Thursday, September 3, 2026

Apple in talks with Khosla Ventures-backed PrismML to shrink AI models for iPhone: Report

Date:

Apple is reportedly evaluating Khoslaventures-backed startup, PrismML’s technology, which the startup says can shrink powerful AI models enough to run directly on an iPhone while using up to 15x less memory, CNBC reported on Tuesday (14 July).

How PrismML claims to shrink AI models

PrismML, which grew out of research at the California Institute of Technology, unveiled compressed versions of Alibaba’s open-source Qwen model on Tuesday. The company said it cut the model’s size from around 54 GB to under 4 GB, enabling all 27 billion of its parameters to operate on an iPhone 15 or newer, according to CNBC.

Chief executive Babak Hassibi told CNBC that Apple and several other firms are currently testing the startup’s models for speed, energy use and overall performance. “They’re really evaluating our technology right now,” Hassibi said. He described the talks as very preliminary but added that “things are progressing nicely.”

Why on-device AI matters for Apple

The development lands a day after Apple opened public beta testing for iOS 27, which includes its long-awaited redesign of Siri.

Apple has been working to make the assistant more competitive with rivals from OpenAI and Anthropic, while keeping as much data and processing as possible on the device itself rather than in the cloud.

Running larger AI models locally could ease one of Apple’s biggest technical constraints, since the most capable systems typically demand more memory and processing power than a smartphone can normally provide. Doing so on-device would cut latency, reduce cloud costs and reinforce Apple’s privacy positioning, while also allowing some features to function offline.

According to CNBC report, PrismML said its method works by simplifying how a model’s internal values are stored, reducing each figure from 16 bits down to as few as one or three possible values. Hassibi compared the approach to the semiconductor industry’s shift from eight-bit to four-bit computing, adding that PrismML “takes it a step further.”

PrismML says its compressed models use up to 15 times less memory, run six to eight times faster and consume up to six times less energy, though Hassibi acknowledged a modest drop in performance, particularly in factual recall rather than reasoning or coding ability.

What it could mean for chip demand

The announcement arrives amid growing debate over whether such efficiency gains might dent demand for memory chips and datacentre hardware. Morgan Stanley has projected Apple’s memory costs could rise sharply in fiscal 2027, potentially pushing iPhone prices higher.

Hassibi said Google’s Gemma model is next for compression, followed eventually by larger frontier models that currently require datacentre-scale hardware. “It’s very important that the intelligence be local and that it can run fast,” he said.

Source link

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Share post:

Subscribe

spot_imgspot_img

Popular

More like this
Related

India urged to review green finance rules, build workforce for nuclear expansion

India needs to review some of its financing rules...

Bank of Japan governor Ueda hints at September rate hike as bets on move mount

Bank of Japan Governor Kazuo Ueda said the bank’s...

RBI appoints Suman Ray as executive director

Ray has also served as Secretary to the Western...

India’s corporate borrowing picks up as capex gains pace: Citi’s Neeraj Kumar

India is seeing a broad-based revival in corporate borrowing,...