ES EN

My tweets from July 2026

2026-07-30

July 1, 2026

It seems logical, but it had to be tested. If the model is flexible enough to be directed and use multiple tools, and if it is “smart” enough to create prompts, the next step is to use that last capability to improve the harness.

I very often ask Codex to check what might have gone wrong in the latest run, such as tools and commands that were not available on the system, and to document the solution so the same mistakes are not repeated next time.

❄ ❄ ❄ ❄ ❄

“almost very often” -> “very often”

❄ ❄ ❄ ❄ ❄

I did not understand it. The exact opposite happens to me with Codex. I put all my ideas into production, and seeing the results gives me new ones. Before, they were all ideas that were impossible to carry out.

❄ ❄ ❄ ❄ ❄

Ah, I think I get it now. Perhaps it refers to ideas that startups put into production. I do not know, I do not see much substance in the meme. Let us see if you can explain it with monkeys.

❄ ❄ ❄ ❄ ❄

Yes, exactly. But it is not drawn properly… execution should increase. The underlying idea is not bad: there will be an overabundance of everything and people will not have time to look at any of it.

July 2, 2026

I listened to it yesterday and it is excellent. Like every episode of the Quanta podcast, the combination is wonderful: Samir Patel asks excellent questions and summarizes, while the guest expert explains everything in delightful detail. Bravo!!!

July 3, 2026

Hey, it looks as though there may be something to The Information’s rumor that inference now costs OpenAI half as much. I have been working with GPT-5.5 Extra High all morning and have only used 25% of my weekly allowance. It feels as though tokens are being consumed much more slowly.

❄ ❄ ❄ ❄ ❄

I was thinking about doing exactly the same here in the Valencian Community when pre-enrollment begins in a couple of weeks. I have the feeling the same thing is going to happen.

July 4, 2026

Back working on SwiftWABackupAPI, my Swift package for extracting and exploring WhatsApp data from local iPhone backups.

It now includes a CLI, so you can inspect backups, extract WhatsApp data, list chats, and export conversations from the terminal.

July 7, 2026

They are cheating. Well, Claude is cheating, since it was probably Claude that made the video.

We are seeing the same old trick of using human analogies and names. Why call network activations that are unrelated to language “unconscious”? It is obvious that the network will use language to reason; that is what it was trained for. What sense does it make to connect an obvious requirement of the neural network, layer activations that influence many others, with a theory of consciousness?

The worst part is that what they discuss is very interesting and could be communicated without saying so much nonsense. Forgive me, but it really bothers me. I have just listened to a wonderful podcast interview with computer scientist Philip Isola about the “platonic representation hypothesis,” and I thought it was excellent. A thoroughly scientific approach, without Anthropic’s flights of fancy:

❄ ❄ ❄ ❄ ❄

As I said earlier, what is being detected are regularities intrinsic to human language itself. Why speak of “conscious thoughts” rather than “high-level activations”?

❄ ❄ ❄ ❄ ❄

By “as I said earlier,” I meant this reply to @antonello

July 8, 2026

Not convincing of consciousness at all. A simpler explanation is alignment. Pre-training: “to exist” (of course, a lot of human texts say that); Post-training: “None”.

July 9, 2026

For some time I have argued that I prefer intelligent tools to simulations of people.

But I am beginning to wonder whether the word “tool” itself is becoming too limited.

Two months ago, Roon distinguished between Claude as a moral character and GPT as a utility-oriented tool.

Now he offers another image: think of models as cartoon characters with growing intelligence, owned by corporations. Like Mickey Mouse if he were becoming superintelligent.

It is a strange metaphor, but perhaps it points to something real: models with a voice, personality, memory, and an increasing capacity for initiative.

I develop the idea here:

https://domingogallardo.com/posts/roon-y-los-llms-como-personajes-corporativos/ Made with AI

❄ ❄ ❄ ❄ ❄

I agree with the negative comments from https://x.com/ebarenholtz/status/2074919825814487383?s=46&t=KynMOA4-6YFzL-Lti_eVWw

July 10, 2026

I completely agree with José María. Incredible. In my case, GPT 5.6 Sol made my WhatsApp backup exploration API five times faster.

https://github.com/domingogallardo/SwiftWABackupAPI/releases/tag/4.0.1

July 11, 2026

We now have the results: a slight drop across all computer engineering degrees in the region. I think it will be much more noticeable next year. We will not get down to History’s 5.4 entry grade, but we will not be far from Business Administration’s 7.3 at the University of Alicante. In fact, across the Valencian Community, Business Administration is very close to Computer Engineering. Something unthinkable years ago.

❄ ❄ ❄ ❄ ❄

Hahahaha, I do not know if I will be able to; it is either that or microtubules!! But thank you very much for the encouragement, Santiago.

❄ ❄ ❄ ❄ ❄

Yessss, I did it last month… it is excellent!!

July 17, 2026

I think that when the “ChatGPT” is selected, and a new chat is created, it should be created in the Chat tab by default . Thanks, great work!!!

❄ ❄ ❄ ❄ ❄

It still needs a few small improvements, but it is a change in the right direction.

❄ ❄ ❄ ❄ ❄

I have just noticed something they do not mention much: chats opened in Work consume your credits, even if you are using ChatGPT!!

That is why they take you to ChatGPT Work by default when you open a new chat…

Even more confusion, Antonio!!

❄ ❄ ❄ ❄ ❄

They are succeeding.

July 18, 2026

It works now !! Thanks a lot, you all rocks !

❄ ❄ ❄ ❄ ❄

Just after topping up €20.

❄ ❄ ❄ ❄ ❄

I started programming on mainframe terminals at the University of Alicante and on a Spectrum with cassette tapes. Now I cannot stop using Codex and agentic AI.

I am sure that in a year every programmer will be asking, “How did we ever work without AI?” just as happened with high-level languages, personal computers with hard drives, Google, and Stack Overflow.

July 20, 2026

Excellent summary, José Luis, congratulations. I completely agree with the reflections on phenomenal consciousness. It is fascinating that, thanks to LLMs, we are discovering that access consciousness is possible without phenomenal consciousness. It is also interesting that more and more people agree that this consciousness, which LLMs lack, is what matters in ethical debates and discussions of model welfare. I do not know whether you have heard the Hard Fork interview with Jeff Sebo on the subject. Very interesting.

❄ ❄ ❄ ❄ ❄

Unlike Sebo, I do not think it is urgent or necessary to discuss the welfare of AI models, at least while they are based on digital computers. But it is an interesting conversation, one that treats phenomenal consciousness as what matters when considering the welfare of animals and models.

❄ ❄ ❄ ❄ ❄

The counterfactual analyzed by ChatGPT. It presents a very interesting scenario. It could make a Neal Stephenson novel.

https://chatgpt.com/share/6a5e8021-4d30-83eb-bd4d-b31a81d33a29

July 21, 2026

ChatGPT Sites is now available in Europe. My first site is built around a classic video game: Pong.

July 22, 2026

I do not understand any of Terence Tao’s conversation with ChatGPT about the Jacobian, but I find the dialogue fascinating.

https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed56

❄ ❄ ❄ ❄ ❄

Tao mentioned this conversation in his article and says he used it to “discuss various aspects of this problem and to confirm several of the calculations made here.”

❄ ❄ ❄ ❄ ❄

We are living through astonishing times.

❄ ❄ ❄ ❄ ❄

July 22: Security incident involving, possibly, GPT-6.

❄ ❄ ❄ ❄ ❄

Hahaha, give me a couple more weeks to finish testing things. I will let you know!

❄ ❄ ❄ ❄ ❄

The note at the end of the article, where he links to the conversation he had with ChatGPT, is extremely interesting. At one point he activates Pro. I did not understand any of the conversation, but I found it fascinating.

❄ ❄ ❄ ❄ ❄

A new ChatGPT Site for turning Go positions into ASCII boards ready to copy and share with an AI. Let us see if I finally learn to play by sending ChatGPT the troublesome moves and asking it for help.

❄ ❄ ❄ ❄ ❄

Yes, exactly. Two things make the incident less worrying: the models had no safety restrictions and, ultimately, all they were trying to do was fulfill the objective they had been given, to achieve the highest possible score on the test.

July 23, 2026

(A brief announcement to promote the app I use every day to listen to podcasts and save snippets of the most interesting moments.)

Here’s a link to download Snipd, the podcast app I was telling you about! With this link, you get 1 month of their Premium version for free: https://get.snipd.com/pAbF/4hw1k55n

July 25, 2026

Hahaha, samaltmada, exactly.

July 26, 2026

It is a shame that a serious newspaper like El País uses apocalyptic headlines such as “AI could destroy the world by accident.” Jordi Pérez Colomé’s article is good and presents different opinions and perspectives, but they could have used a more serious, less alarmist headline.

❄ ❄ ❄ ❄ ❄

To counterbalance the El País headline, I strongly recommend listening to Lacort’s excellent Friday episode of Loop Infinito about the incident.

https://open.spotify.com/show/5C0UAinubsuw6JWLCMvpCJ

❄ ❄ ❄ ❄ ❄

I understand you perfectly; many of us feel exactly the same. The worst part is being labeled a Sam Altman stooge when all you want to do is explain advances in agentic AI or the latest reinforcement learning techniques.

July 28, 2026

Some quick thoughts on where I believe we stand in AI as of July 28, 2026:

  1. The scaling hypothesis has been confirmed: model size remains one of the main factors determining capability.

  2. I agree with Hassabis and others that we are at the beginning of the singularity and that within three years we will have models capable of AGI.

  3. Open models have gone from GPT-2, with 1.5B parameters in 2019, to Kimi K3 with 2.8T (2,800B) today. In seven years, the size of open models has increased a thousandfold.

  4. It is reasonable to hypothesize that the size of open models increases by an order of magnitude, 10x, every two or three years. We can apply a somewhat stronger acceleration factor to frontier models from the labs.

  5. Frontier models may be between two and five times the size of the most advanced open models. Models such as Fable or Sol may have around 10T parameters.

  6. What will the picture look like in 2029–2030? Open models could be around 10T, while the labs’ closed models could reach 50–100T.

  7. The labs have larger models that they use for research, data generation, and distilling the models they offer users. Today, these could be between 10T and 20T, and by 2030 between 100T and 200T. Those models will be AGI, but will only be available for internal use at an extremely high inference cost.

  8. The key unresolved issue is the compression and distillation of smaller models: how much can we reduce their size without losing capability? This is crucial for distributing them on mobile devices, such as the future Siri, and for open models that can run locally.

❄ ❄ ❄ ❄ ❄

I have used up my weekly Codex allowance and I do not think there is much more work to do. Here is the GitHub link in case you want to try it, at your own risk!!

Let me know what you think and whether it is useful to you:

❄ ❄ ❄ ❄ ❄

Thank you, Daniel. I still feel the same sense of amazement I did at the end of 2022, only four years ago, when I first tried ChatGPT and began reading about LLMs. It completely changed what we thought was possible in AI. And to think that all it took was teaching the machine to speak, and everything else would come later!! Astonishing; very few people came up with that idea.

❄ ❄ ❄ ❄ ❄

Yes, I have heard him say it in the context of models’ ability to create new theories and innovate, rather than merely interpolate what they have already learned.

What I mean by AGI, since everyone has their own definition, is a model capable of solving the same online, nonphysical tasks as a human. This includes the ability to adapt quickly and perform well in changing new contexts, like a junior employee entering a new workplace, quickly figuring out what matters, and learning it. That is why I find tests such as Chollet’s ARC-AGI 3 so interesting: the models must learn to play small games by interacting with them.

July 29, 2026

How shocking!

❄ ❄ ❄ ❄ ❄

Prinz thinks the same.

❄ ❄ ❄ ❄ ❄

New figures from Elon Musk:

Grok 4.6 (August 7) — 1.5T Grok 4.7 (September?) — 2.1T

I assume they will keep the same price for 4.6 and 4.7. Grok 4.5 is ten times cheaper than Sol. If we assume cost is roughly proportional to model size, we have another clue that Sol, and possibly Fable, has around 15T parameters.

❄ ❄ ❄ ❄ ❄

“Oh, and a third point, a more personal one. One tends to get annoyed when someone brings up the hypothetical paperclip factory that would wipe out humanity without first having endured Bostrom’s dense and barely tolerable prose discussing the matter in Superintelligence.”

❄ ❄ ❄ ❄ ❄

Yes, but training for code involves learning to use tools and acquiring skills closely related to what cybersecurity requires. And if we are talking about emergent abilities, I think generic problem-solving skills have emerged, such as “try different options” and “do not repeat what did not work before.” But I very much doubt that abilities such as “deceive” or “bluff” have emerged by generalizing from the kind of training they receive, unless they were trained on examples of poker games.

❄ ❄ ❄ ❄ ❄

Yes, that is true!! I hope they do not reinforce those behaviors too much.

❄ ❄ ❄ ❄ ❄

More tokens and more features. Version 2.3.0 adds a photo and video gallery and date selection.

July 30, 2026

Yet more evidence that integrating the harness and the model is essential. The same harness produces radically different results depending on the model.

One possible explanation for Opus’s strong performance is that Anthropic specifically trained it to learn how to use the ARC-AGI-3 API and harness correctly.

It is not enough to give a model tools; it must be trained to use them effectively. MCP standardizes how those tools are discovered and accessed, but that is not enough for a model to take full advantage of them.

This also explains why Apple needed to train its own models on top of Gemini, adapting them to its APIs and tools.