Notes
VI. Coding Agents and Second Reality
Keep calm and be positive. Let music and coffee guide you to the future.
Coding Agents and Second Reality
Autoregressive-transformer LLMs (AT-LLMs) is the start of the bootstrap process to singularity. It’s funny how a mere automatic stream of token generation can be so… idk… going from a Euclidean space to a non-Euclidean space?? Going back to the first note, I made an assumption that in our lifetime, we will reach singularity (of some form) and we will want to integrate ourselves with the silicons. We will go off of that assumption. Beyond that, it’s pointless to think about, so let’s walk back.
First, we will see the collapse of the cyberspace (e.g., dead internet theory.) The human interaction of cyber objects (e.g, virtual objects) cease to be meaningful, except for the art of interacting the cyber objects themselves. What is inherently machine must return to machine. We will see the short return of the reality, albeit a different reality (but reality nonetheless.) Covid was an interesting time, what I thought would be fun (e.g., online schooling) was totally opposite. Eventually, apps like BeReal became popular, so it wasn’t just me who want the reality back. AT-LLMs just adds salt to an open wound—the already distaste for the cyberspace.
In this second reality, AT-LLMs will automate the implementation process, the engineering process (with human guidance), the research process (with human guidance), the commercial writing process (with human guidance), etc. Namely, AT-LLMs, with human guidance, will greatly accelerate the engineering and research process, augmenting itself with visual reasoning, spatial reasoning, and verbal reasoning on top of existing logical/natural-text reasoning to find some suitable architecture X beyond the standard AT-LLMs. An architecture that supports complete replication of human sensory experience that will bring X to singularity. We will not realize it first, but the transition to the next reality, the third reality, would be a quick one.
These are mere abstraction where the borders between stages of reality is very blurred and complicated, filled with nuances. And, these thoughts are practical since it represents what will likely happen during our (my) lifetime.
Cyberspace point-of-view:
- 1st reality: the creation of cyberspace
- 2nd reality: the collapse of cyberspace
- 3rd reality: the biotech integration of cyberspace
Reality point-of-view:
- 1st reality: the escape from reality
- 2nd reality: the return to reality
- 3rd reality: the biotech integration of reality
This is still too abstract. A better way to think about this is cyberspace is constructed from software, firmware, and relevant hardware. As the cyberspace grows and surround our life, so, too, the people who handle the software, firmware, and hardware—1st reality. With AT-LLMs, the entirety of the cyberspace could easily just be a mix of automated self-prompting (it’s alive) or human-guided prompting—the beginning of the 2nd reality. Any cyberspace-derived input or output (i.e., excluding input or output from the reality that includes the world itself and us) would be simplified into a long chains of virtual data passing (e.g., nested function calling at an internet scale.) 2nd reality is defined by a time where cyberspace <-> reality interfaces are expanded to accomodate the accelerated growing complexity and flexibility of the cyberspace. Once the cyberspace of the 2nd reality discovers singularity, assuming we would even know about it by then, would be followed by debates, AI war (???), etc. that would end with the simplification of cyberspace <-> reality interface. In other words, cyberspace and reality into a single, unified space.
Let’s focus on the cyberspace <-> reality interface. We have only talked about AT-LLMs since that’s currently the best modality and architecture current “AI” has (on top of accessibility, integrability, and other factors.) DiT-LIMs (Diffusion transformer and image generators as large image models) or CLIP-VLMs haven’t really reached the quality of AT-LLMs since the problem in vision is higher-dimensional in every way possible. We have even yet to talk about generalized spatial understanding, and VGGT(-Omega) does not count. Because of the lack of scarily good VLMs, let alone spatial understanding, the following interface is safe for now to explore:
- Inverse rendering (reality -> cyberspace)
- Forward rendering (cyberspace -> reality)
- Simulation (cyberspace -> cyberspace)
- Geometric structures (c -> c, r -> c, c -> r)
- and other of these natures
But the following is gone:
- Linear line-by-line coding (not like spatial coding exists?)
- Paper writing, text-based peer-review process, text-based literature review
- Zotero is the next thing that would be gone, or at least adapt to a highly automated workflow
Overall, cyberspace would be more visual and spatial oriented (for now.) Any large text would just be gone, handled by AT-LLMs. Emphasis on any. Maybe because it’s the curse of dimensionality? However, progress is exponential. So, we can expect linear progress towards these higher-dimensional modalities if this is even useful insight.
Future Work
This was a revisit to the first note. I’m satisfied with the conclusion through this excursion: cyberspace will be more visual and spatial oriented going forward as textual objects are mostly automated. Other non-text modalities such as graphs, temporal dimension, and other could further be investigated and whether they relate with the visual/spatial paradigm in the context of AT-LLMs and cyberspace. Third reality and biotech could further be explored. Despite my belief that it will likely occur in our lifetime, it is still too far to be meaningful—yet.