Introducing Latam-GPT: Latin America’s Open-Source AI

Priyadharshini S September 02, 2025 | 12:50 PM Technology

“This work cannot be carried out by a single group or country in Latin America; it requires collective participation,” says Álvaro Soto, director of CENIA, in an interview with WIRED en Español. “Latam-GPT is an open, free, and, above all, collaborative AI project. For the past two years, we’ve followed a bottom-up approach, bringing together citizens from various countries who want to contribute. Recently, we’ve also seen more top-down involvement, with governments taking an interest and joining the initiative.”

Figure 1. Latam-GPT: Open-Source AI Powered by Latin America.

The project’s collaborative ethos sets it apart. “We’re not trying to compete with OpenAI, DeepSeek, or Google. Our goal is a model tailored to Latin America and the Caribbean, one that understands the region’s diverse dialects, history, and cultural nuances,” Soto explains. Figure 1 shows Latam-GPT: Open-Source AI Powered by Latin America.

Through 33 strategic partnerships with institutions across Latin America and the Caribbean, the team has compiled a dataset exceeding eight terabytes of text—the equivalent of millions of books. This foundation has enabled the creation of a language model with 50 billion parameters, comparable in scale to GPT-3.5, capable of performing complex tasks like reasoning, translation, and associations.

Latam-GPT is trained on a regional database containing information from 20 Latin American countries and Spain, totaling 2,645,500 documents. Data distribution reflects regional size and digital development, with Brazil leading at 685,000 documents, followed by Mexico (385,000), Spain (325,000), Colombia (220,000), and Argentina (210,000).

“Initially, we’ll launch a language model. Its performance on general tasks should be similar to large commercial models, but it will excel in topics specific to Latin America. When asked about regional issues, its knowledge will be much deeper,” Soto says.

The first model is just the beginning, paving the way for future technologies, including image and video models, and larger-scale systems. “As an open project, we want other institutions to adapt it. For example, a group in Colombia could tailor it for education, or one in Brazil for healthcare. The goal is to enable organizations to develop specialized models for areas like agriculture, culture, and more,” adds the CENIA director.

A key pillar of Latam-GPT is the supercomputing infrastructure at the University of Tarapacá (UTA) in Arica, Chile. With a $10 million investment, the new center features a cluster of 12 nodes, each equipped with eight NVIDIA H200 GPUs. This unprecedented capacity allows large-scale model training in Chile for the first time while promoting decentralization and energy efficiency.

The first version of Latam-GPT is expected to launch this year, with ongoing refinements and expansions as new partners and richer datasets join the effort.

We must prioritize our own needs; we cannot wait for others to consider what matters to us. Since these are new and disruptive technologies, it is crucial that our region harness their benefits while understanding their risks. Gaining firsthand experience is essential to guide technology use along the right path.

This also opens opportunities for local researchers. At present, Latin American academics have limited access to interact deeply with these models. It’s like wanting to study MRI technology without having a resonator. Latam-GPT aims to be that fundamental tool, enabling the scientific community to experiment and advance.

We’ve focused on creating high-quality data. It’s not just about volume—it’s about composition. We analyze regional diversity to ensure no single country dominates the dataset. If, for example, Nicaragua is underrepresented, we actively seek collaborators there.

We also balance topics—politics, sports, art, and more—and, importantly, cultural diversity. In this first version, we’ve emphasized the cultural knowledge of ancestral peoples like the Aztecs and Incas, rather than focusing on the languages themselves. In the future, we plan to incorporate indigenous languages. At CENIA, we are developing translators for Mapuche and Rapa Nui, while other groups are working on Guaraní. This is something we must do ourselves, because no one else will.

Between 2017 and 2018, a group of experts, including myself, developed Chile’s National Artificial Intelligence Policy. One conclusion was the need for an institution to foster a healthy AI ecosystem, combining science, industry technology transfer, and social responsibility. That institution became CENIA.

While it started in Chile, our vision is regional. We believe that together we are stronger. We’ve promoted initiatives like the Latin American Artificial Intelligence Index, a collaborative study that tracks AI progress across the region.

In cognitive robotics, the “cognitive” part refers to intelligence. My career has focused on developing intelligence for physical machines. Today, language models and foundational models are at the forefront of AI—they’re the most powerful tools available. My work is about understanding and contributing to the scientific and applied development of these technologies.

We face many challenges, but also many strengths, such as openness and a strong capacity for collaboration, as demonstrated in Latam-GPT. One key focus is education. These technologies will change the skills young people need. Memorization will matter less; knowing how to use AI-driven knowledge will matter more. We must prepare our youth while also promoting social sciences and critical thinking. If I had to choose where to apply these technologies first, it would be in education, because it addresses the root of many regional challenges.

Reference:

  1. https://www.wired.com/story/latam-gpt-the-free-open-source-and-collaborative-ai-of-latin-america/

Cite this article:

Priyadharshini S (2025), Introducing Latam-GPT: Latin America’s Open-Source AI, AnaTechMaz, pp. 128

Recent Post

Blog Archive