GENIAC

Challengers Taking on Generative AI to Shape the Future

2025/07/04

GENIAC brings together selected companies that are taking on the challenge of advancing domestically developed generative AI. What kind of individuals are the key drivers behind these efforts? In this article, we introduce Mr. Quan Kong of Woven by Toyota, Inc., who is leading the development of “City-LLM,” a foundation model aimed at understanding “time” and “space” at a real-world urban scale. We spoke with Mr. Kong, who works in Japan far from his hometown in China, about his journey and his vision for developing AI that supports people’s everyday lives.


Profile

Quan Kong
Born in 1987 in Xi’an, China. Staff Research Scientist at Woven by Toyota, Inc. After graduating from Xi’an Jiaotong University in 2011, he obtained a Ph.D. from the Graduate School of Information Science and Technology at Osaka University. In 2016, he joined the Central Research Laboratory of Hitachi, Ltd., where he worked in a division related to media information processing. Since 2022, he has worked in his current role, engaging in research on computer vision and multimodal understanding.


Inspired by Japan’s Gadget Culture, He Chose the Path of a Developer

──You are originally from China. What first sparked your interest in information science?

Kong: Yes, I was born and raised in Xi’an, in Shaanxi Province, China, and I spent my undergraduate years there. I have loved science fiction and gadgets since childhood, and a turning point came when my parents gave me a Walkman when I was in elementary school. That sparked my interest in machines and engineering. There was also a shopping district in my hometown with a vibe similar to Akihabara, where I enjoyed collecting and exploring various devices.

Xi’an Jiaotong University, where I enrolled, had strong programs in communications and electronics, but at the time, interest in information science was rapidly growing. I became particularly interested in areas such as web application development and data mining, and wanted to gain a deeper understanding of how these systems work, which led me to choose an information science major.

During my studies, I worked not only on software but also on projects combining hardware and software. For example, around 2010, I quickly obtained Microsoft’s Kinect with a friend and developed an application that enabled touch-like interaction on a projected screen by combining it with a projector. This project was highly regarded, and I received a top award at my university.

──You later pursued graduate studies at Osaka University and earned your Ph.D. Why did you choose Japan, and Osaka University in particular?

Kong: Universities in Japan were conducting a great deal of interesting research in areas I was interested in, such as human-computer interaction (HCI). That motivated me to see and experience it firsthand.

At the time, I did not have much information about Japanese universities, so I searched using keywords such as “multimedia” and “data mining,” and decided to pursue graduate studies at Osaka University, where research in these areas was particularly active.

──Could you briefly describe your research during your doctoral studies?

Kong: I was affiliated with the laboratories of Professor Takuya Maekawa and Professor Yasuyuki Matsushita, where I worked mainly on wearable computing and ubiquitous computing. For example, by using sensors embedded in devices such as smartwatches and smart glasses, we could capture human behavior and location data and use it to provide seamless services in everyday environments.

A key concept was enabling services to be delivered automatically in a “passive” manner, without requiring users to consciously operate systems. Based on behavioral patterns and location data, we worked on designing and evaluating systems that could deliver optimal user experiences. Through this research, I was able to steadily build a track record, including publishing papers at leading conferences such as Ubicomp on a continuous basis.

──I understand that after earning your Ph.D., you joined a research institute at Hitachi. What led you to choose a manufacturer?

Kong: As I mentioned earlier, many products created in Japan from the 1990s through the 2000s—such as the Walkman, PlayStation, and Vocaloid—have captivated young people in China. I was strongly interested in the “source of innovation” behind these products, namely Japan’s unique culture and development environments.

While I was able to experience some of this through academic research, I wanted to enter a corporate setting to better understand how ideas are formed and how products are actually developed. That led me to join Hitachi’s R&D division, where I worked on developing multimedia technologies by leveraging my research background.

──Could you share, to the extent possible, the projects you worked on at Hitachi?

Kong: In the latter half of my doctoral studies, I worked on a project using Google Glass to control home appliances. The system recognized appliances using the camera on the Glass device and enabled users to control them through their gaze. Through this research, I applied deep learning to image recognition and was amazed by its accuracy, which made me strongly aware of its potential.

At Hitachi, I belonged to a research division that handled multimodal data such as images, audio, and text. I was involved in research on cutting-edge technologies, including video understanding using deep learning, applications in the security domain, and self-supervised learning.

Developing a Foundation Model for Spatiotemporal Understanding of Cities

──What led you to move to Woven by Toyota in 2022?

Kong: At Hitachi, I had the opportunity to work broadly on research related to video understanding, but much of it was B2B-oriented. Over time, I developed a stronger desire to deliver results closer to end users.

Around that time, the concept of “Woven City” was announced at CES, and I felt it was a project closely connected to people’s everyday lives. I saw it as an opportunity to utilize data in real-world living environments, which was very appealing to me.

──At that time, Woven City was still largely at the conceptual stage. Were you also drawn to it because of an interest in mobility?

Kong: Yes. I consider mobility to be part of the broader “environment.” Since my university days, I have consistently worked on research related to smart homes and smart environments, and I see mobility as an extension of that.

Beyond autonomous driving, I believe there is still significant potential in areas such as in-vehicle experiences and information integration. The opportunity to explore the fusion of daily life and technology at the scale of an entire city was extremely appealing to me.

──Could you tell us about the projects you are currently working on?

Kong: In our division, multiple teams collaborate across a wide range of activities—from research and development to engineering—including foundation model development, teams working on enabling future AI services to be easily deployed via urban cloud infrastructure, and teams conducting applied research in machine learning and computer vision.

Within this framework, I am leading the research and development of a multimodal foundation model called “City-LLM” for spatiotemporal understanding of urban environments.

Specifically, we are building a model that integrates multiple modalities such as video, images, and language to understand spatiotemporal information in urban settings. For example, the goal is to capture traffic conditions, pedestrian movements, and patterns of space utilization in real time using cameras and various sensors, and to enable services that respond accordingly.

──So your AI research goes beyond vehicles to encompass infrastructure as a whole?

Kong: Exactly. For instance, achieving “traffic safety” requires close coordination between vehicles and infrastructure. It is not sufficient to rely only on cameras and sensors installed in vehicles; it is also necessary to integrate with infrastructure-side systems, such as cameras at intersections, to analyze pedestrian behavior and provide appropriate information to drivers in real time.

The foundation model we are currently developing is intended to serve as a core component supporting such integration between mobility and infrastructure.

──Can this be considered a kind of “world model”?

Kong: Not exactly. Rather than aiming to “understand everything,” as in a general world model, we place greater emphasis on building models tailored to the real-world domain of cities.

A particularly important aspect is the understanding of “spatiotemporal dynamics.” While many large language models (LLMs) and multimodal models are strong in spatial understanding, accurately handling temporal changes and sequences remains a challenge. For example, determining whether a car is turning left or right can be difficult from a single frame; it requires understanding temporal context from a sequence of frames.

──Indeed, handling “time” is a major challenge for generative AI.

Kong: Exactly. That is why we design our systems with temporal information in mind from the data collection stage. In addition to still images, we incorporate video data and architectures that can accommodate time-series sensor data in the future.

We also start with a relatively manageable model size of around 7–8B parameters, allowing us to iterate quickly through trial and error. This approach also enables flexible deployment to edge devices and cloud environments in the future.

Tackling Societal Challenges Through City-Scale AI Demonstrations

──From the perspective of an AI developer, what do you see as Woven by Toyota’s key strengths?

Kong: I believe the greatest strength lies in the ability to use not only digital tools but also a real urban environment like “Woven City” as a kind of testbed. In other words, we can conduct research and development based not only on data from the internet, but also on real-world data—such as human movement patterns, vehicle behavior, and environmental changes.

In addition, it is not just about collecting data; the infrastructure is also in place to implement and validate technologies using that data. Another key strength is the promotion of open innovation through collaboration with both internal and external partners.

──So participation in GENIAC is part of that approach?

Kong: Yes. We are working on the development of a foundation model called “City-LLM” under GENIAC, and we find it highly valuable to collaborate with leading AI developers and share expertise through the GENIAC development community, which is one of the largest of its kind in Japan.

For example, participating companies bring diverse strengths, such as applying multimodal LLMs to autonomous driving or building large language models from scratch. We are inspired by these efforts and continue to pursue our own approach while taking on new challenges.

──How do you view Japan’s AI development community compared to those overseas?

Kong: This is just my personal impression, but I feel that Japan’s development community has a real sense of speed. It is often said that overseas markets are more active in investing in new trends, while Japan is a more mature market.

However, from my perspective, it is precisely because of Japan’s unique culture and environment that developers are taking on challenges from different perspectives. Rather than simply following approaches already tried by global big tech companies, there is potential to create unique generative AI by building on data and cultural contexts that are specific to Japan. I find these kinds of challenges both meaningful and exciting.

I believe that global big tech companies are also likely to be interested in generative AI developed from Japanese data. That is why I hope that AI developers in Japan will continue to value their own uniqueness and perspectives.

──City-LLM aims to utilize data at an urban scale. What kind of future do you envision?

Kong: What we aim to achieve is a “searchable world” in real urban environments. For example, if someone asks, “Which cafés have available seating right now?”, AI could generate answers in real time based on data from cameras and sensors across the city. This is something that has been difficult to achieve with conventional web search.

Of course, realizing such AI requires careful consideration of privacy and safety in the design of sensors and infrastructure. One of the key strengths of our project is that it goes beyond technological development—we can use the city itself as a testbed to validate functions and business models.

An environment where data can be utilized and algorithms can be tested at a city scale is extremely rare even globally. We aim to achieve outcomes that directly address real-world challenges, such as improving traffic safety and optimizing services through integration with infrastructure.

──Finally, both as an individual and as a member of Woven by Toyota, what kind of challenges would you like to continue taking on with AI going forward?

Kong: Until now, generative AI has often been used as a tool to improve human efficiency. However, what we truly aim for is the creation of new value. Rather than competing for market share within existing markets, we want to create entirely new markets and ways of living. In that sense, I believe AI should be a “tool for exploration.”

To achieve this, we will continue working with real-world data at the scale of entire cities to create new experiences. Believing in that future, I hope to continue taking on new challenges.

GENIAC Top Page
Back to page top