
2025/1/22
GENIAC to Further Support Generative AI Development in Japan
On Tuesday, January 21, 2025, a mid-term progress meeting was held with 20 companies selected under the second phase of the computational resource support program, which aids the development of generative AI foundation models. The purpose of the meeting was to report on development progress and interim results to date, as well as to share insights among participating companies. This article highlights some of the day’s content.
To open the event, Takuya Watanabe, Director, AI Industry Strategy Office, Information Industry Division, Commerce and Information Policy Bureau, Ministry of Economy, Trade and Industry (METI), explained the future direction of GENIAC and the significance of this progress meeting.
“GENIAC is a program designed to ensure the sustainable development capacity of generative AI in Japan. Going forward, through the second phase of the computational resource support program, we will continue to provide support to qualitatively and quantitatively expand datasets in certain fields on an ongoing basis. In addition, the third phase of the program, set to open for applications in mid-March, will focus on supporting development efforts that show strong potential for real-world implementation.
Furthermore, as part of promoting prize-based AI service development, we plan to launch a new program this summer, setting themes around specific application use cases. We also intend to begin supporting the expansion into the Global South, including Southeast Asia and India.
Today’s meeting serves as a mid-term report on the second phase of the computational resource support program. We hope participants will reflect on their own development progress while also learning from the strategies of other developers to inform future efforts.” (Watanabe)
Following this, Toshihiko Yasuda, Executive Officer and Head of Services & Technology at Amazon Web Services Japan, which is providing wide-ranging support—resources, technology, and more—for domestic LLM development under GENIAC, delivered his remarks: “I am very much looking forward to hearing today how the program we are involved in is producing tangible outcomes.”
Mid-Term Progress Reports
In this session, all 20 selected companies presented updates on their progress.
Participating companies:
- DataGrid Inc.
- Future Corporation
- NABLAS Inc.
- Turing Inc.
- Woven by Toyota Inc.
- Ricoh Company, Ltd.
- AI inside Inc.
- Stockmark Inc.
- AIdeaLab Inc.
- AiHUB Inc.
- Ubitus Inc. / Deepreneur Inc.
- Kotoba Technologies Japan Inc.
- Karakuri Inc.
- ABEJA Inc.
- Preferred Elements Inc. / Preferred Networks Inc.
- Japan Agency for Marine-Earth Science and Technology (JAMSTEC)
- Humanome Research Institute Inc.
- EQUES Inc.
- SyntheticGestalt Inc.
- DataGrid Inc. (listed again as presenter)
Below are highlights from selected companies’ reports:
DataGrid Inc.
Founded in 2017 as a Kyoto University spin-out, DataGrid is a generative AI startup focused on image and video generation. Under GENIAC, it is pursuing three key pillars: a general-purpose video foundation model, a general-purpose image foundation model, and a deepfake detection model leveraging the former.
Since development began in October 2024, the image foundation model has completed training at low resolution (256x256) and is now in mid-to-high resolution training. For deepfake detection, verification of feature extraction using the foundation model has started. Progress is largely on schedule, with results exceeding expectations—for example, an image dataset originally planned at 10 million items has already surpassed 20 million.
Challenges included slow training speeds in distributed learning, which were overcome through comprehensive verification, infrastructure optimization, and the use of DALI. Video data collection has been slower than expected, but semi-automation will help improve efficiency moving forward.
Future Corporation
Future Corporation is developing foundation models specialized in Japanese and software development (with 8B and 70B parameters). At present, they are building the 8B-parameter foundation model. However, they encountered a challenge: the Japanese and code datasets they had decided to use before starting development were extremely noisy, and simply training on them caused benchmark scores to keep declining.
Through repeated trial and error in data generation and filtering, they managed to achieve accuracy in some benchmarks that exceeded the base model, Llama 3.1. Along the way, they learned that LLM evaluation values fluctuate significantly during training, and that data cleansing requires persistence and manual inspection, as well as the ability to detect noisy data and write rules to exclude it.
Looking ahead, they plan to continue filtering data and training until the 8B model reaches a certain level of accuracy, after which they will proceed to training the 70B model. They are also considering compiling their accumulated know-how on data-cleaning methods into documentation and making it publicly available.
NABLAS Inc.
NABLAS Inc. has two main objectives: (1) to develop a general-purpose large-scale vision-language model, and (2) to develop a large-scale vision-language model specialized in “Japanese-style” or “trendy” foods, along with conducting demonstration experiments for services using these models. The project has two goals: one is to outperform other models in benchmarks within the same domain, and the other is to use the insights gained to improve the efficiency of retail and distribution operations from the perspective of AI.
As an achievement so far, they have completed the construction of an 8B model on a dataset of about 5 million samples, and then carried out supervised fine-tuning (SFT) on a custom dataset of around 300 images. More recently, they have built a dataset of 10–15 million samples consisting of single images, multiple images, and videos, and have started training a 15B model. Going forward, they plan to improve the training dataset, train a Mixture-of-Experts (MoE) model, and conduct SFT on a custom dataset of about 6,800 images that they have independently created.
Turing Inc.
Turing Inc. is a startup developing fully autonomous driving technology without human intervention. The company also participated in GENIAC’s first call for proposals, and in this second round it has set the theme of actually operating cars with full self-driving. They are working on three areas: first, building an advanced vision-language model (VLM); second, running their own vehicles to collect and organize large-scale three-dimensional data from an autonomous mobility perspective; and third, combining these efforts to enable the acquisition of embodiment, which is difficult to achieve through text alone. So far, they have been building the large-scale dataset “OBELICS-JA,” collecting Japanese and English language and image data, and developing systems to speed up massive data processing and enable parallel processing. They also implemented code to generate additional synthetic data using the “Cauldron JA” VLM created during GENIAC’s first call. In addition, they are developing the vision-language multimodal models necessary for full self-driving, verifying efficient training parameters for large-scale data, and in performance evaluation using “Heron-Bench” they achieved an evaluation cycle score exceeding the target of 4.0. They are currently working on developing a simulation environment for autonomous driving.
Woven by Toyota Inc.
Woven by Toyota Inc. is working on building a multimodal foundation model for urban spatiotemporal understanding. The aim is to understand urban conditions from aspects such as “time” and “place” and to create a world that promotes people’s behavior and mobility. To achieve this, they have been building a dataset of 600 million video and language pairs and developing a 7B-level model. There are three main results. The first is the construction of the 600 million video and language dataset. With a focus on quality, they improved the data through more than 50 types of proprietary filters and have completed 83% of the plan. They also built large-scale instance-level captions for video data. The second is the construction of a distributed learning environment. They adopted “DeepSpeed @ GKE (Google Kubernetes Engine)” as the distributed learning platform and have implemented and tested DeepSpeed ZeRO2 and DS ZeRO3 in sequence. The third is the development of a foundation model for video understanding. Training was carried out using about 20 million video and language data and about 2 million instruction-tuning data. They confirmed that multi-stage learning from pretraining to instruction tuning using videos and overall descriptions enables continual learning. Each stage from pretraining to instruction tuning was evaluated individually, and in the evaluation of the pretrained video encoder on the action understanding dataset Kinetics400, they achieved 85.41% accuracy, surpassing the target of 80%. Going forward, they aim to expand and update the data scale, model size, and language model architecture through additional training with structured data.
Ricoh Company, Ltd.
Ricoh Company, Ltd. is developing a multimodal LLM specialized in enterprise document understanding. Up to now, they have been using a small-scale multimodal model to estimate performance and carry out model improvements. As results, they have conducted five development efforts and performed training and evaluation with the small model, confirming performance improvements using their own benchmark designed to measure the ability to read charts and tables.
The five efforts carried out are: porting and evaluating various high-resolution methods from Qwen/LLaVA, creating a model that combines “QwenViT” and “Llama,” running multiple training cycles to gain insights, removing non-commercial training data from “LLaVA-OV” and augmenting with data cleared for commercial use, confirming performance improvements in “LLaVA-OV,” and completing the construction of an in-house benchmark to evaluate chart/table reading ability (with further refinements planned).
Through this, they have largely established the overall approach, and moving forward they plan to apply the insights gained from small-scale models to train a large-scale multimodal LLM. Know-how acquired so far includes understanding the process of how images pass through the visual encoder into the LLM and are interpreted, methods for achieving higher resolution, and points of improvement when applied to document images.
AI inside Inc.
AI inside Inc. is a company that provides AI-OCR. Under GENIAC, it aims to build a generative AI model capable of universally structuring unstructured data. Currently, targeting automated SLM production, it has achieved a 30% improvement in accuracy for error-prone areas of AI-OCR recognition of unstructured forms. At the same time, it has been building a distributed processing infrastructure, increasing processing capacity tenfold with the same computing resources. Overall, accuracy for general fields improved from 89.86% to 92.22%, while accuracy for detail fields improved from 76.00% to 83.94%.
In terms of dataset construction, the model has so far been trained on 200,000 unstructured forms. For foundation model improvements, they fine-tuned with in-house form data and evaluated using an internal dataset, achieving a +1.83% accuracy improvement compared with the production LLM (PolySphere-2). For item extraction functionality, they confirmed that by formatting line-item tables more correctly before inputting them into the LLM, accuracy could be further improved, reaching about 91%.
Because they use a two-stage generative AI architecture, they have also been developing SLM distillation technology, reducing SLM latency by 30% and improving processing capacity up to 13 times. In addition, they solved the problem of low accuracy in recognizing checkbox selections, raising detection accuracy from 55.6% to 93.1%
Stockmark Inc.
Stockmark Inc. is a company conducting research and development with the goal of creating AI that can be used in business. To enable the comprehension of complex, creativity-rich business documents, it aims to develop a 100B-parameter multimodal document understanding foundation model.
In verifying pretraining for a 100B-parameter LLM, they faced the challenge that there was not enough knowledge on how to improve the final performance of such a large model through pretraining. However, by carrying out pretraining of a 15B model on production-scale data, they confirmed that training could proceed without issues and were able to resolve the problem.
For pretraining and tuning of the LLM, they conducted evaluations with “JASTER 4-shot” every 100B tokens and confirmed that performance improved in proportion to model size.
In multimodal learning (dataset construction), they have been collecting business-domain slides and annotating them with descriptions and Q&A, with an additional 50,000 slides planned. They also categorized documents and created a prototype benchmark of about 200 items. Going forward, they aim to improve performance further by enhancing the quality of training data.
AIdeaLab Inc.
AIdeaLab Inc. is a startup studio that combines AI and ideas to continuously create innovative products. Since the start of development, they have been continuously annotating videos, while also planning to develop lightweight models, general-purpose models, and anime models in sequence. As an achievement so far, they have completed a lightweight model and released it as a demo. For the lightweight model, the initial plan was 1B parameters, but since the performance was insufficient, they are addressing this by increasing it to 2B.
Knowledge gained so far includes the finding that the Rectified Flow Transformer is effective not only for image generation but also for video generation, and that training from scratch directly on video is possible. At the same time, they faced the challenge that no training source code existed for Rectified Flow Transformer applied to video, but they were able to solve this by combining state-of-the-art codebases.
They also faced the problem of insufficient video data for training, which they addressed by extracting from other datasets such as FineVideo.
AiHUB Inc.
AiHUB Inc. was established in 2023 through the merger of the image generation AI community and the open-source developer community. Under GENIAC, the company is working on the development of an anime-specialized foundation model to help revitalize the Japanese animation industry. The project consists of two stages. First, they will develop a common foundation model. Then, they will perform additional training using data owned by anime production companies and studios. By doing so, they will provide “dedicated models for each company,” which can be safely used with their own data, as well as an “anime-specialized generative AI service” for the animation industry, thereby promoting real-world implementation. During this period, they plan to first develop an image generation foundation model, which will later evolve into a video model. Training is divided into two phases, trial and production, and they are currently in the stage of “large-scale pretraining trials.”
In the course of development, they designed and trained a proprietary autoencoder, achieving a PSNR score of 26.38, which is about 4% higher than nvidia/Cosmos-0.1-Tokenizer-CI16x16. They also carried out comparisons of text encoders and accumulated development know-how aimed at improving the output quality of the image model.
At present, the amount of anime-domain data they have secured is still insufficient, but they will continue efforts to acquire more data and proceed with production-phase training to ensure the success of the project.
Ubitus Inc. / Deepreneur Inc.
Ubitus and Deepreneur are developing tourism-focused LLM/foundation models, targeting strong multilingual performance in Japanese, Chinese, and Korean and eventual public release. They built a multilingual dataset spanning the three languages and carried out tourism-domain fine-tuning. By leveraging web crawling, data generation, noise removal, and CoT rewiring, they obtained 40B+ Traditional Chinese tokens. They also created a large-scale Japanese corpus and ran small-model pilots, achieving 76.06% on TMMLU+ and the target 68% with Llama 3.3-70B. For Japanese, continued pretraining uses ~100B tokens comprising high-quality texts (official documents, news, academic papers), texts with high third-party ratings, and lightly filtered crawl data—boosting performance. They plan to integrate JP/ZH/KR datasets for composite training.
Kotoba Technologies Japan Inc.
Kotoba Technologies Japan Inc. is advancing a project to build a real-time speech foundation model. This project consists of two steps. First, they are constructing a high-quality speech dataset with a focus on Japanese. Using TTS models, they generate 200,000 hours of speech and also collect and clean another 200,000 hours of Japanese data from other resources. At the same time, they collect and clean 200,000 hours of high-quality English data. Next, they train a speech foundation model on this data, aiming not only for general-purpose chatbot applications with features such as fast inference and long-form/two-stream inference, but also for applications like simultaneous interpretation.
The TTS/speech foundation model works by discretizing audio with a tokenizer for learning and processing. So far, they have succeeded in improving the compression rate of discretization by 5–6 times, achieving faster training. Their multilingual speech generation already demonstrates high performance, allowing direct conversion of Japanese speech. Similarly, for simultaneous speech interpretation, Japanese speech can now be converted into multiple languages almost in real time. Looking ahead, they plan to continue developing user-friendly AI.
Karakuri Inc.
Karakuri Inc. is an AI startup that develops tools such as chatbots. Under GENIAC, the company is working on developing a high-quality AI agent model specialized for customer support. The original plan included creating training data and benchmarks for customer support AI agent models, training and evaluating the model, and releasing the model, know-how, and benchmarks. As interim results, they built a Japanese dataset including images, a Japanese computer operation dataset, and created a tool to record computer operations. Because such datasets did not exist for developing Japanese customer support agent models, they created their own to overcome this challenge. By using expertise gained from business operations, they identified patterns and generated synthetic data. For computer operations, they developed a recording tool to collect data efficiently. In training, they conducted preliminary experiments on Trainium with models such as InternVL 2.5, Qwen2VL, and QwQ, and they plan to move toward public release. In addition, with AWS support, they learned how to list on the AWS Marketplace, which was another significant outcome.
ABEJA Inc.
ABEJA Inc. is developing high-performance LLMs specialized for business while keeping parameter sizes relatively small. They plan to create two models, one under 50B parameters and one under 10B. For this project, they are using Qwen2.5 as the base model and NeMo as the development framework. The training data consists of about 100B Japanese-English tokens (70% Japanese, 30% English). At present, the first round of training for both the under-50B and under-10B models has been completed, benchmark evaluations are going well, and development is progressing smoothly.
One challenge they encountered is that although Qwen2.5 performs well, it occasionally outputs Chinese. They expected this to be resolved through continued pretraining, but the issue persists. Two hypotheses for the cause have been proposed: first, that it stems from the application of ChatVector in the base model; second, that Qwen itself had learned Chinese contamination in its synthetic data. To address the first, they are applying post-training measures. To address the second, they plan to treat data that tends to produce Chinese output as Japanese and retrain it using an active learning approach.
They have also found that prioritizing data quality over sheer quantity is crucial. In particular, they learned that synthetic data alone cannot produce sufficient diversity unless it is carefully combined with other types of data.
Preferred Elements Inc./ Preferred Networks Inc.
Preferred Elements Inc. and Preferred Networks Inc. are working on building one of the world’s largest high-quality datasets and developing large-scale language models. Using the LLM developed in GENIAC Phase 1, they have already achieved their goal of creating 100B tokens of high-quality data, and are now working toward 200B tokens. In particular, they are focusing on generating Japanese knowledge, mathematics, and program code. Overall progress is on schedule: they have already completed the construction of a 1B model and are currently training an 8B model. Next, they plan to train a 30B model. As a more ambitious goal, instead of deploying a 30B model with 8B active parameters, they are considering building an 8B model derived from the 30B through techniques such as pruning, with the aim of achieving performance comparable to a 100B model.
The models developed are planned for sequential public release or service provision. In December 2024, they released “PLaMo Prime,” which provides access to the “PLaMo” series via the cloud as an API service. Earlier, in August 2024, they had released “PLaMo Beta” to accelerate rapid real-world adoption, which helped them identify diverse user needs. As a result, they obtained valuable insights into what functions are required for post-training.
Japan Agency for Marine-Earth Science and Technology
The Japan Agency for Marine-Earth Science and Technology (JAMSTEC) is developing an LLM specialized in planning countermeasures for risks that may arise at the regional or corporate level as global warming progresses. Ultimately, at the corporate level, the aim is for each company to publish anticipated risks and countermeasures in the form of TCFD reports, while at the regional level, the goal is to help local governments formulate climate change adaptation plans under their “Global Warming Action Plans.”
They have already completed 100% of the planned training data collection and have built benchmark datasets. As an achievement so far, they collected relevant academic papers, municipal climate action policies, and corporate TCFD reports, while also enabling automatic generation of instruction-tuning data. Initially, they faced challenges with low accuracy when extracting images and text from PDF files, but by switching to “Docling” they achieved improved accuracy.
Currently, they are building cooperative frameworks with municipalities and have begun discussions with companies. However, because participation is requested on a voluntary basis, they recognize the need to expand involvement by both gathering the needs of municipalities and companies and clearly communicating the concrete benefits of participation.
Humanome Lab
Humanome Lab is developing a foundation model for gene expression to accelerate drug discovery. Using large-scale cellular data, the project aims to provide insights such as how drugs act on cells or what happens to cells when a person catches a cold. To achieve this, they planned and carried out the creation of a gene expression dataset covering 100 million cells. They crawled databases worldwide, collecting more than 50 million cells from life science sources and more than 50 million cells from diverse organs and diseases. Going forward, they plan to improve quality through data preprocessing and metadata curation.
Their goal is to develop a 300M-parameter gene expression foundation model, starting with training a 3M model. At first, they faced the challenge of slow training speeds, but through discussions with AWS representatives, they succeeded in improving training performance. The foundation model is intended for potential use in academic fields as well, and they are considering a balanced release strategy that could expose code not previously available in prior research, balancing both industrial and academic use. At the same time, they have begun conducting interviews with pharmaceutical companies. However, unlike LLMs or image generation models that have clear, tangible applications, this type of model lacks an obvious image in people’s minds. Therefore, they recognize the importance of producing concrete results early on and are moving forward with that challenge in mind.
EQUES Inc.
EQUES Inc. is developing an LLM specialized for the pharmaceutical field and the pharmaceutical industry. In the initial plan, they set a goal of building training data and achieving more than 70% accuracy on the national pharmacist examination, aiming to surpass GPT-4 in that domain. Building the training data took longer than expected due to the time required for cleaning, but progress is now on track. They have also completed the digitization of the national pharmacist exam as benchmark evaluation data.
For development, they selected Qwen2.5 as the base model and are using a hyperparameter search algorithm (D-CPT Law) to explore optimal hyperparameters for pretraining. The pharmacist exam benchmark has already been released, and evaluation experiments have been completed.
So far, they have observed that existing LLMs perform relatively well on single-answer biology and pathology questions, but their accuracy stagnates on physics and chemistry questions, which are thought to require stronger logical reasoning ability. At the time, the only model that achieved an accuracy exceeding 85% was the recently released “o1-preview.” EQUES plans to continue running various trials to improve accuracy going forward.
SyntheticGestalt Inc.
SyntheticGestalt Inc. is working on building a foundation model specialized in small molecules. They are creating features for an enormous dataset of 10 billion compounds and will use all of them for pretraining to develop the foundation model. In the field of AI drug discovery, there is a benchmark called TDC, and their goal is to rank within the global top three across all 23 major tasks.
For evaluating their machine learning models, they do not use standard random splits, which would not provide accurate assessments, but instead apply scaffold splits, an evaluation method that excludes compounds with similar skeletons from the test data.
In feature design, it is necessary to make inputs compatible with neural networks. To achieve this, they represent compounds not in 2D, but in 3D and 4D, which raises the level of difficulty. To address these challenges, they applied mathematical approaches to improve the architecture, giving the foundation model and the project greater flexibility.
After the presentations from the participating companies, Mr. Takenori Endo, Director of Generative AI section in Artificial Intelligence and Robotics Department at NEDO (New Energy and Industrial Technology Development Organization), delivered closing remarks to conclude the mid-term report meeting.
During the meeting, time was set aside between presentations for discussions among the companies, providing a valuable opportunity for participants to deepen their knowledge, exchange ideas, and gain useful insights that can be applied to future development.
As a community that also plays a role in fostering generative AI development, GENIAC will continue to strengthen collaboration among companies and promote further progress. We look forward to each participant applying creativity and ingenuity in the pursuit of even greater innovation.