Multimodal data labeling and large model thinking chain training method based on visual interaction

By using eye-tracking technology and multimodal data acquisition, and combining CNN and BERT models to train large models, the shortcomings of large models in understanding complex images are solved, multimodal data fusion and self-iterative optimization are realized, and annotation and analysis capabilities are improved.

CN120562596BActive Publication Date: 2025-11-21XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511053014.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-21
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

Existing large models struggle to understand the relationships between image elements when processing complex and specialized images. The labeled data differs significantly from actual needs, and the lack of cross-modal data fusion capabilities results in insufficient analytical capabilities.

Method used

By employing eye-tracking technology combined with multimodal data acquisition, multimodal data is displayed on a monitor and expert eye movement trajectories and voice information are recorded to form multimodal labeled data. After preprocessing using CNN and BERT models, the data is input into a large model for training, forming a large model thought chain.

Benefits of technology

It enhances the ability of large models to understand complex images, enables comprehensive and in-depth analysis of multimodal data, forms a self-iterative optimization closed loop, and improves annotation accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120562596B_ABST
    Figure CN120562596B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of large model training, and particularly relates to a multi-modal data labeling and large model thinking chain training method based on visual interaction. The method comprises the following steps: 1, multi-modal data acquisition; 2, displaying the multi-modal data to a display, showing the multi-modal data to an expert through the display, recording the eye movement track of the expert when watching the multi-modal data on the display by using an eye movement data acquisition device, and collecting the voice information of the expert by using a voice acquisition device; 3, forming multi-modal labeling data according to the eye movement track and the voice information; 4, after pre-processing the multi-modal labeling data, inputting the same to a large model for training to obtain a trained large model. The present application can collect the multi-modal labeling data of the expert in the labeling process, explicitly process the thinking chain and thinking process of the expert, and logically fuse the same, thereby improving the information breadth and accuracy of the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of large model training, specifically involving a method for multimodal data annotation and large model thought chain training based on visual interaction. Background Technology

[0002] While large-scale image models have made significant progress in image understanding, easily performing basic tasks such as classification and object detection for common images, they still exhibit numerous shortcomings when faced with complex and information-rich professional images. Large-scale models also have weaknesses in capturing the relationships between image elements. The spatial positions and logical relationships between multiple objects in an image play a decisive role in the comprehensive understanding of the image content. However, large-scale models struggle to grasp these complex relationships from a holistic perspective, unlike humans, severely limiting their ability to deeply understand and apply images.

[0003] Training large-scale models requires a large amount of labeled data; however, current data labeling tools have significant shortcomings. Firstly, the professionalism of labelers is insufficient. Currently, data labeling is mainly undertaken by IT personnel or outsourced teams. Because these personnel lack professional industry knowledge and a deep understanding of specific application scenarios, the labeling results deviate significantly from actual needs. Secondly, single-modal and correlation analysis are lacking. Most existing data labeling tools only support the labeling of single-modal data and lack the ability to fuse cross-modal data. This makes it impossible to establish effective correlations between different types of data, hindering the formation of a comprehensive and multi-dimensional data support system. In medical scenarios, image data, medical record text data, and doctors' diagnostic voice data should corroborate each other, but due to the limitations of existing tools, these data cannot be effectively fused and analyzed, resulting in incomplete labeling information. Furthermore, the existing labeling process suffers from logical process and efficiency issues. The current labeling process lacks a data-driven expert-level judgment chain, significantly reducing the accuracy and reliability of the labeling results. Even when multiple experts' opinions are gathered, the lack of an effective integration mechanism makes it difficult to fully leverage collaborative advantages. Summary of the Invention

[0004] The purpose of this invention is to address the technical problem that existing large model training methods rely on labeled data that are far from meeting actual needs, and that they often depend on single-modal data and lack cross-modal data fusion capabilities, resulting in large models that lack comprehensive, in-depth, and accurate analytical capabilities. The invention provides a method for multimodal data annotation and large model thought chain training based on visual interaction.

[0005] The design concept of this invention is as follows:

[0006] The successful application of thought chains (CoT) in natural language processing has provided new insights for research in the image domain. CoT significantly improves the performance of large models in tasks such as text generation, question answering systems, and machine translation by decomposing complex problems into a series of logically related sub-problems and guiding the model through step-by-step reasoning. Introducing the concept of thought chains into the image domain can similarly enhance the ability of large models to understand and process images. Meanwhile, the development of eye-tracking technology provides new technical means for this research. Eye-tracking technology can accurately record information such as the observer's gaze point, gaze duration, and saccade path when viewing an image. This information intuitively reflects the observer's visual attention distribution and thought process. For example, in industrial inspection scenarios, the gaze trajectory of a skilled technician reveals their process of identifying product faults. Combining eye-tracking technology with multimodal data opens up new research directions for improving the performance of large models in specialized fields.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] A method for multimodal data annotation and large-scale model thought chain training based on visual interaction is characterized by the following steps:

[0009] Step 1: Multimodal data acquisition, including image data acquisition, text data acquisition, and sensor monitoring data acquisition;

[0010] Step 2: Display the collected multimodal data in different areas of the monitor according to the set arrangement. Show the multimodal data to the expert through the monitor. During this process, use an eye-tracking data acquisition device to record the eye movement trajectory of the expert when viewing the multimodal data on the monitor, and use a voice acquisition device to collect the expert's voice information simultaneously.

[0011] Step 3: Based on the collected eye movement trajectory data, determine the coordinates of the expert's gaze point; then, based on the expert's gaze point coordinates, determine the order in which the expert gazes at each area on the display screen through time series analysis, and crop the corresponding areas of gaze into images in sequence according to this order to form image annotation data containing the expert's thought process.

[0012] The collected voice information is converted into voice-text information, and then the voice-text information is matched with the images in the image annotation data in chronological order. Corresponding voice-text labels are added to each image to form multimodal annotation data.

[0013] Step 4: Preprocess the multimodal labeled data, and then input the preprocessed multimodal labeled data into the large model for training to obtain the trained AI large model, thus completing the training of the large model's thought chain.

[0014] Furthermore, in step 2, the multimodal data is displayed to multiple experts via a monitor, and an eye-tracking data acquisition device is used to record the eye movement trajectories of multiple experts as they view the multimodal data on the monitor. Additionally, a voice acquisition device is used to simultaneously acquire the voice information of multiple experts.

[0015] In step 3, the coordinates of each expert's gaze point are determined based on the collected eye movement trajectory data of multiple experts; then, based on the coordinates of each expert's gaze point, the order in which each expert gazes at different areas on the display screen is determined, and the gaze areas are cropped into images in this order to form multiple image annotation data containing the expert's thought process.

[0016] The collected voice information from multiple experts is converted into voice-text information, and then the voice-text information of each expert is matched with the images in the corresponding image annotation data to form multiple multimodal annotation data.

[0017] In step 4, multiple multimodal labeled data are input into the large model for training.

[0018] Furthermore, it also includes step 5:

[0019] The current AI model is used to update and iterate the current multimodal data. Then, the updated and iterated multimodal data is used to obtain the updated and iterated AI model in the same way as steps 2 to 4, thus realizing the closed-loop update and iteration of the AI ​​model.

[0020] Furthermore, the preprocessing described in step 4 specifically includes:

[0021] Y1. Use a CNN model to convert images in multimodal labeled data into image feature vectors, and use a BERT model to convert speech and text information in multimodal labeled data into text feature vectors.

[0022] Y2. Using the image-text contrast learning method, image feature vectors and text feature vectors are aligned to the same space;

[0023] Y3. The image feature vector and the text feature vector are added by an adder to obtain the fused feature vector, which is the preprocessed multimodal labeled data.

[0024] Furthermore, in step 3, after cropping the gaze area into images in sequence, the following steps are also included: labeling the cropped images, including gaze time, image category, position information in the original image, and semantic description; extracting visual features of the images, including color, texture, and shape; and combining the visual features with the labeled content to form image labeled data that contains expert thought processes and features.

[0025] The collected voice information is converted into voice-text information, and then the voice-text information is matched with images in image annotation data that contain expert thought processes and feature representations in chronological order.

[0026] The beneficial effects of this invention are:

[0027] 1. Construction of Multimodal Data and Establishment of Multimodal Data Correlation: Collect multimodal data from experts during the annotation process, make the experts' thought processes and thought chains explicit, and logically integrate them to effectively improve the information breadth and final accuracy of the large model. This integration method can gather the rich knowledge and valuable experience of experts, enabling the trained AI large model to have more comprehensive and in-depth analytical capabilities, and to demonstrate superior performance and more accurate judgment in solving complex problems.

[0028] 2. By extracting the thought processes of multiple experts, it is possible to create a large-scale AI model that surpasses the level of a single expert.

[0029] 3. Optimize the data annotation process to form a self-iterative optimization loop: A multimodal thinking chain facilitates data annotation, improving annotation accuracy and efficiency. High-quality annotated data feeds back into the training of large AI models, enhancing their performance. Improved model performance further optimizes the thinking chain, forming a closed loop of "annotation optimization—model improvement—thinking chain improvement," achieving self-iteration, reducing costs, improving the quality of multimodal data, and driving the continuous progress of large AI models. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of an expert multimodal annotation data acquisition device according to an embodiment of the present invention;

[0031] Figure 2 This is a schematic diagram of the expert multimodal annotation data acquisition interface according to an embodiment of the present invention;

[0032] Figure 3 This is a flowchart illustrating the construction process of expert multimodal annotation data according to an embodiment of the present invention;

[0033] Figure 4 This is a schematic diagram of expert multimodal thinking chain data acquisition and model training according to an embodiment of the present invention;

[0034] Figure 5 yes Figure 4 A schematic diagram of the information to be labeled displayed on the monitor;

[0035] Figure 6 yes Figure 4 Diagram illustrating eye-tracking information from Chinese experts;

[0036] Figure 7 yes Figure 4 Chinese text information illustration;

[0037] Figure 8 This is a schematic diagram illustrating the fusion of multi-expert thought chains according to an embodiment of the present invention;

[0038] Figure 9 This is a schematic diagram illustrating the iterative process of the multimodal dataset and expert model according to an embodiment of the present invention;

[0039] In the diagram: 101-monitor, 102-host, 103-eye-tracking data acquisition device, 104-voice acquisition device, 105-keyboard and mouse. Detailed Implementation

[0040] To make the objectives, advantages, and features of this invention clearer, the following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a method for multimodal data annotation and large-scale model thought chain training based on visual interaction. The advantages and features of this invention will become clearer according to the following specific embodiments.

[0041] See Figure 1 To implement the multimodal data annotation and large model thinking chain training method based on visual interaction proposed in this embodiment, the equipment used includes a display 101, a host 102, an eye-tracking data acquisition device 103, a voice acquisition device 104, and a keyboard and mouse 105.

[0042] The display 101 is used to display multimodal data, which includes image information, text information, and other modal information to be labeled, and arranges this information in a certain order. For example, this embodiment follows... Figure 2 The arrangement shown is used to arrange multimodal data.

[0043] The host 102 is responsible for processing the data collected by the information acquisition devices (eye-tracking data acquisition device 103 and voice acquisition device 104) and storing multimodal data.

[0044] The eye-tracking data acquisition device 103 is used to collect the expert's eye-tracking trajectory, while the voice acquisition device 104 is used to collect the expert's voice information.

[0045] The keyboard and mouse (105) are operating devices, and the experts are experts in a specific field.

[0046] See Figure 3 This embodiment of a method for multimodal data annotation and large model thinking chain training based on visual interaction specifically includes the following steps:

[0047] Step 1: Display multimodal data, including image data, sensor monitoring data, and text data, on display 101, and categorize the multimodal data according to... Figure 2 The arrangement shown is displayed on monitor 101.

[0048] Image data includes medical images, industrial inspection images, and natural scene videos. This image data can be obtained from medical testing instruments such as CT images and MRI images, as well as surveillance footage captured by security cameras or pictures taken by video cameras.

[0049] Text data: Collect textual materials related to images or videos, such as diagnostic reports from medical images and technical documents from industrial testing. Simultaneously, convert collected audio information into speech-to-text information using speech-to-text algorithms, enriching the sources of text data.

[0050] Sensor monitoring data: Monitoring data obtained by various sensors, such as temperature sensors, current sensors, pressure sensors, etc.

[0051] Step 2: Have multiple experts in the relevant fields view the multimodal data displayed on monitor 101. During this process, an eye-tracking data acquisition device 103 is used to record the eye movements of the experts as they view the multimodal data. Experts include professionals (such as doctors and engineers) and members of the general public. The experimental environment should be kept quiet and with uniform lighting. The eye-tracking data acquisition device 103 should be calibrated before each experiment to ensure data accuracy.

[0052] Step 3: Use an eye-tracking algorithm to identify and calibrate the gaze points of each expert to improve recognition accuracy, obtaining the gaze points (t, x, y) of each expert on the screen. Where: t represents time information, and x and y represent the coordinate values ​​of the expert's gaze point. For example, if the resolution of display 101 is 1920×1080, when an expert gazes at the multimodal data displayed on display 101, then x and y are values ​​from 1 to 1920 and from 1 to 1080, respectively.

[0053] Step 4: Determining the gaze region. Image data is extracted using gaze point information. For example, at time t, the gaze point is focused on position (x, y) of display 101. Therefore, a square area centered at (x, y) with a side length of 50 pixels is preferably constructed and cropped from the screen area as the gaze region at time t, denoted as (t, img_t). Through time series analysis, the order in which each expert gazes at each region is determined. Following this order, the gaze regions are sequentially cropped into smaller images, forming multiple image annotation data containing the expert's thought process.

[0054] Step 5: Expert Voice Acquisition. When different experts view the multimodal data on the display 101, they are required to simultaneously describe the information they are viewing and their own judgment using voice, and the voice information is collected by the voice acquisition device 104.

[0055] Step 6: Convert the collected voice information from multiple experts into voice-to-text information using existing speech-to-text algorithms.

[0056] Step 7: Align the converted voice and text information of each expert with the images in the corresponding image annotation data in terms of time.

[0057] Step 8: After aligning the time, save the voice and text information of each expert in the form of (t, txt_t).

[0058] Step 9: Multimodal Data Fusion. The speech-text information (t, txt_t) from each expert is fused with the corresponding images (t, img_t) in chronological order. A corresponding speech-text label is added to each image to form multiple multimodal labeled datasets.

[0059] To further illustrate the above steps, we will use mechanical fault labeling in an industrial setting as an example. (See [link to documentation]). Figures 4 to 7 The display 101 displays an image of the faulty object in its image display window; it displays the time-domain signal and spectral analysis waveform of the bearing vibration in its sensor data window; and it displays corresponding text information in its text information window. When the expert views the information on the display 101, the eye-tracking data acquisition device 103 records the expert's eye movement trajectory, such as... Figure 6 The green circle in to These represent the expert's gaze points. When the expert observes unusual information, they will verbally describe it. For example, when the expert's gaze falls on the gaze point... When the expert is located in a specific area, they verbally describe a frequency peak at 108.6 Hz, which is close to the theoretical value of 107.91 Hz for bearing BPFO, and there is a significant overtone. At this point, the voice acquisition device 104 will capture the expert's voice and convert it into text. Figure 4 Image information containing time sequence is Figure 6 The green circle in to After time alignment, multimodal data fusion can be performed to obtain multimodal labeled data.

[0060] Step 10: Since data from different modalities, such as images, speech, and text, have different features and formats, they need to be converted into a form that can be uniformly processed by large models. For example, a CNN model can be used to convert images into image feature vectors, and a BERT model can be used to convert speech and text information into text feature vectors. Then, alignment techniques, such as image-text contrastive learning (ITC), are used to align the image feature vectors and text feature vectors to the same space. Finally, an adder is used to add the image feature vectors and text feature vectors to obtain a fused feature vector, which is the preprocessed multimodal labeled data, so that large models can better understand and fuse this multimodal labeled data.

[0061] After inputting multimodal labeled data into a large model, the model needs to be trained to learn how to predict the output from the input. This process involves adjusting and optimizing model parameters and requires selecting a suitable large model architecture, such as Transformer, CNN-LSTM, etc.

[0062] At the same time, because each expert's thought process differs to some extent, by learning from the thought processes of different experts through a large model, and integrating the strengths of different experts, the goal can be achieved to surpass the work of a single expert. Figure 8 As shown, this demonstrates a process that makes tacit knowledge from the expert diagnostic process explicit and utilizes artificial intelligence technology for fault diagnosis. Figure 8 The image information window can be seen in [the image information window]. Figure 6 The image information window is used to illustrate this. First, eye-tracking reveals the intuitive judgments of experts based on years of experience, transforming this tacit knowledge into a structured decision-making process, quantifiable attention time, and a visualized experience transfer system. Next, the diagnostic processes of different types of experts are compared; for example, experienced experts rely on sensory information for broad but brief fixations, while analytical experts focus on data for concentrated and sustained fixations. AI eye-tracking annotation technology combines the strengths of both types of experts, achieving more efficient and accurate annotation. Finally, by fusing eye-tracking annotation data from multiple experts, multiple multimodal annotation datasets can be obtained. These multimodal annotation datasets are input into a large model, training the large model to form a large AI model, namely an AI intelligent fault diagnosis system. This system can optimize the diagnostic path and improve the accuracy and efficiency of diagnosis.

[0063] To achieve iterative iteration of large-scale AI models, an iterative process is designed between industry big data and large-scale AI models, such as... Figure 9As shown, this diagram depicts the iterative process between industry big data and AI big models. The diagram illustrates three main components: industry big data, thought chains, and AI big models. The industry big data component emphasizes the accumulation of high-quality multimodal data, including image data, text data, and sensor monitoring data. This industry big data can be updated annually in real time, such as Dataset 2025, Dataset 2026, ..., Dataset N, demonstrating the continuous enrichment and optimization of the data. Thought chains represent the expert diagnostic thinking process. Through training and optimization, industry big data is used to construct thought chains, while the application and validation of thought chains drive the iteration of AI big models. With the iteration and optimization of industry big data, AI big models also continuously iterate to achieve the explicit expression of expert knowledge and the improvement of AI predictive capabilities. Figure 9 Arrows in the diagram illustrate the interactions between the various components: industry big data, after training and optimization, enters the thought chain; the application and validation of the thought chain drives the iteration of the AI ​​big model; and the iteration of the AI ​​big model, in turn, promotes the iteration of industry big data. The entire process forms a closed loop, demonstrating the dynamic relationship between data richness and knowledge iteration.

[0064] Through the above steps, the AI ​​large model can effectively learn the thinking patterns of experts from multimodal data and make predictions, providing strong support for various applications.

Claims

1. A method for multimodal data annotation and large model thought chain training based on visual interaction, characterized in that, Includes the following steps: Step 1: Multimodal data acquisition, including image data acquisition, text data acquisition, and sensor monitoring data acquisition; Step 2: Display the collected multimodal data in different areas of the monitor according to the set arrangement. Show the multimodal data to the expert through the monitor. During this process, use an eye-tracking data acquisition device to record the eye movement trajectory of the expert when viewing the multimodal data on the monitor, and use a voice acquisition device to collect the expert's voice information simultaneously. Step 3: Based on the collected eye movement trajectory data, determine the coordinates of the expert's gaze point; then, based on the expert's gaze point coordinates, determine the order in which the expert gazes at each area on the monitor through time series analysis, and crop the gaze area into images in this order to form image annotation data containing the expert's thought process. The collected voice information is converted into voice-text information, and then the voice-text information is matched with the images in the image annotation data in chronological order. Corresponding voice-text labels are added to each image to form multimodal annotation data. Step 4: Preprocess the multimodal labeled data, and then input the preprocessed multimodal labeled data into the large model for training to obtain the trained AI large model and complete the training of the large model's thought chain. The preprocessing specifically includes: Y1. Use a CNN model to convert images in multimodal labeled data into image feature vectors, and use a BERT model to convert speech and text information in multimodal labeled data into text feature vectors. Y2. Using the image-text contrast learning method, image feature vectors and text feature vectors are aligned to the same space; Y3. The image feature vector and the text feature vector are added by an adder to obtain the fused feature vector, which is the preprocessed multimodal labeled data.

2. The method for multimodal data annotation and large model thinking chain training based on visual interaction according to claim 1, characterized in that: In step 2, the multimodal data is displayed to multiple experts via a monitor, and the eye movement trajectories of the multiple experts as they view the multimodal data on the monitor are recorded using an eye-tracking data acquisition device. Additionally, the voice information of the multiple experts is collected synchronously using a voice acquisition device. In step 3, the coordinates of each expert's gaze point are determined based on the collected eye movement trajectory data of multiple experts; then, based on the coordinates of each expert's gaze point, the order in which each expert gazes at different areas on the display screen is determined, and the areas corresponding to the gazes are cropped into images in sequence according to this order, forming multiple image annotation data containing the expert's thought process. The collected voice information from multiple experts is converted into voice-text information, and then the voice-text information of each expert is matched with the images in the corresponding image annotation data to form multiple multimodal annotation data. In step 4, multiple multimodal labeled data are input into the large model for training.

3. The method for multimodal data annotation and large model thinking chain training based on visual interaction according to claim 1 or 2, characterized in that, It also includes step 5: The current AI model is used to update and iterate the current multimodal data. Then, the updated and iterated multimodal data is used to obtain the updated and iterated AI model in the same way as steps 2 to 4, thus realizing the closed-loop update and iteration of the AI ​​model.

4. The method for multimodal data annotation and large model thinking chain training based on visual interaction according to claim 1, characterized in that: In step 3, after cropping the gaze area into images in sequence, the following steps are also included: labeling the cropped images, including gaze time, image category, location information in the original image, and semantic description; extracting visual features of the images, including color, texture, and shape; and combining the visual features with the labeled content to form image labeled data that contains expert thought processes and feature representations. The collected voice information is converted into voice-text information, and then the voice-text information is matched with images in image annotation data that contain expert thought processes and feature representations in chronological order.

Citation Information

Patent Citations

  • Multi-modal generative large model training method and device and computer equipment

    CN117011686A

  • Multi-modal data acquisition, storage and labeling system

    CN118155288A