Multi-modal data intelligent labeling system based on computer vision
Through the coordinated work of reference point alignment, gaze point prediction and modal generalization modules, the problems of difficulty in reference point alignment and insufficient labeling quality in multimodal data labeling are solved, and efficient and accurate multimodal data labeling are achieved.
Patent Information
- Application Number
- CN202510682421.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing multimodal data annotation technology, the reference points are difficult to automatically align, and the spatial consistency of cross-modal annotation is insufficient. The focus point prediction in dynamic scenarios relies on single-modal timing analysis to lack multimodal collaboration, and the modal generalization ability is weak, making it difficult to achieve cross-modal annotation knowledge transfer.
The reference point alignment module is used to obtain feature points and align coordinates through the deep learning model. The gaze point prediction module uses time series analysis and optical flow method to predict dynamic gaze points. The modal generalization module uses cross-modal learning framework to map feature space, and the labeling optimization module optimizes the labeling results through statistical models and consistency checks.
The spatial consistency and dynamic optimization of multimodal data annotation are realized, the labeling efficiency and accuracy are improved, and the consistency and reliability of the annotation results are ensured.
Smart Images

Figure CN120580692A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer processing, and in particular relates to a multimodal data intelligent labeling system based on computer vision. Background Art
[0002] With the rapid development of computer vision and artificial intelligence technologies, intelligent annotation of multimodal data has gradually become a research hotspot. This technology can integrate visual data from multiple modalities, such as images, videos, and point clouds, and automatically extract features and generate high-quality annotations through deep learning models, significantly improving the efficiency and accuracy of data annotation.
[0003] Traditionally, multimodal data annotation typically utilizes a single-modality independent processing approach, requiring separate feature extraction and annotation algorithms designed for each modality. For example, image data utilizes 2D keypoint detection, video data relies on optical flow to track target motion, and point cloud data utilizes 3D geometric feature matching. Furthermore, traditional methods often require manual annotation or the assistance of semi-automatic annotation tools, with the annotation results then undergoing post-processing for cross-modal alignment and optimization.
[0004] However, there are several problems with the current annotation method: the reference points of data from different modalities are difficult to automatically align, resulting in insufficient spatial consistency in cross-modal annotation; gaze point prediction in dynamic scenes relies on the temporal analysis of a single modality and lacks the precise positioning capability of multi-modal collaboration; the modal generalization capability of traditional annotation systems is weak, making it difficult to achieve cross-modal annotation knowledge transfer. Summary of the Invention
[0005] Based on this, it is necessary to provide a multimodal data intelligent labeling system based on computer vision that can accurately label the above technical problems.
[0006] In a first aspect, the present application provides a multimodal data intelligent annotation system based on computer vision, comprising:
[0007] The reference point alignment module is used to obtain the feature extraction results of multimodal data, perform reference point alignment on the feature extraction results, and generate a consistent reference point set;
[0008] A gaze point prediction module is used to dynamically predict the gaze point in the visual data based on the reference point set and generate a dynamic gaze point position;
[0009] The modality generalization module is used to perform feature space mapping processing based on the dynamic gaze point position through a cross-modal learning framework to obtain cross-modal annotation results;
[0010] The annotation optimization module is used to evaluate the quality of annotation results through statistical models and consistency checks to generate optimized annotation results.
[0011] In one embodiment, the reference point alignment module includes the following units:
[0012] A data acquisition unit is used to acquire key feature points in images, videos, and point cloud data through a deep learning model to obtain a feature point set;
[0013] A feature matching unit, configured to perform cross-modal feature point matching based on a feature point set using a feature matching algorithm to generate an original reference point set;
[0014] The coordinate alignment unit is used to align the original reference point set into the same coordinate system through the coordinate transformation matrix to obtain a spatially consistent reference point set.
[0015] In one embodiment, the gaze point prediction module includes the following units:
[0016] The motion trajectory prediction unit is used to predict the motion trajectory of objects in the video based on continuous frame images through time series analysis and optical flow method, and calculate the dynamic gaze point position using the following formula:
[0017] G t+1 =G t +F(I t ,I t+1 )
[0018] Among them, G t+1 is the dynamic gaze point position, G t is the gaze point position of the current frame, F is the optical flow calculation function, I t and I t+1 It is two consecutive frames of images;
[0019] The gaze point optimization unit is used to perform key area priority labeling processing based on the dynamic gaze point position through the gaze point prediction model to generate an optimized gaze point position.
[0020] In one embodiment, the modality generalization module includes the following units:
[0021] The feature extraction unit is used to obtain different modal data, extract modal invariant features through feature extraction functions, and map different modal data to the same feature space. The extraction formula of modal invariant features is:
[0022] Z=E(X)+E(Y)
[0023] Among them, Z is the modal invariant feature, X and Y are different modal data, and E is the feature extraction function;
[0024] A feature decoupling unit is used to separate modality-specific information and general semantic information based on modality-invariant features through feature decoupling technology to obtain general semantic features;
[0025] The annotation transfer unit is used to perform cross-modal annotation knowledge transfer based on common semantic features and annotation information of the annotated data using the annotation propagation algorithm and calculate the cross-modal annotation results using the following formula:
[0026] Y p =D(Z g )+S(Y l )
[0027] Among them, Y p is the cross-modal annotation result, D is the decoder, S is the annotation propagation function, and Z g is a general semantic feature, Y l It is the annotation information of the labeled data.
[0028] In one embodiment, the annotation optimization module includes the following units:
[0029] Anomaly detection unit, which is used to detect outliers based on the cross-modal annotation results and the annotations that do not match the data distribution through the statistical model to obtain a preliminary anomaly annotation set;
[0030] The consistency check unit is used to perform multi-view or multi-modal annotation consistency verification based on the preliminary abnormal annotation set through the consistency check function to generate the final consistent annotation result; wherein, the calculation formula of the consistency check function is:
[0031] A c =A⊙M(A)
[0032] Where A is the consistency annotation result, M is the consistency check function, and ⊙ is the element-by-element multiplication;
[0033] The manual correction unit is used to conduct expert review and correction on the detected preliminary anomaly annotation set through an interactive annotation interface to generate the final optimized annotation data.
[0034] In one embodiment, the system further includes a vision center module, including the following units:
[0035] A data management unit is used to store and manage multimodal visual data. Multimodal visual data supports multiple visual data formats including images, videos, and point clouds to obtain raw visual data.
[0036] The pipeline processing unit is used to perform preprocessing, feature extraction, and annotation generation based on the original visual data through the data pipeline to obtain standardized intermediate processing results;
[0037] The distributed computing unit is used to perform parallel computing and load balancing processing based on the intermediate processing results through a distributed computing framework to generate the final traceable large-scale annotation data.
[0038] In one embodiment, the system further includes a multimodal data enhancement and fusion annotation module, including the following units:
[0039] The data enhancement unit is used to enhance the data diversity of the original visual data through geometric transformation, illumination adjustment and noise injection technology to generate enhanced data samples;
[0040] The feature fusion unit is used to perform feature alignment and fusion processing based on the enhanced data samples through cross-modal feature extraction technology to generate comprehensive annotation results;
[0041] The unified annotation unit is used to generate the same annotation through the multimodal fusion network based on the comprehensive annotation results to obtain the final fusion annotation results.
[0042] In a second aspect, the present application also provides a method for intelligently labeling multimodal data based on computer vision, comprising:
[0043] Obtain feature extraction results of multimodal data, perform benchmark alignment on the feature extraction results, and generate a consistent benchmark set;
[0044] Perform dynamic prediction processing on the fixation point in the visual data based on the reference point set to generate the dynamic fixation point position;
[0045] Based on the dynamic gaze point position, feature space mapping is performed through a cross-modal learning framework to obtain annotation results applicable to multiple modalities.
[0046] The annotation results are evaluated for quality through statistical models and consistency checks to generate optimized annotation data.
[0047] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements any module of the above-mentioned embodiment when executing the computer program.
[0048] In a fourth aspect, the present application further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, any module of the above-mentioned embodiment is implemented.
[0049] The above-mentioned computer vision-based multimodal data intelligent labeling system, through the collaborative work of modules such as reference point alignment, gaze point prediction, modal generalization and labeling optimization, can effectively solve the problems of reference point alignment difficulties, low labeling efficiency, and difficulty in ensuring labeling quality in multimodal data labeling, significantly improve the efficiency and accuracy of multimodal data labeling, and ensure the consistency and reliability of the labeling results. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 This is a structural diagram of the computer vision-based multimodal data intelligent annotation system of the present invention;
[0052] Figure 2 This is a structural diagram of the data processing module for intelligent annotation of multimodal data based on computer vision of the present invention;
[0053] Figure 3 This is a structural diagram of a unified annotation module for intelligent annotation of multimodal data based on computer vision of the present invention;
[0054] Figure 4 This is a flow chart of the computer vision-based multimodal data intelligent labeling method of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0056] In one embodiment, Figure 1 As shown, a multimodal data intelligent annotation system 10 based on computer vision is provided. This embodiment uses the system applied to a terminal as an example. It is understandable that the system can also be applied to a server, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the system includes the following modules:
[0057] The reference point alignment module 11 is used to obtain feature extraction results of multimodal data, perform reference point alignment processing on the feature extraction results, and generate a consistent reference point set.
[0058] For example, key feature points in multimodal data are extracted through a deep learning model, and feature matching algorithms are used to match feature points in different modal data. The reference points are aligned to the same coordinate system through a coordinate transformation matrix to ensure the spatial consistency of multimodal data.
[0059] The gaze point prediction module 12 is used to perform dynamic prediction processing on the gaze point in the visual data according to the reference point set to generate a dynamic gaze point position.
[0060] Optionally, by analyzing the reference point set and key features in the visual data, combined with time series analysis and optical flow, the motion trajectory of objects in the video is predicted, thereby generating dynamic gaze point locations. The reference point set is generated by feature extraction and alignment of multimodal data, and serves as a consistent set of reference points to guide gaze point prediction. The dynamic gaze point location is predicted using a deep learning model and optical flow, and can dynamically reflect the gaze point location in key areas of the visual data.
[0061] The modality generalization module 13 is used to perform feature space mapping processing based on the dynamic gaze point position through a cross-modal learning framework to obtain cross-modal annotation results.
[0062] Specifically, based on the dynamic gaze point location, a cross-modal learning framework is used to map data from different modalities into a unified feature space. Feature decoupling and alignment techniques are then used to extract modality-invariant semantic features, thereby enabling the generation of cross-modal annotation results. The cross-modal learning framework aligns and shares information between modalities through feature space mapping; feature space mapping involves converting features from different modal data into the same feature space through a specific mapping method, enabling cross-modal comparison and annotation.
[0063] The annotation optimization module 14 is used to perform quality assessment on the annotation results through statistical models and consistency checks to generate optimized annotation results.
[0064] Statistical models are used to assess the quality of annotation results, identifying outliers and inconsistencies in the annotations. Consistency checks are also used to ensure consistency between annotations from different annotators or from different perspectives, ultimately generating optimized annotation results. Statistical models are used to assess the quality of annotation results, such as calculating accuracy and consistency. Consistency checks verify the consistency of annotation results from different annotators or from different perspectives and are often used to detect annotation deviations and errors.
[0065] In the above-mentioned computer vision-based multimodal data intelligent labeling system, the spatial consistency of multimodal data is achieved through the reference point alignment module, the gaze point prediction module dynamically optimizes the labeling area, the modal generalization module improves the cross-modal labeling capability, and the labeling optimization module ensures the labeling quality, so as to achieve the efficient and accurate generation of high-quality multimodal labeling results, significantly improving the labeling efficiency and data reliability.
[0066] In one embodiment, the reference point alignment module includes the following units:
[0067] A data acquisition unit is used to acquire key feature points in images, videos, and point cloud data through a deep learning model to obtain a feature point set;
[0068] Optionally, a deep learning model is used to automatically identify and extract a set of key feature points from multimodal data, such as images, videos, and point clouds, for subsequent feature matching and alignment. A deep learning model is a neural network-based model that automatically learns and extracts features from data; key feature points are representative and distinguishing points in the data, such as corners and edges in images, and geometric feature points in point clouds.
[0069] The feature matching unit is used to perform cross-modal feature point matching based on the feature point set through a feature matching algorithm to generate an original reference point set.
[0070] Specifically, by extracting a set of feature points from multimodal data and comparing them across different modalities using a feature matching algorithm, a set of unprocessed original reference points is generated. This ensures consistency across modalities by addressing the feature differences between the modalities. Feature point matching algorithms, such as SIFT, SURF, and ORB, are used to compare and match feature points across different modalities.
[0071] The coordinate alignment unit is used to align the original reference point set into the same coordinate system through the coordinate transformation matrix to obtain a spatially consistent reference point set.
[0072] Preferably, the points in the original reference point set are transformed from the coordinate systems of their respective modalities to a unified coordinate system by calculating a coordinate transformation matrix, thereby obtaining a spatially consistent reference point set. A coordinate transformation matrix is a mathematical tool used to transform points from one coordinate system to another, typically including transformations such as translation, rotation, and scaling.
[0073] In the above-mentioned computer vision-based multimodal data intelligent annotation system, a set of original reference points is generated by matching cross-modal feature points, and these reference points are further aligned to the same coordinate system to ensure the spatial consistency of multimodal data and provide an accurate spatial positioning basis for subsequent annotation work.
[0074] In an exemplary embodiment, the gaze point prediction module includes the following units:
[0075] The motion trajectory prediction unit is used to predict the motion trajectory of objects in the video based on continuous frame images through time series analysis and optical flow method, and calculate the dynamic gaze point position using the following formula:
[0076] G t+1 =C t +F(I t ,I t+1 )
[0077] Among them, G t+1 is the dynamic gaze point position, G t is the gaze point position of the current frame, F is the optical flow calculation function, I t and I t+1 It is two consecutive frames of images.
[0078] For example, by analyzing consecutive frames, time series analysis and optical flow methods are used to track the movement of objects in real time, dynamically adjusting the gaze point position to adapt to dynamic scene changes in the video. Consecutive frames are two consecutive frames in a video sequence and are used to analyze object motion. Time series analysis, which predicts future trends by studying how data changes over time, is used to analyze the movement of objects in consecutive frames. Optical flow is used to estimate the direction and speed of pixel movement between consecutive image frames and is commonly used in motion analysis.
[0079] The gaze point optimization unit is used to perform key area priority labeling processing based on the dynamic gaze point position through the gaze point prediction model to generate an optimized gaze point position.
[0080] Specifically, based on the dynamic gaze point position, a gaze point prediction model is used to prioritize labeling of key areas in the visual data. The model optimizes and adjusts the gaze point position to ensure accurate and efficient labeling. The gaze point prediction model is used to predict and optimize the gaze point position, improving labeling efficiency and accuracy. Prioritized labeling of key areas prioritizes labeling of the most informative or important areas in the visual data to improve labeling efficiency.
[0081] In the above-mentioned multimodal data intelligent labeling system based on computer vision, the motion trajectory of objects in the video is predicted through time series analysis and optical flow method, and the position of the gaze point is dynamically adjusted; further, the key areas are prioritized for labeling according to the dynamic gaze point position, and the gaze point position is optimized, so that the key areas in the visual data can be efficiently and accurately located and labeled, significantly improving the labeling efficiency and accuracy.
[0082] In one embodiment, the modality generalization module includes the following units:
[0083] The feature extraction unit is used to obtain different modal data, extract modal invariant features through feature extraction functions, and map different modal data to the same feature space. The extraction formula of modal invariant features is:
[0084] Z=E(X)+E(Y)
[0085] Among them, Z is the modal invariant feature, X and Y are different modal data, and E is the feature extraction function.
[0086] Optionally, by acquiring data from different modalities and processing each modal data using a feature extraction function E, modality-invariant features Z are extracted and mapped into the same feature space for subsequent cross-modal processing and annotation. Module-invariant features are a feature representation that reflects the essential characteristics of the data and is unaffected by modal differences, and are used for unified processing of cross-modal data. A feature space is an abstract mathematical space used to represent and process the extracted features, facilitating comparison and fusion of data from different modalities. Different modal data X and Y refer to data from different data sources or different types, such as images, videos, or point clouds.
[0087] The feature decoupling unit is used to separate the modality-specific information and the general semantic information based on the modality-invariant features through feature decoupling technology to obtain the general semantic features.
[0088] Using modality-invariant features, decoupling ensures that semantic information in data from different modalities is independent of modality-specific information, thereby improving the generalization and consistency of annotation results. Modality-specific information is feature information related to a particular modality, such as the color of an image or the grammatical information of text. Universal semantic information is feature information related to the semantic content of the data and independent of the modality, such as the shape or movement of an object. Decoupling involves breaking down features into independent, unrelated parts using specific algorithms.
[0089] The annotation transfer unit is used to perform cross-modal annotation knowledge transfer based on common semantic features and annotation information of the annotated data using the annotation propagation algorithm and calculate the cross-modal annotation results using the following formula:
[0090] Y p =D(Z g )+S(Y l )
[0091] Among them, Y p is the cross-modal annotation result, D is the decoder, S is the annotation propagation function, and Z g is a general semantic feature, Y l It is the annotation information of the labeled data.
[0092] For example, by utilizing existing annotation information and universal semantic features, annotation knowledge is transferred from the annotated modality to the unannotated modality through an annotation propagation algorithm, thereby achieving cross-modal annotation. Universal semantic features are modality-independent semantic features extracted by a feature decoupling unit and are used for cross-modal annotation. The annotation information of the annotated data is the annotation result of the annotated data, which is used to guide the annotation process of the unlabeled data. The decoder is a model or function used to decode universal semantic features into specific annotation results. The annotation propagation function is used to propagate the annotation information of the annotated data to the unlabeled data, achieving the transfer of annotation knowledge. The cross-modal annotation results are calculated by the annotation migration unit and are applicable to the annotation results of data in different modalities.
[0093] In this computer vision-based multimodal data intelligent annotation system, decoupling technology is used to extract common semantic features, and an annotation propagation algorithm is used to transfer annotation knowledge across modalities, resulting in cross-modal annotation results. This process efficiently transfers annotation knowledge from annotated modalities to unannotated modalities, significantly improving annotation efficiency and generalization capabilities.
[0094] In one embodiment, the annotation optimization module includes the following units:
[0095] The anomaly detection unit is used to detect outliers based on the cross-modal annotation results and the annotations that do not match the data distribution through the statistical model to obtain a preliminary anomaly annotation set.
[0096] Optionally, by analyzing the statistical characteristics of the annotation results and comparing them with the normal data distribution, anomaly annotations that deviate from the normal distribution can be effectively identified. A statistical model is a model based on the statistical characteristics of data, used to analyze the data distribution and identify outliers; the preliminary anomaly annotation set is the set of annotations that do not conform to the normal data distribution.
[0097] The consistency check unit is used to perform multi-view or multi-modal annotation consistency verification based on the preliminary abnormal annotation set through the consistency check function to generate the final consistent annotation result; wherein, the calculation formula of the consistency check function is:
[0098] A c =A⊙M(A)
[0099] Among them, A is the consistency annotation result, M is the consistency check function, and ⊙ is the element-by-element multiplication.
[0100] Specifically, element-by-element multiplication, combined with a consistency check function, ensures that annotation results remain consistent across different viewpoints or modalities. A consistency check function is a function used to verify the consistency of annotation results, implemented using a specific algorithm or model. Element-by-element multiplication is a mathematical operation used to perform element-by-element multiplication on two arrays or matrices of the same dimension.
[0101] The manual correction unit is used to conduct expert review and correction on the detected preliminary anomaly annotation set through an interactive annotation interface to generate the final optimized annotation results.
[0102] Specifically, the automated annotation results are corrected based on the expert's expertise and experience to ensure their accuracy and reliability. The interactive annotation interface is a user interface that allows experts to interact with the annotation system, reviewing and modifying the annotation results. Expert review and correction is the process by which domain experts review and modify the automated annotation results to improve their quality. The final optimized annotation results are high-quality, accurate annotation data after expert review and correction.
[0103] The above-mentioned multimodal data intelligent labeling system based on computer vision can automatically identify and correct outliers in the labeling, ensure the consistency and accuracy of multimodal data labeling, and significantly improve the labeling quality.
[0104] In one embodiment, Figure 2 As shown, the system further includes a visual center module 15, including the following units:
[0105] The data management unit 151 is used to store and manage multimodal visual data. The multimodal visual data supports multiple visual data formats including images, videos, and point clouds to obtain original visual data.
[0106] Efficient data storage and management mechanisms ensure that different types of data can be effectively read and processed by the system. Multimodal visual data is visual data in various formats, including images, videos, and point clouds, which come from different sensors or data sources. Raw visual data is unprocessed visual data obtained directly from the data source and serves as the basis for subsequent data processing and analysis.
[0107] The pipeline processing unit 152 is used to perform preprocessing, feature extraction and annotation generation processing based on the original visual data through the data pipeline to obtain a standardized intermediate processing result.
[0108] Optionally, automated processing can be used to ensure efficient data flow and processing at each stage, providing standardized data support for subsequent annotation tasks. Data pipelines are automated processes that sequentially pass data through multiple processing stages to ensure efficient processing and flow. Standardized intermediate processing results are intermediate data that meet system requirements after pipeline processing and are used for subsequent annotation optimization and storage.
[0109] The distributed computing unit 153 is used to perform parallel computing and load balancing processing based on the intermediate processing results through a distributed computing framework to generate final traceable large-scale annotation data.
[0110] For example, parallel processing and load balancing within a distributed computing framework ensure even distribution of computing tasks across multiple nodes, ultimately generating traceable, large-scale annotated data, significantly improving the efficiency and reliability of large-scale data processing. Distributed computing frameworks, such as MapReduce or Spark, are used to decompose computing tasks into multiple subtasks and execute them in parallel across multiple computing nodes. Load balancing evenly distributes workloads across multiple computing resources to optimize resource utilization, maximize throughput, and avoid system overload. Traceable, large-scale annotated data is generated after distributed computing processing, with detailed processing records, facilitating subsequent query and verification.
[0111] The above-mentioned multimodal data intelligent labeling system based on computer vision can efficiently process and manage large-scale multimodal visual data, ensure the efficiency and scalability of labeling work, and at the same time ensure the traceability of data processing.
[0112] In one embodiment, Figure 3 As shown, the system also includes a multimodal data enhancement and fusion annotation module 16, which includes the following units:
[0113] The data enhancement unit 161 is used to perform data diversity enhancement processing on the original visual data through geometric transformation, illumination adjustment and noise injection technology to generate enhanced data samples.
[0114] Specifically, raw visual data is processed through geometric transformation, lighting adjustment, and noise injection to generate diverse augmented data samples. These techniques increase data diversity and improve the model's adaptability to different conditions. Geometric transformations, including operations such as rotation, flipping, and scaling, are used to change the geometric properties of an image, increasing data diversity. Lighting adjustment simulates different lighting conditions by changing image parameters such as brightness and contrast, improving the model's robustness to lighting changes. Noise injection adds noise, such as Gaussian noise, to the data to enhance the model's ability to resist interference.
[0115] The feature fusion unit 162 is used to perform feature alignment and fusion processing based on the enhanced data samples through cross-modal feature extraction technology to generate a comprehensive annotation result.
[0116] Feature alignment and fusion integrate feature information from different modalities, improving the model's ability to understand and label multimodal data. Augmented data samples are generated through data augmentation and contain more variants. Cross-modal feature extraction extracts features from data of different modalities and aligns them into the same feature space for fusion. Comprehensive annotation results are generated through feature fusion, integrating multimodal information to improve annotation accuracy and reliability.
[0117] The unified annotation unit 163 is used to perform a unified annotation generation process through a multimodal fusion network according to the comprehensive annotation result to obtain a final fusion annotation result.
[0118] For example, a multimodal fusion network is used to integrate feature information from different modalities to generate a unified annotation result, ensuring the accuracy and consistency of the annotation. A multimodal fusion network is a neural network architecture used to integrate feature information from different modalities (such as images, text, and video) to improve information processing and comprehension capabilities. The final fused annotation result is a unified and accurate annotation result obtained after processing by the multimodal fusion network.
[0119] In the above-mentioned computer vision-based multimodal data intelligent labeling system, the multimodal fusion network is used to perform the same labeling to generate the final fused labeling result, which can effectively improve the accuracy and reliability of multimodal data labeling, while enhancing the model's adaptability to different conditions, ensuring the consistency and high quality of the labeling results.
[0120] In order to further illustrate the solution of the embodiment of the present application, a specific example is given below.
[0121] Taking autonomous driving R&D Project A as an example, the system needs to automatically annotate multimodal road environment data collected by the vehicle. The camera installed on the test vehicle collects a 1920×1080 resolution video stream (30fps), and the lidar generates point cloud data at 100,000 points per second, which together constitute the raw visual data input.
[0122] A pretrained ResNet-50 network was used to extract traffic sign feature points from video keyframes, while PointNet++ was used to extract 3D feature points of the same signs from the point cloud. The feature matching unit used the SIFT algorithm to match feature points from both modalities, initially obtaining 23 pairs of reference points. The coordinate alignment unit calculated the transformation matrix T = [R|t], where the rotation matrix R was obtained through SVD decomposition. This was used to map the point cloud reference points to the image coordinate system. After optimization using the LM algorithm, the average alignment error was reduced from an initial 4.7 pixels to 1.2 pixels.
[0123] When the pedestrian is detected at the image coordinate (320, 480) in the video frame t, the motion trajectory prediction unit calculates the optical flow field F(I t ,I t+1 ))Get the displacement vector (12, -5). Substitute it into formula G t+1 =(320,480)+(12,-5)=(332,475), predicting the next frame's gaze point location. When the gaze point optimization unit detects occlusion in this area, it automatically adjusts it to the adjacent high-confidence area (330,472).
[0124] The features of the pedestrian bounding box (X) in the video and the corresponding point cluster (Y) in the point cloud are extracted respectively. After E(X) = [0.7, 1.2, 0.3] and E(Y) = [0.6, 1.1, 0.4], Z = [1.3, 2.3, 0.7] is obtained. After providing feature decoupling to remove the reflection intensity feature unique to the point cloud, Zg = [1.2, 2.1, _] is retained as the general semantic feature. The annotation transfer unit adds the "pedestrian" label Y in the annotated video l By the formula Yp=softmax(Zg)+0.3Y l Migrated to point cloud data, the final output cross-modal annotation confidence is 0.87.
[0125] The vehicle dimensions annotated in one frame of the point cloud were found to be outside the 3σ range (6.2m in length, 3.8±0.9m in mean), marking it as an anomaly. After comparing the multi-view data, the consistency check unit corrected the annotation to 5.1m using Ac=0.9A. The manual correction unit displayed this anomaly to the annotator, who confirmed it was a special engineering vehicle and retained the original annotation.
[0126] 10TB of raw data was organized by timestamp. The pipeline processing unit decoded the video using FFmpeg and voxelized the point cloud (at 0.1m resolution), generating standardized feature vectors. Parallel processing on an 8-node cluster reduced the 100-hour single-machine annotation task to 14 hours.
[0127] Simulated raindrop noise (PSNR = 28dB) was added to the rainy video, and the point cloud was randomly rotated (±5°). The feature fusion unit projected the enhanced point cloud onto the image plane and calculated the fusion weights of 0.73:0.27 using a cross-attention mechanism. The unified annotation unit outputs the final annotation with a multimodal confidence score (0.92 for video and 0.85 for point cloud). The 3D bounding box achieved an Intersection over Union (IoU) of 0.81, a 19% improvement over unimodal annotation.
[0128] Through the above steps, the system can efficiently process and annotate multimodal visual data, provide accurate environmental perception information for autonomous driving vehicles, and ensure the reliability and safety of the autonomous driving system.
[0129] Based on the same inventive concept, the present application also provides a computer vision-based multimodal data intelligent annotation method. The implementation solution provided by this method is similar to the implementation solution described in the above system. Therefore, the specific limitations of one or more computer vision-based multimodal data intelligent annotation method embodiments provided below can be found in the above-mentioned limitations of a computer vision-based multimodal data intelligent annotation system, and will not be repeated here.
[0130] In an exemplary embodiment, Figure 4 As shown, a multimodal data intelligent labeling method based on computer vision is provided, including:
[0131] S111, obtaining feature extraction results of multimodal data, performing reference point alignment processing on the feature extraction results, and generating a consistent reference point set;
[0132] S112, performing dynamic prediction processing on the fixation point in the visual data according to the reference point set to generate a dynamic fixation point position;
[0133] S113, performing feature space mapping processing based on the dynamic gaze point position through a cross-modal learning framework to obtain annotation results applicable to multiple modalities;
[0134] S114, perform quality assessment on the annotation results through statistical models and consistency checks to generate optimized annotation data.
[0135] In one embodiment, obtaining feature extraction results of multimodal data, performing reference point alignment processing on the feature extraction results, and generating a consistent reference point set include:
[0136] S211, based on the feature point set, performing cross-modal feature point matching through a feature matching algorithm to generate an original reference point set;
[0137] S212: Align the original reference point set to the same coordinate system through a coordinate transformation matrix to obtain a spatially consistent reference point set.
[0138] In one embodiment, dynamically predicting a gaze point in visual data based on a reference point set to generate a dynamic gaze point position includes:
[0139] S311, based on the continuous frame images, the motion trajectory of the object in the video is predicted through time series analysis and optical flow method, and the dynamic gaze point position is calculated using the following formula:
[0140] G t+1 =G t +F(I t ,I t+1 )
[0141] Among them, G t+1 is the dynamic gaze point position, G t is the gaze point position of the current frame, F is the optical flow calculation function, I t and I t+1 It is two consecutive frames of images;
[0142] S312: Based on the dynamic gaze point position, the key area is prioritized and labeled using the gaze point prediction model to generate an optimized gaze point position.
[0143] In one embodiment, based on the dynamic gaze point position, feature space mapping is performed through a cross-modal learning framework to obtain annotation results applicable to multiple modalities, including:
[0144] S411, obtain different modal data, extract modal invariant features through feature extraction function, and map different modal data to the same feature space. The extraction formula of modal invariant features is:
[0145] Z=E(X)+E(Y)
[0146] Among them, Z is the modal invariant feature, X and Y are different modal data, and E is the feature extraction function;
[0147] S412, based on the modality-invariant features, separating the modality-specific information from the general semantic information using a feature decoupling technique to obtain a general semantic feature;
[0148] S413: Using the following formula, based on the common semantic features and the annotation information of the annotated data, the annotation propagation algorithm is used to perform cross-modal annotation knowledge transfer and calculate the cross-modal annotation results:
[0149] Y p =D(Z g )+S(Y l )
[0150] Among them, Y p is the cross-modal annotation result, D is the decoder, S is the annotation propagation function, and Z g is a general semantic feature, Y l It is the annotation information of the labeled data.
[0151] In one embodiment, the quality of the annotation results is assessed using statistical models and consistency checks to generate optimized annotation data, including:
[0152] S511, based on the cross-modal annotation results, outlier detection is performed on the annotations that do not match the data distribution using a statistical model to obtain a preliminary anomaly annotation set;
[0153] S512: Based on the preliminary abnormal annotation set, a consistency check function is used to perform multi-view or multi-modal annotation consistency verification to generate a final consistent annotation result; wherein the calculation formula of the consistency check function is:
[0154] A c =A⊙M(A)
[0155] Where A is the consistency annotation result, M is the consistency check function, and ⊙ is the element-by-element multiplication;
[0156] S513, the detected preliminary anomaly annotation set is reviewed and revised by experts through an interactive annotation interface to generate the final optimized annotation data.
[0157] In one embodiment, the method further comprises:
[0158] S611, storing and managing multimodal visual data, where the multimodal visual data supports multiple visual data formats including images, videos, and point clouds, to obtain original visual data;
[0159] S612, based on the original visual data, preprocessing, feature extraction and annotation generation are performed through the data pipeline to obtain a standardized intermediate processing result;
[0160] S613, based on the intermediate processing results, parallel computing and load balancing processing are performed through a distributed computing framework to generate the final traceable large-scale annotation data.
[0161] In one embodiment, the method further comprises:
[0162] S711, performing data diversity enhancement processing on the original visual data through geometric transformation, illumination adjustment and noise injection technology to generate enhanced data samples;
[0163] S712, based on the enhanced data samples, feature alignment and fusion processing is performed through cross-modal feature extraction technology to generate a comprehensive annotation result;
[0164] S713: Based on the comprehensive annotation results, the same annotation generation process is performed through the multimodal fusion network to obtain the final fusion annotation result.
[0165] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0166] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, a module of a multimodal data intelligent annotation system based on computer vision as described above is implemented.
[0167] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the modules in the above-mentioned system embodiments are implemented.
[0168] For the device embodiments, since they basically correspond to the system embodiments, the relevant parts can be referred to the partial description of the system embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. Those of ordinary skill in the art can understand and implement it without expending creative work.
[0169] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.
Claims
1. A multimodal data intelligent annotation system based on computer vision, characterized by: The system comprises: A reference point alignment module is used to obtain feature extraction results of multimodal data, perform reference point alignment processing on the feature extraction results, and generate a consistent reference point set; A gaze point prediction module, configured to dynamically predict the gaze point in the visual data based on the reference point set to generate a dynamic gaze point position; The modality generalization module is used to perform feature space mapping processing based on the dynamic gaze point position through a cross-modal learning framework to obtain cross-modal annotation results; The annotation optimization module is used to perform quality assessment on the annotation results through statistical models and consistency checks to generate optimized annotation results.
2. The system according to claim 1, wherein: The reference point alignment module includes: A data acquisition unit is used to acquire key feature points in images, videos, and point cloud data through a deep learning model to obtain a feature point set; A feature matching unit, configured to perform cross-modal feature point matching based on the feature point set using a feature matching algorithm to generate an original reference point set; The coordinate alignment unit is used to align the original reference point set into the same coordinate system through a coordinate transformation matrix to obtain the spatially consistent reference point set.
3. The system according to claim 1, wherein: The gaze point prediction module includes: The motion trajectory prediction unit is used to predict the motion trajectory of objects in the video based on continuous frame images through time series analysis and optical flow method, and calculate the dynamic gaze point position using the following formula: G t+1 =G t +F(I t ,I t+1 ) Among them, G t+1 is the dynamic gaze point position, G t is the gaze point position of the current frame, F is the optical flow calculation function, I t and I t+1 It is two consecutive frames of images; A gaze point optimization unit is used to perform key area priority labeling processing based on the dynamic gaze point position through a gaze point prediction model to generate an optimized gaze point position.
4. The system according to claim 1, wherein: The modality generalization module includes: The feature extraction unit is used to obtain different modal data, extract modal invariant features through a feature extraction function, and map the different modal data to the same feature space. The extraction formula of the modal invariant features is: Z=E(X)+E(Y) Among them, Z is the modal invariant feature, X and Y are different modal data, and E is the feature extraction function; a feature decoupling unit, configured to separate modality-specific information and general semantic information based on the modality-invariant features by using a feature decoupling technique to obtain general semantic features; The annotation transfer unit is configured to use the following formula to perform cross-modal annotation knowledge transfer based on the common semantic features and the annotation information of the annotated data through an annotation propagation algorithm to calculate the cross-modal annotation result: Y p =D(Z g )+S(Y l ) Among them, Y p is the cross-modal annotation result, D is the decoder, S is the annotation propagation function, and Z g is a general semantic feature, Y l It is the annotation information of the labeled data.
5. The system according to claim 1, wherein: The annotation optimization module includes: an anomaly detection unit, configured to detect outliers based on the cross-modal annotation results and the annotations that do not match the data distribution using a statistical model, to obtain a preliminary anomaly annotation set; A consistency checking unit is configured to perform multi-view or multi-modal annotation consistency verification based on the preliminary anomaly annotation set using a consistency checking function to generate a final consistent annotation result; wherein the calculation formula of the consistency checking function is: A c =A⊙M(A) Where A is the consistency annotation result, M is the consistency check function, and ⊙ is the element-by-element multiplication; The manual correction unit is used to perform expert review and correction on the detected preliminary abnormal annotation set through an interactive annotation interface to generate final optimized annotation data.
6. The system according to any one of claims 1 to 5, characterized in that The system also includes a visual center module, which includes: A data management unit, configured to store and manage multimodal visual data in various visual data formats including images, videos, and point clouds, to obtain raw visual data; A pipeline processing unit, configured to perform preprocessing, feature extraction, and annotation generation processing on the raw visual data through a data pipeline to obtain a standardized intermediate processing result; The distributed computing unit is used to perform parallel computing and load balancing processing based on the intermediate processing results through a distributed computing framework to generate final traceable large-scale annotation data.
7. The system according to claim 6, characterized in that The system also includes a multimodal data enhancement and fusion annotation module, which includes: a data enhancement unit, configured to perform data diversity enhancement processing on the original visual data through geometric transformation, illumination adjustment, and noise injection techniques to generate enhanced data samples; A feature fusion unit, configured to perform feature alignment and fusion processing based on the enhanced data samples through a cross-modal feature extraction technology to generate a comprehensive annotation result; The unified annotation unit is used to perform a unified annotation generation process through a multimodal fusion network according to the comprehensive annotation result to obtain a final fusion annotation result.
8. A multimodal data intelligent labeling method based on computer vision, characterized in that: The method comprises: Acquiring feature extraction results of multimodal data, performing reference point alignment processing on the feature extraction results, and generating the consistent reference point set; Performing dynamic prediction processing on the gaze point in the visual data according to the reference point set to generate a dynamic gaze point position; Based on the dynamic gaze point position, feature space mapping is performed through a cross-modal learning framework to obtain annotation results applicable to multiple modalities. The annotation results are quality assessed through statistical models and consistency checks to generate optimized annotation data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the system according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the system according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Automatic driving scene labeling method and system for multi-modal data fusion
CN120954000A
Automatic driving scene labeling method and system for multi-modal data fusion
CN120954000B