Labeling method and device based on multi-modal model

By using multimodal models in data annotation to process text, images and point cloud data, the problem of model customization and cross-domain data in the prior art is solved, and more efficient and accurate data annotation is achieved.

CN119939241AActive Publication Date: 2025-05-06ZHEJIANG WUWEN ZHIXING TECHNOLOGY CO LTD

Patent Information

Application Number
CN202411897602.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

The existing AI-based data annotation methods have problems such as serious model customization and difficult to solve cross-domain data business requirements, which leads to the inability to meet the requirements of labeling efficiency and accuracy.

Method used

The annotation method based on multimodal model is adopted, and data preprocessing is performed by receiving the client terminal data set, including text word segmentation, image segmentation and feature extraction, point cloud feature extraction, to build matching space, generate multimodal models, and to perform prediction and annotation.

Benefits of technology

It improves the efficiency and accuracy of data annotation, can adapt to different customer needs, solve cross-domain data problems, and improves the flexibility and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939241A_ABST
    Figure CN119939241A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an annotation method and device based on a multi-modal model, and the method comprises the steps: inputting a client terminal data set into a preset initial model, segmenting text data according to a set adaptive word segmentation algorithm, obtaining a text vector, segmenting image data according to a preset image segmentation algorithm, and obtaining an image vector; feature extraction is carried out on the target image area, the foreground image area and the background image area obtained after segmentation to obtain image vectors, multi-scale point cloud features of the three-dimensional point cloud data are extracted according to a preset multi-scale feature extraction network to obtain point cloud vectors, and the point cloud vectors are subjected to point cloud feature extraction; and inputting the text vector, the image vector and the point cloud vector into a backbone network of a preset initial model to perform matching space construction operation, determining a corresponding multi-modal model, performing prediction labeling operation on the client terminal data according to the multi-modal model, determining a prediction result of the client terminal data output by the multi-modal model, and outputting the prediction result of the client terminal data. According to the invention, the efficiency and accuracy of data annotation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to a labeling method and device based on a multimodal model. Background Art

[0002] In the field of data annotation, with the rapid development of big data and artificial intelligence technologies, the requirements for data annotation efficiency and accuracy are increasing. Traditional data annotation methods rely on manual annotation, which is inefficient and costly. To improve annotation efficiency, automatic annotation methods based on AI technology have emerged.

[0003] The core of existing AI-based methods for improving annotation efficiency lies in designing corresponding AI models based on the application scenarios proposed by customers and using the service provider's existing data to train the models. The trained AI models are then used in the annotation business to improve annotation efficiency. However, this approach has the following major problems:

[0004] 1. Model customization is critical: Different customers have vastly different needs, requiring the design of unique models and training based on diverse data. This places high demands on service providers' infrastructure, algorithm development capabilities, and data reserves.

[0005] 2. Cross-domain data business needs are difficult to address: If the data in the customer's business needs is cross-domain with the supplier's data (e.g., different sensors, different countries, cities, or seasons of data collection), the performance of the model trained based on existing stock data on the customer data will be severely degraded, and AI capabilities will be difficult to provide significant assistance to labeling efficiency.

[0006] Based on the above problems, the existing AI-based labeling methods cannot meet the requirements of data labeling efficiency and accuracy in the field of data labeling. Summary of the Invention

[0007] In response to the problems in the prior art, the present application provides a labeling method and device based on a multimodal model, which can improve the efficiency and accuracy of data labeling.

[0008] In order to solve at least one of the above problems, the present application provides the following technical solutions:

[0009] In a first aspect, the present application provides a labeling method based on a multimodal model, comprising:

[0010] receiving a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data;

[0011] Input the client terminal data set into the data preprocessing module of the preset initial model, perform a segmentation operation on the text data according to the set adaptive word segmentation algorithm to determine the corresponding text vector, perform a segmentation operation on the image data according to the preset image segmentation algorithm to determine the corresponding target image area, foreground image area and background image area, perform a high-definition feature extraction operation on the foreground image area to determine the corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area to determine the corresponding background feature, perform a feature extraction operation on the target image area to determine the corresponding target feature, perform a vectorization operation on the target feature, the foreground feature and the background feature to determine the corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to the preset multi-scale feature extraction network to determine the corresponding point cloud vector, input the text vector, the image vector and the point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation to determine the corresponding multimodal model;

[0012] A predictive annotation operation is performed on the client terminal data according to the multimodal model to determine a prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes predicted text annotation, predicted image annotation, and predicted three-dimensional point cloud annotation.

[0013] Furthermore, before segmenting the text data according to the set word segmentation algorithm and determining the corresponding text vectors, the method further includes:

[0014] Collecting a historical text data set, performing a word segmentation operation on the historical text data set, and determining a corresponding minimum text subunit;

[0015] Inputting the minimum text subunits into a preset long-short term network for model training, and determining the dependency relationship between the minimum text subunits output by the long-short term network;

[0016] According to the dependency relationship, the preset verification set dependency relationship and the preset cross entropy loss function, the corresponding loss function is determined, and the parameters of the long-term and short-term network are updated according to the optimal loss function and the reverse update algorithm to determine the corresponding adaptive word segmentation algorithm.

[0017] Furthermore, segmenting the text data according to a set word segmentation algorithm to determine corresponding text vectors includes:

[0018] Analyze the text data according to a set word segmentation algorithm to determine corresponding text data segmentation points, segment the text data according to the text data segmentation points, and determine corresponding optimal text subunits;

[0019] The optimal text sub-unit is vectorized according to a preset word vector model to determine a corresponding text vector.

[0020] Furthermore, performing a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to a preset multi-scale feature extraction network to determine the corresponding point cloud vector includes:

[0021] Performing a sampling operation on the three-dimensional point cloud data according to a preset far-point sampling algorithm to determine a corresponding centroid point, and performing a neighborhood construction operation on each of the centroid points according to a preset multi-scale radius to determine a corresponding multi-scale neighborhood;

[0022] A feature extraction operation is performed on the multi-scale neighborhood, and a feature fusion operation is performed on the multi-scale neighborhood features obtained after the feature extraction operation to determine the multi-scale features corresponding to the centroid point, and a vectorization operation is performed on the multi-scale features to determine the corresponding point cloud vector.

[0023] Furthermore, the step of inputting the text vector, the image vector, and the point cloud vector into a backbone network of a preset initial model to perform a matching space construction operation and determine a corresponding multimodal model includes:

[0024] Inputting the text vector, the image vector, and the point cloud vector into a backbone network of a preset initial model, performing a distance measurement operation on the text vector, the image vector, and the point cloud vector according to a preset shared space, and determining corresponding alignment;

[0025] A matching space construction operation is performed on the text vector, the image vector, and the point cloud vector according to the alignment degree to determine a corresponding multimodal vector mapping relationship, and a corresponding multimodal model is determined according to the multimodal vector mapping relationship.

[0026] Furthermore, after performing the prediction and labeling operation on the client terminal data according to the multimodal model and determining the prediction result of the client terminal data output by the multimodal model, the method further includes:

[0027] Perform a difference comparison operation based on the predicted result and the set actual result to determine the corresponding deviation value.

[0028] The multimodal model is updated according to the deviation value and the back propagation algorithm to determine the updated multimodal model, wherein the set true result is determined by the annotator end.

[0029] Furthermore, before performing a difference comparison operation based on the predicted result and the set actual result to determine the corresponding deviation value, the method includes:

[0030] The annotator manually annotates the dynamic data in the client terminal data at the annotator end to determine the corresponding dynamic data envelope;

[0031] The annotator manually annotates the static data in the client terminal data on the annotator side to determine the corresponding static data polyline;

[0032] The corresponding real result is determined according to the dynamic data envelope and the static data broken line.

[0033] In a second aspect, the present application provides a labeling device based on a multimodal model, comprising:

[0034] A data receiving module, configured to receive a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data;

[0035] a multimodal model training module, configured to input the client terminal data set into a data preprocessing module of a preset initial model, perform a segmentation operation on the text data according to a set adaptive word segmentation algorithm to determine a corresponding text vector, perform a segmentation operation on the image data according to a preset image segmentation algorithm to determine a corresponding target image area, a foreground image area, and a background image area, perform a high-definition feature extraction operation on the foreground image area to determine a corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area to determine a corresponding background feature, perform a feature extraction operation on the target image area to determine a corresponding target feature, perform a vectorization operation on the target feature, the foreground feature, and the background feature to determine a corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to a preset multi-scale feature extraction network to determine a corresponding point cloud vector, input the text vector, the image vector, and the point cloud vector into a backbone network of a preset initial model to perform a matching space construction operation to determine a corresponding multimodal model;

[0036] A prediction and annotation module is used to perform a prediction and annotation operation on the client terminal data according to the multimodal model, and determine the prediction results of the client terminal data output by the multimodal model, wherein the prediction results include predicted text annotation, predicted image annotation and predicted three-dimensional point cloud annotation.

[0037] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the multimodal model-based labeling method when executing the program.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal model-based annotation method.

[0039] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the multimodal model-based annotation method.

[0040] It can be seen from the above technical solution that the present application provides a labeling method and device based on a multimodal model, which inputs a client terminal data set into a preset initial model, segments the text data according to a set adaptive word segmentation algorithm to obtain a text vector, segments the image data according to a preset image segmentation algorithm, performs feature extraction on the target image area, foreground image area and background image area obtained after segmentation to obtain an image vector, extracts multi-scale point cloud features of the three-dimensional point cloud data according to a preset multi-scale feature extraction network to obtain a point cloud vector, inputs the text vector, image vector and point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation, determines the corresponding multimodal model, performs a predictive labeling operation on the client terminal data according to the multimodal model, and determines the prediction result of the client terminal data output by the multimodal model, thereby improving the efficiency and accuracy of data labeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] Figure 1 This is a flowchart of a multimodal model-based annotation method according to an embodiment of the present application;

[0043] Figure 2 This is a second flow chart of the multimodal model-based annotation method in an embodiment of the present application;

[0044] Figure 3 This is a third flow chart of the multimodal model-based annotation method in an embodiment of the present application;

[0045] Figure 4 This is a fourth flow chart of the multimodal model-based annotation method in an embodiment of the present application;

[0046] Figure 5 This is a fifth flowchart of the multimodal model-based annotation method in an embodiment of the present application;

[0047] Figure 6 This is a sixth flowchart of the multimodal model-based annotation method in an embodiment of the present application;

[0048] Figure 7 FIG7 is a flowchart of a multimodal model-based annotation method according to an embodiment of the present application;

[0049] Figure 8 This is a structural diagram of a multimodal model-based annotation device in an embodiment of the present application;

[0050] Figure 9 Schematic diagram of the structure of the electronic device in the embodiment of the present application.

[0051] Reference numerals:

[0052] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION

[0053] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0054] The acquisition, storage, use, and processing of data in this application's technical solution comply with relevant national laws and regulations.

[0055] Considering the problem that the existing AI-based annotation method cannot meet the requirements of data annotation efficiency and accuracy in the field of data annotation. The present application provides a annotation method and device based on a multimodal model, which inputs a client terminal data set into a preset initial model, segments the text data according to a set adaptive word segmentation algorithm to obtain a text vector, segments the image data according to a preset image segmentation algorithm, performs feature extraction on the target image area, foreground image area and background image area obtained after segmentation to obtain an image vector, extracts multi-scale point cloud features of the three-dimensional point cloud data according to a preset multi-scale feature extraction network to obtain a point cloud vector, inputs the text vector, image vector and point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation, determines the corresponding multimodal model, performs a predictive annotation operation on the client terminal data according to the multimodal model, and determines the prediction result of the client terminal data output by the multimodal model, thereby improving the efficiency and accuracy of data annotation.

[0056] In order to improve the efficiency and accuracy of data annotation, this application provides an embodiment of an annotation method based on a multimodal model. Figure 1 The multimodal model-based annotation method specifically includes the following contents:

[0057] Step S101: receiving a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data;

[0058] Optionally, in this embodiment, the client terminal data set is used to train a multimodal model. Therefore, the data set needs to have multiple data forms, and the annotated data includes text, images (2D data), 3D point cloud data, etc.

[0059] It is understandable that compared with existing technologies, traditional models tend to only process a specific type of single data (such as images or text) and lack a comprehensive understanding of different data modalities. Therefore, when dealing with complex and diverse annotation tasks, for example, when there is text data in a picture, the performance of traditional models is relatively limited. The multimodal capability obtained by this technical solution through multimodal data training enables the model to better adapt to various different data sources and application scenarios.

[0060] Step S102: Input the client terminal data set into the data preprocessing module of the preset initial model, perform a segmentation operation on the text data according to the set adaptive word segmentation algorithm to determine the corresponding text vector, perform a segmentation operation on the image data according to the preset image segmentation algorithm to determine the corresponding target image area, foreground image area and background image area, perform a high-definition feature extraction operation on the foreground image area to determine the corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area to determine the corresponding background feature, perform a feature extraction operation on the target image area to determine the corresponding target feature, perform a vectorization operation on the target feature, the foreground feature and the background feature to determine the corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to the preset multi-scale feature extraction network to determine the corresponding point cloud vector, input the text vector, the image vector and the point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation to determine the corresponding multimodal model;

[0061] Optionally, in this embodiment, this step innovates in the data preprocessing part by converting different types of data (text, pictures, point clouds, etc.) into a unified token (unit element) format and sending it into the Transformer (deep learning model) architecture for processing, so that data of different modalities can be interactively inferred in the same model, increasing the flexibility and adaptability of the model.

[0062] Optionally, in this embodiment, data collected by the terminal enters the preprocessing module of the Multimodal Large Model (MMLM) for preprocessing. Different types of data (such as images, text, and point clouds) are converted into tokens in a unified format through different preprocessing processes. For example, text is tokenized, and images and point cloud data are converted into tokens through patch segmentation and feature extraction.

[0063] Optionally, in this embodiment, the preprocessed Token is input into the inference module of the multimodal large model (MMLM). By matching the space, the AI ​​model will output the inference result. In the subsequent step S103, the model training is performed through the feedback mechanism.

[0064] Optionally, in this embodiment, in a multimodal AI model (MMLM model based on the Transformer architecture), the preprocessing step is a key stage for converting input data into a form that the model can process. For data of different modalities, the core task of preprocessing is tokenization to ensure that different types of data can be uniformly represented and input into the neural network for processing.

[0065] Optionally, in this embodiment, the text data is tokenized.

[0066] Optionally, natural language processing (NLP) techniques, such as WordPiece (a word segmentation algorithm), Byte Pair Encoding (BPE) (a machine translation algorithm), or SentencePiece (a subword matching algorithm), can be used to tokenize text data. These methods break down the input text into words or subwords, which are then converted into corresponding tokens. The main purpose of tokenization is to convert text content into a series of digital tokens, which are then fed into the model as an input sequence.

[0067] Preferably, for long or compound words in the text, an adaptive word segmentation algorithm is used to segment the words into more appropriate subword units. This method can better handle polysemous words and rare words, and improve the model's adaptability to different languages.

[0068] Specifically, adaptive word segmentation algorithms use contextual models like LSTM to process text, extracting semantic information about each word or character within its context. They then dynamically adjust the word segmentation points based on this information. When processing sequential data, LSTMs remember information about the preceding and following sequences. Therefore, they can analyze the context of preceding and following words to determine whether the current word requires further segmentation. For example, when a word is a compound word in a specific context, LSTM can use its memory mechanism to adjust the segmentation strategy.

[0069] Another example is an adaptive word segmentation algorithm that uses a contextual model like the Transformer to process text. Through its self-attention mechanism, the Transformer can simultaneously consider the context of the entire sentence at each position in the text. This self-attention mechanism allows the model to adjust the position of word segments based on the global context. After the model captures the relationship between each word in the context, it dynamically adjusts the word segmentation boundaries based on this contextual information.

[0070] It's worth noting that text data first undergoes word segmentation. After word segmentation, the text data is converted into tokens and fed into the model for processing. At this point, the model adjusts word boundaries through contextual analysis to ensure more accurate text processing. During joint training, preprocessed data from different modalities, such as text, images, and point clouds, is integrated into the model. This joint training enables the model to understand the relationships between these different data types. During this process, the text's word segmentation is also optimized and adjusted based on its context within the overall multimodal task.

[0071] Finally, the text that has undergone adaptive word segmentation processing is converted into a vector model through word embedding, laying the foundation for the subsequent construction of a matching space with a unified token input model backbone network.

[0072] Optionally, in this embodiment, the image data is tokenized.

[0073] Optionally, during image data preprocessing, different processing strategies are applied according to different regions of the image (foreground, background, target region, etc.).

[0074] Specifically, an image is first divided into different regions, including foreground, background, and target regions, using image segmentation algorithms (such as U-Net and FCN). Then, based on the different characteristics of these regions, appropriate preprocessing methods are dynamically selected, including contrast enhancement and noise removal for the target region and specific processing (such as high definition, blurring, and noise reduction) for the background region.

[0075] For example, a higher resolution is used in the target and foreground regions to preserve more detail, while a lower resolution is used in the background regions to reduce computational effort. Information from different regions in the image is extracted as feature vectors, and regional adaptation is performed within the model based on the features of these regions. This means the model can identify and process different regions of the image differently. For example, for the foreground region of an image, the model might select a more detailed feature representation, while for the background region, it might select a more simplified feature representation.

[0076] It's understandable that this region-based dynamic adjustment method can flexibly adjust preprocessing methods based on the characteristics of different image regions, thereby improving the model's processing efficiency and accuracy. During subsequent model applications, the model can adjust its processing methods for different regions in real time. For example, when annotators annotate a new region, the model can learn and adapt to the characteristics of this new area, dynamically optimizing the processing method.

[0077] Finally, the preprocessed image data is converted into a unified vector format, laying the foundation for the subsequent construction of the matching space using a unified token input model backbone network.

[0078] Optionally, in this embodiment, the three-dimensional point cloud data is tokenized.

[0079] Optionally, point cloud data usually contains multi-scale information. In order to process multi-scale point cloud data, the point cloud data needs to be decomposed through a hierarchical structure first.

[0080] Specifically, when processing point cloud data, the farthest point sampling (FPS) strategy is first used to select a set of key points from the input point cloud. Then, for each point, it not only considers the surrounding local point cloud data, but also extracts features through different sampling scales. For example, you can start with a small neighborhood to extract detail information, and then expand to a larger neighborhood to capture larger contextual information. Then, from multiple levels, each layer uses the feature extraction method of the local point cloud to effectively extract details from the point cloud data. By aggregating features at different scales, the network can comprehensively consider information at multiple levels from local details to global structures.

[0081] In terms of effect, the multi-scale feature extraction capability can significantly improve the adaptability of the model in complex point cloud scenes. Including:

[0082] 1) Capturing diverse object forms: The representations of different objects in point clouds sometimes vary greatly in scale and density. PointNet++ can handle this difference and capture the various representations of objects through a hierarchical structure.

[0083] 2) Enhanced recognition of small objects: For small objects, the network can focus on details through a fine local feature extraction module without being disturbed by a large range of background.

[0084] Finally, the preprocessed 3D point cloud data is converted into a unified vector format, laying the foundation for the subsequent construction of the matching space using a unified token input model backbone network.

[0085] Optionally, in this embodiment, in the training of a large multimodal model, the construction of a matching space is a crucial step. Its main purpose is to ensure that the model can efficiently perform multimodal reasoning and achieve rapid convergence by effectively matching the features of different modal data (such as text, images, 3D point clouds, etc.).

[0086] Optionally, a neural network based on the Transformer architecture is used, which is able to process different types of data.

[0087] Text data: The text is processed through tokenization to generate a series of tokens, which are then input into the Transformer model.

[0088] Image data: The image is split into multiple small patches (Patch), and each patch is converted into a vector representation (Token) through feature extraction method.

[0089] 3D point cloud data: The coordinates of each point cloud data (such as x, y, z) are processed as tokens, and spatial features are extracted through a specific network structure.

[0090] Optionally, next, similarities between different modal data are calculated to establish a mapping relationship.

[0091] Specifically, by calculating the alignment between different modalities, different modalities (such as images and text, or 3D point clouds and text) are transformed into a shared space. This shared space effectively unifies data from different modalities within the same feature space for comparison and matching. Using cross-modal embedding techniques, data from different modalities are embedded into a common feature space. This allows the model to understand the relationships and interdependencies between different modalities. In this shared space, distance metrics (such as Euclidean distance or cosine similarity) are used to calculate the matching between different modalities. Only when the distance between the features of two modalities in the shared space is sufficiently small are they considered matched. Then, for the matched data from different modalities, a self-attention mechanism is used to perform a weighted fusion of multiple modalities, such as images, text, and point clouds. This allows the model to dynamically adjust the contribution of data based on the importance of each modality. This construction of a matching space improves the model's adaptability to complex datasets, not only improving model accuracy but also providing strong support for the effective integration of multimodal learning frameworks.

[0092] Understandably, the matching space is not fixed due to the varying needs of different customers and application scenarios. The matching space is dynamically adjusted based on new annotation data to ensure that each device's annotation task receives appropriate training and inference support. Each time annotators reject an annotation result, a new training cycle is triggered to update the model and optimize the matching space.

[0093] After the matching space is constructed, a multimodal model is obtained, which is trained and tuned based on a feedback mechanism.

[0094] Step S103: performing a predictive annotation operation on the client terminal data according to the multimodal model to determine a prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes predicted text annotation, predicted image annotation, and predicted three-dimensional point cloud annotation.

[0095] Optionally, this step predicts and labels the client terminal data by building a multimodal model of the matching space to obtain the labeling results. The intelligent agent platform is deployed on the annotator side, and the annotator side and the multimodal model perform labeling simultaneously.

[0096] Subsequently, the annotators will inspect and accept the multimodal model predictions. The annotators’ feedback on the inspection will be compared with the inference results of the multimodal large model. The differences will be used for feedback model iteration and training to achieve rapid convergence of the neural network model and training of the multimodal large model.

[0097] This example demonstrates how this embodiment trains a model to process multimodal data and performs model self-training based on a feedback mechanism to implement the terminal agent intelligent labeling paradigm.

[0098] From the above description, it can be seen that the multimodal model-based labeling method provided in the embodiment of the present application can obtain text vectors by inputting the client terminal data set into a preset initial model, segmenting the text data according to the set adaptive word segmentation algorithm, and segmenting the image data according to the preset image segmentation algorithm. Feature extraction is performed on the target image area, foreground image area, and background image area obtained after segmentation to obtain image vectors. Multi-scale point cloud features of three-dimensional point cloud data are extracted according to a preset multi-scale feature extraction network to obtain point cloud vectors. The text vector, image vector, and point cloud vector are input into the backbone network of the preset initial model for matching space construction operations, and the corresponding multimodal model is determined. The client terminal data is predicted and labeled according to the multimodal model to determine the prediction results of the client terminal data output by the multimodal model, thereby improving the efficiency and accuracy of data labeling.

[0099] In one embodiment of the multimodal model-based annotation method of the present application, see Figure 2 , and can also include the following:

[0100] Step S201: collecting a historical text dataset, performing a word segmentation operation on the historical text dataset, and determining a corresponding minimum text sub-unit;

[0101] Step S202: inputting the minimum text subunits into a preset long-short term network for model training, and determining the dependency relationship between the minimum text subunits output by the long-short term network;

[0102] Step S203: Determine the corresponding loss function according to the dependency relationship, the preset validation set dependency relationship and the preset cross entropy loss function, and update the parameters of the long-term and short-term network according to the optimal loss function and the reverse update algorithm to determine the corresponding adaptive word segmentation algorithm.

[0103] Optionally, in this embodiment, the purpose of this step is to train an adaptive word segmentation model and implement the data preprocessing process in the multimodal model through this model. The core goal of the adaptive word segmentation strategy is to select the most suitable word segmentation method based on the different characteristics of the text data (such as register, content, context, etc.).

[0104] Optionally, in this embodiment, a large amount of historical text data is first collected and basic text cleaning operations are performed, including removing useless symbols, standardizing punctuation marks, and processing whitespace characters. For multilingual text data, language recognition is required to classify texts in different languages and select appropriate tokenizers for processing. Stop words (such as "of", "is", "in", etc.) are removed to reduce noise and improve the tokenization effect.

[0105] Secondly, a tokenization strategy is selected and tokenization annotation and processing are performed according to the tokenization strategy.

[0106] Rule-based tokenization: For some structured texts (such as legal documents, technical documents, etc.), rule-driven tokenization methods can be used to perform tokenization based on dictionaries, regular expressions, etc.

[0107] Subword Tokenization: The BPE algorithm is used to segment out-of-vocabulary words, breaking words into smaller units to improve the robustness of the model.

[0108] Context-aware tokenization: In some scenarios with rich context information, pre-trained language models (such as BERT) can be used to determine word boundaries. For example, context information is used to handle the segmentation of some polysemous words or compound words.

[0109] After tokenization annotation of the words, they are input into a long short-term network for model training so that the long short-term network can recognize the dependencies between words and output the text tokenization predicted by the model according to the dependencies.

[0110] Finally, the cross-entropy loss is calculated between the true value results in the validation set and the text tokenization results predicted by the model. The calculated loss function dynamically adjusts the parameters of the long short-term network through backpropagation until the optimal parameters are obtained, so that the long short-term network can extract the semantic information of each word or character in the text in context, and then dynamically adjust the tokenization split points based on this information.

[0111] Through step S203, this embodiment successfully trains an adaptive tokenization algorithm, laying a foundation for the subsequent text preprocessing process of the multimodal model.

[0112] In an embodiment of the annotation method based on a multimodal model of the present application, referring to Figure 3 , it may specifically include the following content:

[0113] Step S301: Analyze the text data according to the set tokenization algorithm to determine the corresponding text data split points, and perform a splitting operation on the text data according to the text data split points to determine the corresponding optimal text sub-units;

[0114] Step S302: Perform vectorization on the optimal text sub-unit according to a preset word vector model to determine the corresponding text vector.

[0115] Optionally, in this embodiment, the purpose of this step is to preprocess the text data to obtain a text vector, laying a foundation for subsequent joint learning and reasoning with image vectors and point cloud vectors in the same feature space.

[0116] Optionally, in this embodiment, the text data in the obtained customer terminal data is input into a data preprocessing module, and the text data is segmented according to the adaptive word segmentation algorithm obtained in step S203. Specifically, the adaptive word segmentation algorithm will identify the optimal segmentation points of the text data and segment the text accordingly.

[0117] For example, assume the input text is "The annotator manually annotated the text data". After tokenization, it may be decomposed into the following tokens: ["annotator", "manual", "annotated", "the", "text", "data"].

[0118] In terms of effect, adaptive word segmentation dynamically adjusts the word segmentation method according to the needs of different fields. For example, when processing technical documents, it can identify and split professional terms, while when processing social media text, it can identify and correctly segment some slang, abbreviations, etc.

[0119] Adaptive word segmentation not only depends on the dictionary but also combines context information to avoid word segmentation errors caused by polysemy or context dependence and improve accuracy. For example, for the word "apple", whether it is "fruit" or "brand" in the context, the word segmentation result can adapt automatically.

[0120] Adaptive word segmentation effectively processes out-of-vocabulary (OOV) words through sub-word tokenization (such as BPE, WordPiece, etc.). Even when encountering new words, the algorithm can generate appropriate word segmentation results based on existing vocabulary and context.

[0121] Finally, through the method of word embedding, these tokens obtained by adaptive word segmentation are further transformed into a numerical representation that the model can process, such as the corresponding word vector (word embedding).

[0122] Through step S302, this embodiment realizes segmenting the text data in the customer terminal data through the adaptive word segmentation algorithm, laying a solid foundation for subsequent joint training of the input multi-modal model.

[0123] In an embodiment of the annotation method based on a multi-modal model of the present application, refer to Figure 4 , and it may specifically include the following content:

[0124] Step S401: performing a sampling operation on the three-dimensional point cloud data according to a preset far-point sampling algorithm to determine a corresponding centroid point, and performing a neighborhood construction operation on each of the centroid points according to a preset multi-scale radius to determine a corresponding multi-scale neighborhood;

[0125] Step S402: performing a feature extraction operation on the multi-scale neighborhood, and performing a feature fusion operation on the multi-scale neighborhood features obtained after the feature extraction operation, determining the multi-scale features corresponding to the centroid point, performing a vectorization operation on the multi-scale features, and determining the corresponding point cloud vector.

[0126] Optionally, in this embodiment, the purpose of this step is to preprocess the three-dimensional point cloud data of the client terminal data to obtain a three-dimensional point cloud vector, laying the foundation for subsequent joint learning and reasoning with the image vector and text vector into the same feature space.

[0127] Optionally, in this embodiment, a set of key points is first selected from the input point cloud using the Farthest Point Sampling (FPS) strategy. For each key point, its neighborhood range is determined according to multiple different radii (i.e., different scales). For example, for a center point, neighborhoods with radii of 1 unit, 2 units, and 3 units may be considered at the same time. For each neighborhood defined by a different radius, a PointNet structure is used to extract local features. Finally, the features extracted from different neighborhoods are connected to form a multi-scale feature representation of the center point.

[0128] Optionally, in this embodiment, multiple such feature extraction layers are used. In each layer of the network, features of local areas are extracted, and then multiple local features are aggregated at a higher level to capture information at different scales.

[0129] Finally, the extracted point cloud data is converted into point cloud vectors, laying a solid foundation for subsequent input into the multimodal model for joint training.

[0130] Through step S402, this embodiment successfully obtains high-precision three-dimensional point cloud features through multi-scale feature extraction, laying the foundation for subsequent training of a multimodal model.

[0131] In one embodiment of the multimodal model-based annotation method of the present application, see Figure 5 , and can also include the following:

[0132] Step S501: Inputting the text vector, the image vector, and the point cloud vector into a backbone network of a preset initial model, performing a distance measurement operation on the text vector, the image vector, and the point cloud vector according to a preset shared space, and determining corresponding alignment;

[0133] Step S502: performing a matching space construction operation on the text vector, the image vector, and the point cloud vector according to the alignment degree, determining a corresponding multimodal vector mapping relationship, and determining a corresponding multimodal model according to the multimodal vector mapping relationship.

[0134] Optionally, in this embodiment, the purpose of this step is to construct a matching space by converting different types of data (text, pictures, point clouds, etc.) into a unified token format and sending them into the Transformer architecture for processing in a shared space, so that the trained model has stronger multimodal data processing capabilities.

[0135] Optionally, in this embodiment, the key to a multimodal AI model is to unify data from different modalities (such as text, images, and 3D point clouds) into a tokenized form that the model can process. Through tokenization, all data (whether text, images, or point clouds) is converted into a standardized sequence of tokens, ensuring that different data sources can interact and perform joint reasoning within the same model.

[0136] After unified Tokenization, text, image, and point cloud data will all be converted into Tokens and input into the model in a specific order or structure. For example, a text Token sequence, an image Patch Token sequence, and a point cloud coordinate Token sequence will be combined and input into the input layer of the Transformer model. In terms of effect, Tokenization not only converts the data into a unified form, but also extracts potential features in the data (for example, semantics in text, local features in images, spatial relationships in point cloud data, etc.) for subsequent model reasoning and learning. After Tokenization, all modal data can be processed through the same mechanism in the neural network, ensuring that the model's reasoning ability can be effectively exerted on different types of data.

[0137] Optionally, in this embodiment, an effective matching space is constructed, and a mapping relationship is established by calculating the similarity between different modal data. That is, through cross-modal embedding technology, different modalities (such as images and text, or 3D point clouds and text) are converted into a shared space to calculate their alignment, which enables the model to understand the relationship and interdependence between different modalities. The shared space can effectively unify different modal data in the same feature space for comparison and matching. In this shared space, a specific distance metric (such as Euclidean distance or cosine similarity) is used to calculate the matching degree between different modalities. Only when the distance between the features of the two modalities in the shared space is small enough can they be considered to be matched.

[0138] Then, for the matched data of different modalities, a self-attention mechanism is used to perform a weighted fusion of multiple modalities, such as images, text, and point clouds. This allows the model to dynamically adjust the contribution of data based on the importance of different modalities. This construction of a matching space improves the model's adaptability to complex datasets, not only improving model accuracy but also providing strong support for the effective integration of multimodal learning frameworks.

[0139] Through step S502, this embodiment successfully obtains a multimodal model by constructing a matching space, laying a foundation for subsequent prediction of client terminal multimodal data through the model and continuous optimization.

[0140] In one embodiment of the multimodal model-based annotation method of the present application, see Figure 6 , and can also include the following:

[0141] Step S601: performing a difference comparison operation based on the predicted result and the set actual result to determine the corresponding deviation value;

[0142] Step S602: performing an update operation on the multimodal model according to the deviation value and the back propagation algorithm to determine an updated multimodal model, wherein the set true result is determined by the annotator end.

[0143] Optionally, in this embodiment, the purpose of this step is to iteratively optimize the model through a feedback mechanism to ensure that the model can learn quickly based on real-time data and annotator feedback.

[0144] Optionally, in this embodiment, the multimodal model is deployed on the annotation terminal used by the annotator. The intelligent agent platform (AIAgent) of each terminal observes the annotator's annotation behavior in real time, and continuously compares it with the multimodal model inference results, continuously calibrates the model prediction results, and achieves the effect of model training.

[0145] Optionally, in this embodiment, the prediction results of the model are compared with the annotation results of the annotators one by one, and the differences are calculated.

[0146] Optionally, for example, taking image 2D data and point cloud 3D data as examples, the annotator manually annotates in 2D / 3D space respectively, and extracts the true value of the target in the corresponding space (for dynamic targets, the data result is generally a 2D / 3D bounding box (for 2D bounding boxes, the format of the Bounding Box is the center point coordinates and the corresponding width and height, x1, y1, w, l); for static map elements, it is generally a polyline). The inference result of the model matches the output format of the manual annotation result of the annotator. The result output by the annotator (x_l, y_l, w_l, h_l) and the result of model inference (x_m, y_m, w_m, h_m) are directly subtracted to obtain the difference between the two (x_l-x_m, y_l-y_m, w_l-w_m, h_l-h_m). The annotation results of the annotators are used as the true value, and the deviation (x_l-x_m, y_l-y_m, w_l-w_m, h_l-h_m) after the model's prediction results are compared with the annotators' annotation results in real time is passed to the model as feedback for model training and optimization based on the backpropagation algorithm of the neural network model training.

[0147] Optionally, in this optimization step, after the model prediction, the annotator will inspect the model's prediction results. If the inspection fails, the annotations will be manually adjusted. Each manual adjustment by the annotator will trigger model training on the intelligent agent platform. Feedback iteration can accelerate the performance improvement of the model.

[0148] It is understandable that for multimodal data (such as text data, image data, and point cloud 3D data), the AI ​​Agent (intelligent agent platform) also compares the manual annotations of the annotator with the inference results of the model. For the difference feedback of multimodal data, all difference calculations will be used as training samples for cross-modal model iteration, thereby optimizing the predictive ability of the multimodal large model.

[0149] Through step S602 , this embodiment successfully performs iterative tuning on the modal model through the feedback mechanism, ensuring rapid convergence and optimization of the model performance.

[0150] In one embodiment of the multimodal model-based annotation method of the present application, see Figure 7 , and can also include the following:

[0151] Step S701: a labeler manually labels the dynamic data in the client terminal data at the labeler terminal to determine the corresponding dynamic data envelope;

[0152] Step S702: The annotator manually annotates the static data in the client terminal data at the annotator terminal to determine the corresponding static data polyline;

[0153] Step S703: Determine the corresponding true result according to the dynamic data envelope and the static data polyline.

[0154] Optionally, the actual result is obtained by the annotator annotating the client terminal data at the annotation terminal.

[0155] Specifically, the annotator manually annotates data using a terminal device, and the annotated data includes text, images (2D data), 3D point cloud data, etc. Every annotation action of the annotator will be recorded in real time by the AIAgent.

[0156] Specifically, for various types of client terminal data (text, images, and 3D point clouds), there are two forms: dynamic and static. For dynamic objects, the data result is generally a 2D / 3D bounding box (for 2D bounding boxes, the format of the Bounding Box is the center point coordinates and the corresponding width and height, x1, y1, w, l); for static map elements, it is generally a polyline.

[0157] Through step S703, this embodiment successfully obtains the actual results of the annotators' annotations, which lays the foundation for subsequent comparison with the model's prediction results. By sending a feedback mechanism, the model is helped to quickly adapt to the new data distribution and requirements, ensuring the rapid convergence and optimization of the model performance.

[0158] In order to improve the efficiency and accuracy of data annotation, the present application provides an embodiment of a multimodal model-based annotation device for implementing all or part of the content of the multimodal model-based annotation method, see Figure 8 The multimodal model-based annotation device specifically includes the following contents:

[0159] A data receiving module 10 is configured to receive a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data;

[0160] A multimodal model training module 20 is configured to input the client terminal data set into a data preprocessing module of a preset initial model, perform a segmentation operation on the text data according to a set adaptive word segmentation algorithm to determine a corresponding text vector, perform a segmentation operation on the image data according to a preset image segmentation algorithm to determine a corresponding target image area, a foreground image area, and a background image area, perform a high-definition feature extraction operation on the foreground image area to determine a corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area to determine a corresponding background feature, perform a feature extraction operation on the target image area to determine a corresponding target feature, perform a vectorization operation on the target feature, the foreground feature, and the background feature to determine a corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to a preset multi-scale feature extraction network to determine a corresponding point cloud vector, input the text vector, the image vector, and the point cloud vector into a backbone network of a preset initial model to perform a matching space construction operation to determine a corresponding multimodal model;

[0161] The prediction and annotation module 30 is used to perform a prediction and annotation operation on the client terminal data according to the multimodal model, and determine the prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes predicted text annotation, predicted image annotation and predicted three-dimensional point cloud annotation.

[0162] From the above description, it can be seen that the multimodal model-based labeling device provided in the embodiment of the present application can input the client terminal data set into a preset initial model, segment the text data according to the set adaptive word segmentation algorithm to obtain a text vector, segment the image data according to a preset image segmentation algorithm, and perform feature extraction on the target image area, foreground image area and background image area obtained after segmentation to obtain an image vector, extract the multi-scale point cloud features of the three-dimensional point cloud data according to a preset multi-scale feature extraction network to obtain a point cloud vector, input the text vector, image vector and point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation, determine the corresponding multimodal model, perform a predictive labeling operation on the client terminal data according to the multimodal model, and determine the prediction result of the client terminal data output by the multimodal model, thereby improving the efficiency and accuracy of data labeling.

[0163] From a hardware perspective, in order to improve the efficiency and accuracy of data annotation, the present application provides an embodiment of an electronic device for implementing all or part of the multimodal model-based annotation method. The electronic device specifically includes the following:

[0164] A processor, a memory, a communications interface, and a bus; wherein the processor, the memory, and the communications interface communicate with each other via the bus; the communications interface is used to implement information transmission between the multimodal model-based annotation method and related devices such as the core business system, user terminals, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiment of the multimodal model-based annotation method and the embodiment of the multimodal model-based annotation method in the embodiment, and their contents are incorporated herein, and repeated parts are not repeated.

[0165] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0166] In practical applications, portions of the multimodal model-based annotation method can be performed on the electronic device side as described above, or all operations can be performed on the client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not impose any restrictions on this. If all operations are performed on the client device, the client device may also include a processor.

[0167] The client device may include a communication module (i.e., a communication unit) that can establish a communication connection with a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0168] Figure 9 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 9 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0169] In one embodiment, the multimodal model-based annotation method function may be integrated into the central processing unit 9100. The central processing unit 9100 may be configured to perform the following control:

[0170] Step S101: receiving a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data;

[0171] Step S102: Input the client terminal data set into the data preprocessing module of the preset initial model, perform a segmentation operation on the text data according to the set adaptive word segmentation algorithm to determine the corresponding text vector, perform a segmentation operation on the image data according to the preset image segmentation algorithm to determine the corresponding target image area, foreground image area and background image area, perform a high-definition feature extraction operation on the foreground image area to determine the corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area to determine the corresponding background feature, perform a feature extraction operation on the target image area to determine the corresponding target feature, perform a vectorization operation on the target feature, the foreground feature and the background feature to determine the corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to the preset multi-scale feature extraction network to determine the corresponding point cloud vector, input the text vector, the image vector and the point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation to determine the corresponding multimodal model;

[0172] Step S103: performing a predictive annotation operation on the client terminal data according to the multimodal model to determine a prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes predicted text annotation, predicted image annotation, and predicted three-dimensional point cloud annotation.

[0173] From the above description, it can be seen that the electronic device provided by the embodiment of the present application inputs a client terminal data set into a preset initial model, segments the text data according to a set adaptive word segmentation algorithm to obtain a text vector, segments the image data according to a preset image segmentation algorithm, performs feature extraction on the target image area, foreground image area and background image area obtained after segmentation to obtain an image vector, extracts multi-scale point cloud features of the three-dimensional point cloud data according to a preset multi-scale feature extraction network to obtain a point cloud vector, inputs the text vector, image vector and point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation, determines the corresponding multimodal model, performs a prediction and annotation operation on the client terminal data according to the multimodal model, and determines the prediction result of the client terminal data output by the multimodal model, thereby improving the efficiency and accuracy of data annotation.

[0174] In another embodiment, the multimodal model-based annotation method can be configured separately from the central processing unit 9100. For example, the multimodal model-based annotation method can be configured as a chip connected to the central processing unit 9100, and the function of the multimodal model-based annotation method can be realized through the control of the central processing unit.

[0175] like Figure 9 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 9 In addition, the electronic device 9600 may also include all components shown in Figure 9 For components not shown, reference may be made to the prior art.

[0176] like Figure 9 As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.

[0177] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.

[0178] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.

[0179] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), or a SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is provided with more data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 by the central processing unit 9100.

[0180] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0181] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 is coupled to the central processing unit 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0182] Based on different communication technologies, multiple communication modules 9110 can be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module. The communication module 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby implementing common telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is also coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.

[0183] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the multimodal model-based annotation method in the above-mentioned embodiment, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the multimodal model-based annotation method in the above-mentioned embodiment, where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented:

[0184] Step S101: receiving a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data;

[0185] Step S102: Input the client terminal data set into the data preprocessing module of the preset initial model, perform a segmentation operation on the text data according to the set adaptive word segmentation algorithm to determine the corresponding text vector, perform a segmentation operation on the image data according to the preset image segmentation algorithm to determine the corresponding target image area, foreground image area and background image area, perform a high-definition feature extraction operation on the foreground image area to determine the corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area to determine the corresponding background feature, perform a feature extraction operation on the target image area to determine the corresponding target feature, perform a vectorization operation on the target feature, the foreground feature and the background feature to determine the corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to the preset multi-scale feature extraction network to determine the corresponding point cloud vector, input the text vector, the image vector and the point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation to determine the corresponding multimodal model;

[0186] Step S103: performing a predictive annotation operation on the client terminal data according to the multimodal model to determine a prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes predicted text annotation, predicted image annotation, and predicted three-dimensional point cloud annotation.

[0187] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application inputs a client terminal data set into a preset initial model, segments the text data according to a set adaptive word segmentation algorithm to obtain a text vector, segments the image data according to a preset image segmentation algorithm, performs feature extraction on the target image area, foreground image area, and background image area obtained after segmentation to obtain an image vector, extracts multi-scale point cloud features of the three-dimensional point cloud data according to a preset multi-scale feature extraction network to obtain a point cloud vector, inputs the text vector, image vector, and point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation, determines the corresponding multimodal model, performs a prediction and annotation operation on the client terminal data according to the multimodal model, and determines the prediction result of the client terminal data output by the multimodal model, thereby improving the efficiency and accuracy of data annotation.

[0188] The embodiments of the present application also provide a computer program product capable of implementing all steps of the multimodal model-based annotation method in the above-mentioned embodiment, where the execution subject is a server or a client. When the computer program / instructions are executed by a processor, the computer program / instructions implement the steps of the multimodal model-based annotation method. For example, the computer program / instructions implement the following steps:

[0189] Step S101: receiving a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data;

[0190] Step S102: Input the client terminal data set into the data preprocessing module of the preset initial model, perform a segmentation operation on the text data according to the set adaptive word segmentation algorithm to determine the corresponding text vector, perform a segmentation operation on the image data according to the preset image segmentation algorithm to determine the corresponding target image area, foreground image area and background image area, perform a high-definition feature extraction operation on the foreground image area to determine the corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area to determine the corresponding background feature, perform a feature extraction operation on the target image area to determine the corresponding target feature, perform a vectorization operation on the target feature, the foreground feature and the background feature to determine the corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to the preset multi-scale feature extraction network to determine the corresponding point cloud vector, input the text vector, the image vector and the point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation to determine the corresponding multimodal model;

[0191] Step S103: performing a predictive annotation operation on the client terminal data according to the multimodal model to determine a prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes predicted text annotation, predicted image annotation, and predicted three-dimensional point cloud annotation.

[0192] From the above description, it can be seen that the computer program product provided by the embodiment of the present application inputs a client terminal data set into a preset initial model, segments the text data according to a set adaptive word segmentation algorithm to obtain a text vector, segments the image data according to a preset image segmentation algorithm, performs feature extraction on the target image area, foreground image area and background image area obtained after segmentation to obtain an image vector, extracts multi-scale point cloud features of the three-dimensional point cloud data according to a preset multi-scale feature extraction network to obtain a point cloud vector, inputs the text vector, image vector and point cloud vector into the backbone network of the preset initial model to perform a matching space construction operation, determines the corresponding multimodal model, performs a prediction and annotation operation on the client terminal data according to the multimodal model, and determines the prediction result of the client terminal data output by the multimodal model, thereby improving the efficiency and accuracy of data annotation.

[0193] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0194] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0195] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0196] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0197] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. A labeling method based on a multimodal model, characterized in that: Applied to a proxy annotation platform, the proxy annotation platform is communicatively connected with a preset annotator terminal, and the method includes: Receiving a client terminal data set, wherein the client terminal data set includes text data, image data, and three-dimensional point cloud data; Input the client terminal data set into the data preprocessing module of the preset initial model, perform segmentation operation on the text data according to the set adaptive word segmentation algorithm, determine the corresponding text vector, perform segmentation operation on the image data according to the preset image segmentation algorithm, determine the corresponding target image area, foreground image area and background image area, perform high-definition feature extraction operation on the foreground image area, determine the corresponding foreground feature, perform fuzzy feature extraction operation on the background image area, determine the corresponding background feature, perform feature extraction operation on the target image area, determine the corresponding target feature, perform vectorization operation on the target feature, the foreground feature and the background feature, determine the corresponding image vector, perform multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to the preset multi-scale feature extraction network, determine the corresponding point cloud vector, input the text vector, the image vector and the point cloud vector into the backbone network of the preset initial model to perform matching space construction operation, and determine the corresponding multimodal model; A predictive annotation operation is performed on the client terminal data according to the multimodal model to determine a prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes a predicted text annotation, a predicted image annotation, and a predicted three-dimensional point cloud annotation.

2. The labeling method based on the multimodal model according to claim 1, characterized in that: Before segmenting the text data according to the set word segmentation algorithm and determining the corresponding text vector, the method includes: Collecting a historical text data set, performing a word segmentation operation on the historical text data set, and determining a corresponding minimum text subunit; Inputting the minimum text subunit into a preset long-short term network for model training, and determining the dependency relationship between the minimum text subunits output by the long-short term network; According to the dependency, the preset validation set dependency and the preset cross entropy loss function, the corresponding loss function is determined, and the parameters of the long-term and short-term network are updated according to the optimal loss function and the reverse update algorithm to determine the corresponding adaptive word segmentation algorithm.

3. The labeling method based on a multimodal model according to claim 1, characterized in that: The step of segmenting the text data according to a set word segmentation algorithm to determine a corresponding text vector includes: Analyze the text data according to the set word segmentation algorithm to determine the corresponding text data segmentation points, segment the text data according to the text data segmentation points to determine the corresponding optimal text sub-units; The optimal text sub-unit is vectorized according to a preset word vector model to determine a corresponding text vector.

4. The labeling method based on a multimodal model according to claim 1, characterized in that: The performing a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to a preset multi-scale feature extraction network to determine a corresponding point cloud vector includes: Performing a sampling operation on the three-dimensional point cloud data according to a preset far-point sampling algorithm to determine a corresponding centroid point, and performing a neighborhood construction operation on the centroid point according to a preset multi-scale radius to determine a corresponding multi-scale neighborhood; A feature extraction operation is performed on the multi-scale neighborhood, and a feature fusion operation is performed on the multi-scale neighborhood features obtained after the feature extraction operation to determine the multi-scale features corresponding to the centroid point, and a vectorization operation is performed on the multi-scale features to determine the corresponding point cloud vector.

5. The labeling method based on a multimodal model according to claim 1, characterized in that: The step of inputting the text vector, the image vector, and the point cloud vector into a backbone network of a preset initial model to perform a matching space construction operation and determine a corresponding multimodal model includes: Inputting the text vector, the image vector and the point cloud vector into a backbone network of a preset initial model, performing a distance measurement operation on the text vector, the image vector and the point cloud vector according to a preset shared space, and determining a corresponding alignment degree; A matching space construction operation is performed on the text vector, the image vector and the point cloud vector according to the alignment degree to determine a corresponding multimodal vector mapping relationship, and a corresponding multimodal model is determined according to the multimodal vector mapping relationship.

6. The labeling method based on a multimodal model according to claim 1, characterized in that: After performing a prediction and labeling operation on the client terminal data according to the multimodal model to determine a prediction result of the client terminal data output by the multimodal model, the method further comprises: Perform a difference comparison operation based on the predicted result and the set actual result to determine the corresponding deviation value; The multimodal model is updated according to the deviation value and the back propagation algorithm to determine an updated multimodal model, wherein the set true result is determined by the annotator end.

7. The labeling method based on a multimodal model according to claim 6, characterized in that: Before performing a difference comparison operation based on the predicted result and the set actual result to determine the corresponding deviation value, the method includes: The annotator manually annotates the dynamic data in the client terminal data at the annotator end to determine the corresponding dynamic data envelope frame; The annotator manually annotates the static data in the client terminal data at the annotator end to determine the corresponding static data polyline; The corresponding real result is determined according to the dynamic data envelope frame and the static data polyline.

8. A labeling device based on a multimodal model, characterized in that: The device comprises: A data receiving module, used for receiving a client terminal data set, wherein the client terminal data set includes text data, image data and three-dimensional point cloud data; A multimodal model training module is used to input the client terminal data set into a data preprocessing module of a preset initial model, perform a segmentation operation on the text data according to a set adaptive word segmentation algorithm, determine a corresponding text vector, perform a segmentation operation on the image data according to a preset image segmentation algorithm, determine a corresponding target image area, a foreground image area, and a background image area, perform a high-definition feature extraction operation on the foreground image area, determine a corresponding foreground feature, perform a fuzzy feature extraction operation on the background image area, determine a corresponding background feature, perform a feature extraction operation on the target image area, determine a corresponding target feature, perform a vectorization operation on the target feature, the foreground feature, and the background feature, determine a corresponding image vector, perform a multi-scale point cloud feature extraction operation on the three-dimensional point cloud data according to a preset multi-scale feature extraction network, determine a corresponding point cloud vector, input the text vector, the image vector, and the point cloud vector into a backbone network of a preset initial model to perform a matching space construction operation, and determine a corresponding multimodal model; A prediction and annotation module is used to perform a prediction and annotation operation on the client terminal data according to the multimodal model, and determine the prediction result of the client terminal data output by the multimodal model, wherein the prediction result includes predicted text annotation, predicted image annotation and predicted three-dimensional point cloud annotation.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the multimodal model-based annotation method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal model-based annotation method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Cross-modal single-sample three-dimensional point cloud segmentation method

    CN114529757A

  • Point cloud instance segmentation method based on comparative language image pre-training technology

    CN116152267A

  • Multi-modal model training method, device and equipment and readable storage medium

    CN116561570A

  • 3D self-supervised pre-training method based on large-scale language-image model guidance

    CN116681107A

  • Multimodal unsupervised pedestrian pixel-level semantic labeling method and system

    WO2022141721A1

Cited By

  • Spanish multimodal dictionary construction method and system based on large language model

    CN120448558A

  • Self-closed-loop optimized point cloud fusion data labeling method and system

    CN121330685A