Multi-modal information tagging method, apparatus and device, and storage medium and product

Through the multimodal information marking system, the problem of poor single-modal marking effect is solved, and the precise marking of diverse content, especially the effective identification of inferior content is achieved.

WO2025148651A1PCT designated stage expired Publication Date: 2025-07-17BIGO TECH PTE LTD +1

Patent Information

Application Number
PCT/CN2024/140760
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-12-19
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing single-modal information marking method has poor results when facing diverse user content experience needs, making it difficult to accurately identify inferior content of abstract and confrontational nature.

Method used

The multimodal information marking method is adopted, and the multimodal marking system is used to integrate visual information, identify text information and describe text information to determine visual features and graphic correlation, and the cross-modal correlation is used to identify and process inferior content.

Benefits of technology

It improves the accuracy and coverage of information marking, especially the recognition ability of inferior content with abstract and confrontational properties, and enhances the multi-dimensional effect of information marking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140760_17072025_PF_FP_ABST
    Figure CN2024140760_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A multi-modal information tagging method, apparatus and device, and a storage medium and a product. An information tagging effect is effectively improved by means of determining visual feature information of visual information by means of a multi-modal tagging system, performing multi-modal fusion processing on the basis of the visual feature information, the visual information, recognized text information and descriptive text information to obtain image-text fusion features, determining image-text correlation information between the visual information and both the recognized text information and the descriptive text information, on the basis of the image-text fusion features and the image-text correlation information, determining a tagging result for information to be tagged, and performing multidimensional information tagging on the basis of multi-modal information fusion and cross-modal correlations.
Need to check novelty before this filing date? Find Prior Art

Description

A multimodal information marking method, device, equipment, storage medium and product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 8, 2024, with application number 202410026073.1, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a multimodal information marking method, apparatus, device, storage medium, and product. Background Art

[0003] With the rapid development of internet and mobile communications technologies, cultural exchanges around the world are becoming increasingly frequent, and cultural expressions are becoming increasingly diverse. Against this backdrop, traditional video / live streaming products rely on coarse-grained tagging and profiling systems in a single visual or textual dimension to generate massive user traffic and build strong user engagement.

[0004] Existing methods for tagging video / live video content typically rely on single-modal computer vision technology or natural language processing methods, extracting visual features from videos or images for label classification or classifying labels based on image text and description text. However, with the continued increase in users and growing cultural diversity, traditional single-modal coarse-grained tagging and profiling methods are increasingly unable to meet the diverse content experience needs of users, resulting in poor information tagging results. Summary of the Invention

[0005] The embodiments of the present application provide a multimodal information marking method, apparatus, device, storage medium and product to solve the technical problem of poor marking effect of single-modal information marking methods in related technologies, and effectively improve the information marking effect.

[0006] In a first aspect, an embodiment of the present application provides a multimodal information marking method, comprising:

[0007] Acquiring information to be marked, wherein the information to be marked includes visual information, identification text information, and description text information;

[0008] The information to be labeled is input into a trained multimodal labeling system, the visual feature information of the visual information is determined by the multimodal labeling system, multimodal fusion processing is performed based on the visual feature information, the visual information, the recognition text information and the description text information to obtain image-text fusion features, and image-text correlation information of the visual information, the recognition text information and the description text information is determined, and the labeling result of the information to be labeled is determined based on the image-text fusion features and the image-text correlation information.

[0009] In a second aspect, an embodiment of the present application provides a multimodal information marking device, including an information acquisition module and a marking processing module, wherein:

[0010] The information acquisition module is configured to acquire information to be marked, wherein the information to be marked includes visual information, identification text information and description text information;

[0011] The marking processing module is configured to input the information to be marked into a trained multimodal marking system, determine the visual feature information of the visual information through the multimodal marking system, perform multimodal fusion processing based on the visual feature information, the visual information, the recognition text information and the description text information to obtain image-text fusion features, and determine the image-text correlation information of the visual information, the recognition text information and the description text information, and determine the marking result of the information to be marked based on the image-text fusion features and the image-text correlation information.

[0012] In a third aspect, an embodiment of the present application provides a multimodal information marking device, comprising: a memory and one or more processors;

[0013] The memory is used to store one or more programs;

[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal information marking method as described in the first aspect.

[0015] In a fourth aspect, an embodiment of the present application provides a non-volatile storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to execute the multimodal information marking method as described in the first aspect.

[0016] In the fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads and executes the computer program from the computer-readable storage medium, so that the device performs the multimodal information marking method described in the first aspect.

[0017] The embodiment of the present application determines the visual feature information of the visual information through a multimodal labeling system, performs multimodal fusion processing based on the visual feature information, the visual information, the recognition text information and the description text information to obtain the image-text fusion feature, and determines the image-text correlation information of the visual information, the recognition text information and the description text information, and determines the labeling result of the information to be labeled based on the image-text fusion feature and the image-text correlation information, and performs multi-dimensional information labeling based on multimodal information fusion and cross-modal correlation, effectively improving the information labeling effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a flow chart of a multimodal information marking method provided in an embodiment of the present application;

[0019] FIG2 is a schematic diagram of a flow chart of determining visual feature information based on a live broadcast image according to an embodiment of the present application;

[0020] FIG3 is a schematic diagram of the structure of a live stream visual feature extraction network based on live images provided in an embodiment of the present application;

[0021] FIG4 is a schematic diagram of a process for determining visual feature information based on a video image according to an embodiment of the present application;

[0022] FIG5 is a schematic diagram of the structure of a video frame visual feature extraction network based on video images provided in an embodiment of the present application;

[0023] FIG6 is a schematic diagram showing the principle of a temporal visual query vector learner provided in an embodiment of the present application;

[0024] FIG7 is a schematic diagram showing the principle of a static visual query vector learner provided in an embodiment of the present application;

[0025] FIG8 is a schematic diagram of the interaction principle of a visual text fusion network provided in an embodiment of the present application;

[0026] FIG9 is a schematic diagram of a flow chart of determining image-text relevance information provided by an embodiment of the present application;

[0027] FIG10 is a schematic diagram of a cross-modal calculation principle of image-text relevance provided by an embodiment of the present application;

[0028] FIG11 is a schematic diagram of an alignment model structure provided in an embodiment of the present application;

[0029] FIG12 is a schematic structural diagram of a multimodal information marking system provided in an embodiment of the present application;

[0030] FIG13 is a schematic structural diagram of a multimodal information marking device provided in an embodiment of the present application;

[0031] FIG14 is a schematic structural diagram of a multimodal information marking device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only some, but not all, of the contents related to the present application are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The above process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The above process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0033] The multimodal information labeling method provided in this application can be applied to the labeling of videos, live broadcasts, pictures and other content. It aims to perform multi-dimensional information labeling on the information to be labeled based on multimodal information fusion and cross-modal correlation, thereby improving the information labeling effect.

[0034] The labeling results can be used to assist in stratifying recommendations for users with different interests and defining communities of interest. They can also provide personalized, continuous push notifications tailored to specific users' interests. Existing information labeling solutions typically leverage computer vision or natural language processing methods, extracting visual features from videos / images to classify labels or analyzing relevant content based on image text and descriptions. To balance development costs and label scalability, existing labeling solutions typically incorporate a hierarchical labeling structure, dividing labels into several subsets with minimal overlap. Each subset is assigned a specific "expert" network, and the outputs of all expert networks are combined to produce a complete labeling result. While this type of labeling technology can address most obvious labeling issues, it typically extracts features from only a single modality, which can lead to misclassification of low-relevance labels. For example, when a labeling system is preparing to label a series of videos in the "car review" category, while cars are one of the visual elements of "car review," the model in the labeling system only uses features from the visual modality to make labeling decisions, which can easily misclassify videos containing cars as having that label.

[0035] Based on this, a multimodal information labeling method of an embodiment of the present application is provided to solve the technical problem that the existing single-modal information labeling method has a poor labeling effect. The multimodal information labeling method provided by this solution, based on the labeling technology of multimodal information, solves the problem that the existing labeling system is not good at labeling abstract content by selecting and fusing the features of different modalities. Similarly, for "car review" videos, the labeling technology based on multimodal fusion can not only focus on visual elements, but also extract key information from the voice, and give a labeling judgment based on the context of the video. The same abstract content has a more obvious improvement in the effect of low-quality labeling. For example, for the identification of obscure label content such as fraudulent traffic, dirty jokes, political metaphors, etc., the existing labeling system cannot label these contents because of the single input modal information and the overly single judgment dimension. Especially when these contents have adversarial characteristics, it will further bring adverse effects to the content ecology. The multimodal information labeling method provided by this solution uses a cross-modal correlation mechanism to enhance the identification of such low-quality content. Based on the image-text correlation and audio-text correlation, multimodal information is integrated to accurately mine and process more obscure low-quality content, thereby improving the information labeling effect.

[0036] Figure 1 shows a flowchart of a multimodal information marking method provided in an embodiment of the present application. The multimodal information marking method provided in an embodiment of the present application can be executed by a multimodal information marking device, which can be implemented in hardware and / or software and integrated into a multimodal information marking device.

[0037] The following description will be made by taking a multimodal information marking device executing a multimodal information marking method as an example. Referring to FIG1 , the multimodal information marking method includes:

[0038] S110: Acquire information to be marked, where the information to be marked includes visual information, recognition text information, and description text information.

[0039] Exemplarily, the information to be marked provided by this solution can be a video image (for example, a video obtained by video editing, which includes multiple video image frames), a live image (for example, a live stream obtained in real time, a recorded video obtained by recording a live broadcast, etc., which includes multiple live video frames), or a picture (which can be a single image or multiple images).

[0040] Exemplarily, information to be marked that needs to be marked is obtained, and the information to be marked provided by this solution includes visual information, recognition text information and description text information. Among them, visual information can be video information, live broadcast information, pictures, etc. The recognition text can be text information obtained by performing text recognition on visual information or information related to visual information (such as voice information corresponding to video information or live broadcast information), such as performing text recognition (OCR text recognition) on video image frames in video information, live image frames in live broadcast information, or pictures to obtain recognition text information (the recognition text information can be the text recognized in the image frame or picture), or performing voice recognition (ASR voice recognition) on audio information corresponding to video information or live broadcast information to obtain corresponding recognition text information (the recognition text information can be the text corresponding to the speech in the audio information). The description text information can be understood as text that describes the information to be marked or the visual information therein, such as a video introduction, a live broadcast introduction, a picture introduction, etc. Optionally, the identified text information can be the audio information corresponding to the video information or live broadcast information, and the multimodal information labeling system identifies the audio information to obtain the corresponding text information. Alternatively, the audio information corresponding to the video information or live broadcast information is first identified to obtain the corresponding text information and then transmitted to the multimodal information labeling system.

[0041] In one embodiment, the information to be marked obtained by this scheme can be a combination of one or more of visual information, recognition text information and descriptive text information. For the information that is not obtained (visual information, recognition text information, descriptive text information, etc.), it can be replaced by pre-set default information to ensure the normal progress of information marking.

[0042] S120: Input the information to be labeled into the trained multimodal labeling system, determine the visual feature information of the visual information through the multimodal labeling system, perform multimodal fusion processing based on the visual feature information, visual information, recognition text information and description text information to obtain image-text fusion features, and determine the image-text correlation information between the visual information and the recognition text information and the description text information, and determine the labeling result of the information to be labeled based on the image-text fusion features and the image-text correlation information.

[0043] Exemplarily, after obtaining the information to be marked, the information to be marked is input into a trained multimodal marking system, and the multimodal marking system analyzes and processes the information to be marked to obtain a marking result corresponding to the information to be marked.

[0044] Among them, after receiving the information to be marked, the multimodal marking system determines the visual feature information corresponding to the visual information in the information to be marked (visual feature information of multiple video frames or pictures), and performs multimodal fusion processing based on the above-mentioned visual feature information, as well as the visual information, recognition text information and description text information in the information to be marked, to obtain the image-text fusion feature. At the same time, the multimodal marking system determines the image-text correlation information between the visual information and the recognition text information and the description text information. After obtaining the image-text fusion feature and the image-text correlation information, the multimodal marking system can determine the marking result of the information to be marked based on the image-text fusion feature and the image-text correlation information.

[0045] Optionally, the multimodal fusion processing of visual feature information, visual information, recognition text information and description text information can be achieved by aligning the above-mentioned multiple information, or by performing nonlinear mapping in the inference process of the multimodal labeling system, or by performing information fusion through a trained visual text fusion network. The image-text correlation information can be understood as the degree of correlation between visual information and recognition text information and / or description text information, or whether the degree of correlation between visual information and recognition text information and / or description text information reaches a set correlation threshold. Optionally, when determining the marking result of the information to be marked based on the image-text fusion features and the image-text correlation information, this scheme can be obtained by processing through a trained neural network model.

[0046] In one possible embodiment, as shown in FIG2 , a schematic diagram of a flow chart for determining visual feature information based on a live image, the multimodal marking system provided in this solution includes:

[0047] S211: In the case where the visual information is a live image, each live image frame in the visual information is input into the first visual processing model in chronological order, and the first picture token mapping information of each live image frame is determined by the first visual processing model.

[0048] S212: stacking the first image token mapping information, and performing temporal context association fusion on the stacked first image token mapping information through a pseudo three-dimensional convolutional model to obtain feature maps corresponding to each first image token mapping information.

[0049] S213: Pooling the feature maps corresponding to the first image token mapping information into a first feature map of a set dimension, and mapping the first feature map into visual feature information through a fully connected layer.

[0050] For example, after obtaining the information to be labeled, if the visual information in the information to be labeled is a live image, there is a strong temporal relationship between the live image frames in the visual information. In this case, each live image frame in the visual information can be input into the trained first visual processing model according to the corresponding time sequence of each live image frame in the visual information (which can be determined based on the timestamp of the live image frame), and the first image token mapping information (Token Map) corresponding to each live image frame is determined by the first visual processing model.

[0051] In one embodiment, after obtaining the first image token mapping information corresponding to each live image frame, each first image token mapping information is stacked, and the stacked first image token mapping information is temporally context-fused using a pseudo three-dimensional convolutional model to obtain a feature map corresponding to each first image token mapping information. In one embodiment, after obtaining the feature map corresponding to each first image token mapping information, the feature map corresponding to each first image token mapping information is pooled into a first feature map of a set dimension (e.g., one dimension, two dimensions, etc.), and the first feature map is mapped into visual feature information through a fully connected layer (FC layer).

[0052] As shown in the structural diagram of a live stream visual feature extraction network based on live images provided in FIG3 , this solution can determine the visual feature information of visual information through the live stream visual feature extraction network. The live stream visual feature extraction network provided by this solution includes a first visual processing model (Vision Encoder in FIG3 ), a pseudo three-dimensional convolutional model (P3D-Res15 unit in FIG3 ) and a fully connected layer (FC layer in FIG3 ). Optionally, the visual processing model provided by this solution (including the first visual processing model and the second visual processing model) can be a large visual model (LVM, Large Vision Model). Optionally, the first visual processing model can extract token features (first image token mapping information) containing local spatial visual feature interactions through an encoder-based deep ViT (Vision Transformer, a Transformer module used in a visual deep neural network model, usually an encoder structure, with a token vector encoded by a picture block as the input source of the attention operation) network while retaining details as much as possible. Compared with traditional deep convolutional networks, it effectively avoids the relative information loss caused by downsampling and strengthens the connection between each local visual information. Considering that the visual encoder based on the Transformer (a neural network module that operates through the self-attention mechanism and the cross-attention mechanism, with two modules, encoder and decoder, which is good at modeling the correlation relationship and details of the context) neural network only considers the visual attention relationship on the spatial scale, the lack of temporal information aggregation for video / live streaming will lead to too many, complex and redundant features, and even lack of contextual interpretation. In order to solve this defect, this solution supplements the visual encoder based on the first visual processing model with a lightweight timing module to aggregate temporal features. Among them, for live streaming screenshots with strong correlation between visual information before and after the time sequence, pseudo 3D convolution is used here to process the fusion of visual information in a specific window period.

[0053] Exemplarily, each live image frame in the visual information (e.g., a screenshot of a live stream) will enter the live stream visual feature extraction network one by one in sequence. Each live image frame will output a first image token mapping information (Image1 Token Map~ImageN Token Map) through the first visual processing model in the live stream visual feature extraction network. The first image token mapping information is encoded by the local image of each grid in the live image frame combined with the position information, where the number of grids is equal to the number of rows of the live image frame. The live stream visual feature extraction network stacks these first image token mapping information and performs temporal context association fusion on the stacked first image token mapping information through a pseudo-three-dimensional convolutional model to obtain the feature map corresponding to each first image token mapping information. The live stream visual feature extraction network pools the feature map corresponding to each first image token mapping information into a first feature map (Image1 Token Map Pooling~ImageN Token Map Pooling) of a set dimension (e.g., one dimension), and uses a fully connected layer to map the first feature map to the final feature (i.e., the visual feature information of the visual information) output.

[0054] Optionally, the first visual processing model provided by this solution can adopt the open source 48-layer large parameter model ViT-L. The pseudo 3D convolution model can be a 15-layer pseudo 3D convolution residual network (ResNet). The pseudo 3D convolution model effectively speeds up the calculation speed without losing performance by splitting the spatiotemporal convolution into two types of convolution layers with shared parameters. Among them, the spatial convolution layer only changes the size and number of channels of the spatial level of the feature map, and multiple feature maps share the convolution kernel for parallel calculation, while the temporal convolution layer groups the feature maps at the same position on the channels at different time points for calculation, but does not change the spatial size and number of channels of the feature map. It can effectively achieve the effect of approximating 3D convolution while improving the calculation speed.

[0055] In one embodiment, the pseudo 3D convolution model provided by this solution can be trained in an unsupervised manner at low cost. The training of the pseudo 3D convolution model does not require labeled data. Instead, the goal is to maximize the number of screenshots within a window of size k to predict the next screenshot, and converge autoregressively. The objective function for training the pseudo 3D convolution model can be: obj live =max(P(im n |im n-1 ,im n-2 ,...,im n-k )),n=1,2,...,T

[0056] Among them, im nThe next image frame needs to be predicted by the first k images in the window, T is the number of image frames, and the objective function needs to achieve the maximum likelihood of this prediction probability. When the pseudo three-dimensional convolution model performs inference, a vector is used as the unit of operation. The autoregressive training method of the pseudo three-dimensional convolution model can still learn the context association pattern without manual annotation, while ensuring the temporal context association fusion effect, and effectively improve the training efficiency of the pseudo three-dimensional convolution model. When the visual information is a live image, this solution inputs each live image frame in the visual information into the first visual processing model in chronological order to obtain the first image token mapping information of each live image frame, stacks the first image token mapping information, and performs temporal context association fusion on the stacked first image token mapping information through the pseudo three-dimensional convolution model to obtain the feature map corresponding to each first image token mapping information, and pools the feature map corresponding to each first image token mapping information into a first feature map of a set dimension, and maps the first feature map to visual feature information through a fully connected layer, so as to efficiently and accurately obtain the visual feature information of the visual information and improve the multimodal information labeling effect.

[0057] In one possible embodiment, as shown in FIG4 , a schematic diagram of a process for determining visual feature information based on a video image is provided. The multimodal marking system provided by this solution includes:

[0058] S214: In the case where the visual information is a video image, each video image frame in the visual information is input into a plurality of second visual processing models respectively, and second picture token mapping information corresponding to each video image frame is determined by the second visual processing model.

[0059] S215: Pooling the second image token mapping information into an image semantic feature vector of a set dimension, and performing an attention operation on the image semantic feature vector through a multi-query attention mechanism network to obtain visual feature information.

[0060] Exemplarily, after obtaining the information to be labeled, if the visual information in the information to be labeled is a video image, the temporal relationship between the video image frames in the visual information is a non-strongly correlated temporal relationship. At this time, each video image frame in the visual information can be input into multiple second visual processing models, and the second visual processing models can be used to determine the second image token mapping information corresponding to the received video image frames, and these second image token mapping information are pooled into image semantic feature vectors of a set dimension (for example, one dimension, two dimensions, etc.). The multi-query attention mechanism network performs attention operation processing on the image semantic feature vector to obtain visual feature information of the visual information.

[0061] In one embodiment, as shown in a structural diagram of a video frame visual feature extraction network based on video images provided in FIG5 , the present solution can determine the visual feature information of visual information through the video frame visual feature extraction network. The video frame visual feature extraction network provided by the present solution includes a multi-query attention mechanism (MQA) network and multiple parallel second visual processing models. Since the temporal correlation between video image frames may not be strong due to excessive jumps, the present solution may not use three-dimensional convolution to model their temporal relationship, but instead use a parallel second visual processing model (visual encoder) to extract features from each frame of the image, and then model the temporal relationship between multiple video image frames through a lightweight multi-query attention mechanism network.

[0062] In one embodiment, the second visual processing model provided by this solution is configured with 15 parallel second visual processing models. The second visual processing model can be obtained based on the various expert sub-networks (20-layer ViT) in the open source hybrid expert system (VMoE, Vision Model of Experts), but the expert network gate selection mechanism is cancelled, and the output results of each expert sub-network are directly used. Each sub-network is a second visual processing model, which is similar to the first visual processing model of the live stream screenshot, but because it only extracts features from a certain frame of the video frame sequence and has no shared weights, that is, it only works on the local problem space of the visual elements, so each sub-network does not need too many weight parameters to fit the overly long context.

[0063] Exemplarily, after visual information is input into the video frame visual feature extraction network, each second visual processing model in the video frame visual feature extraction network will process the input video image frame and output a second picture token mapping information (picture Token Map), which can reflect the correlation between each part of each picture. The second picture token mapping information will be pooled into a picture semantic feature vector of a set dimension (for example, one dimension) through a logistic regression operation (LR). Among them, the logistic regression operation is adopted because the features of each picture have been divided into feature subspaces, and the logistic regression operation has the advantages of fast convergence and accurate fitting for the subspace. Conventional neural network structures generally construct query vectors (Queries) with words / picture blocks. When performing self-attention operations or cross-attention operations, although the query vectors will calculate attention scores with each other according to the masking strategy, the query vectors are essentially from the same local domain and cannot model long context information in time series well. Therefore, this scheme adopts a multi-query attention mechanism network to learn temporal association relationships. Compared with traditional neural networks, the multi-query attention mechanism network adopts a two-level self-attention mechanism, that is, after performing the attention score operation on multiple query vectors in the local domain, they are first weighted, and then the same attention operation is performed on the long query vectors from the remaining local domains, that is, the temporal attention score is used for secondary weighting, which can effectively realize the aggregation of temporal visual information.

[0064] In one embodiment, the multi-query attention mechanism network provided by this solution can be composed of a 5-layer multi-query neural network (multi-query Transformer), which is trained in an unsupervised autoregressive manner. Unlike the optimization goal of live stream screenshots, the training of the multi-query attention mechanism network needs to predict the missing frame based on the context of the missing frame, rather than just predicting the current frame based on the previous k frames in the window, so as to fully utilize the bidirectional temporal information. The objective function of the multi-query attention mechanism network is as follows:

[0065] where im k is the kth frame that is set as the missing frame, N is the number of video image frames, and the goal of multi-query attention mechanism network training is to maximize the probability of predicting the current frame through the previous and next frames, w kis an importance factor, which can be obtained through training. The pseudo three-dimensional convolutional model and multi-query attention mechanism network provided by this solution can be trained unsupervisedly using a set number of video frame sequence clusters / live stream screenshot clusters, and both do not require manual labeling. In this solution, when the visual information is a video image, each video image frame in the visual information is input into multiple second visual processing models respectively, and the second image token mapping information corresponding to each video image frame is determined by the second visual processing model, and the second image token mapping information is pooled into a picture semantic feature vector of a set dimension, and the picture semantic feature vector is processed by the multi-query attention mechanism network to obtain visual feature information by performing attention operation, so as to efficiently and accurately obtain the visual feature information of the visual information and improve the multimodal information labeling effect.

[0066] In one possible embodiment, the multimodal marking system provided by this solution performs multimodal fusion processing based on visual feature information, visual information, recognition text information, and description text information to obtain image-text fusion features, which may be:

[0067] S221: Obtain query vectors for each image frame in the visual information.

[0068] S222: Input the visual feature information, query vector, recognition text information and description text information into the trained visual-text fusion network. The visual-text fusion network performs multimodal fusion processing based on the cross-attention interaction mechanism based on the visual feature information, query vector, recognition text information and description text information to obtain image-text fusion features.

[0069] Exemplarily, after determining the visual feature information of the visual information, the multimodal tagging system further obtains the query vector of each image frame (live image frame, video image frame or picture) in the visual information, and inputs the above-determined visual feature information, query vector, recognition text information and description text information into the trained visual-text fusion network (VT Q-Former, Vision-Text Query Transformer). The visual-text fusion network performs multimodal fusion processing based on the cross-attention interaction mechanism based on the visual feature information, query vector, recognition text information and description text information to obtain image-text fusion features. This solution extracts the query vector of each image frame in the visual information and connects the visual-text fusion network that can realize the fusion of context or previous and next time node information after the visual processing model (including the first visual processing model and the second visual processing model), thereby reducing the output of erroneous information caused by time sequence fragmentation.

[0070] Optionally, the multimodal labeling system provided by this solution can obtain the query vector of the image frame based on the query vector learner of different visual types of images. Optionally, the query vector learner can be a temporal visual query vector learner that obtains query vectors for video images and live broadcast images, and a static visual query vector learner that obtains query vectors for single-image images (single pictures). Among them, the visual-text fusion network is configured with a visual feature end for receiving visual-related information and a text feature end for receiving text-related information. The visual feature end of the visual-text fusion network is directly input as an abstracted query vector. These query vectors remove redundant information as much as possible and model the surface spatial characteristics and temporal characteristics. This solution divides the query vector learner into a temporal visual query vector learner and a static visual query vector learner, which can effectively reduce the computational overhead when operating a single-image task and reduce the computational cost.

[0071] In one possible embodiment, the multimodal marking system provided by the present solution may, when obtaining the query vector of each image frame in the visual information, be: when the visual information is a time-series image, the visual information is input into a trained time-series visual query vector learner, and the time-series visual query vector learner uses a multi-scale convolutional neural network to determine the multi-scale visual association features corresponding to each image frame in the visual information, and perform feature extraction and fusion operations on the multi-scale visual association features to obtain the first query vector of each image frame in the visual information.

[0072] Exemplarily, when the visual information in the acquired information to be labeled is a time-series image (video image or live broadcast image), the visual information is input into a trained time-series visual query vector learner, and the time-series visual query vector learner is used to obtain a first query vector corresponding to each image frame in the visual information. The time-series visual query vector learner can use a multi-scale convolutional neural network (CNN) to determine the multi-scale visual correlation features corresponding to each image frame in the visual information, and perform feature extraction and fusion operations on the multi-scale visual correlation features to obtain the first query vector for each image frame in the visual information.

[0073] In one embodiment, as shown in the principle diagram of a temporal visual query vector learner provided in FIG6 , the visual information will first pass through a series of multi-scale convolutional neural networks (the multi-scale convolutional neural network ResNext15 in the figure), and the feature maps of the visual information at different scales are refined by neural network modules (Transformer modules) of different scales in the multi-scale convolutional neural network to extract local visual correlation features, and the local visual correlation features corresponding to each image frame are spliced ​​to obtain multi-scale visual correlation features. Then, a two-stage feature extraction and fusion operation (Mixer operation) is performed on the spliced ​​multi-scale visual correlation features from multiple image frames. For example, a group of multi-layer perceptrons (MLPs) are used to model the correlation of visual information across time, and then another group of multi-layer perceptrons MLPs are used to realize the self-modeling of the multi-scale information. After completing the two-stage feature extraction and fusion operation, the original input (multi-scale visual correlation features) will be passed to the subsequent steps through cross-layer connections. Considering that the temporal visual query vector learner proposed in this solution does not use a gating strategy to control the fusion of multiple sub-networks, and to reduce the number of parameters, the multi-scale convolutional neural network and the multi-layer perceptron in the two-stage feature extraction fusion operation all share weight parameters within the same group. The multi-scale convolutional neural network exploits the invariance of the convolution kernel to translation, rotation, and scaling to refine feature maps, mitigating the impact of positional information loss in the full neural network (Transformer network). Using neural network modules of different sizes on feature maps at different scales not only strengthens the connections between local regions, but also provides representations for image patches of different sizes, increasing the richness of the local attention mechanism and the diversity of associations. The combination of the multi-scale convolutional neural network and neural networks (CNN and Transformer) combines the advantages of both while further reducing computational overhead. The multi-scale feature construction model also reduces the inappropriate loss of semantic information. The subsequent feature extraction fusion operation models the temporal information of multiple image frames without using expensive attention strategies. Because the multi-layer perceptron has a good fit for sequential information, it also avoids the overuse of the full-scale matching attention mechanism.

[0074] Optionally, the temporal visual query vector learner provided by this solution does not need to be trained using manually labeled data. It only requires a loss function constructed with a small amount of prior information to complete the training, thereby improving the training efficiency of the temporal visual query vector learner.

[0075] In one embodiment, the loss function of the temporal visual query vector learner training can be expressed as: L tsq =λ1CE ts +λ2CE reg +λ3CE contr CEts =CE(In(sort(video i _frames),sort'(video i _frames)),In(sort(video i _frames),shuffle(video i _frames))) CE reg =CE(In(sort(video i _frames),sort'(video i _frames)),In(sort(video i _frames),shuffle(video j _frames))) CE contr =CE(In(sort(video i _frames),sort'(video i _frames)),In(sort(video i _frames),sort(video j _frames)))

[0076] Among them, the loss function L tsq Including CE ts ,CE reg and CE contr Three cross entropies, where In() can be understood as calculating the "feature difference" obtained by performing simple image transformation on a sequence of images from the same video / live stream screenshot, sort() can be understood as referring to a collection of image frames sorted in chronological order, and shuffle() can be understood as a collection of image frames that are randomly shuffled in order; CE ts CE is the cross entropy of the “feature difference” and the “feature difference” within the same cluster when the input is a sequence of sorted images and a randomly shuffled sequence of images from the same video / live stream screenshots, respectively. reg and CE contrSimilarly, the "feature differences" used for classification operations come from different video / live stream screenshots. This loss function hopes that the model will maximize the correlation within the same video / live stream screenshot cluster and model the temporal correlation, that is, the feature differences of the same video / live stream screenshot evolution cluster should be identified as the smallest. The regularization coefficients λ1, λ2 and λ3 can take values ​​of 0.5, 1.0 and 0.8 respectively, which are used to perform regularization of different strengths on different classification differences. This scheme targets the visual information of time-series images. Through the time-series visual query vector learner, a multi-scale convolutional neural network is used to determine the multi-scale visual association features corresponding to each image frame in the visual information, and feature extraction and fusion operations are performed on the multi-scale visual association features to accurately obtain the first query vector of each image frame in the visual information, thereby improving the accuracy of multimodal information labeling.

[0077] In one possible embodiment, the multimodal marking system provided by the present solution may, when obtaining the query vector of each image frame in the visual information, be: when the visual information is a single-image image, the visual information is input into a trained static visual query vector learner, and the static visual query vector learner uses a dual-task neural network to analyze and process the visual information to obtain a second query vector of the image frame corresponding to the visual information.

[0078] For example, when the visual information in the acquired information to be labeled is a single-image image (picture), the visual information is input into a trained static visual query vector learner, and the static visual query vector learner obtains a second query vector corresponding to each image frame in the visual information. The static visual query vector learner can analyze and process the visual information using a dual-task neural network to obtain the second query vector for the image frame corresponding to the visual information.

[0079] In one embodiment, as shown in the principle diagram of a static visual query vector learner provided in FIG7 , the static visual query vector learner is designed based on the idea of ​​a hybrid expert system (MoE, Mixture of Experts). Among them, the visual information of a single image first passes through a dual-task neural network (a two-task Transformer network). The dual-task neural network can be used as a routing network for the expert sub-network, that is, it generates a gating signal for each sub-network to control the importance of each sub-network to the final output. At the same time, the dual-task neural network will also output the third image token mapping information (image Token Map) corresponding to the image. The tokens (Token) of different rows of the third image token mapping information are grouped, and each group of tokens passes through a multi-layer perceptron (MLP) expert network with exclusive parameters to deeply abstract the information. Unlike general hybrid expert systems, the output after the multi-layer perceptron reasoning will be reassembled into the fourth image token mapping information (Token Map) and randomly shuffled (shuffle) in units of behavior, and regrouped to enter the next module. Among them, since the grouped token mapping information comes from non-overlapping local visual information, the gating signal generated by the routing network is exerted through pixel-by-pixel multiplication, and the weight parameters of the expert sub-network are not shared and require independent sub-networks to process. The randomly shuffled token mapping information effectively improves the generalization ability of each sub-network.

[0080] In one embodiment, the training of the static visual query vector learner provided by this solution does not require manually labeled data. The training of the static visual query vector learner needs to be completed through two stages of training. The loss function of the two-stage training is as follows:

[0081] Among them, L2 is the L2 loss based on distance metric (a loss function based on distance metric), BCE i is the binary cross entropy loss between different image clusters, each image cluster (im_clu i ) are considered to be independent categories, and a binary cross entropy loss is calculated between them. The loss function of the first stage of training aims to minimize the distance between the same image and its small variants while maximizing the distance between different images. The loss function of the second stage of training aims to maximize the distance between each cluster of images. The purpose of both is to enable the network to extract the general semantic features of the image while mining the detailed features to support a wider range of general tasks. This scheme uses a dual-task neural network to analyze and process visual information through a static visual query vector learner, accurately obtains the second query vector of the image frame corresponding to the visual information, and improves the accuracy of multimodal information labeling.

[0082] In one possible embodiment, as shown in a schematic diagram of the interaction principle of a visual-text fusion network provided in FIG8 , the visual-text fusion network provided by this solution includes a visual decoder and a text decoder. The visual decoder of the visual-text fusion network includes a first set number of stacked first neural network modules, the first neural network module includes a first self-attention subunit for processing query vectors and a first cross-attention subunit for processing visual feature information, and the text decoder of the visual-text fusion network includes a first set number of stacked second neural network modules, the second neural network module includes a second self-attention subunit for processing recognition text information and description text information, and the visual decoder and the text decoder interact based on a cross-attention interaction mechanism.

[0083] The visual-text fusion network is used to integrate visual and linguistic features, enabling subsequent text processing models (such as the Large Language Model (LLM)) to fully utilize information from both images and text for discrimination and generation. Therefore, the visual-text fusion network includes two decoders of different modalities: a visual decoder and a text decoder. Both decoders are composed of stacked neural network modules (Transformer modules). To ensure synchronization, both decoders use the same number (e.g., 12) of neural network modules. Because visual features contain rich information at both spatial and temporal levels, relying solely on the query vector learner and the visual decoder in the visual-text fusion network can easily miss or ignore features that are critical for task discrimination. Compared to the text decoder, the visual decoder requires an additional operation to interact with the visual processing model, which accurately supplements this information at multiple levels of abstraction. The LVM visual encoder (the visual encoder in the visual processing model) interacts with the visual decoder in the visual-text fusion network through a cross-attention mechanism. Specifically, the feature vector output by the LVM visual encoder is first converted to a query vector, and then an attention operation is performed and weighted. Because the visual processing model (visual decoder) lacks the ability to extract features from temporal information and super-resolution static images, while the text processing model (text transformer) already has the ability to model contextual information, the inputs of the visual decoder and the text decoder differ. The visual decoder receives an abstracted query vector, while the text decoder directly extracts features along with positional encoding information. It is important to note that the cross-attention layer of each visual decoder receives features from different levels of the LVM visual encoder, rather than from the same level, to fully leverage the abstract nature of features at different levels of the LVM. Information from the visual and text processing models primarily interacts in the self-attention layer, treating the visual and text query vectors as query vectors from the same domain and performing attention operations together. At different training stages, some different vector information can be masked to enable the visual-text fusion network to learn to handle different multimodal tasks. When providing query vectors for the visual-text fusion network, two specialized query vector learners are provided for temporal association and temporal context, solving the problem of previous query vector learners that only extract features from a single image and avoiding temporal gaps in the output results.

[0084] In one possible embodiment, the training of the visual-text fusion network provided by this solution includes:

[0085] S201: Input a first sample image-text pair, determine a first sample visual embedding vector of a sample image and a first sample text embedding vector of a sample text in the first sample image-text pair, make a visual query vector and a text query vector in a visual-text fusion network invisible to each other, and update a network weight of the visual-text fusion network according to distance information between the first sample visual embedding vector and the first sample text embedding vector.

[0086] S202: Input a second sample image-text pair, determine a second sample visual embedding vector of the sample image and a second sample text embedding vector of the sample text in the second sample image-text pair, make the visual query vector in the visual-text fusion network invisible to the text query vector, make the text query vector visible to the visual query vector and invisible to the text query vector of the text postscript, and update the network weight of the visual-text fusion network according to the distance information between the second sample visual embedding vector and the second sample text embedding vector.

[0087] S203: Input a third sample image-text pair, determine a third sample visual embedding vector of the sample image and a third sample text embedding vector of the sample text in the third sample image-text pair, make the visual query vector and the text query vector in the visual-text fusion network visible to each other, and update the network weight of the visual-text fusion network according to the distance information between the third sample visual embedding vector and the third sample text embedding vector.

[0088] S204: Input a fourth sample image-text pair, determine a fourth sample visual embedding vector of the sample image and a fourth sample text embedding vector of the sample text in the fourth sample image-text pair, randomly mask a set number of visual query vectors in the visual-text fusion network so that the text query vectors in the visual-text fusion network are invisible to the masked visual query vectors, and update the network weights of the visual-text fusion network according to the distance information between the fourth sample visual embedding vector and the fourth sample text embedding vector.

[0089] The training of the visual text fusion network provided by this solution includes four training stages (i.e., steps S201-S204), and the training of each stage is used to enable the visual text fusion network to learn the processing capability of a multimodal task. The first training stage is used to enable the visual text fusion network to learn the semantic level image-text comparison capability. First, the first sample image-text pair (the image-text pair is the matching image information and text information) is input into the visual text fusion network, wherein the image information is first extracted with high-quality features through the LVM visual encoder (i.e., the text processing model provided by this solution) while using the query vector learner (including the temporal visual query vector learner and the static visual query vector learner) to learn the query vector as input, and then add type tags (such as the corresponding image) to the text information according to the input type of the visual information (including single image, video and live stream screenshots). <independent>, and corresponding video and live stream screenshots <time>), then uses the visual decoder and text decoder to calculate the visual embedding vector (emb, embedding) and text embedding vector, respectively, and calculates the cosine distance between them. This cosine distance is used to optimize the L2 loss based on image-text similarity and update the network weights of the visual-text fusion network. To prevent the image encoder from directly learning text information and causing overfitting, the self-attention layers of the visual decoder and text decoder mask the query vectors. This first training phase enables the visual-text fusion network to compare semantic information between images and text.

[0090] In one embodiment, after the first training stage converges, the visual-text fusion network will learn the ability to generate text from images. The first two steps use pre-processing similar to the first training stage. The difference is that the text decoder does not mask the query vector from the visual decoder, but the text decoder will mask the query vector corresponding to the word after the current word to enable the model to learn the ability to generate subsequent text. At the same time, this step also uses the maximum likelihood method to optimize the objective function, and no additional labeling is required.

[0091] In one embodiment, the third training phase enables the visual-text fusion network to achieve image-text matching capabilities. Neither decoder class will obscure the other's query vector, allowing image features to match text features in detail. This training phase calculates the image-text matching degree feature by feature and optimizes a binary loss function based on the feature-by-feature proximity. After completing the first three phases of training, the visual-text fusion network has essentially acquired the ability to handle relatively complete image-text modality tasks.

[0092] In one embodiment, since there is often a considerable amount of adversarial data in video / live broadcast products, traditional multimodal networks are prone to "answering the wrong question" on these adversarial data. Therefore, this solution adds a training stage based on the first three types of training. This stage will randomly mask a certain number of visual query vectors, and then make the text decoder invisible to these query vectors. The loss function is the loss function of the third stage supplemented with the corresponding regularization term. After the fourth stage of training is completed, the visual text fusion network can also accurately extract features and fuse data that is deliberately masked or covers up specific frames or live broadcast screenshots. This solution draws on the ideas of unsupervised training deep neural networks such as autoregression, and designs four training methods for the visual text fusion network based on contextual association and the importance of image association in the same time interval, which greatly reduces the cost of manual labeling and training, and enables the visual text fusion network to handle strong adversarial data such as covering or concealing, thereby improving the security of the model and reducing the production of unsafe content.

[0093] In one possible embodiment, as shown in a schematic diagram of a process for determining image-text relevance information provided in FIG9 , the multimodal marking system provided in this solution, when determining image-text relevance information between visual information and recognized text information and descriptive text information, includes:

[0094] S231: Determine a visual feature vector corresponding to each image frame in the visual information.

[0095] S232: Determine text feature vectors corresponding to the recognition text information and the description text information.

[0096] S233: Determine the image-text similarity between each visual feature vector and text feature vector, and determine image-text association information based on the image-text similarity.

[0097] Exemplarily, after acquiring the information to be labeled, the multimodal labeling system determines the visual feature vectors corresponding to each image frame in the visual information, as well as the text feature vectors corresponding to the identification text information and description text information in the visual information. After determining the visual feature vectors and text feature vectors, the image-text similarity between each visual feature vector and text feature vector is determined, and based on the image-text similarity, the image-text association information corresponding to the current information to be labeled is determined. The image-text similarity can be represented by Hamming distance, cosine distance, Euclidean distance, Manhattan distance, Chebyshev distance, Minkowski distance, Mahalanobis distance, etc.

[0098] Optionally, the visual feature vectors of the image frames, as well as the text feature vectors for the identified and descriptive text information, can be extracted using a pre-set feature vector extraction algorithm or a trained feature vector extraction model. This solution determines image-text correlation information based on the image-text similarity between the visual feature vectors corresponding to each image frame in the visual information and the text feature vectors corresponding to the identified and descriptive text information. This accurately reflects the degree of correlation between the visual and textual aspects of the information to be identified, thereby improving the accuracy of multimodal information labeling.

[0099] Specifically, for the calculation of cross-modal correlation (i.e., the calculation of image-text correlation information), this solution is based on the principle of language image contrast learning (CLIP, Contrast Language Image Pretrain) and designs a cross-modal calculation method for image-text correlation, which is used to mark some low-quality content that requires cross-modal information to be associated for judgment. As shown in Figure 10, a schematic diagram of the principle of cross-modal calculation of image-text correlation is provided. Unlike ordinary language image contrast learning that only calculates correlation for a single static image and a single text description, considering that videos and live streams need to combine temporal visual correlation, this solution extracts visual feature vectors one by one from the frame sequence after video decoding (video image frame sequence) or the screenshot sequence of the live stream (live image frame sequence), and also extracts text feature vectors from the text transcribed from the audio segment corresponding to each image frame. After calculating the similarity (for example, cosine similarity) of multiple sets of features (visual feature vectors and text feature vectors) pairwise, they are integrated according to the similarity ratio distribution to obtain the image-text correlation information, effectively alleviating the situation where excessive local correlation leads to false suppression. In addition, both audio and images are divided in time sequence, reducing the amount of calculation.

[0100] Optionally, the visual encoder used in this solution to extract visual feature vectors is the pre-trained Res101 network from the original CLIP (Language-Image Contrastive Learning) combination, with the layer immediately preceding the output layer used as the encoded feature vector. The text feature encoder, meanwhile, utilizes a 12-layer neural network (Transformer network) based on an encoder-decoder architecture, also using an open-source pre-trained version. Both the visual encoder and text feature encoder extract abstract semantic features from images and text while retaining strong generalization capabilities. Because CLIP (Language-Image Contrastive Learning) is trained using contrastive learning, the visual and text feature vectors from different modalities are aligned and can directly participate in various aggregation operations. The proposed cross-modal image-text relevance computation also utilizes the visual text information in the image and the conversational text information obtained by converting the audio into the conversational text. The visual text information is extracted using an optical character recognition model (OCR) based on the text-to-image ratio, while the conversational text information is extracted using an audio-to-text speech recognition model (ASR) based on invalid speech filtering. The visual text information and the dialogue text information correspond to the video frames or live screenshots at the corresponding moments, effectively improving the alignment between the text feature vector and the visual feature vector, and also solving the problem of insufficient details when only using text information such as video descriptions and live broadcast theme descriptions.

[0101] In one possible embodiment, the multimodal tagging system provided by this solution, when determining image-text relevance information based on image-text similarity, includes:

[0102] S2331: When the number of visual feature vectors is an integer multiple of the number of text feature vectors, determine the image-text association information according to the proportion of the number of image-text similarities that reaches a set similarity threshold in the total number of feature vectors of the visual feature vectors and the text feature vectors.

[0103] S2332: When the number of visual feature vectors does not divide the number of text feature vectors, determine image-text association information according to the proportion of the number of image-text similarities that reach a set similarity threshold in the number of feature vectors of the visual feature vectors.

[0104] Exemplarily, after determining the visual feature vectors and the text feature vectors, it is determined whether the number of the visual feature vectors is divisible by the number of the text feature vectors.

[0105] In one embodiment, if the number of visual feature vectors is evenly divisible by the number of text feature vectors, it can be understood that the features of the visual information and the text information are evenly sampled in the same time period. At this time, each visual feature vector only needs to find the number of text feature vectors with a high correlation with it and calculate the proportion in the collection, that is, the image-text correlation information is determined based on the proportion of the number of image-text similarities that reach the set similarity threshold in the total number of feature vectors of the visual feature vector and the text feature vector.

[0106] In one embodiment, if the number of visual feature vectors cannot be divided evenly into the number of text feature vectors, it can be understood that the visual feature vector needs to be matched with all text feature vectors, and the correlation with the greatest impact is taken as the image-text correlation information, that is, the image-text correlation information is determined based on the proportion of the number of image-text similarities that reach the set similarity threshold in the number of feature vectors in the visual feature vector.

[0107] In one embodiment, the calculation formula for the image-text relevance information provided by this solution can be expressed as:

[0108] Among them, d(IE i ,TE j ) is the i-th visual feature vector IE i and the jth text feature vector TE j The cosine similarity of N im is the number of eigenvectors of the visual feature vector, N te is the number of feature vectors of the text feature vector, thd is the set similarity threshold, and the value range of the set similarity threshold can be 0.8 to 0.95. For example, the set similarity threshold can be 0.9. In one embodiment, when the text information in the input information to be labeled is a single video description and a live broadcast theme description, the image-text correlation information can be determined based on the proportion of the number of image-text similarities that reach the set similarity threshold in the number of feature vectors of the visual feature vector. This solution improves the accuracy of determining the image-text correlation information and improves the accuracy of multimodal information labeling by respectively determining the image-text correlation information based on the integer divisibility of the number of visual feature vectors and the number of text feature vectors.

[0109] In one possible embodiment, when the multimodal labeling system provided by the present solution determines the labeling result of the information to be labeled based on the image-text fusion features and the image-text correlation information, it can be: inputting the image-text fusion features and the image-text correlation information into a text processing model, and analyzing and processing the image-text fusion features and the image-text correlation information through the text processing model to obtain the content description information and / or open set keywords of the information to be labeled.

[0110] The text processing model provided by this solution can be built based on a large language model (LLM). In one embodiment, the labeling results provided by this solution can be content description information and / or an open set of keywords. For example, the image-text fusion features and image-text relevance information are input into the text processing model, and the text processing model analyzes and processes the image-text fusion features and image-text relevance information to obtain the content description information and / or an open set of keywords for the information to be labeled.

[0111] Among them, open-set keywords can be used as a supplement to the topK keywords when the labels are too coarse during recommendation. New labels can also be summarized by the multi-task synonym summarizer and added to the label system to improve the coverage of the labeling system. Due to the characteristics of the open set of keywords, the design and expansion of the label tree basically realize automatic circulation, which can effectively reduce the cost of human evaluation and increase the research and development of labels. For example, the trained multi-task synonym summarizer can be used to summarize the open set keyword set, and the labels (including high-quality labels and low-quality labels) can be expanded according to the summary results. This solution uses a text processing model to analyze and process the image-text fusion features and image-text correlation information to obtain content description information and / or open-set keywords of the information to be labeled, accurately output the labeling results of the information to be labeled, and improve the labeling effect of multimodal information.

[0112] In a possible embodiment, the multimodal labeling system provided by the present solution further includes: inputting the image-text fusion features into a trained alignment model, and aligning the image-text fusion features through the alignment model before determining the labeling results of the information to be labeled based on the image-text fusion features and the image-text correlation information, so that the image-text fusion features after alignment conform to the output preferences of the text processing model.

[0113] Exemplarily, after obtaining the image-text fusion feature, the image-text fusion feature is input into the trained alignment model, and the image-text fusion feature is aligned by the alignment model so that the aligned image-text fusion feature conforms to the output preference of the text processing model. Among them, after completing the fusion of the image-text bimodal features to obtain the image-text fusion feature, if the embedded feature vector corresponding to the image-text fusion feature is directly input into the text processing model and generates an output, since the open source LLM generally adjusts the output preference through reinforcement learning, its encoded input vector also needs to conform to the corresponding output preference, so as to avoid security issues. There are situations where some security issues are caused by the unlimited generation of content. Based on this, this solution can perform an alignment operation on the image-text fusion feature before inputting it into the text processing model, so that the aligned image-text fusion feature conforms to the output preference of the text processing model, thereby ensuring the accuracy of multimodal information labeling.

[0114] In one embodiment, the alignment model provided by this solution includes:

[0115] S241: Convert the previously obtained image-text fusion features into a context feature group through a fully connected layer, and determine the routing value corresponding to the context feature group through a gate selection network.

[0116] S242: Perform weighted processing on the routing value and the conversion vector obtained through training to obtain a weighted conversion vector.

[0117] S243: Through the cascaded decoder network and multi-layer perceptron, the current image-text fusion features and the weighted transformation vector are aligned and summed feature by feature.

[0118] As shown in Figure 11, the alignment model provided in this solution includes multiple connected layers (FC layers in Figure 11), a decoder network (Decoder-Transformer network in Figure 11), and a multi-layer perceptron (MLP). The decoder network and the MLP are cascaded. This alignment model is based on the concept of a hybrid expert system, with a series of fully connected layers acting as expert sub-networks, and a decoder network cascaded with a multi-layer perceptron acting as the gate selector for the expert network. During actual inference, if it is determined that previous context information from previous inputs needs to be considered, the image-text fusion features obtained from the previous encoding fusion are cached and used as the corpus input for each new inference. Each image-text fusion feature obtained from the previous encoding fusion is input into a fully connected layer, where it is converted into a contextual feature group. The concatenated contextual feature group is then input into a gate selector network to calculate a routing value. The routing value and the output transformation vector are weighted, and then the sum of the weights is aligned and summed to form the input to the text processing model. Optionally, multiple fully connected layers in the alignment model can share weight parameters to save parameters.

[0119] Exemplarily, previously obtained image-text fusion features are input into the connection layer. A fully connected layer converts these features into a contextual feature group, and a gated network determines the routing value corresponding to the contextual feature group. The routing value and the trained transformation vector are weighted to produce a weighted transformation vector. The current image-text fusion features are then aligned and summed with the weighted transformation vector through a cascaded decoder network and a multi-layer perceptron, yielding the aligned image-text fusion features.

[0120] Optionally, the training of the alignment model module requires first fixing the parameters of the visual processing model, the visual-text fusion network, and the text processing model, using the feature vector inferred by the visual-text fusion network as input and the output of the text processing model as the predicted value matching the label, for integrated parameter optimization. The label uses the result generated by the third-party large model interface as the candidate set, and the matching text is manually selected as the real label judgment. This solution converts the previously obtained image-text fusion features into a context feature group through a fully connected layer, and determines the routing value corresponding to the context feature group through a gate selection network. The routing value and the trained conversion vector are weighted to obtain a weighted conversion vector, and the current image-text fusion feature is aligned and summed with the weighted conversion vector through a cascaded decoder network and a multi-layer perceptron to obtain the aligned image-text fusion feature, so that the aligned image-text fusion feature meets the output preference of the text processing model and ensures the accuracy of multimodal information labeling.

[0121] In one possible embodiment, because the alignment of image-text fusion features takes into account previous context, the multimodal labeling system proposed in this solution also exhibits a certain degree of "self-correction" capability, allowing for continuous correction through prompts (instructional text) to achieve personalized task settings. When determining the labeling result for the information to be labeled based on the image-text fusion features and image-text relevance information, the multimodal labeling system provided in this solution can also: input pre-configured instructional text, image-text fusion features, and image-text relevance information into a text processing model, and then analyze and process the image-text fusion features and image-text relevance information to obtain the labeling result for the information to be labeled. The instructional text is used to activate the text processing model's reasoning ability for the task of identifying a set label range.

[0122] Exemplarily, a pre-configured instruction text ( <prefix>In-context Mode < / prefix> ) and image-text relevance information ( <prefix> Inferior Reg< / prefix> ), the aligned image-text fusion features are input to the text decoder of the text processing model, and the text processing model analyzes and processes the image-text fusion features and the image-text correlation information to obtain the labeling results of the information to be labeled, and the text processing model can activate the text processing model's reasoning ability for the set label range recognition task based on the instruction text, thereby improving the sensitivity to the set label range recognition. This solution increases the prefix prompt for the recognition of low-quality labels by adding instruction text, thereby achieving range-controllable low-quality label recognition and improving the multimodal information labeling effect. Compared with the conventional text processing model that only includes a decoder, the text encoder of the text processing model can also use instruction text (prefix prompt) to maximize the combination of contextual image and text information to improve the accuracy of the output material. Among them, the embedded prompt text is encoded into an embedding vector with few-shot prompt capability through a high-quality feature abstraction process, which can effectively activate the text processing model's reasoning ability for specific mode tasks.

[0123] In a possible embodiment, the multimodal information marking method provided by this solution further includes:

[0124] S251: Inputting the image-text fusion features and the image-text relevance information into the text processing model, and analyzing and processing the image-text fusion features and the image-text relevance information through the text processing model to obtain the contextual text features of the information to be marked.

[0125] S252: Input the contextual text features, and a combination of one or more of the voice information, recognition text information, and description text information corresponding to the visual information into the trained first label recognition model, and analyze and process the contextual text features, and a combination of one or more of the voice information, recognition text information, and description text information corresponding to the visual information through the first label recognition model to obtain the first label information.

[0126] For example, the image-text fusion features and image-text relevance information are input into a text processing model, which then analyzes and processes the image-text fusion features and image-text relevance information to obtain contextual text features (In-Context) of the information to be labeled. Optionally, the contextual text features of the information to be labeled can be used as a form of representation of the labeling result.

[0127] In one embodiment, after obtaining the contextual text features of the information to be labeled, the contextual text features, and a combination of one or more of the voice information, recognition text information, and description text information corresponding to the visual information are input into a trained first label recognition model (e.g., a unimodal text / speech small model). The first label recognition model analyzes and processes the contextual text features, and a combination of one or more of the voice information, recognition text information, and description text information corresponding to the visual information to obtain first label information. This solution utilizes the first label recognition model to output the first label information corresponding to the information to be labeled, and utilizes the intermediate features produced by the large model (visual processing model and text processing model) as the input of the existing label system small model to assist in achieving classification labeling and improve the multimodal labeling effect.

[0128] In a possible embodiment, the multimodal information labeling method provided by this solution also includes: inputting the image-text fusion features and / or visual information into a trained second label recognition model, and analyzing and processing the image-text fusion features and / or visual information through the second label recognition model to obtain second label information.

[0129] Exemplarily, the obtained image-text fusion features and / or visual information are input into a trained second label recognition model (e.g., a small unimodal image model). The second label recognition model then analyzes and processes the image-text fusion features and / or visual information to obtain the second label information. This solution utilizes the second label recognition model to output the second label information corresponding to the information to be labeled, and uses the intermediate features generated by the large model as input to the small model of the existing label system to assist in classification labeling and improve the multimodal labeling effect.

[0130] As shown in the structural diagram of a multimodal information labeling system provided in Figure 12, the multimodal information labeling system provided by this solution includes a visual processing model (including a first visual processing model Vision Encoder and a second visual processing model Vision Encoder cluster), a visual text fusion network (VT Q-Former), a text processing model (including a text encoder Language Encoder and a text decoder Language Decoder), an alignment model (a multimodal LLM alignment module based on context association), a first label recognition model (a single-modal text / speech small model set), and a second label recognition model (a single-modal image small model set). The labeling result of the information to be labeled provided by this solution can be a combination of one or more of image-text fusion features, aligned image-text fusion features, content description information, open-set keywords, contextual text features, first label information, and second label information.

[0131] The multimodal information labeling system is based on an open-source visual processing model (LVM) and a text processing model (LLM) pre-trained using ultra-large-scale parameters. It matches and fuses image and text information through a visual-text fusion network, combining high-quality image features extracted by the visual processing model with prompts. Finally, based on the alignment, the text processing model generates keywords and a description with specific semantic understanding. Because the large-parameter model learns patterns at a very large scale, it possesses strong generalization and complex multimodal understanding capabilities, adapting to a wide range of tasks without any adjustments, eliminating the need for specialized multi-task training. Furthermore, since it utilizes an open-source model, it eliminates the need for extensive manual annotation and expensive training. However, the original visual processing model and text processing model learn from different data distributions, resulting in significant feature differences. This solution uses a visual-text fusion network to match and fuse image and text features and aligns the models using multimodality, enabling the text model to process the fused image and text features.

[0132] To learn key contextual information about videos / live streams in time series, this solution inputs query vectors (learned queries) extracted from temporally correlated visual features into the visual side of the visual-text fusion network. It also integrates a contextual or preceding and following time node information fusion module behind two open-source vision encoders in the visual processing model to reduce the "hallucination" output caused by temporal fragmentation. After focusing on the fusion of visual and language information, the core dual-model can produce several feature vectors with different emphases, such as multimodal fusion feature vectors, text-aligned multimodal feature vectors, and image-text feature vectors incorporating prefix language features. These feature vectors can be output as labeling results for the information to be labeled and used as input for downstream tasks. The multimodal fusion feature vectors can be directly used for machine review and data mining of image-text associations, while the text-aligned multimodal feature vectors and the image-text feature vectors incorporating prefix language features can be embedded in the inference process of content generation to enhance user experiences such as AI drawing and AI question-answering. In addition to the high-quality content mining and labeling that conventional labeling systems focus on, this solution can also determine the relevance of text and image information for content such as irrelevant image and text diversion / like fraud and obscure political metaphors, quickly determine the scope of low-quality content, and add prefix prompts for low-quality label identification through an optional text encoder, achieving controllable range of low-quality label identification.

[0133] In one embodiment, in addition to using the intermediate features produced by the large model as input to the existing small model of the labeling system to assist in classification and labeling, the original small model set can also directly accept single-modal information such as images, text, or speech, or multimodal features produced by the multimodal information labeling system as input. Small models customized for different scenarios can be used to complete different tasks, complementing the large model and improving the effectiveness of multimodal information labeling. Of course, the large model itself can also output keywords, descriptions, and other information required for specific tasks. At the same time, because the large model has learned the ability to reason about multiple tasks from a large amount of visual and linguistic materials, it can smoothly switch tasks by simply changing the prompt statement. The multimodal labeling system provided by this solution can also produce multi-granular information and apply this information to a wide range of downstream tasks, further facilitating the refined delivery of downstream content. Compared to traditional labeling systems, it is not only easier to scale but also reduces the development costs increased by the lack of commonality between tasks. The number of labels and keywords provided by this solution can reach hundreds of times that of traditional labeling systems. It is more adept at processing abstract and adversarial data. In addition to providing a basis for targeted distribution and recommendation of high-quality content, it can also suppress low-quality content. This approach, which uses a large model as the core and incorporates smaller models for correction, significantly overcomes the slow scalability of mainstream tagging systems. It adds a keyword set of hundreds of thousands to the existing multi-level tagging system. These keywords can be easily added to new tags in response to changes in the content ecosystem, eliminating the need to collect data and train models specifically for these new tags. Furthermore, these keywords are open sets, capped at the total number of image / text patterns contained in the large model. This allows for rapid adaptation and adjustment to changes in the content ecosystem, significantly reducing the costs of data collection and model training.

[0134] Taking content information from different modalities, such as vision, audio, and text, as raw input, this approach generates intermediate feature maps and embedded feature vectors through cross-modal correlation matching and multimodal fusion. Based on this intermediate information, a multi-layered tagging system, built around a "large model" and using multiple small models and intelligent processing strategies as "gate units," serves as the inference module. This system outputs multiple signals, including features, keywords, tags, and even descriptions. These signals are used for user segmentation, customized user recommendations, enhanced content production experience, and content quality management. Because this solution uses multimodal information as metadata for model system inference, it can accurately extract content characteristics from videos / live broadcasts, providing accurate tagging results, even in situations where current tagging systems struggle with abstract or niche content. This is particularly true for low-quality content, such as obscure or adversarial content, where single-modal tagging systems often fail due to ignoring detailed features. The multi-granularity labeling model proposed in this solution can still find obscure details from the details, and provide keywords and supplementary descriptions. Adversarial content can reduce the probability of distortion over time through the support of massive image and text patterns by "big models" and the convenient generalization of multiple tasks, which is beneficial to the governance of the content ecosystem. In addition, in terms of architecture, this solution also builds a related technology stack with common deep learning tasks as the pre-foundation support, effectively reducing overhead and ensuring that the most basic video / live multimodal information can be effectively reused by the multi-granularity labeling system. In addition, the output of the layered multi-granularity multimodal labeling system can cover multiple dimensions such as high-quality / low-quality content, providing comprehensive support for various downstream product applications, such as recommendation systems, advertising algorithms, content production and order systems.

[0135] As described above, the visual feature information of the visual information is determined by a multimodal tagging system. Multimodal fusion processing is performed based on the visual feature information, visual information, recognition text information, and description text information to obtain the image-text fusion feature, as well as the image-text correlation information between the visual information, recognition text information, and description text information. The tagging result of the information to be tagged is determined based on the image-text fusion feature and the image-text correlation information. Multi-dimensional information tagging is performed based on multimodal information fusion and cross-modal correlation, effectively improving the information tagging effect, effectively improving the tagging coverage and tagging accuracy of abstract content and adversarial content, and considering the correlation between different modalities, overcoming the shortcomings of single-modal tagging, which is single-minded and easy to judge based on non-core elements. It also considers the calculation of the correlation between visual information and recognition text information and description text information for multiple time segments, providing the ability to identify low-quality labels of obscure and adversarial nature, and supplementing the defect that conventional tagging systems can only cover the tagging of high-quality content.

[0136] FIG13 is a schematic diagram of the structure of a multimodal information marking device provided in an embodiment of the present application. Referring to FIG13 , the multimodal information marking device includes an information acquisition module 31 and a marking processing module 32 .

[0137] Among them, the information acquisition module 31 is configured to acquire information to be labeled, which includes visual information, recognition text information and description text information; the labeling processing module 32 is configured to input the information to be labeled into the trained multimodal labeling system, determine the visual feature information of the visual information through the multimodal labeling system, perform multimodal fusion processing based on the visual feature information, visual information, recognition text information and description text information, obtain image-text fusion features, and determine the image-text correlation information between the visual information and the recognition text information and the description text information, and determine the labeling result of the information to be labeled based on the image-text fusion features and the image-text correlation information.

[0138] In the above, the visual feature information of the visual information is determined by the multimodal labeling system, and multimodal fusion processing is performed according to the visual feature information, visual information, recognition text information and description text information to obtain the image-text fusion feature, and the image-text correlation information of the visual information and the recognition text information and the description text information is determined, and the labeling result of the information to be labeled is determined according to the image-text fusion feature and the image-text correlation information, and multi-dimensional information labeling is performed based on multimodal information fusion and cross-modal correlation, which effectively improves the information labeling effect.

[0139] In one possible embodiment, when determining the visual feature information of visual information, the multimodal labeling system is configured as follows: when the visual information is a live image, each live image frame in the visual information is input into the first visual processing model in chronological order, and the first image token mapping information of each live image frame is determined through the first visual processing model; the first image token mapping information is stacked, and the stacked first image token mapping information is subjected to temporal context association fusion through a pseudo three-dimensional convolutional model to obtain a feature map corresponding to each first image token mapping information; the feature map corresponding to each first image token mapping information is pooled into a first feature map of a set dimension, and the first feature map is mapped to visual feature information through a fully connected layer.

[0140] In one possible embodiment, when determining the visual feature information of visual information, the multimodal labeling system is configured as follows: when the visual information is a video image, each video image frame in the visual information is input into multiple second visual processing models respectively, and the second image token mapping information corresponding to each video image frame is determined by the second visual processing model; the second image token mapping information is pooled into an image semantic feature vector of a set dimension, and the image semantic feature vector is subjected to attention operation processing through a multi-query attention mechanism network to obtain visual feature information.

[0141] In one possible embodiment, when the multimodal tagging system performs multimodal fusion processing based on visual feature information, visual information, recognition text information and description text information to obtain image-text fusion features, the system is configured to: obtain the query vector of each image frame in the visual information, input the visual feature information, query vector, recognition text information and description text information into the trained visual-text fusion network, and perform multimodal fusion processing based on the visual feature information, query vector, recognition text information and description text information through the visual-text fusion network based on the cross-attention interaction mechanism to obtain image-text fusion features.

[0142] In one possible embodiment, when obtaining the query vector of each image frame in the visual information, the multimodal labeling system is configured as follows: when the visual information is a time-series image, the visual information is input into a trained time-series visual query vector learner, and the time-series visual query vector learner uses a multi-scale convolutional neural network to determine the multi-scale visual association features corresponding to each image frame in the visual information, and performs feature extraction and fusion operations on the multi-scale visual association features to obtain the first query vector of each image frame in the visual information.

[0143] In one possible embodiment, when obtaining the query vector of each image frame in the visual information, the multimodal labeling system is configured as follows: when the visual information is a single-image image, the visual information is input into a trained static visual query vector learner, and the static visual query vector learner uses a dual-task neural network to analyze and process the visual information to obtain a second query vector of the image frame corresponding to the visual information.

[0144] In one possible embodiment, the visual-text fusion network includes a visual decoder and a text decoder, the visual decoder includes a first set number of stacked first neural network modules, the first neural network module includes a first self-attention subunit for processing query vectors and a first cross-attention subunit for processing visual feature information, the text decoder includes a first set number of stacked second neural network modules, the second neural network module includes a second self-attention subunit for processing recognition text information and description text information, and the visual decoder and the text decoder interact based on a cross-attention interaction mechanism.

[0145] In one possible embodiment, the multimodal information labeling device includes a visual text fusion network training model, and the visual text fusion network training model is configured as follows: inputting a first sample image-text pair, determining a first sample visual embedding vector of a sample image and a first sample text embedding vector of a sample text in the first sample image-text pair, so that a visual query vector and a text query vector in the visual text fusion network are invisible to each other, and updating the network weights of the visual text fusion network according to the distance information between the first sample visual embedding vector and the first sample text embedding vector; inputting a second sample image-text pair, determining a second sample visual embedding vector of a sample image and a second sample text embedding vector of a sample text in the second sample image-text pair, so that the visual query vector in the visual text fusion network is invisible to the text query vector, the text query vector is visible to the visual query vector and is invisible to the text query vector of the subsequent text, and updating the network weights of the visual text fusion network according to the distance information between the second sample visual embedding vector and the second sample text embedding vector. The distance information of the text embedding vector is used to update the network weight of the visual text fusion network; a third sample image-text pair is input to determine the third sample visual embedding vector of the sample image and the third sample text embedding vector of the sample text in the third sample image-text pair, so that the visual query vector and the text query vector in the visual text fusion network are visible to each other, and the network weight of the visual text fusion network is updated according to the distance information of the third sample visual embedding vector and the third sample text embedding vector; a fourth sample image-text pair is input to determine the fourth sample visual embedding vector of the sample image and the fourth sample text embedding vector of the sample text in the fourth sample image-text pair, a set number of visual query vectors in the visual text fusion network are randomly masked, so that the text query vector in the visual text fusion network is invisible to the masked visual query vector, and the network weight of the visual text fusion network is updated according to the distance information of the fourth sample visual embedding vector and the fourth sample text embedding vector.

[0146] In one possible embodiment, when determining the graphic-text correlation information between visual information and identification text information and description text information, the multimodal marking system is configured to: determine the visual feature vector corresponding to each image frame in the visual information; determine the text feature vector corresponding to the identification text information and description text information; determine the graphic-text similarity between each visual feature vector and the text feature vector, and determine the graphic-text correlation information based on the graphic-text similarity.

[0147] In one possible embodiment, when the multimodal tagging system determines the image-text correlation information based on the image-text similarity, it is configured as follows: when the number of visual feature vectors is evenly divisible by the number of text feature vectors, the image-text correlation information is determined based on the proportion of the number of image-text similarities that reach a set similarity threshold in the total number of feature vectors of the visual feature vector and the text feature vector; when the number of visual feature vectors is not evenly divisible by the number of text feature vectors, the image-text correlation information is determined based on the proportion of the number of image-text similarities that reach a set similarity threshold in the total number of feature vectors of the visual feature vector.

[0148] In one possible embodiment, when the multimodal tagging system determines the tagging result of the information to be tagged based on the image-text fusion features and the image-text correlation information, it is configured as follows: the image-text fusion features and the image-text correlation information are input into a text processing model, and the image-text fusion features and the image-text correlation information are analyzed and processed by the text processing model to obtain the content description information and / or open set keywords of the information to be tagged.

[0149] In one possible embodiment, before determining the labeling result of the information to be labeled based on the image-text fusion features and the image-text correlation information, the multimodal labeling system is further configured to: input the image-text fusion features into a trained alignment model, and align the image-text fusion features through the alignment model so that the aligned image-text fusion features meet the output preferences of the text processing model.

[0150] In one possible embodiment, when the alignment model performs alignment processing on the image-text fusion features, it is configured as follows: the image-text fusion features previously obtained are converted into a context feature group through a fully connected layer, and the routing value corresponding to the context feature group is determined through a gate selection network; the routing value and the trained conversion vector are weighted to obtain a weighted conversion vector; and the current image-text fusion features and the weighted conversion vector are aligned and summed feature by feature through a cascaded decoder network and a multi-layer perceptron.

[0151] In one possible embodiment, when determining the marking result of the information to be marked based on the image-text fusion feature and the image-text relevance information, the multimodal marking system is configured as follows:

[0152] The pre-configured instruction text, image-text fusion features, and image-text correlation information are input into the text processing model. The image-text fusion features and image-text correlation information are analyzed and processed by the text processing model to obtain the labeling results of the information to be labeled. The instruction text is used to activate the reasoning ability of the text processing model for the set label range recognition task.

[0153] In a possible embodiment, multimodal information labeling also includes a first labeling module, which is configured to: input the image-text fusion features and the image-text correlation information into the text processing model, and analyze and process the image-text fusion features and the image-text correlation information through the text processing model to obtain the contextual text features of the information to be labeled; input the contextual text features, and a combination of one or more of the voice information, recognition text information, and description text information corresponding to the visual information into the trained first label recognition model, and analyze and process the contextual text features, and a combination of one or more of the voice information, recognition text information, and description text information corresponding to the visual information through the first label recognition model to obtain the first label information.

[0154] In a possible embodiment, multimodal information labeling also includes a second label module, which is configured to: input the image-text fusion features and / or visual information into a trained second label recognition model, and analyze and process the image-text fusion features and / or visual information through the second label recognition model to obtain second label information.

[0155] It is worth noting that in the embodiment of the above-mentioned multimodal information marking device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of this application.

[0156] The embodiment of the present application also provides a multimodal information marking device, which can integrate the multimodal information marking device provided in the embodiment of the present application. Figure 14 is a structural diagram of a multimodal information marking device provided in the embodiment of the present application. Referring to Figure 14, the multimodal information marking device includes: an input device 43, an output device 44, a memory 42 and one or more processors 41; the memory 42 is used to store one or more programs; when one or more programs are executed by one or more processors 41, the one or more processors 41 implement the multimodal information marking method provided in the above embodiment. The multimodal information marking device, equipment and computer provided above can be used to execute the multimodal information marking method provided in any of the above embodiments, and have corresponding functions and beneficial effects.

[0157] The embodiments of the present application also provide a non-volatile storage medium that stores computer-executable instructions, which are used to execute the multimodal information marking method provided in the above embodiments when executed by a computer processor. Of course, the non-volatile storage medium that stores computer-executable instructions provided in the embodiments of the present application, whose computer-executable instructions are not limited to the multimodal information marking method provided above, can also execute the related operations in the multimodal information marking method provided in any embodiment of the present application. The multimodal information marking device, equipment and storage medium provided in the above embodiments can execute the multimodal information marking method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the multimodal information marking method provided in any embodiment of the present application.

[0158] On the basis of the above embodiments, the embodiments of the present application also provide a computer program product. The essence of the technical solution of the present application or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes a number of instructions for enabling a computer device, a mobile terminal or a processor therein to execute all or part of the steps of the multimodal information marking method provided in each embodiment of the present application.< / time> < / independent>

Claims

1. A multimodal information marking method, wherein, Including: Obtain the information to be labeled, where the information to be labeled includes visual information, recognized text information, and descriptive text information; Input the information to be labeled into the trained multi-modal labeling system. Determine the visual feature information of the visual information through the multi-modal labeling system. Perform multi-modal fusion processing based on the visual feature information, the visual information, the recognized text information, and the descriptive text information to obtain the graphic-text fusion feature, and determine the graphic-text correlation degree information between the visual information and the recognized text information and the descriptive text information, and determine the labeling result of the information to be labeled according to the graphic-text fusion feature and the graphic-text correlation degree information.

2. The multimodal information marking method according to claim 1, wherein, The multi-modal labeling system determines the visual feature information of the visual information, including: When the visual information is a live image, input each live image frame in the visual information into the first visual processing model in chronological order, and determine the first picture token mapping information of each live image frame through the first visual processing model; Perform stacking processing on the first picture token mapping information, and perform temporal context correlation fusion on the stacked first picture token mapping information through a pseudo-three-dimensional convolution model to obtain a feature map corresponding to each first picture token mapping information; Pool the feature map corresponding to each first picture token mapping information into a first feature map of a set dimension, and map the first feature map to visual feature information through a fully connected layer.

3. The multimodal information tagging method according to claim 1, wherein, The multi-modal labeling system determines the visual feature information of the visual information, including: When the visual information is a video image, input each video image frame in the visual information into multiple second visual processing models respectively, and determine the second picture token mapping information corresponding to each video image frame through the second visual processing model; Pool the second picture token mapping information into a picture semantic feature vector of a set dimension, and perform attention operation processing on the picture semantic feature vector through a multi-query attention mechanism network to obtain visual feature information.

4. The multimodal information marking method according to claim 1, wherein, The multi-modal labeling system performs multi-modal fusion processing based on the visual feature information, the visual information, the recognized text information, and the descriptive text information to obtain the graphic-text fusion feature, including: Obtain the query vector of each image frame in the visual information, input the visual feature information, the query vector, the recognized text information, and the descriptive text information into the trained visual-text fusion network, and perform multi-modal fusion processing based on the visual feature information, the query vector, the recognized text information, and the descriptive text information through the visual-text fusion network based on the cross-attention interaction mechanism to obtain the graphic-text fusion feature.

5. The multimodal information marking method according to claim 4, wherein, The obtaining of the query vector of each image frame in the visual information includes: In the case where the visual information is sequential images, input the visual information into the trained sequential visual query vector learner. Through the sequential visual query vector learner, use a multi-scale convolutional neural network to determine the multi-scale visual correlation features corresponding to each image frame in the visual information, and perform a feature extraction and fusion operation on the multi-scale visual correlation features to obtain the first query vector for each image frame in the visual information.

6. The multimodal information marking method according to claim 4, wherein, Obtaining the query vectors for each image frame in the visual information includes: In the case where the visual information is a single-image, input the visual information into the trained static visual query vector learner. Through the static visual query vector learner, use a dual-task neural network to analyze and process the visual information to obtain the second query vector for the image frame corresponding to the visual information.

7. The multimodal information marking method according to claim 4, wherein, The visual-text fusion network includes a visual decoder and a text decoder. The visual decoder includes a first set number of stacked first neural network modules. The first neural network module includes a first self-attention sub-unit for processing the query vector and a first cross-attention sub-unit for processing the visual feature information. The text decoder includes a first set number of stacked second neural network modules. The second neural network module includes a second self-attention sub-unit for processing the recognized text information and the descriptive text information. The visual decoder and the text decoder interact based on a cross-attention interaction mechanism.

8. The multimodal information tagging method according to claim 4, wherein, The training of the visual-text fusion network includes: Input a first sample image-text pair, determine the first sample visual embedding vector of the sample image and the first sample text embedding vector of the sample text in the first sample image-text pair, make the visual query vector and the text query vector in the visual-text fusion network mutually invisible, and update the network weights of the visual-text fusion network according to the distance information between the first sample visual embedding vector and the first sample text embedding vector; Input a second sample image-text pair, determine the second sample visual embedding vector of the sample image and the second sample text embedding vector of the sample text in the second sample image-text pair, make the visual query vector in the visual-text fusion network invisible to the text query vector, make the text query vector visible to the visual query vector and invisible to the text query vectors of the subsequent text, and update the network weights of the visual-text fusion network according to the distance information between the second sample visual embedding vector and the second sample text embedding vector; Input a third sample image-text pair, determine the third sample visual embedding vector of the sample image and the third sample text embedding vector of the sample text in the third sample image-text pair, make the visual query vector and the text query vector in the visual-text fusion network mutually visible, and update the network weights of the visual-text fusion network according to the distance information between the third sample visual embedding vector and the third sample text embedding vector; A fourth sample image-text pair is input, a fourth sample visual embedding vector of the sample image and a fourth sample text embedding vector of the sample text in the fourth sample image-text pair are determined, a set number of visual query vectors in the visual-text fusion network are randomly masked so that the text query vector in the visual-text fusion network is invisible to the masked visual query vector, and a network weight of the visual-text fusion network is updated according to the distance information between the fourth sample visual embedding vector and the fourth sample text embedding vector.

9. The multimodal information marking method according to claim 1, wherein, The multimodal marking system determines the graphic-text association information between the visual information and the identification text information and the description text information, including: Determining a visual feature vector corresponding to each image frame in the visual information; Determine text feature vectors corresponding to the identification text information and the description text information; The image-text similarity between each of the visual feature vectors and the text feature vector is determined, and image-text association information is determined according to the image-text similarity.

10. The multi-modal information tagging method according to claim 9, wherein, The determining of the image-text relevance information according to the image-text similarity includes: When the number of the visual feature vectors is an integer divider of the number of the text feature vectors, determining the image-text association information according to the proportion of the number of image-text similarities reaching a set similarity threshold in the total number of feature vectors of the visual feature vector and the text feature vector; When the number of the visual feature vectors does not divide the number of the text feature vectors, the image-text association information is determined according to the proportion of the number of image-text similarities reaching a set similarity threshold in the number of feature vectors of the visual feature vectors.

11. The multimodal information marking method according to claim 1, wherein, The multimodal marking system determines the marking result of the information to be marked according to the image-text fusion feature and the image-text relevance information, including: The image-text fusion feature and the image-text relevance information are input into a text processing model, and the image-text fusion feature and the image-text relevance information are analyzed and processed by the text processing model to obtain content description information and / or open set keywords of the information to be marked.

12. The multimodal information marking method according to claim 11, wherein, Before the multimodal marking system determines the marking result of the information to be marked according to the image-text fusion feature and the image-text association information, the multimodal marking system includes: The image-text fusion features are input into a trained alignment model, and the image-text fusion features are aligned by the alignment model so that the image-text fusion features after alignment meet the output preference of the text processing model.

13. The multimodal information tagging method according to claim 12, wherein, The alignment model performs alignment processing on the image-text fusion features, including: The previously obtained image-text fusion features are converted into a context feature group through a fully connected layer, and the routing value corresponding to the context feature group is determined through a gate selection network; Performing weighted processing on the routing value and the conversion vector obtained through training to obtain a weighted conversion vector; Through the cascaded decoder network and the multi-layer perceptron, the current image-text fusion feature and the weighted conversion vector are aligned and summed feature by feature.

14. The multimodal information marking method according to claim 1, wherein, The multimodal marking system determines the marking result of the information to be marked according to the image-text fusion feature and the image-text relevance information, including: Input the pre-configured indication text, the graphic-text fusion feature, and the graphic-text correlation information into a text processing model. Through the text processing model, analyze and process the graphic-text fusion feature and the graphic-text correlation information to obtain the marking result of the information to be marked. The indication text is used to activate the inference ability of the text processing model for the recognition task of the set label range.

15. The multimodal information marking method according to claim 1, wherein, The multimodal information marking method further includes: Input the graphic-text fusion feature and the graphic-text correlation information into a text processing model. Through the text processing model, analyze and process the graphic-text fusion feature and the graphic-text correlation information to obtain the context text feature of the information to be marked. Input the context text feature, and a combination of one or more of the speech information corresponding to the visual information, the recognition text information, and the description text information into the trained first label recognition model. Through the first label recognition model, analyze and process the context text feature, and a combination of one or more of the speech information corresponding to the visual information, the recognition text information, and the description text information to obtain the first label information.

16. The multimodal information marking method according to claim 1, wherein, The multimodal information marking method further includes: Input the graphic-text fusion feature and / or the visual information into the trained second label recognition model. Through the second label recognition model, analyze and process the graphic-text fusion feature and / or the visual information to obtain the second label information.

17. A multimodal information marking device, wherein, It includes an information acquisition module and a marking processing module, where: The information acquisition module is configured to acquire the information to be marked, and the information to be marked includes visual information, recognition text information, and description text information. The marking processing module is configured to input the information to be marked into the trained multimodal marking system. Through the multimodal marking system, determine the visual feature information of the visual information, perform multimodal fusion processing based on the visual feature information, the visual information, the recognition text information, and the description text information to obtain the graphic-text fusion feature, and determine the graphic-text correlation information between the visual information and the recognition text information and the description text information, and determine the marking result of the information to be marked according to the graphic-text fusion feature and the graphic-text correlation information.

18. A multimodal information marking device, wherein, It includes: A memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal information marking method according to any one of claims 1-16.

19. A non-volatile storage medium storing computer-executable instructions, wherein, The computer-executable instructions are used to execute the multimodal information marking method according to any one of claims 1-16 when executed by a computer processor.

20. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the multimodal information marking method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Video tag classification method and system and computer readable storage medium

    CN114758283A

  • Multi-modal false news identification method and system based on correlation information extension

    CN115100664A

  • Multi-modal rumor detection method and system

    CN115545039A

  • Social media false information detection method based on multi-modal entity fusion and alignment

    CN116452939A

  • Multi-modal information marking method and device, equipment, storage medium and product

    CN118038125A

Cited By

  • Abnormity detection method and system for traffic infrastructure, medium and equipment

    CN120635612A

  • Pancreatic operation area inflammation existence discrimination method fusing fat related image features

    CN120713554A

  • On-demand content recommendation method and device, storage medium and electronic equipment

    CN120751203A

  • Online anomaly detection method, device and equipment for traditional Chinese medicine pelleting machine and storage medium

    CN120850172A

  • Multimodal large model-based multimedia material label management method and system

    CN120873214A