Data enhancement method and device, computer equipment, storage medium and program product
By performing data fusion, feature fusion and tag information fusion on multimodal data, enhancing data is generated, and the problem of poor enhancement effect in the prior art is solved, and a more efficient enhancement effect in multimodal data is achieved.
Patent Information
- Application Number
- CN202510213285.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has poor results in multimodal data enhancement, especially the enhancement scheme for multimodal data, which mainly relies on image processing methods, resulting in low effects.
By obtaining the multimodal data and tag information to be enhanced, data fusion, feature fusion and tag information fusion are carried out to generate enhanced data. Specific steps include: data fusion, feature fusion and tag information fusion, and obtain enhanced data based on the fusion data, feature and tag information.
It effectively improves the effect of multimodal data enhancement, realizes the enhancement of multimodal data in the original data space and feature space, and improves model performance.
Smart Images

Figure CN120217280A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a data augmentation method, device, computer device, storage medium, and program product. Background Art
[0002] With the rise of artificial intelligence technology, data augmentation has become particularly important in improving model performance. For example, in the field of vision, data augmentation helps the model generalize better by increasing the diversity of training data.
[0003] In traditional data augmentation schemes, augmentation is usually performed on single-modal data. For example, for image data augmentation, image processing operations such as rotation, cropping, scaling, and color transformation are mainly carried out on the image data. For multi-modal data, the above image processing methods are mainly used for sampling to augment the image data in the multi-modal data, resulting in low effectiveness of multi-modal data augmentation. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a data augmentation method, device, computer device, storage medium, and program product that can effectively improve the effect of multi-modal data augmentation.
[0005] In a first aspect, this application provides a data augmentation method, which includes:
[0006] Obtain data to be augmented; the data to be augmented includes multi-modal data to be augmented and label information;
[0007] Perform data fusion processing on the data belonging to the same modality in the multi-modal data according to a preset ratio to obtain the fusion data of each modality;
[0008] Extract the features of each modality from the multi-modal data, and perform feature fusion processing on at least two features belonging to the same modality according to the preset ratio to obtain the fusion features of each modality;
[0009] Fuse the label information of each modality according to the preset ratio to obtain fused label information;
[0010] Based on the fusion data, the fusion features, and the fused label information, obtain augmented data.
[0011] In a second aspect, this application also provides a data augmentation device, which includes:
[0012] An acquisition module, configured to acquire data to be augmented; the data to be augmented includes multi-modal data to be augmented and label information;
[0013] The first fusion module is used to perform data fusion processing on the data belonging to the same modality in the multi-modal data according to a preset ratio to obtain the fusion data of each modality;
[0014] The second fusion module is used to extract the features of each modality from the multi-modal data, and perform feature fusion processing on at least two features belonging to the same modality according to the preset ratio to obtain the fusion features of each modality;
[0015] The third fusion module is used to fuse the label information of each modality according to the preset ratio to obtain the fused label information;
[0016] The enhancement module is used to obtain enhanced data based on the fusion data, the fusion features, and the fused label information.
[0017] In one embodiment, the first fusion module is further used to divide the data belonging to the same modality in the multi-modal data; perform data fusion processing on the data belonging to the same modality according to a preset ratio to obtain the fusion data of each modality; or, replace a preset size of partial data in the data belonging to the same modality with other data that matches the same modality and is of the preset size to obtain the fusion data of each modality.
[0018] In one embodiment, the multi-modal data includes visual data, and the visual data includes first visual data and second visual data;
[0019] The first fusion module is further used to perform data fusion processing on the first visual data and the second visual data according to a preset ratio to obtain visual fusion data;
[0020] The second fusion module is further used to extract first visual features and second visual features from the first visual data and the second visual data, and perform feature fusion processing on the first visual features and the second visual features according to the preset ratio to obtain visual fusion features.
[0021] In one embodiment, the preset ratio includes a first preset ratio and a second preset ratio;
[0022] The first fusion module is further used to crop a first visual data block from the first visual data according to the first preset ratio; crop a second visual data block from the second visual data according to the second preset ratio; splice the first visual data block and the second visual data block to obtain visual fusion data; or, perform weighted processing on the first visual data and the second visual data according to the first preset ratio and the second preset ratio to obtain visual fusion features.
[0023] In one embodiment, the first fusion module is further configured to determine a to-be-covered area with a target size in the first visual data according to a preset ratio; intercept a target tile with the target size from the second visual data, and respectively cover the target tile on the to-be-covered area to obtain visual fusion data; or, obtain a mask tile with the target size, and respectively cover the mask tile on the to-be-covered area of the first visual data and the to-be-covered area of the second visual data to obtain visual fusion data.
[0024] In one embodiment, the first visual data is first video data or first image data, and the second visual data is second video data or second image data;
[0025] The first fusion module is further configured to perform video frame fusion processing on the video frames in the first video data and the video frames in the second video data according to a preset ratio to obtain video fusion data; or, perform tile fusion processing on the first image data and the second image data according to the preset ratio to obtain image fusion data.
[0026] In one embodiment, the multimodal data includes text data, the text data includes first text data and second text data, and the preset ratio includes a first preset ratio and a second preset ratio;
[0027] The first fusion module is further configured to splice the first text data and the second text data according to a preset ratio to obtain text fusion data; or, truncate the first text data according to the first preset ratio to obtain a first text data segment; truncate the second text data according to the second preset ratio to obtain a second text data segment; and splice the first text data segment and the second text data segment to obtain text fusion data.
[0028] In one embodiment, the second fusion module is further configured to respectively extract a first text feature and a second text feature from the first text data and the second text data; and perform weighted summation on the first text feature and the second text feature according to the first preset ratio and the second preset ratio to obtain a text fusion feature.
[0029] In one embodiment, the multimodal data includes audio data, and the audio data includes first audio data and second audio data;
[0030] The first fusion module is further configured to splice the first audio data and the second audio data according to a preset ratio to obtain audio fusion data; or perform parallel superposition processing on the first audio data and the second audio data according to a preset ratio to obtain audio fusion data.
[0031] In one embodiment, the preset ratio is the ratio of a first ratio length to a second ratio length;
[0032] The second fusion module is further configured to respectively extract a first audio feature and a second audio feature from the first audio data and the second audio data; perform weighted summation on the first audio feature and the second audio feature according to the first ratio length and the second ratio length to obtain an audio fusion feature; or intercept a first sub-audio feature of the first ratio length and a second sub-audio feature of the second ratio length from the first audio feature and the second audio feature; perform splicing processing on the first sub-audio feature and the second sub-audio feature to obtain an audio fusion feature.
[0033] In one embodiment, the device further includes:
[0034] The third fusion module is further configured to obtain first enhanced data based on the fusion data and the fusion label information; obtain second enhanced data based on the fusion feature and the fusion label information;
[0035] The training module is configured to use the first enhanced data and the data to be enhanced as a training set, and train a multi-modal classification model based on the training set to obtain a trained multi-modal classification model; or classify the fusion feature in the second enhanced data through a classifier of the multi-modal classification model, determine a classification loss based on the obtained classification result and the label information; optimize the parameters of the multi-modal classification model according to the classification loss to obtain a trained multi-modal classification model; wherein, the fusion feature in the second enhanced data is obtained by the feature extraction network of the multi-modal classification model through extraction and feature fusion processing.
[0036] In one embodiment, the training module is further configured to an adjustment module, which is configured to dynamically adjust the preset ratio during the process of training the multi-modal classification model, so as to perform data fusion processing and feature fusion processing using the adjusted preset ratio to obtain dynamically adjusted enhanced data for model training; wherein, there is a difference between the dynamically adjusted enhanced data and the enhanced data before dynamic adjustment; or adjust the sample ratio, group the training set according to the adjusted sample ratio to obtain at least two training subsets; respectively train the multi-modal classification model based on each of the training subsets.
[0037] In a third aspect, the present application further provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the data enhancement method are implemented.
[0038] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data enhancement method are implemented.
[0039] In a fifth aspect, the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the data enhancement method are implemented.
[0040] For each modality data in the multi-modal data, the above data enhancement method, device, computer device, storage medium and program product can all perform data fusion processing on the data belonging to the same modality in the multi-modal data according to a preset ratio to obtain the fusion data of each modality, so as to enhance the multi-modal data in the original data space, which is beneficial to improving the effect of multi-modal data enhancement; in addition, the features of each modality can be extracted from the multi-modal data, and at least two features belonging to the same modality are subjected to feature fusion processing according to a preset ratio to obtain the fusion features of each modality, so as to enhance the multi-modal data in the feature space and further improve the effect of multi-modal data enhancement; finally, the label information of each modality is fused according to a preset ratio, and enhanced data is obtained based on the obtained fusion label information, fusion data and fusion features, so as to enhance the multi-modal data in the original data space and feature space, effectively improving the effect of multi-modal data enhancement; moreover, by enhancing the data of each modality in the multi-modal data, when performing model training, the model performance can be effectively improved without increasing training resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is an application environment diagram of the data enhancement method in an embodiment;
[0042] Figure 2 It is a schematic flowchart of the data enhancement method in an embodiment;
[0043] Figure 3 It is a schematic flowchart of the data enhancement method in another embodiment;
[0044] Figure 4 It is a schematic diagram of video data fusion to obtain video fusion data in an embodiment;
[0045] Figure 5It is a schematic diagram of obtaining video fusion data by fusing video data by splicing or linear interpolation in one embodiment;
[0046] Figure 6 is a schematic diagram of three image data fusion methods in one embodiment;
[0047] Figure 7 It is a schematic diagram of obtaining a video fusion feature by fusing video features of video data by splicing or linear interpolation in one embodiment;
[0048] Figure 8 A schematic diagram of obtaining text fusion data by fusing text data in one embodiment;
[0049] Figure 9 It is a schematic diagram of obtaining text fusion features by fusing text features of text data by splicing or linear interpolation in one embodiment;
[0050] Figure 10 A schematic diagram of obtaining audio fusion data by fusing audio data in one embodiment;
[0051] Figure 11 It is a schematic diagram of obtaining an audio fusion feature by fusing audio features of audio data by splicing or linear interpolation in one embodiment;
[0052] Figure 12 is a schematic diagram of the structure of a multimodal classification model in one embodiment;
[0053] Figure 13 A schematic diagram of performing data enhancement on two samples to obtain enhanced data in one embodiment;
[0054] Figure 14 is a schematic diagram of a beta probability density function in one embodiment;
[0055] Figure 15 It is a schematic diagram of data enhancement of two samples in feature space in one embodiment;
[0056] Figure 16 A schematic diagram of data enhancement, model training, and content distribution in one embodiment;
[0057] Figure 17 A schematic diagram of prompting a wonderful clip during video playback in one embodiment;
[0058] Figure 18 is a structural block diagram of a data enhancement device in one embodiment;
[0059] Figure 19 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0060] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0061] It should be noted that in the following description, the terms "first", "second" and "third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second" and "third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0062] The data augmentation method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other network servers.
[0063] The server 104 can obtain the data to be augmented, perform augmentation processing on the multi-modal data in the data to be augmented in the original data space and the feature space, so as to obtain the fusion data and fusion features of each modality; in addition, it can also obtain fusion label information, so as to obtain augmented data based on the fusion data, fusion features and fusion label information, and use it in model training, which can effectively increase data diversity and model performance; therefore, when content recommendation is required, the trained multi-modal classification model can be used for classification, so as to share the content that the user may be interested in with the terminal 102 of the user. It should be noted that the terminal 102 can also implement the above multi-modal data augmentation solution.
[0064] Combined with Figure 2 it can be seen that for the multi-modal data augmentation in the original data space, mainly the data belonging to the same modality in the multi-modal data are subjected to data fusion processing according to a preset ratio to obtain the fusion data of each modality; for the multi-modal data augmentation in the feature space, mainly the features of each modality are extracted from the multi-modal data, and at least two features belonging to the same modality are subjected to feature fusion processing according to a preset ratio to obtain the fusion features of each modality.
[0065] Among them, the terminal 102 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, an Internet of Things device and a portable wearable device, and the Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner and a smart vehicle-mounted device, etc.
[0066] Server 104 can be an independent physical server or a service node in a blockchain system. A peer-to-peer network is formed among the service nodes in the blockchain system. The peer-to-peer protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In addition, server 104 can also be a server cluster composed of multiple physical servers, and can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), as well as big data and artificial intelligence platforms.
[0067] In one embodiment, as Figure 3 shown, a data enhancement method is provided. This method can be executed by the Figure 1 server or terminal in Figure 1 or jointly executed by the server and the terminal. Taking the execution of this method by the
[0068] server in
[0069] as an example, the method includes the following steps:
[0070] S302, obtain the data to be enhanced; the data to be enhanced includes the multi-modal data to be enhanced and label information.
[0071] Among them, the multi-modal data can be data of multiple modalities, including at least two of visual modality data (abbreviated as visual data), audio modality data (abbreviated as audio data), or text modality data (abbreviated as text data). The visual modality can be a video modality or an image modality, so the visual data can be video data or image data.
[0070] In addition, the multi-modal data can be multi-modal data in different application scenarios. For example, it can include: multi-modal data on a film and television platform (such as data of TV dramas + text + audio), multi-modal data published on a short video platform through a video application (such as data of video + text), multi-modal data for social interaction published on an interactive platform through a social application, and multi-modal data of user life sharing published on a sharing platform through a sharing application.
[0071] It should be noted that the number of the above-mentioned multimodal data can be n. For example, there are two multimodal data, namely multimodal data a and multimodal data b. In addition, the visual data, text data, and audio data in each multimodal data can be at least 1. Thus, the number of the visual data, text data, and audio data in the multimodal data can be at least n respectively, where n is a positive integer greater than or equal to 2. For example, visual data a and visual data b. Thus, the visual data a and visual data b can be processed for data fusion according to a preset ratio to achieve data enhancement.
[0072] The label information can be information used to represent the category to which the multimodal data belongs, such as the category information of the multimodal data, or the multimodal data containing annotation information obtained by annotating the multimodal data. In addition, the label information can also be used to represent whether there is key data in the multimodal data and the position corresponding to the key data. For example, in the field of film and television, the label information can be whether there is a popular segment in a certain film and television video (such as a certain episode of a TV drama) and the corresponding position. The popular segment can be determined according to the interaction data. For example, if the comments (such as bullet screen information) of a certain video segment exceed a preset number, then this video segment can be determined as a popular segment.
[0073] In one embodiment, the server can directly obtain the data to be enhanced including multimodal data and label information from the database; or, the server can collect multimodal data without label information, classify or annotate the multimodal data to obtain result information, and obtain the label information of the multimodal data based on the result information.
[0074] Among them, when performing classification or annotation, a classification model can be used, such as the CLIP (Contrastive Language-Image Pre-Training) model. The CLIP model is a multimodal machine learning model that can understand the relationship between visual data (such as image data or video data) and speech information (such as text data and audio data).
[0075] For example, obtain multimodal data, which includes text data (such as the title information of video data), video data, and audio data, as described in detail below:
[0076] Title information: "The Story of XX", introducing the growth process of the protagonist, getting better every day;
[0077] Video data: <video1.mp4>;
[0078] Audio data: <audio1.wav>;
[0079] Input the above-mentioned title data, video data, and audio data into a classification model. The classification model uses these input data for classification processing to obtain a classification result, and uses this classification result as the label information of the multimodal data.
[0080] It should be noted that for the multimodal data in the above-mentioned film and television field, the text data can be not only title information, but also lines or interactive data. Among them, the above-mentioned lines can be read from the corresponding line file; the interactive data can be the comment information of the user on the video data or audio data during or after watching or listening, such as the comment information published in the form of bullet screens during the process of watching the video data (i.e., bullet screen information), or the comment information published in the comment area of the playback page.
[0081] In one embodiment, when the multimodal data is video data, audio data, and text data (including title information and bullet screen information) in a film and television scene, the classification or annotation of the multimodal data can be implemented in the following way: The terminal can first determine the playback time of the bullet screen information, determine the number of bullet screen information belonging to each video segment according to the playback time, and select the video segments with the number of bullet screen texts exceeding the preset number from these video segments as popular segments, so as to mark and obtain the corresponding label information.
[0082] S304, perform data fusion processing on the data belonging to the same modality in the multimodal data according to a preset ratio to obtain the fusion data of each modality.
[0083] Among them, the preset ratio can be used to determine the data size (or data length) required for different data during fusion. For example, if the preset ratio is 0.5:0.5, it means that half of the data a and data b of a certain modality are intercepted for fusion to obtain the fusion data of this modality.
[0084] The data belonging to the same modality can refer to at least two different data belonging to the same modality, such as visual data a and visual data b that are both in the visual modality.
[0085] The above data fusion processing method can be concatenation (concat) or linear interpolation (interpolation) in data dimensions. For example, at least two data are directly concatenated one by one, or a part is intercepted from at least two data according to a preset ratio for concatenation, or a mask data of a target size (such as a mask image of a target size) is intercepted according to a preset ratio for data covering, or at least two data are weighted and summed according to a preset ratio (such as weights).
[0086] In one embodiment, when the multimodal data includes data of at least two modalities among visual data, text data, or audio data, the server may perform data fusion of at least two modalities according to a preset ratio. Among them, the process of data fusion may refer to the following method:
[0087] 1) Fusion of visual data: The server fuses at least every two visual data in the multimodal data according to a preset ratio to obtain fused data of the visual modality (referred to as visual fused data). For example, half of the first video frame from video data a and half of the second video frame from video data b are cropped and spliced to obtain video fused data;
[0088] In one embodiment, before fusing the visual data, the server may also perform at least one visual data enhancement process on each visual data, such as rotation, cropping, scaling, color transformation, or noise addition.
[0089] 2) Fusion of text data: The server fuses at least every two text data in the multimodal data according to a preset ratio to obtain fused data of the text modality (referred to as text fused data). For example, half of the text data a and half of the text data b are intercepted and spliced to obtain text fused data;
[0090] 3) Fusion of audio data: The server fuses at least every two audio data in the multimodal data according to a preset ratio to obtain fused data of the audio modality (referred to as audio fused data). For example, half of the audio data a and half of the audio data b are intercepted and spliced, or the audio data a and the audio data b are processed in parallel and superimposed in a 1:1 ratio to obtain audio fused data.
[0091] In one embodiment, the server may respectively use the fused data of each modality of at least every two multimodal data as a combination to obtain corresponding multimodal fused data.
[0092] For example, assuming that data fusion is performed pairwise among multimodal data a, multimodal data b, and multimodal data c, the data of each modality included in multimodal data a, multimodal data b, and multimodal data c and the fusion results are shown in Table 1, and thus the multimodal fused data shown in Table 1 can be obtained.
[0093] Table 1
[0094]
[0095] S306, extract the features of each modality from the multimodal data, and perform feature fusion processing on at least two features belonging to the same modality according to a preset ratio to obtain the fused features of each modality.
[0096] Among them, the features of each modality can be at least two of the visual features of the visual modality, the audio features of the audio modality, or the text features of the text modality.
[0097] The above feature fusion processing methods can be concatenation or linear interpolation in the feature dimension. For example, at least two features are directly concatenated one by one, or a part is intercepted from at least two features respectively according to a preset ratio for concatenation, or a mask feature of a target size is intercepted according to a preset ratio for feature coverage, or at least two features are weighted and summed according to a preset ratio (such as weight).
[0098] In one embodiment, when the multi-modal data includes at least two types of data among visual data, text data, or audio data, the extracted features can include at least two of visual features, text features, or audio features. At this time, the server can perform feature fusion of at least two modalities according to a preset ratio. Among them, the process of feature fusion can refer to the following methods:
[0099] 1) Fusion of visual features: The server fuses the visual features of at least every two visual data according to a preset ratio to obtain visual fusion features. For example, after extracting video feature a and video feature b from the first video frame in video data a and the second video frame in video data b respectively, perform linear interpolation or concatenation on video feature a and video feature b, such as weighted summation using two preset interpolation coefficients, or select partial video features according to two preset ratio values, and then perform vector concatenation processing to obtain video fusion features;
[0100] 2) Fusion of text features: The server fuses the text features of at least every two text data according to a preset ratio to obtain text fusion features. For example, after extracting text feature a and text feature b from text data a and text data b, perform concatenation processing on text feature a and text feature b, select partial text features according to two preset ratio values, and then perform vector concatenation processing to obtain text fusion features;
[0101] 3) Fusion of audio features: The server fuses the audio features of at least every two audio data according to a preset ratio to obtain audio fusion features. For example, after extracting audio feature a and audio feature b from audio data a and audio data b respectively, perform linear interpolation or concatenation on audio feature a and audio feature b, such as weighted summation using two preset interpolation coefficients, or extract partial audio features from audio feature a and audio feature b respectively according to two preset ratio lengths, and then perform feature concatenation to obtain audio fusion features.
[0102] For example, for the splicing and fusion of audio features, let's assume that two audio features with a feature dimension of [32, 1024] are extracted from audio data a and audio data b. Take p and 1-p proportion lengths of these two audio features respectively. Assume p is 0.5, that is, audio features with a feature dimension of [16, 1024] are taken from both audio features. Then, the two extracted audio features are spliced, and the resulting audio fusion feature is still [32, 1024]. In this way, the fusion of the two audio features is realized. Therefore, this audio fusion feature contains p proportion of the features in audio feature a and 1-p proportion of the features in audio feature b.
[0103] In one embodiment, the server can take the fusion features of each modality of each multimodal data as a combination respectively, so as to obtain multimodal fusion features.
[0104] For example, assume that data fusion is performed pairwise among multimodal data a, multimodal data b, and multimodal data c. The features of each modality corresponding to multimodal data a, multimodal data b, and multimodal data c respectively, as well as the fusion results, are shown in Table 2. Thus, the multimodal fusion features shown in Table 2 can be obtained.
[0105] Table 2
[0106]
[0107] S308, fuse the label information of each modality according to a preset ratio to obtain fused label information.
[0108] Among them, the label information of multimodal data can be that one kind of label information corresponds to the data of each modality. In addition, it can also be that one kind of label information corresponds to the data of each modality respectively.
[0109] For example, for a film and television scene, the label information of this multimodal data can be label information of the TV drama category, such as at least one label information of the TV drama category like TV drama, emotional category, or powerful actor, etc.
[0110] In one embodiment, the server fuses the label information corresponding to at least every two data of each modality according to a preset ratio to obtain fused label information.
[0111] For example, assume that video data a and video data b have undergone data fusion processing according to a preset ratio. For instance, the video frames of video data a and video frames of video data b are fused according to the ratio of p:(1 - p). Then, continue to fuse the label information of video data a and the label information of video data b according to this preset ratio. For example, fuse the label information of video data a and the label information of video data b according to the ratio of p:(1 - p). Thus, the resulting fused label information is (TV drama p, technology 1 - p), where p in the fused label information represents the weight of the label being a TV drama is p, and 1 - p represents the weight of the label being technology is 1 - p.
[0112] S310. Obtain enhanced data based on the fused data, fused features, and fused label information.
[0113] Among them, considering that this application performs data enhancement in two different spaces, that is, data enhancement is performed in the original data space and the feature space. Therefore, enhanced data in two spatial dimensions can be obtained, namely the first enhanced data in the original data space and the second enhanced data in the feature space.
[0114] In one embodiment, the server can use the fused data and fused label information as the first enhanced data in the original data space, and use the fused features and fused label information as the second enhanced data in the feature space.
[0115] In the above embodiment, for each modality data in the multi-modal data, data fusion processing can be performed on the data belonging to the same modality in the multi-modal data according to a preset ratio to obtain the fused data of each modality. Thus, multi-modal data enhancement can be achieved in the original data space, which is beneficial to improving the effect of multi-modal data enhancement. In addition, features of each modality can be extracted from the multi-modal data, and at least two features belonging to the same modality are fused according to a preset ratio to obtain the fused features of each modality. Thus, multi-modal data enhancement can be achieved in the feature space, further improving the effect of multi-modal data enhancement. Finally, fuse the label information of each modality according to a preset ratio, and obtain enhanced data based on the obtained fused label information, fused data, and fused features. Thus, multi-modal data enhancement is achieved in the original data space and the feature space, effectively improving the effect of multi-modal data enhancement. Moreover, data enhancement is performed on each modality data in the multi-modal data, and during model training, without increasing training resources, the model performance can be effectively improved.
[0116] In one embodiment, the server may partition the multimodal data to identify data belonging to the same modality; perform data fusion processing on the data belonging to the same modality according to a preset ratio to obtain fused data for each modality; or replace a preset-sized portion of the data belonging to the same modality with other data that matches the same modality and is of a preset size to obtain fused data for each modality.
[0117] Among them, the above replacement may be that the original partial data is deleted and supplemented with data of the same size; or, it may also be that the same-sized data overwrites the original partial data, such as a certain rectangular area in an image being covered by a mask image (i.e., a black mask).
[0118] The preset size may be a preset data size, target size, and proportional length, which can be calculated based on the preset ratio and the data size of the corresponding modality of the multimodal data.
[0119] Next, the processes of fusing visual data, text data, and audio data in the multimodal data, and the processes of fusing visual features, text features, and audio features will be described separately as follows:
[0120] (1) Fusion of visual data and visual features.
[0121] A. Fusion of visual data:
[0122] In one embodiment, the multimodal data includes visual data, and the visual data includes first visual data and second visual data; thus, the fusion of visual data includes: the server performs data fusion processing on the first visual data and the second visual data according to a preset ratio to obtain visual fusion data.
[0123] Among them, the first visual data and the second visual data may be two different data of the visual modality, such as visual data a and visual data b.
[0124] Considering that visual data may include video data and image data, the first visual data may be first video data, the second visual data may be second video data, or the first visual data may be first image data, and the second visual data may be second image data. In particular, the first visual data may be a video frame in the video data (such as a video frame sampled at regular intervals), and the second visual data may be image data.
[0125] In one embodiment, the first visual data is the first video data or the first image data, and the second visual data is the second video data or the second image data; thus, the fusion of visual data may include: for data fusion in the video modality, the server may perform video frame fusion processing on the video frames in the first video data and the video frames in the second video data according to a preset ratio to obtain video fusion data. In addition, for data fusion in the image modality, the server may directly perform tile fusion processing on the first image data and the second image data according to a preset ratio to obtain image fusion data.
[0126] For example, for data fusion in the video modality, the video frame sequences a and b can be decoded from video data a and video data b respectively; corresponding video frames a (denoted as cv1) and video frames b (denoted as cv2) are extracted from the video frame sequences a and b according to a preset ratio, and then the extracted video frames a and b are combined to obtain video fusion data concat(cv1, cv2), as Figure 4 and Figure 5 shown.
[0127] In addition, after obtaining the video frame sequences a and b, corresponding video frames can also be extracted from the video frame sequences a and b to obtain a target number of video frames a and a target number of video frames b, and then linear interpolation is performed on the video frames a and b according to a preset ratio to obtain video fusion data interpolation(cv1, cv2), as Figure 5 shown; alternatively, image block stitching can also be performed to obtain video fusion data. Similarly, image data can also be fused by referring to this method of linear interpolation or image block stitching.
[0128] In one embodiment, the server may perform data fusion processing by means of data stitching. Specifically: the server crops a first visual data block from the first visual data according to a first preset ratio; crops a second visual data block from the second visual data according to a second preset ratio; and performs stitching processing on the first visual data block and the second visual data block to obtain visual fusion data.
[0129] For example, the preset ratio is 0.5:0.5, and the sizes of the two image data are 1024×1024. At this time, the first half can be cropped from image data a, the second half can be cropped from image data b, and then these two parts are stitched together to obtain image fusion data.
[0130] In another embodiment, the server may perform data fusion processing in a linear interpolation manner. Specifically, the server may perform weighted processing on the first visual data and the second visual data according to a first preset ratio and a second preset ratio to obtain visual fusion features.
[0131] For example, if the preset ratio is 0.5:0.5, then when performing data fusion, the two image data can be linearly interpolated according to the ratio of 0.5:0.5. As Figure 5 shown, by performing weighted summation on the two image data according to the ratio of 0.5:0.5, the image fusion data of the two image data can be obtained. As Figure 6 shown, for a dog image, a cat image can be used to perform linear interpolation in the original data space with the dog image, so as to obtain image data that combines the dog and the cat. As Figure 6 shown in Figure (b).
[0132] In another embodiment, the server may perform data fusion processing using an image block covering method. Specifically, the server determines a to-be-covered area with a target size in the first visual data according to a preset ratio; intercepts a target tile with the target size from the second visual data, and respectively covers the target tile on the to-be-covered area to obtain visual fusion data.
[0133] Among them, the target size can be calculated according to a preset ratio and the size of the visual data (such as a certain video frame or image data). For example, if the preset ratio is 0.6:0.4 and the size of the image data is 1024×1024, then the target size can be 410×410 = (1024×0.4)×(1024×0.4).
[0134] For example, if the preset ratio is 0.6:0.4 and the size of the image data is 1024×1024, it means that image data a requires 60% of the image data size, and image data b requires 40% of the data size. At this time, a rectangular area of 410×410 can be extracted from image data b to obtain an image data block with 40% of the original image size, and it is covered on image data a to obtain image fusion data. As Figure 6 shown, for a dog image, a rectangular area of 410×410 in a cat image can be covered on the dog image to obtain image data that combines the dog and the cat. As Figure 6 shown in Figure (d).
[0135] In another embodiment, the server can perform data fusion processing by covering with a mask image (i.e., a black mask). Specifically: In the first visual data, the server determines a to-be-covered area of a target size according to a preset ratio; obtains a mask tile of the target size, and covers the mask tile on the to-be-covered area of the first visual data and the to-be-covered area of the second visual data respectively to obtain visual fusion data.
[0136] For example, if the preset ratio is 0.6:0.4 and the size of the image data is 1024×1024, then 60% of the area in the image data will be retained, and the other 40% of the area will be replaced by the mask image. That is, a rectangular area of 410×410 in the image data is covered by a mask image of the same size. As Figure 6 shown, for a dog image, a mask image of 410×410 can be covered on the dog image to obtain image data that combines the dog and noise, as shown in Figure 6 Figure (c).
[0137] B. Fusion of visual features:
[0138] In one embodiment, the server can extract first visual features and second visual features from the first visual data and the second visual data, and perform feature fusion processing on the first visual features and the second visual features according to a preset ratio to obtain visual fusion features.
[0139] Among them, the above feature fusion processing can be feature splicing or linear interpolation in the feature space.
[0140] The first visual data can be the first video data or the first image data, and the second visual data can be the second video data or the second image data.
[0141] In one embodiment, the server can first decode the first video data and the second video data to obtain a first video frame sequence and a second video frame sequence, and then extract first video features and second video features from each video frame (or partially extracted video frames) of the first video frame sequence and each video frame (or partially extracted video frames) of the second video frame sequence; then, the server can perform fusion processing in the way of feature splicing, that is: from each of the first video features and each of the second video features, video features of corresponding dimensions are extracted according to a preset ratio, and the extracted video features are spliced respectively to obtain video fusion features.
[0142] For example, video frames a and b pass through a visual feature extractor to obtain video features (CV embedding), with a feature dimension of [N, 1024]. Then, p and 1-p ratios are taken for the video features of video frame a and video frame b respectively, that is, video features with feature dimensions of [p*N, 1024] and [(1-p)*N, 1024] respectively. Then, they are concatenated to obtain a video fusion feature, and the feature dimension of this video fusion feature is still [N, 1024]. For reference, see Figure 7 , it should be noted that Figure 7 the ratios in
[0143] are not shown.
[0144] In addition, the server can perform fusion processing in a linear interpolation manner, that is: weighted sum the first video features and the second video features according to a preset ratio to obtain a video fusion feature. Figure 7 For example, video frames a and b pass through a visual feature extractor to obtain video features (CV embedding, abbreviated as CV emb), with a feature dimension of [N, 1024]. Then, linear interpolation is performed on the video features of video frame a (i.e., CV emb1) and the video features of video frame b (i.e., CV emb2). For reference, see
[0145] (2) Fusion of text data and text features.
[0146] A. Fusion of text data:
[0147] In one embodiment, the multimodal data includes text data, the text data includes first text data and second text data, and the preset ratio includes a first preset ratio and a second preset ratio; therefore, the server can concatenate the first text data and the second text data according to the preset ratio to obtain text fusion data; or, truncate the first text data according to the first preset ratio to obtain a first text data segment; truncate the second text data according to the second preset ratio to obtain a second text data segment; concatenate the first text data segment and the second text data segment to obtain text fusion data. For reference, see Figure 8 .
[0148] For example, the text concatenation of two text data (i.e., text data a and text data b) is directly concatenated at a ratio of 1:1 to obtain text fusion data. For example, the title information of the above-mentioned "The Story of XX" is concatenated with the title information of a certain fruit at a ratio of 1:1 to obtain new title information, namely: "The Story of XX", which introduces the growth process of the protagonist, and every day is better than yesterday; Big news! A certain fruit announced a cooperation with OpenAI, and its mobile phones, computers and other systems were fully updated! Smart glasses will also be on sale, with a starting price of nearly 30,000.
[0149] In addition, a proportional truncation and re-joining method may also be adopted, that is, the text data a is truncated at a ratio of p, such as truncating the title information of the above-mentioned "The Story of XX" from the middle to obtain the first half of the title information, and the text data b is truncated at a ratio of 1-p, such as truncating the title information of a certain fruit from the middle to obtain the second half of the truncation, and then splicing is performed, such as Figure 8 As shown in the figure, direct concatenation can ensure the integrity of each sample text, while the proportional truncation method is more aligned with the fusion methods of other modalities.
[0150] B. Fusion of text features:
[0151] In one embodiment, the server can extract the first text feature and the second text feature from the first text data and the second text data respectively; perform weighted summation on the first text feature and the second text feature according to the first preset ratio and the second preset ratio to obtain a text fusion feature, which can be referred to as Figure 9 .
[0152] For example, linear interpolation is performed on the feature dimension (feature dimension + interpolation), such as using a feature extractor to extract text features from text data a and text data b, and then linear interpolation is performed on the extracted text features, such as using a preset ratio to perform weighted summation on the text features of text data a and the text features of text data b; or, text features of corresponding dimensional size are extracted from the text features of text data a according to the first-meaning preset ratio, and text features of corresponding dimensional size are extracted from the text features of text data b according to the second-meaning preset ratio, and then the extracted text features are spliced to obtain text fusion features with the same feature dimension as the original text features.
[0153] (3) Fusion of audio data and audio features.
[0154] A. Fusion of audio data:
[0155] In one embodiment, the multimodal data includes audio data, and the audio data includes first audio data and second audio data; thus, the server can splice the first audio data and the second audio data according to a preset ratio to obtain audio fusion data; alternatively, the first audio data and the second audio data are subjected to parallel superposition processing according to a preset ratio to obtain audio fusion data, for reference Figure 10 .
[0156] For example, the audio data a and the audio data b are serially spliced together in a 1:1 manner, or the audio data a and the audio data b are superposed in parallel in a 1:1 manner. Tools such as python or FFmpeg are used to implement the superposition of audio data, such as Figure 10 shown.
[0157] B. Fusion of audio features:
[0158] In one embodiment, the preset ratio is the ratio of the first ratio length to the second ratio length; thus, the server can respectively extract the first audio feature and the second audio feature from the first audio data and the second audio data; perform weighted summation on the first audio feature and the second audio feature according to the first ratio length and the second ratio length to obtain an audio fusion feature; alternatively, intercept the first sub-audio feature of the first ratio length and the second sub-audio feature of the second ratio length from the first audio feature and the second audio feature; perform splicing processing on the first sub-audio feature and the second sub-audio feature to obtain an audio fusion feature.
[0159] For example, an audio feature extractor is used to extract audio feature 1 (audioembedding1) and audio feature 2 (audio embedding2) from the audio data a and the audio data b, and the feature dimension is [32, 1024]. Then, the p and 1-p ratio lengths are respectively taken for the audio features of the two audio data. Assuming p is 0.5, that is, [16, 1024]. Then, the audio features extracted from the two audio data are spliced, and the feature dimension of the obtained audio fusion feature is still [32, 1024]. In this way, the audio feature fusion of the two audio data is realized, and the audio fusion feature contains the audio feature of the audio data a (the proportion is p) and the audio feature of the audio data b (the proportion is 1-p).
[0160] In one embodiment, considering that enhancement is performed in the original data space and the feature space, after obtaining the fusion data, fusion features, and fusion label information of the multimodal data, the first enhanced data in the original data space and the second enhanced data in the feature space can be obtained.
[0161] Specifically, for obtaining the first augmented data in the original data space and the second augmented data in the feature space, the following methods can be referred to: The server can obtain the first augmented data based on the fused data and the fused label information; and obtain the second augmented data based on the fused features and the fused label information.
[0162] After obtaining the first augmented data in the original data space and the second augmented data in the feature space, they can be applied to the model training process, and the training process is as follows: The server can use the first augmented data and the data to be augmented as the training set, and train the multimodal classification model based on the training set to obtain the trained multimodal classification model.
[0163] Alternatively, the server can also classify the fused features in the second augmented data through the classifier of the multimodal classification model, determine the classification loss based on the obtained classification results and the label information; optimize the parameters of the multimodal classification model according to the classification loss to obtain the trained multimodal classification model; wherein, the fused features in the second augmented data are obtained by the feature extraction network of the multimodal classification model through extraction and feature fusion processing.
[0164] Among them, after the multimodal classification model is trained, it can be applied to application scenarios of content distribution, such as video recommendation scenarios and video content operation scenarios. For example, when a user is swiping videos, the multimodal classification model can use the multimodal data of candidate videos and the user's relevant information set for video recommendation.
[0165] In one embodiment, the server can also dynamically adjust the preset ratio during the process of training the multimodal classification model, so as to perform data fusion processing and feature fusion processing using the adjusted preset ratio to obtain the dynamically adjusted augmented data for model training; wherein, there are differences between the dynamically adjusted augmented data and the augmented data before dynamic adjustment; or, adjust the sample ratio, group the training set according to the adjusted sample ratio to obtain at least two training subsets; and train the multimodal classification model based on each training subset respectively.
[0166] For example, the p during each sample augmentation can be dynamically adjusted, such as adjusting the size of p using the beta function. It is also possible to adjust the ratio of fused samples in each batch, without the need to read twice the amount of data for fusion, so it will not increase the burden of data reading and will not increase the consumption of training resources.
[0167] As an example, combined with Figure 12 A simple description of the solution of this application is as follows:
[0168] Such as Figure 12As shown, the multi-modal classification model has a main network Blender, which is a network model similar to Bert and adopts the Transformer Encoder structure. Its input features include text data, video data, and audio data in these three modalities.
[0169] For text features (text embedding, abbreviated as text emb), text information such as title information, information from automatic speech recognition (ASR), and information from optical character recognition (OCR) can be used. Then, the text information is passed through a Tokenizer to obtain a series of tokens, and then converted into text emb.
[0170] For video data, first extract video frames from the video data, and then use a visual feature extractor (such as ViT, Swin Transformer, etc.) to extract CV embedding (abbreviated as CV emb). Specifically, extract n frames from the video data, treat each frame as an image and send it into the visual feature extractor, and then take the cls token embedding of each frame as the video feature of that frame (i.e., CV emb). Then, a total of n frames of CV emb are obtained for this video data.
[0171] For audio data, use an audio feature extractor (such as the Vggish model) to extract audio features from the audio data to obtain an audio embedding (abbreviated as audio emb) with a length of 32. Concatenate text emb, CV emb, and audio emb together, and at the same time add a cls token, and send them into the main network Blender to finally obtain the classification result.
[0172] Among them, the Tokenizer is an important component in natural language processing (NLP), mainly used to split text data into a series of tokens.
[0173] As another example, the following two multi-modal data (referred to as sample 1 and sample 2 for convenience) are combined for illustration, as follows:
[0174] (1) Use the following two different samples:
[0175] Sample 1:
[0176] Title information (text): "The Story of XX", which introduces the growth process of the protagonist, getting better every day;
[0177] Video data: <video1.mp4>;
[0178] Audio data: <audio1.wav>;
[0179] Tag information: TV series;
[0180] Sample 2:
[0181] Title information: Big news! Apple announced a partnership with OpenAI, and its mobile phones, computers and other systems have been fully updated! Smart glasses will also be on sale, with a starting price of nearly 30,000 yuan;
[0182] Video data:<video2.mp4> ;
[0183] Audio Data:<audio2.wav> ;
[0184] Tag information: technology;
[0185] (2) For the text part (i.e., title information), there are multiple fusion methods and fusion dimensions, as follows:
[0186] a) Concatenate the text of the two samples (original dimension + concat). You can directly concatenate the text of the two samples and turn it into:
[0187] The merged title information is: "The Story of XX", which introduces the protagonist's growth process, and every day is better than yesterday; Big news! Apple announced a partnership with OpenAI, and its mobile phones and computers and other systems have been fully updated! Smart glasses will also be on sale, with a starting price of nearly 30,000 yuan.
[0188] In addition, a proportional truncation and splicing method can also be used, such as truncating the title information of sample 1 according to the ratio p, and truncating the title information of sample 2 according to the ratio 1-p, and then splicing. Among them, direct splicing can ensure the integrity of each sample text, while the proportional truncation method is more aligned with the fusion method of other modalities.
[0189] b) Linear interpolation on the feature dimension (feature dimension + interpolation), that is, after extracting text features (text embedding) from the texts of the two samples, linear interpolation is performed, such as Figure 9 shown.
[0190] (3) For the audio part (i.e., audio data), there can also be multiple fusion methods and fusion dimensions, as follows:
[0191] a) concatenating the audio data of the two samples (original dimension + concat), that is, concatenating the audio data of the two samples serially together to obtain a new audio data (ie, fused audio data); in addition, the obtained fused audio data can be input into an audio feature extractor (Vggish).
[0192] b) Overlay the audio data of the two samples (original dimension + interpolation), that is, overlay the two audio data in parallel. Tools such as Python or FFmpeg can be used to achieve the overlay of audio data.
[0193] c) Concatenation in the feature dimension (feature dimension + concat), that is, each sample uses an audio feature extractor to obtain audio features with a feature dimension of [32, 1024]. Then, take p and 1 - p proportion lengths of the audio features of the two samples respectively. Assuming p is 0.5, that is, [16, 1024]. Then concatenate the audio features of the two samples, and the feature dimension of the resulting audio fusion features is still [32, 1024]. In this way, the audio feature fusion of the two samples is achieved, and the audio fusion features contain the audio features of sample 1 (p proportion) and the audio features of sample 2 (1 - p proportion).
[0194] d) Linear interpolation in the feature dimension (feature dimension + interpolation). The specific method can refer to the linear interpolation in the text part, such as fusing in the way of weighted summation in the feature dimension.
[0195] (4) For the visual part (i.e., video data), there are also various fusion methods and fusion dimensions, as follows:
[0196] a) Concatenate the video frames of the two samples (original dimension + concat), that is, extract N frames from the video data of each sample, and then take p and 1 - p proportions of the extracted video frames of the two samples respectively. Assuming p is 0.5, then take N / 2 frames from the video data of sample 1 and N / 2 frames from the video data of sample 2, and then combine them together, still N frames.
[0197] b) Perform linear interpolation on the video frames of the two samples (original dimension + interpolation), that is, directly perform weighted summation on the video frames of sample 1 and the video frames of sample 2 in the pixel dimension to obtain p * video frames of sample 1 + (1 - p) * video frames of sample 2.
[0198] c) Concatenation in the feature dimension (feature dimension + concat). The video frames of the two samples obtain CV embeddings through a visual feature extractor with a feature dimension of [N, 1024]. Then, take p and 1 - p proportions of the features of the two samples respectively, that is, [p * N, 1024] and [(1 - p) * N, 1024]. The feature dimension after concatenation is still [N, 1024].
[0199] d) Linear interpolation in the feature dimension (feature dimension + interpolation), and linear interpolation is achieved by means of weighted summation in the feature dimension.
[0200] In addition to the fusion methods introduced above, the cutout method can also be used to fuse the above various modalities of data.
[0201] After enhancement by the above methods, enhanced multi-modal data can be obtained, such as Figure 13 shown, where Figure 13 shows the fusion results of the data of each modality in two samples. These fusion results include the fused data of each modality in the original data space; in addition, these fusion results also include the fused features in the feature space (not shown in the figure).
[0202] (5) For the label part (i.e., label information), it can be fused according to a ratio. For example, if the label information of sample 1 is TV drama and the label information of sample 2 is technology, then the fused label information is (TV drama p, technology 1 - p).
[0203] After obtaining the enhanced data in the original data space, this enhanced data and the corresponding label information are used as new training samples to train the multi-modal classification model together with the original samples 1 and 2.
[0204] In actual training, p for each sample enhancement can be dynamically adjusted. For example, the beta function is used to obtain p, as Figure 14 . In addition, the ratio of the fused samples in each batch can also be adjusted. It should be noted that the solution of this application does not require reading twice the amount of data for fusion. Only the batch needs to be dynamically adjusted and then fused with the original batch. Therefore, it will not increase the burden of data reading or the consumption of training resources.
[0205] In the solution of this application, data enhancement is performed in the original data space, and after the multi-modal model extracts features, data enhancement is performed in the feature space. Specifically, input samples 1 and 2 into the multi-modal classification model. The multi-modal classification model performs feature extraction to obtain the features of the corresponding modalities in the feature space, then the features of the two samples are fused, and finally sent to a classifier (such as softmax) to obtain the classification result, as Figure 15 shown. Finally, the loss is calculated between this classification result and the corresponding label information, and the parameters of the multi-modal classification model are optimized using this loss until the model converges.
[0206] After completing the training of the multi-modal classification model, this multi-modal classification model can be used in application scenarios such as recommendation and content operation. Such as Figure 16As shown, the creator creates various multimodal data through relevant visual creation tools and audio creation tools, such as TV dramas, voices, and corresponding lines (or other text data). Then, content processing is performed on this multimodal data (such as data augmentation and model training). Subsequently, the trained multimodal classification model is applied to content distribution application scenarios, such as recommendation and content operation application scenarios.
[0207] Among them, for content recommendation, when a user swipes through videos in a video application, such as swiping on a video page to view videos, the multimodal classification model can use the multimodal data of TV dramas, movies, live videos (which can also be live rooms) or short videos to make content recommendations to the user, such as recommending TV dramas, movies, live videos or short videos that the user may be interested in.
[0208] Or, during the process of a user watching a certain video, the multimodal classification model can also determine the popular segments of the video (such as the highlight segments of the video, that is, the wonderful segments) based on the multimodal data of the video, and then display the position of the popular segment on the progress bar, as Figure 17 shown.
[0209] The solution of this application is simple and general, and can be applied generally in multimodal classification tasks. By enhancing multimodal data in the mixup series manner, the performance of the multimodal classification model is effectively improved without increasing training resources. In addition, the improvement of the multimodal classification model can significantly improve the efficiency of machine review and save the cost of manual review.
[0210] In addition, the solution of this application can be implemented in actual business and has been verified effective in multiple business scenarios. For example, in the browser classification business, when the multimodal classification model has been fully fine-tuned to relatively high metrics, using this method, the classification effect can still be steadily improved, as shown in Table 3.
[0211] Table 3
[0212]
[0213] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0214] Based on the same inventive concept, an embodiment of the present application further provides a data augmentation device for implementing the data augmentation method involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the data augmentation device provided below can refer to the limitations on the data augmentation method in the above text and will not be repeated here.
[0215] In one embodiment, as Figure 18 shown, a data augmentation device is provided, including: an acquisition module 1802, a first fusion module 1804, a second fusion module 1806, a third fusion module 1808, and an augmentation module 1810, where:
[0216] The acquisition module 1802 is configured to acquire data to be augmented; the data to be augmented includes multi-modal data to be augmented and label information;
[0217] The first fusion module 1804 is configured to perform data fusion processing on the data belonging to the same modality in the multi-modal data according to a preset ratio to obtain the fusion data of each modality;
[0218] The second fusion module 1806 is configured to extract the features of each modality from the multi-modal data, and perform feature fusion processing on at least two features belonging to the same modality according to a preset ratio to obtain the fusion features of each modality;
[0219] The third fusion module 1808 is configured to fuse the label information of each modality according to a preset ratio to obtain the fused label information;
[0220] The augmentation module 1810 is configured to obtain augmented data based on the fusion data, the fusion features, and the fused label information.
[0221] In the above embodiments, for each modality data in the multimodal data, the data fusion process can be performed on the data belonging to the same modality in the multimodal data according to a preset ratio to obtain the fusion data of each modality, so that the enhancement of multimodal data can be realized in the original data space, which is beneficial to improving the effect of multimodal data enhancement; in addition, the features of each modality can be extracted from the multimodal data, and at least two features belonging to the same modality are subjected to feature fusion processing according to a preset ratio to obtain the fusion features of each modality, so that the enhancement of multimodal data can be realized in the feature space, further improving the effect of multimodal data enhancement; finally, the label information of each modality is fused according to a preset ratio, and the enhanced data is obtained based on the obtained fusion label information, fusion data and fusion features, so as to realize the enhancement of multimodal data in the original data space and feature space, effectively improving the effect of multimodal data enhancement; moreover, by performing data enhancement on the data of each modality in the multimodal data, the model performance can be effectively improved without increasing training resources during model training.
[0222] In one of the embodiments, the first fusion module 1804 is further configured to partition the data belonging to the same modality in the multimodal data; perform data fusion processing on the data belonging to the same modality according to a preset ratio to obtain the fusion data of each modality; or replace a part of the data with a preset size in the data belonging to the same modality with other data that matches the same modality and has a preset size to obtain the fusion data of each modality.
[0223] In one of the embodiments, the multimodal data includes visual data, and the visual data includes first visual data and second visual data;
[0224] The first fusion module 1804 is further configured to perform data fusion processing on the first visual data and the second visual data according to a preset ratio to obtain visual fusion data;
[0225] The second fusion module 1806 is further configured to extract the first visual feature and the second visual feature from the first visual data and the second visual data, and perform feature fusion processing on the first visual feature and the second visual feature according to a preset ratio to obtain visual fusion features.
[0226] In one of the embodiments, the preset ratio includes a first preset ratio and a second preset ratio;
[0227] The first fusion module 1804 is further configured to crop a first visual data block from the first visual data according to a first preset ratio; crop a second visual data block from the second visual data according to a second preset ratio; splice the first visual data block and the second visual data block to obtain visual fusion data; or perform weighted processing on the first visual data and the second visual data according to the first preset ratio and the second preset ratio to obtain visual fusion features.
[0228] In one embodiment, the first fusion module 1804 is further configured to determine a to-be-covered area with a target size in the first visual data according to a preset ratio; intercept a target tile with the target size from the second visual data, and respectively cover the target tile on the to-be-covered area to obtain visual fusion data; or obtain a mask tile with the target size, and respectively cover the mask tile on the to-be-covered area of the first visual data and the to-be-covered area of the second visual data to obtain visual fusion data.
[0229] In one embodiment, the first visual data is first video data or first image data, and the second visual data is second video data or second image data;
[0230] The first fusion module 1804 is further configured to perform video frame fusion processing on the video frames in the first video data and the video frames in the second video data according to a preset ratio to obtain video fusion data; or perform tile fusion processing on the first image data and the second image data according to a preset ratio to obtain image fusion data.
[0231] In one embodiment, the multimodal data includes text data, the text data includes first text data and second text data, and the preset ratio includes a first preset ratio and a second preset ratio;
[0232] The first fusion module 1804 is further configured to splice the first text data and the second text data according to a preset ratio to obtain text fusion data; or truncate the first text data according to the first preset ratio to obtain a first text data segment; truncate the second text data according to the second preset ratio to obtain a second text data segment; and splice the first text data segment and the second text data segment to obtain text fusion data.
[0233] In one embodiment, the second fusion module 1806 is further configured to respectively extract a first text feature and a second text feature from the first text data and the second text data; perform weighted summation on the first text feature and the second text feature according to the first preset ratio and the second preset ratio to obtain text fusion features.
[0234] In one embodiment, the multimodal data includes audio data, and the audio data includes first audio data and second audio data;
[0235] The first fusion module 1804 is further configured to splice the first audio data and the second audio data according to a preset ratio to obtain audio fusion data; or, perform parallel superposition processing on the first audio data and the second audio data according to a preset ratio to obtain audio fusion data.
[0236] In one embodiment, the preset ratio is the ratio of the first ratio length to the second ratio length;
[0237] The second fusion module 1806 is further configured to respectively extract a first audio feature and a second audio feature from the first audio data and the second audio data; perform weighted summation on the first audio feature and the second audio feature according to the first ratio length and the second ratio length to obtain an audio fusion feature; or, intercept a first sub-audio feature of the first ratio length and a second sub-audio feature of the second ratio length from the first audio feature and the second audio feature; perform splicing processing on the first sub-audio feature and the second sub-audio feature to obtain an audio fusion feature.
[0238] In one embodiment, the third fusion module 1808 is further configured to obtain first enhanced data based on the fusion data and the fusion label information; obtain second enhanced data based on the fusion feature and the fusion label information;
[0239] The training module is configured to use the first enhanced data and the data to be enhanced as a training set, and train the multimodal classification model based on the training set to obtain a trained multimodal classification model; or, classify the fusion feature in the second enhanced data through the classifier of the multimodal classification model, and determine the classification loss based on the obtained classification result and the label information; optimize the parameters of the multimodal classification model according to the classification loss to obtain a trained multimodal classification model; wherein, the fusion feature in the second enhanced data is obtained by the feature extraction network of the multimodal classification model through extraction and feature fusion processing.
[0240] In one embodiment, the training module is further configured as an adjustment module, which is configured to dynamically adjust the preset ratio during the process of training the multimodal classification model, so as to perform data fusion processing and feature fusion processing using the adjusted preset ratio to obtain dynamically adjusted enhanced data for model training; wherein, there is a difference between the dynamically adjusted enhanced data and the enhanced data before dynamic adjustment; or, adjust the sample ratio, group the training set according to the adjusted sample ratio to obtain at least two training subsets; respectively train the multimodal classification model based on each training subset.
[0241] Each module in the above data augmentation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0242] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 19 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store multi-modal data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a data augmentation method.
[0243] In one embodiment, a computer device is provided. The computer device may be a terminal, including a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a data enhancement method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0244] Those skilled in the art can understand that Figure 19 the structure shown in
[0245] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps of the above data enhancement method.
[0246] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of the above data enhancement method.
[0247] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps of the above data enhancement method.
[0248] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0249] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0250] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0251] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A data enhancement method, characterized in that: The method comprises: Acquire data to be enhanced; the data to be enhanced includes multimodal data to be enhanced and label information; Performing data fusion processing on the data belonging to the same modality in the multimodal data according to a preset ratio to obtain fused data of each modality; Extracting features of each modality from the multimodal data, performing feature fusion processing on at least two features belonging to the same modality according to the preset ratio, to obtain fusion features of each modality; The label information of each of the modes is fused according to the preset ratio to obtain fused label information; Enhanced data is obtained based on the fused data, the fused features and the fused label information.
2. The method according to claim 1, characterized in that The step of fusing the data of the same modality in the multimodal data according to a preset ratio to obtain fused data of each modality includes: Dividing data belonging to the same modality from the multimodal data; Performing data fusion processing on the data belonging to the same modality according to a preset ratio to obtain fused data of each modality; or, Part of the data of a preset size in the data belonging to the same modality is replaced with other data matching the same modality and of the preset size, so as to obtain fused data of each of the modalities.
3. The method according to claim 2, characterized in that The multimodal data includes visual data, and the visual data includes first visual data and second visual data; The data belonging to the same modality are subjected to data fusion processing according to a preset ratio to obtain fused data of each modality, including: Performing data fusion processing on the first visual data and the second visual data according to a preset ratio to obtain visual fusion data; The extracting features of each modality from the multimodal data, performing feature fusion processing on at least two features belonging to the same modality according to the preset ratio, and obtaining fusion features of each modality includes: A first visual feature and a second visual feature are extracted from the first visual data and the second visual data, and the first visual feature and the second visual feature are subjected to feature fusion processing according to the preset ratio to obtain a visual fusion feature.
4. The method according to claim 3, characterized in that: The preset ratio includes a first preset ratio and a second preset ratio; The step of fusing the first visual data and the second visual data according to a preset ratio to obtain visual fusion data comprises: cutting out a first visual data block from the first visual data according to the first preset ratio; cutting out a second visual data block from the second visual data according to the second preset ratio; The first visual data block and the second visual data block are spliced to obtain visual fusion data; or, The first visual data and the second visual data are weighted according to the first preset ratio and the second preset ratio to obtain a visual fusion feature.
5. The method according to claim 3, characterized in that: The step of replacing part of the data of the same modality with a preset size of data with other data matching the same modality and having the preset size to obtain fused data of each modality includes: In the first visual data, determining a to-be-covered area of a target size according to a preset ratio; intercepting target blocks of the target size from the second visual data, and covering the target blocks on the areas to be covered respectively, to obtain visual fusion data; or, A mask image block of the target size is obtained, and the mask image block is respectively covered on the area to be covered of the first visual data and the area to be covered of the second visual data to obtain visual fusion data.
6. The method according to claim 3, characterized in that The first visual data is first video data or first image data, and the second visual data is second video data or second image data; The step of fusing the first visual data and the second visual data according to a preset ratio to obtain visual fusion data comprises: Performing video frame fusion processing on the video frames in the first video data and the video frames in the second video data according to a preset ratio to obtain video fusion data; or, The first image data and the second image data are subjected to block fusion processing according to the preset ratio to obtain image fusion data.
7. The method according to claim 2, characterized in that The multimodal data includes text data, the text data includes first text data and second text data, and the preset ratio includes a first preset ratio and a second preset ratio; The data belonging to the same modality are subjected to data fusion processing according to a preset ratio to obtain fused data of each modality, including: The first text data and the second text data are spliced according to a preset ratio to obtain text fusion data; or, The first text data is truncated according to a first preset ratio to obtain a first text data segment; the second text data is truncated according to a second preset ratio to obtain a second text data segment; the first text data segment and the second text data segment are spliced to obtain text fusion data.
8. The method according to claim 7, characterized in that The extracting features of each modality from the multimodal data, performing feature fusion processing on at least two features belonging to the same modality according to the preset ratio, and obtaining fusion features of each modality includes: Extracting a first text feature and a second text feature from the first text data and the second text data respectively; According to the first preset ratio and the second preset ratio, the first text feature and the second text feature are weightedly summed to obtain a text fusion feature.
9. The method according to claim 2, characterized in that: The multimodal data includes audio data, and the audio data includes first audio data and second audio data; The data belonging to the same modality are subjected to data fusion processing according to a preset ratio to obtain fused data of each modality, including: The first audio data and the second audio data are spliced together according to a preset ratio to obtain audio fusion data; or, The first audio data and the second audio data are parallelly superimposed according to a preset ratio to obtain audio fusion data.
10. The method according to claim 9, characterized in that The preset ratio is the ratio of the first ratio length to the second ratio length; The extracting features of each modality from the multimodal data, performing feature fusion processing on at least two features belonging to the same modality according to the preset ratio, and obtaining fusion features of each modality includes: Extracting a first audio feature and a second audio feature from the first audio data and the second audio data, respectively; According to the first proportional length and the second proportional length, weighted summation is performed on the first audio feature and the second audio feature to obtain an audio fusion feature; or, A first sub-audio feature of the first proportional length and a second sub-audio feature of the second proportional length are extracted from the first audio feature and the second audio feature; the first sub-audio feature and the second sub-audio feature are concatenated to obtain an audio fusion feature.
11. The method according to any one of claims 1 to 10, characterized in that: The obtaining enhanced data based on the fused data, the fused features and the fused label information includes: Obtain first enhanced data based on the fused data and the fused label information; Obtaining second enhanced data based on the fusion feature and the fusion label information; The method further includes: using the first enhanced data and the to-be-enhanced data as training sets, training a multimodal classification model based on the training sets, and obtaining a trained multimodal classification model; or, The fused features in the second enhanced data are classified by the classifier of the multimodal classification model, and the classification loss is determined based on the obtained classification result and the label information; the parameters of the multimodal classification model are optimized according to the classification loss to obtain a trained multimodal classification model; wherein the fused features in the second enhanced data are obtained by extraction and feature fusion processing by the feature extraction network of the multimodal classification model.
12. The method according to claim 11, characterized in that The method further comprises: In the process of training the multimodal classification model, the preset ratio is dynamically adjusted to perform data fusion processing and feature fusion processing using the adjusted preset ratio to obtain dynamically adjusted enhanced data for model training; wherein there is a difference between the dynamically adjusted enhanced data and the enhanced data before the dynamic adjustment; or, The sample ratio is adjusted, and the training set is grouped according to the adjusted sample ratio to obtain at least two training subsets; and a multimodal classification model is trained based on each of the training subsets.
13. A data enhancement device, characterized in that: The device comprises: An acquisition module, used to acquire data to be enhanced; the data to be enhanced includes multimodal data to be enhanced and label information; A first fusion module is used to perform data fusion processing on the data belonging to the same modality in the multimodal data according to a preset ratio to obtain fused data of each modality; A second fusion module is used to extract features of each modality from the multimodal data, and perform feature fusion processing on at least two features belonging to the same modality according to the preset ratio to obtain fusion features of each modality; A third fusion module, used to fuse the label information of each of the modalities according to the preset ratio to obtain fused label information; The enhancement module is used to obtain enhanced data based on the fused data, the fused features and the fused label information.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.