Data processing method and device, electronic equipment, storage medium and program product
By using pre-trained audio and video encoders, dual encoding paths and contrastive learning, the semantic and temporal alignment problems of audio and video data are solved, and the processing effect of multimodal tasks is improved.
Patent Information
- Application Number
- CN202510992129.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-24
AI Technical Summary
Existing technologies make it difficult to extract features with high temporal and semantic alignment in audio and video data, resulting in poor multimodal task processing results.
Pre-trained audio encoder and video encoder are used to ensure their identical structures, and dual encoding paths and contrastive learning training are used to improve the semantic and temporal alignment of audio and video features.
It improves the alignment of audio and video features and enhances the data processing effect of multimodal tasks, including the accuracy and efficiency of generation, retrieval, synchronization and recognition.
Smart Images

Figure CN120832632A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computers, and particularly relates to a data processing method and device, electronic equipment, storage medium and program product. BACKGROUND
[0002] With the development of computer technology, multi-modal data processing technology has also developed. In multi-modal data processing technology, multi-modal data can be first converted into a unified representation form, so that the multi-modal data is in a vector space of the same dimension, and then subsequent multi-modal tasks, such as feature fusion, data synchronization, etc., are executed. SUMMARY
[0003] To solve the problem that in related technologies, the processing of audio and video data is difficult to obtain audio and video features with high time and semantic alignment, the present disclosure provides a data processing method, device, electronic equipment, storage medium and program product, which improves the semantic and time alignment of audio and video features extracted based on audio and video data, and further improves the data processing effect of multi-modal tasks involving audio and video data.
[0004] According to a first aspect of an embodiment of the present disclosure, a data processing method is provided, comprising: obtaining audio data and video data; encoding the audio data and the video data through a pre-trained multi-modal model to obtain audio features and video features that are aligned in semantics and time; wherein the pre-trained multi-modal model comprises a pre-trained audio encoder and a pre-trained video encoder, the structure of the pre-trained audio encoder is the same as the structure of the pre-trained video encoder, the pre-trained audio encoder is used to encode the audio data to obtain the audio features, and the pre-trained video encoder is used to encode the video data to obtain the video features.
[0005] The audio data and the video data are encoded through the pre-trained multi-modal model to obtain audio features and video features that are aligned in semantics and time. The pre-trained multi-modal model comprises a pre-trained audio encoder and a pre-trained video encoder, and the structures of the two encoders are the same. In the case that the structures of the two encoders are the same, the semantic and time deviation between the audio features and the video features caused by different structures of the encoders can be reduced, so that the semantic and time alignment effect of the audio features and the video features will be better. Therefore, this technical solution can improve the semantic and time alignment of the audio and video features extracted based on audio and video data, and further improve the data processing effect of multi-modal tasks involving audio and video data.
[0006] In some possible implementation manners, the encoding, by the pre-trained multi-modal model, of the audio data and the video data to obtain audio features and video features that are aligned in semantics and time includes: encoding, by a first audio encoding path and a second audio encoding path of the pre-trained audio encoder, the audio data respectively to obtain the audio features; and encoding, by a first video encoding path and a second video encoding path of the pre-trained video encoder, the video data respectively to obtain the video features.
[0007] The audio encoder and the video encoder both adopt the dual encoding path structure, which can extract features from more dimensions and improve the richness of the features. Therefore, on the basis that the structures of the audio-video encoder are the same and both adopt the dual encoding path structure, the time and semantic alignment of the audio-video features can be improved, and the richness of the audio-video features can also be improved.
[0008] In some possible implementation manners, the sampling rate of the first audio encoding path is less than a preset sampling rate, the running capacity of the first audio encoding path is greater than or equal to a preset running capacity, and the first audio encoding path has a time resolution limit and a time kernel limit, and the sampling rate of the second audio encoding path is greater than or equal to the preset sampling rate, and the running capacity of the second audio encoding path is less than the preset running capacity.
[0009] The sampling rate of the first audio encoding path is low, the running capacity is large, and the first audio encoding path has a time resolution limit and a time kernel limit, which can be regarded as a slow path and can effectively capture spatial semantic information. The sampling rate of the second audio encoding path is high, and the running capacity is small, which can capture rapidly changing motion features. Therefore, the audio encoder can extract more rich audio features.
[0010] In some possible implementation manners, the running frame rate of the first video encoding path is less than a preset running frame rate, the processing time step of the first video encoding path is less than a preset processing time step, the running frame rate of the second video encoding path is greater than or equal to the preset running frame rate, and the processing time step of the second video encoding path is greater than or equal to the preset processing time step.
[0011] The running frame rate of the first video encoding path is small, and the processing time step is large, which can be regarded as a slow path and focuses on the spatial domain and semantics, and can effectively capture spatial semantic information. The running frame rate of the second video encoding path is high, and the processing time step is small, and the sampling is more intensive, which has a time resolution feature and can capture rapidly changing motion features. Therefore, the video encoder can extract more rich video features.
[0012] In some possible implementation manners, the data processing method further includes: obtaining training data, the training data including a plurality of training samples, each training sample including: an audio data sample and a video data sample; encoding, by a to-be-trained multi-modal model, the plurality of training samples respectively to obtain a plurality of feature samples, each feature sample including: an audio feature sample and a video feature sample; and training, according to the plurality of feature samples, the to-be-trained multi-modal model to obtain the pre-trained multi-modal model.
[0013] The training data includes a plurality of training samples, each training sample can be regarded as an audio-video sample pair, and through the audio-video sample pair, the contrastive learning training can be performed on the audio-video encoder in the multi-modal model, so that the audio-video encoder can extract the audio-video features aligned in time and speech.
[0014] In some possible implementation manners, the training, according to the plurality of feature samples, of the to-be-trained multi-modal model to obtain the pre-trained multi-modal model includes: calculating, according to the plurality of feature samples, a contrastive loss of the to-be-trained multi-modal model, the contrastive loss being used to represent a deviation between the plurality of audio feature samples and the plurality of video feature samples; and training, according to the contrastive loss, of the to-be-trained multi-modal model to obtain the pre-trained multi-modal model.
[0015] The model training through the contrastive loss representing the deviation between the plurality of audio data sample features and the plurality of video data sample features can have the effect of contrastive learning, and can constrain the feature extraction manner of the audio-video encoder, so that the audio-video encoder can extract the audio-video features aligned in time and speech.
[0016] In some possible implementation manners, the data processing method further includes: training, according to the audio feature and the video feature, of a to-be-trained generation model to obtain a pre-trained generation model; and the pre-trained generation model is used to generate a sound video according to a soundless video.
[0017] Since the audio feature and the video feature are highly aligned in time and semantics, the generation effect of the generation model can be improved.
[0018] In some possible implementation manners, the obtaining of the audio data and the video data includes: in response to receiving a first retrieval request used to instruct to retrieve a video related to a target sound video, obtaining the target sound video from the first retrieval request; and extracting the audio data and the video data from the target sound video, the video data being soundless video data; and the data processing method further includes: performing retrieval in a preset video database according to the audio feature and the video feature to obtain a retrieval result for the first retrieval request.
[0019] By aligning the audio features and the video features in semantics and time, a video feature matching video feature can be quickly and accurately retrieved from a preset video database, improving retrieval efficiency and accuracy.
[0020] In some possible implementation manners, the obtaining the audio data and the video data includes: in response to receiving a second retrieval request for indicating retrieval of a video related to the audio data, obtaining the audio data from the second retrieval request, and obtaining the video data from a preset video database; and the data processing method further includes: determining a feature similarity between the audio features and the video features; and determining a retrieval result for the second retrieval request according to the feature similarity.
[0021] By aligning the audio features and the video features in semantics and time, it can be quickly and accurately determined whether the video data corresponding to the video feature is data meeting the retrieval requirement, improving retrieval efficiency and accuracy.
[0022] In some possible implementation manners, the obtaining the audio data and the video data includes: in response to receiving a processing request for indicating audio-video synchronization of a to-be-processed voice video, obtaining the to-be-processed voice video from the processing request; and extracting the audio data and the video data from the to-be-processed voice video, the video data being voiceless video data; and the data processing method further includes: determining a feature similarity between the audio features and the video features; and performing audio-video synchronization processing on the to-be-processed voice video according to the feature similarity.
[0023] By mining the audio features and the video features aligned in time and voice, the accuracy and efficiency of audio-video synchronization processing can be improved, and the audio-visual experience of a user can be improved.
[0024] In some possible implementation manners, the obtaining the audio data and the video data includes: in response to receiving an identification request for indicating content identification of a to-be-identified voice video, obtaining the to-be-identified voice video from the identification request; and extracting the audio data and the video data from the to-be-identified voice video, the video data being voiceless video data; and the data processing method further includes: determining an identification result for the identification request according to the audio features and the video features.
[0025] By mining the audio features and the video features aligned in time and voice, the efficiency and accuracy of video identification can be improved.
[0026] According to a second aspect of the embodiments of the present disclosure, a data processing apparatus is provided, configured to implement the data processing method according to the first aspect of the present disclosure.
[0027] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement the data processing method according to the first aspect of the present disclosure.
[0028] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer program instructions, which, when executed by a processor, implement the data processing method according to the first aspect of the present disclosure.
[0029] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the data processing method according to the first aspect of the present disclosure.
[0030] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0031] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure together with the description.
[0032] Figure 1 is a flowchart of a data processing method according to an example embodiment.
[0033] Figure 2 is a structural block diagram of a pre-trained multi-modal model according to an example embodiment.
[0034] Figure 3 is a contrast learning schematic diagram of a multi-modal model according to an example embodiment.
[0035] Figure 4 is a schematic diagram of the discrimination of a multi-modal model of a related technology on audio-video data according to an example embodiment.
[0036] Figure 5 is a schematic diagram of the discrimination of a multi-modal model of an embodiment of the present disclosure on audio-video data according to an example embodiment.
[0037] Figure 6 is a block diagram of a data processing apparatus according to an example embodiment.
[0038] Figure 7 is a block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION
[0039] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements, unless the context of use indicates otherwise. The following description of exemplary embodiments is not representative of all embodiments consistent with the present disclosure. Rather, it is merely an example of apparatus and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0040] It should be noted that all the actions of acquiring signals, information or data in the present disclosure are carried out in accordance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the owner of the corresponding device.
[0041] As mentioned in the background, in the multi-modal data processing technology, the conversion processing of multi-modal data is involved, which can be understood as encoding based on multi-modal data through a multi-modal model to obtain multi-modal data features, and the obtained multi-modal data features have a certain correlation.
[0042] Among them, the multi-modal model needs to be pre-trained to obtain a pre-trained multi-modal model, and then the pre-trained multi-modal model is used to encode the multi-modal data to obtain multi-modal data features with a certain correlation.
[0043] In the related art, for the two kinds of multi-modal data of image and text, through the CLIP (Contrastive Language-Image Pretraining) model, the acquisition of image features and text features with correlation can be realized.
[0044] For the two kinds of multi-modal data of text and audio, through the CLAP (Contrastive Language-Audio Pretraining) model, the acquisition of text features and audio features with correlation can be realized.
[0045] However, for the two kinds of multi-modal data of audio and video, there is still no relatively mature multi-modal model for processing of audio and video data, and it is difficult to obtain audio and video features with high time and semantic alignment (with correlation).
[0046] Based on this, the embodiment of the disclosure provides a technical scheme, through the pre-trained multi-modal model, the audio data and the video data are encoded, and the audio features and the video features that are aligned in semantics and time are obtained. Wherein, the pre-trained multi-modal model includes a pre-trained audio encoder and a pre-trained video encoder, and the structures of the two encoders are the same. In the case that the structures of the two encoders are the same, the semantic and temporal deviation between the audio and video features caused by the different structures of the encoders can be reduced, so that the semantic and temporal alignment effect of the audio features and the video features will be better.
[0047] Therefore, the technical scheme can improve the alignment degree of the audio and video features extracted based on the audio and video data in semantics and time, and further improve the data processing effect of the multi-modal task related to the audio and video data.
[0048] The technical scheme provided by the embodiment of the disclosure can be applied to various multi-modal task scenarios. The specific multi-modal task scenarios are introduced in subsequent embodiments.
[0049] Figure 1 is a flowchart of a data processing method according to an exemplary embodiment. The data processing method can be applied to an electronic device, such as Figure 1 As shown in the figure, the data processing method includes the following steps: Step S11, obtaining audio data and video data.
[0050] Step S12, through the pre-trained multi-modal model, the audio data and the video data are encoded, and the audio features and the video features that are aligned in semantics and time are obtained.
[0051] In step S11, the audio data and the video data can be understood as data that needs to be encoded. In different multi-modal task scenarios, different implementation manners can be adopted to obtain the audio data and the video data. The specific obtaining manner is introduced in subsequent embodiments.
[0052] Figure 2 is a structural block diagram of a pre-trained multi-modal model according to an exemplary embodiment. As shown in the figure, the pre-trained multi-modal model includes a pre-trained audio encoder and a pre-trained video encoder. Wherein, the pre-trained audio encoder and the pre-trained video encoder have the same structure. Figure 2
[0053] Through the pre-trained audio encoder, the audio data can be encoded; through the pre-trained video encoder, the video data can be encoded. In addition, since the audio encoder and the video encoder are both pre-trained, the obtained audio features and video features are features that are aligned in semantics and time.
[0054] As an example, the video data is silent video data, the audio data and the video data are data extracted from the same voiced video, and the pre-trained multi-modal model can output pairs of audio-video features corresponding to the same time and the same content (i.e., semantics).
[0055] As another example, the video data is silent video data, the audio data and the video data are data extracted from different voiced videos, and the pre-trained multi-modal model can output pairs of audio-video features with a correlation between the time and the content.
[0056] Therefore, the alignment in semantics and time can be that the semantics and the time are consistent, or that the semantics and the time have a correlation.
[0057] As an optional implementation, step S12 includes: encoding the audio data through a first audio encoding path and a second audio encoding path of a pre-trained audio encoder to obtain audio features; and encoding the video data through a first video encoding path and a second video encoding path of a pre-trained video encoder to obtain video features.
[0058] In this implementation, the audio encoder and the video encoder both adopt a double-encoding-path structure. This double-encoding-path structure can extract features from more dimensions and improve the richness of the features. Therefore, on the basis of the same structure of the audio-video encoder and the adoption of the double-encoding-path structure, the time and semantic alignment of the audio-video features can be improved, and the richness of the audio-video features can also be improved.
[0059] In some embodiments, for the double-encoding-path encoder, the structures of the two encoding paths can be the same, and different encoding parameters are adopted. The two encoding paths can be fused through lateral connection, and the advantages are complementary.
[0060] In some embodiments, the backbone network of the audio encoder and the video encoder can both be ResNet-50.
[0061] In some embodiments, the double-encoding-path can be a fast-slow double-flow path. The fast-slow double-flow path includes a fast path and a slow path. The encoding parameters of the fast path and the slow path are different, and the different encoding parameters can have different feature extraction effects.
[0062] As an optional implementation, the sampling rate of the first audio encoding path is less than the preset sampling rate, the running capacity of the first audio encoding path is greater than or equal to the preset running capacity, and the first audio encoding path has a time resolution limit and a time kernel limit, and the sampling rate of the second audio encoding path is greater than or equal to the preset sampling rate, and the running capacity of the second audio encoding path is less than the preset running capacity.
[0063] In this implementation, the first audio encoding path has a lower sampling rate, a larger running capacity, and a time resolution limit and a time kernel limit, which can be regarded as a slow path, and can focus on learning frequency semantics, thereby effectively capturing semantic features. The second audio encoding path has a higher sampling rate and a smaller running capacity, and can focus on learning time patterns, thereby effectively capturing time features. Therefore, the audio encoder can extract more abundant audio features.
[0064] In some embodiments, the preset sampling rate can be set according to different scenes, which can serve as a dividing line between high and low sampling rates.
[0065] In some embodiments, the preset running capacity can be set according to different scenes, which can serve as a dividing line between high and low channel capacities.
[0066] In some embodiments, the time resolution limit can limit the size of the time resolution, and the time kernel limit can limit the number of time kernels (convolution kernels), which can be set according to different application scenarios.
[0067] It can be understood that the running capacity, sampling rate, time resolution, and time kernel mentioned above all belong to encoding parameters, and the specific meanings of these encoding parameters can be referred to mature encoder technology in the art, and the related meanings will not be described in detail here.
[0068] As an optional implementation, the running frame rate of the first video encoding path is less than the preset running frame rate, the processing time step of the first video encoding path is less than the preset processing time step, the running frame rate of the second video encoding path is greater than or equal to the preset running frame rate, and the processing time step of the second video encoding path is greater than or equal to the preset processing time step.
[0069] In this implementation, the running frame rate of the first video encoding path is smaller, and the processing time step is larger, which can be regarded as a slow path, and can focus on the spatial domain and semantics, and can effectively capture spatial semantic information. The running frame rate of the second video encoding path is higher, and the processing time step is smaller, and the sampling is more intensive, and has a time resolution feature, and can capture rapidly changing motion features. Therefore, the video encoder can extract more abundant video features.
[0070] In some embodiments, the preset running frame rate can be configured according to different application scenarios, which can serve as a dividing line between high frame rate and low frame rate.
[0071] In some embodiments, the preset time processing step can be configured according to different application scenarios, which can serve as a dividing line between a larger time step and a smaller time step.
[0072] It can be understood that the running frame rate and the time step, etc., all belong to encoding parameters. The specific meanings of these encoding parameters can be referred to mature encoder technology in the art, and the related meanings will not be introduced in detail here.
[0073] In some embodiments, the audio encoder and the video encoder can also adopt other structures, which are not limited to the dual-path encoder structure proposed in the embodiments of the present disclosure. Alternatively, the audio encoder and the video encoder can adopt the dual-path encoder structure, but the encoding parameters configured for the dual paths are not limited to the encoding parameters proposed in the embodiments of the present disclosure.
[0074] It can be understood that the multi-modal model needs to be pre-trained to extract the audio-visual features that are aligned in time and semantics. Next, the training method of the multi-modal model will be introduced.
[0075] As an optional implementation, the training process of the multi-modal model includes: obtaining training data, the training data including a plurality of training samples, each training sample including: an audio data sample and a video data sample; encoding the plurality of training samples respectively by the multi-modal model to be trained to obtain a plurality of feature samples, each feature sample including: an audio feature sample and a video feature sample; training the multi-modal model to be trained according to the plurality of feature samples to obtain a pre-trained multi-modal model.
[0076] In this implementation, the training data includes a plurality of training samples, each training sample can be regarded as a kind of audio-visual sample pair. Through the audio-visual sample pair, the audio-visual encoder in the multi-modal model can be trained by contrastive learning, so that the audio-visual encoder can extract the audio-visual features that are aligned in time and speech.
[0077] In some embodiments, a plurality of original videos can be obtained first, and a plurality of training samples can be obtained by processing the plurality of original videos.
[0078] In some embodiments, the original video can be obtained from a large-scale open-source audio-visual data set VGGSound, which contains 180,000 video training sets and 15,000 video test sets. When selecting the original video, a video with a time length greater than a preset time length can be selected. For example, the preset time length can be 9.5s.
[0079] As an example, the obtaining of the video data sample can include: resampling the original video at a specific frame rate to obtain resampled video data; processing each video frame in the resampled video data into a fixed size video frame to obtain processed video data; and cutting each processed video data into a plurality of video clips of equal time length, and each video clip can be regarded as a video data sample.
[0080] The specific frame rate can be 25fps or other frame rates, which are not limited herein. The fixed size can be [224, 224] or other sizes, which are not limited herein. The time length of each video clip can be 1.28s or other lengths, which are not limited herein.
[0081] In some embodiments, one original video can be cut into a preset number of video clips, and the preset number can be 7 or other values, which are not limited herein. Further, a plurality of video clips can be obtained by cutting a plurality of original videos.
[0082] In some embodiments, the video data sample can be a silent video data, which does not contain audio.
[0083] As an example, the obtaining of the audio data sample can include: audio resampling the original video at a specific sampling rate to obtain resampled audio data; converting the resampled audio data into a mel spectrum; and cutting the mel spectrum into a plurality of mel spectrum clips of equal time length, and each mel spectrum clip can be regarded as an audio data sample.
[0084] The specific sampling rate can be 16k or other values, which are not limited herein. The time length of each mel spectrum clip can be the same as the time length of the video clip, for example, 1.28s.
[0085] In some embodiments, based on one original video, a plurality of audio data samples and a plurality of video data samples can be obtained, and based on a plurality of original videos, a plurality of audio data samples and a plurality of video data samples obtained respectively, sample pair combination can be performed to obtain a plurality of training samples.
[0086] In each contrast learning process, a corresponding number of training samples can be used for model training.
[0087] As an optional implementation, the multi-modal model to be trained is trained according to the plurality of feature samples to obtain the pre-trained multi-modal model, including: calculating, according to the plurality of feature samples, a contrastive loss of the multi-modal model to be trained, the contrastive loss being used to represent a deviation between the plurality of audio feature samples and the plurality of video feature samples; and training, according to the contrastive loss, the multi-modal model to be trained to obtain the pre-trained multi-modal model.
[0088] In this implementation, the model training is performed by the contrastive loss representing the deviation between the plurality of audio feature samples and the plurality of video feature samples, which can have a contrastive learning effect and constrain the feature extraction manner of the audio-video encoder, so that the audio-video encoder can extract audio-video features that are aligned in time and speech.
[0089] In some embodiments, the training of the model to be trained is performed by the contrastive loss for the purpose of maximizing the similarity between the audio-video features from the same video and minimizing the similarity between the audio-video features from different videos or different time periods.
[0090] For the specific implementation of how to train the model to be trained based on the contrastive loss, mature training techniques in the art can be referred to, and no detailed introduction is made herein.
[0091] In some embodiments, each audio feature sample can be a one-dimensional feature vector with a dimension of 2304. Each video feature sample can be a one-dimensional feature vector with a dimension of 2304.
[0092] In some embodiments, the contrastive loss can be an average of the audio-video contrastive loss and the video-audio contrastive loss.
[0093] As an example, in each contrastive learning process, video clips obtained from B different videos can be obtained, and S video clips are obtained for each video. Therefore, the number of audio-video feature sample pairs can be BS. Therefore, the audio feature sample (vector) can be represented as , and the video feature sample (vector) can be represented as , wherein .
[0094] As an example, B can be 16, S can be 7, and BS = 112.
[0095] As an example, the contrastive loss can be represented as: , wherein is the contrastive loss, is the audio-video contrastive loss, is the video-audio contrastive loss.
[0096] As an example, may be represented as:
[0097] where exp represents the natural exponential function, represents the audio feature samples extracted from the audio data samples, corresponding to the video feature samples extracted from the video data samples, represents the audio feature samples not extracted from the audio data samples, corresponding to the video feature samples extracted from the video data samples.
[0098] and, is a trainable temperature parameter, which acts as a regularization, for controlling the contrastive loss from deviating too much.
[0099] It can be understood that, the definition of may be the same, which can be represented as:
[0100] represents the audio feature samples extracted from the audio data samples, corresponding to the video feature samples extracted from the video data samples, represents the audio feature samples not extracted from the audio data samples, corresponding to the video feature samples extracted from the video data samples.
[0101] In some embodiments, in the multimodal model, at the back end of the audio encoder and the video encoder, a pooling layer can also be configured for dimensionality reduction of the features.
[0102] Figure 3 is a contrastive learning schematic diagram of a multimodal model according to an exemplary embodiment, as shown in Figure 3 In each contrastive learning process, the audio data sample-video data sample pair is input into the multimodal model, the audio data sample is encoded by the audio encoder, and then the dimensionality is reduced by the pooling layer to obtain the audio feature sample. The video data sample is encoded by the video encoder to obtain the video feature sample.
[0103] Based on the audio feature sample and the video feature sample, a feature sample pair is obtained, based on the feature sample pair, a contrastive loss is calculated, and by using the contrastive loss, the audio encoder and the video encoder are trained respectively by means of back propagation, and finally the pre-trained audio encoder and the pre-trained video encoder are obtained, i.e. the pre-trained multimodal model is obtained.
[0104] By selecting a test set, the audio-video features are extracted by using the multi-modal model of the related technology and the multi-modal model of the embodiment of the present disclosure respectively. By calculating the similarity of the extracted audio-video features, the discrimination of the multi-modal model on the audio-video data can be reflected, and then the feature alignment performance of the multi-modal model can be reflected.
[0105] Figure 4 is a diagram of the discrimination of the multi-modal model of the related technology on the audio-video data according to an example embodiment, in Figure 4 which the horizontal axis represents the number of segments of the audio data, the vertical axis represents the number of segments of the video data, the number of segments of the audio data and the number of segments of the video data are the same, representing coming from the same video, and the number of segments of the audio data and the number of segments of the video data are different, representing coming from different videos.
[0106] And, in Figure 4 , the lighter (brighter) the color is, the higher the similarity is; the darker the color is, the lower the similarity is. The greater the color difference is, the more obvious the discrimination is.
[0107] As can be seen from Figure 4 , the audio-video features obtained by the multi-modal model of the related technology are not very obvious in discrimination for the audio-video features coming from the same video and different videos, and then the correlation between the features cannot be effectively mined, and the audio-video features aligned in time and semantics are obtained.
[0108] Figure 5 is a diagram of the discrimination of the multi-modal model of the embodiment of the present disclosure on the audio-video data according to an example embodiment, in Figure 5 which the horizontal axis represents the number of segments of the audio data, the vertical axis represents the number of segments of the video data, the number of segments of the audio data and the number of segments of the video data are the same, representing coming from the same video, and the number of segments of the audio data and the number of segments of the video data are different, representing coming from different videos.
[0109] And, in Figure 5 , the lighter (brighter) the color is, the higher the similarity is; the darker the color is, the lower the similarity is. The greater the color difference is, the more obvious the discrimination is.
[0110] As can be seen from Figure 5 , the audio-video features obtained by the multi-modal model of the embodiment of the present disclosure are obvious in discrimination for the audio-video features coming from the same video and different videos, which can effectively mine the correlation between the features, and obtain the audio-video features aligned in time and semantics.
[0111] Specifically, the upper left block and the lower right block are the similarity between audio-video features from the same video, and the color is lighter, representing higher similarity. The lower left block and the upper right block are the similarity between audio-video features from different videos, and the color is darker, representing lower similarity. Moreover, the similarity represented by the upper left block and the lower right block is obviously different from the similarity represented by the lower left block and the upper right block, and the discrimination is obvious.
[0112] The technical solution of the embodiments of the present disclosure can be applied to various multi-modal task scenarios. Next, some multi-modal task scenarios are exemplarily introduced.
[0113] In a multi-modal task scenario, a sound video can be generated according to a silent video. In this scenario, a generation model needs to be trained, and the training data of the generation model includes an audio-video feature pair. Through the scheme of the embodiments of the present disclosure, data for training the generation model can be obtained.
[0114] Therefore, as an optional implementation, the data processing method further includes: training the to-be-trained generation model according to the audio feature and the video feature to obtain a pre-trained generation model; and the pre-trained generation model is used to generate a sound video according to a silent video.
[0115] In this implementation, since the audio feature and the video feature have high alignment in time and semantics, the generation effect of the generation model can be improved.
[0116] In this application scenario, the video data and the audio data in step S11 can be data selected for the generation model for training.
[0117] In this application scenario, the generation model can be a diffusion model or other network model suitable for content generation, which is not limited here.
[0118] As can be seen, in this multi-modal task scenario, the multi-modal model of the embodiments of the present disclosure is used to obtain an audio-video feature pair to train a generation model, so that the generation model can generate audio adapted to the picture content, action change, etc. in the video, add sound effects to the silent video, create background music for the video, etc., play an auxiliary effect for video content creation, and enrich the video content.
[0119] In a multi-modal task scenario, a related video can be retrieved according to a video or audio.
[0120] In the scenario of searching according to the video, the step S11 can include: in response to receiving a first search request for indicating searching for a video related to a target sound video, obtaining the target sound video from the first search request; extracting audio data and video data from the target sound video, the video data being soundless video data.
[0121] In some embodiments, the target sound video can be a video uploaded by a user or a video selected by the user from existing video resources, and different implementation manners can be adopted according to different search scenarios, which are not limited herein.
[0122] In some embodiments, the video data is obtained by performing soundless video data extraction on the target sound video, and the audio data is obtained by performing audio data extraction.
[0123] Further, the data processing method further includes: searching in a preset video database according to the audio feature and the video feature to obtain a search result for the first search request.
[0124] By aligning the audio feature and the video feature in semantics and time, the related video meeting the audio feature and the video feature can be quickly and accurately searched from the preset video database, and the search efficiency and accuracy are improved.
[0125] In some embodiments, the preset video database can be a database configured in a search platform or a search database, and the database includes a large amount of video data.
[0126] In some embodiments, when searching according to the audio feature and the video feature, the audio video feature corresponding to the video data in the preset video database can be obtained, and then the target video meeting the audio feature and the video feature is determined as the search result for the first search request by feature comparison.
[0127] In the scenario of searching according to the audio, the step S11 can include: in response to receiving a second search request for indicating searching for a video related to audio data, obtaining the audio data from the second search request, and obtaining the video data from a preset video database.
[0128] In some embodiments, the audio data can be a video uploaded by a user or an audio selected by the user from existing audio resources, and different implementation manners can be adopted according to different search scenarios, which are not limited herein.
[0129] It can be understood that the preset video database includes a large amount of video data, and the video data can be obtained therefrom.
[0130] Further, the data processing method further includes: determining a feature similarity between the audio feature and the video feature; and determining the search result for the second search request according to the feature similarity.
[0131] By determining the audio feature and the video feature that are aligned in semantics and time, it can be quickly and accurately determined whether the video data corresponding to the video feature is data that meets the search requirement, thereby improving the search efficiency and accuracy.
[0132] In some embodiments, the video feature with a similarity higher than a preset similarity to the audio feature can be determined, and the video data corresponding to the video feature can be taken as the search result for the second search request. The preset similarity can be configured according to different application scenarios, which is not limited herein.
[0133] It can be understood that in such a search scenario, the technical solution of the embodiments of the present disclosure can be used to analyze the matching degree between features.
[0134] In a multi-modal task scenario, audio data and video data can be processed synchronously. In such a scenario, feature extraction can be performed on the audio-video data that needs to be processed synchronously to improve the synchronization processing accuracy.
[0135] In such a scenario, step S11 can include: in response to receiving a processing request for indicating audio-video synchronization of a to-be-processed sound video, obtaining the to-be-processed sound video from the processing request; and extracting audio data and video data from the to-be-processed sound video, the video data being soundless video data.
[0136] In scenarios such as film and television production, online education videos, and live streaming, there can be a demand for audio-video synchronization. Therefore, the to-be-processed sound video can be a produced video, an online education video, a live streaming video, and the like, which is not limited herein.
[0137] Further, the data processing method can further include: determining a feature similarity between the audio feature and the video feature; and performing audio-video synchronization processing on the to-be-processed sound video according to the feature similarity.
[0138] By mining the audio feature and the video feature that are aligned in time and voice, the accuracy and efficiency of the audio-video synchronization processing can be improved, thereby improving the audio-visual experience of users.
[0139] In some embodiments, the audio-video features with a low feature similarity can be located, and then the audio-video data corresponding to the part of the audio-video features of the to-be-processed sound video can be processed synchronously, for example, the order of video frames can be adjusted, video frames can be added or deleted, and the like. For details, refer to mature audio-video synchronization technology in the art, which is not specifically introduced herein.
[0140] In a multi-modal task scenario, audio data and video data in a sound video can be analyzed to implement video content recognition.
[0141] In this scenario, step S11 can include: in response to receiving an identification request for indicating content recognition of a sound video to be identified, obtaining the sound video to be identified from the identification request; extracting audio data and video data from the sound video to be identified, the video data being soundless video data.
[0142] As an example, in a video resource management scenario, it is necessary to classify video resources, in which case, the content of the video needs to be identified, and the video content recognition result can be used as a basis for classification of video resources. Therefore, the sound video to be identified can be a video that needs to be classified.
[0143] Further, the data processing method further includes: determining an identification result for the identification request according to the audio feature and the video feature.
[0144] By mining audio features and video features that are aligned in time and speech, the efficiency and accuracy of video recognition can be improved.
[0145] In some embodiments, the audio features and video features can be input into a pre-trained content recognition model, and the pre-trained content recognition model determines the identification result based on the audio features and video features.
[0146] In some embodiments, the identification result can include the content type of the sound video, etc., which can be used for classifying video resources.
[0147] It can be understood that by obtaining audio features and video features that are aligned in time and speech, the input data of the pre-trained content recognition model is preprocessed, which can make the identification accuracy and efficiency of the pre-trained content recognition model higher.
[0148] The above several multi-modal task scenarios are only examples, and the technical solutions of the embodiments of the present disclosure can also be applied to more multi-modal task scenarios to improve the execution effect of the corresponding multi-modal task, which is not limited here.
[0149] Figure 6 is a block diagram of a data processing apparatus 600 according to an exemplary embodiment. Referring to Figure 6 The apparatus includes: The acquisition module 601 is configured to acquire audio data and video data.
[0150] The processing module 602 is configured to encode the audio data and the video data by using a pre-trained multi-modal model to obtain audio features and video features that are aligned in semantics and time; wherein the pre-trained multi-modal model comprises a pre-trained audio encoder and a pre-trained video encoder, the pre-trained audio encoder has the same structure as the pre-trained video encoder, the pre-trained audio encoder is configured to encode the audio data to obtain the audio features, and the pre-trained video encoder is configured to encode the video data to obtain the video features.
[0151] Optionally, the processing module 602 is further configured to encode the audio data by using a first audio encoding path and a second audio encoding path of the pre-trained audio encoder to obtain the audio features, and encode the video data by using a first video encoding path and a second video encoding path of the pre-trained video encoder to obtain the video features.
[0152] Optionally, the apparatus further comprises a training module configured to obtain training data, the training data comprising a plurality of training samples, each training sample comprising audio data and video data; encode the plurality of training samples by using a multi-modal model to be trained to obtain a plurality of feature samples, each feature sample comprising an audio feature sample and a video feature sample; and train the multi-modal model to be trained according to the plurality of feature samples to obtain the pre-trained multi-modal model.
[0153] Optionally, the training module is further configured to calculate a contrastive loss of the multi-modal model to be trained according to the plurality of feature samples, the contrastive loss being used to represent a deviation between the plurality of audio feature samples and the plurality of video feature samples; and train the multi-modal model to be trained according to the contrastive loss to obtain the pre-trained multi-modal model.
[0154] Optionally, the training module is further configured to train a generative model to be trained according to the audio features and the video features to obtain a pre-trained generative model; wherein the pre-trained generative model is used to generate a sound video from a soundless video.
[0155] Optionally, the obtaining module 601 is further configured to obtain the target sound video from a first retrieval request for indicating retrieval of a video related to the target sound video in response to receiving the first retrieval request; extract the audio data and the video data from the target sound video, the video data being soundless video data; and the processing module 602 is further configured to perform retrieval in a preset video database according to the audio features and the video features to obtain a retrieval result for the first retrieval request.
[0156] Optionally, the obtaining module 601 is further configured to: in response to receiving a second retrieval request for indicating retrieval of a video related to the audio data, obtain the audio data from the second retrieval request, and obtain the video data from a preset video database; and the processing module 602 is further configured to: determine a feature similarity between the audio feature and the video feature; and determine a retrieval result for the second retrieval request according to the feature similarity.
[0157] Optionally, the obtaining module 601 is further configured to: in response to receiving a processing request for indicating audio-video synchronization of a to-be-processed sound video, obtain the to-be-processed sound video from the processing request; and extract the audio data and the video data from the to-be-processed sound video, the video data being soundless video data; and the processing module 602 is further configured to: determine a feature similarity between the audio feature and the video feature; and perform audio-video synchronization processing on the to-be-processed sound video according to the feature similarity.
[0158] Optionally, the obtaining module 601 is further configured to: in response to receiving an identification request for indicating content identification of a to-be-identified sound video, obtain the to-be-identified sound video from the identification request; and extract the audio data and the video data from the to-be-identified sound video, the video data being soundless video data; and the processing module 602 is further configured to: determine an identification result for the identification request according to the audio feature and the video feature.
[0159] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in details in the embodiments of the method, and thus will not be described in details here.
[0160] The present disclosure also provides a computer readable storage medium, having stored thereon computer program instructions, which, when executed by a processor, implement the steps of the data processing method provided by the present disclosure.
[0161] Figure 7 FIG. 7 is a block diagram of an electronic device 700 according to an exemplary embodiment. The electronic device 700 can be, for example, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, or the like.
[0162] Referring to FIG. 7, Figure 7 The electronic device 700 can include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.
[0163] The processing component 702 generally controls the overall operations of the electronic device 700, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 702 can include one or more processors 720 to execute instructions and to complete all or part of steps of the data processing methods described above. In addition, the processing component 702 can include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 can include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.
[0164] The memory 704 is configured to store various types of data to support operations of the electronic device 700. Examples of these data include instructions for any application or methods operating on the electronic device 700, contact data, phonebook data, messages, pictures, videos, and so on. The memory 704 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage devices, flash memory, magnetic disks, or optical disks.
[0165] The power component 706 provides power to the various components of the electronic device 700. The power component 706 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.
[0166] The multimedia component 708 includes a screen providing an output interface between the electronic device 700 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensors can not only sense a boundary of a touching or sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0167] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive an external audio signal when the electronic device 700 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 also includes a speaker for outputting audio signals.
[0168] The input / output interface 712 provides an interface between the processing component 702 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0169] The sensor component 714 includes one or more sensors for providing status assessments of various aspects of the electronic device 700. For example, the sensor component 714 can detect an open / closed position of the electronic device 700, relative positioning of components, such as a display and a keypad of the electronic device 700, a change of position of the electronic device 700 or a component of the electronic device 700, presence or absence of user contact with the electronic device 700, orientation or acceleration / deceleration / g-force and temperature of the electronic device 700. The sensor component 714 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0170] The communication component 716 is configured to facilitate wired or wireless communication between the electronic device 700 and other devices. The electronic device 700 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 716 receives broadcast signals or broadcast-related information from an external broadcasting management system via a broadcast channel. In an example embodiment, the communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0171] In exemplary embodiments, the electronic device 700 can be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements for performing the above-described data processing method.
[0172] In exemplary embodiments, a non-transitory computer-readable storage medium including instructions, such as the memory 704 including instructions, is also provided, which can be executed by the processor 720 of the electronic device 700 to accomplish the above-described data processing method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0173] In another exemplary embodiment, a computer program product including a computer program capable of being executed by a programmable apparatus is also provided, which has a code portion for executing the above-described data processing method when executed by the programmable apparatus.
[0174] It should be understood that the features of various some embodiments of the present disclosure described herein can be combined with each other unless specifically noted otherwise. As used herein, the term "and / or" includes any one of the associated listed items, as well as any combination of any two or more of the associated listed items; similarly, "at least one of" includes any one of the associated listed items, as well as any combination of any two or more of the associated listed items.
[0175] Although terms such as "first", "second", and "third" can be used herein to describe various components, parts, regions, layers, or segments, these components, parts, regions, layers, or segments are not limited to these terms. Rather, these terms are used only to distinguish one component, part, region, layer, or segment from another component, part, region, layer, or segment. Therefore, the first component, part, region, layer, or segment mentioned in the examples described herein can also be referred to as the second component, part, region, layer, or segment without departing from the teachings of the examples. In addition, the terms "first", "second" are used for description purposes only, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description herein, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise specifically limited.
[0176] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, use of the word exemplary is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless specified otherwise, or clear from context, "X employs A or B" is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied under any of the foregoing instances. In addition, the articles "a" and "an" as used in this application and the appended claims should generally be construed to mean "one or more" unless specified otherwise or clear from context to be directed to a singular form. Thus, use of the articles in this application and the following claims is not limiting.
[0177] Also, although the disclosure has been described with respect to only one or more implementations thereof, those skilled in the art will readily appreciate that other alternatives can be used. It is contemplated that the disclosure can be carried out in other specific ways than those expressly disclosed herein. Any and all such changes and modifications other than those already described and claimed are intended to be included within the scope of the disclosure. Specifically with respect to the various functions performed by the components (e.g., elements, resources, etc.) described above, unless otherwise specified, the terms used are intended to encompass any component which performs the specified function for an equivalent result. In addition, it is contemplated that various combinations or sub-combinations of the specific features and / or aspects of the disclosure can be made, and are intended to be within the scope of the disclosure. Further, unless otherwise specified, use of the terms in the description above are intended to be inclusive of the singular and the plural. Additionally, although the disclosure has been described with respect to only a few implementations, it will be appreciated that those skilled in the art will readily apply the principles of the disclosure to other implementations and applications without departing from the spirit and scope of the disclosure. It is intended, therefore, that the disclosure be limited only by the scope of the appended claims.
[0178] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features of the disclosure disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.
[0179] It is understood that the present disclosure is not limited to the precise construction disclosed and as illustrated in the accompanying drawings, and that various modifications can be made by those skilled in the art without departing from the scope of the disclosure. The scope of the disclosure is limited only by the claims appended hereto.
Claims
1. A data processing method, characterized by, The method comprises: obtaining audio data and video data; encoding the audio data and the video data by using a pre-trained multi-modal model to obtain audio features and video features that are aligned in semantics and time; wherein the pre-trained multi-modal model comprises a pre-trained audio encoder and a pre-trained video encoder, the pre-trained audio encoder has the same structure as the pre-trained video encoder, the pre-trained audio encoder is configured to encode the audio data to obtain the audio features, and the pre-trained video encoder is configured to encode the video data to obtain the video features.
2. The data processing method according to claim 1, characterized in that, The encoding of the audio data and the video data by using the pre-trained multi-modal model to obtain the audio features and the video features that are aligned in semantics and time comprises: encoding the audio data by using a first audio encoding path and a second audio encoding path of the pre-trained audio encoder to obtain the audio features; encoding the video data by using a first video encoding path and a second video encoding path of the pre-trained video encoder to obtain the video features.
3. The data processing method according to claim 2, characterized in that, The first audio encoding path has a sampling rate that is less than a preset sampling rate, a running capacity that is greater than or equal to a preset running capacity, and a time resolution limit and a time kernel limit, and the second audio encoding path has a sampling rate that is greater than or equal to the preset sampling rate and a running capacity that is less than the preset running capacity.
4. The data processing method according to claim 2, characterized in that, The first video encoding path has a running frame rate that is less than a preset running frame rate and a processing time step that is less than a preset processing time step, and the second video encoding path has a running frame rate that is greater than or equal to the preset running frame rate and a processing time step that is greater than or equal to the preset processing time step.
5. The data processing method according to any one of claims 1 to 4, characterized in that, The method further comprises: obtaining training data, the training data comprising a plurality of training samples, each training sample comprising audio data and video data; encoding the plurality of training samples by using a to-be-trained multi-modal model to obtain a plurality of feature samples, each feature sample comprising an audio feature sample and a video feature sample; training the to-be-trained multi-modal model according to the plurality of feature samples to obtain the pre-trained multi-modal model.
6. The data processing method according to claim 5, characterized in that, The training of the to-be-trained multi-modal model according to the plurality of feature samples to obtain the pre-trained multi-modal model comprises: calculating a contrastive loss of the to-be-trained multi-modal model according to the plurality of feature samples, the contrastive loss being used to represent a deviation between a plurality of audio feature samples and a plurality of video feature samples; and training the to-be-trained multi-modal model according to the contrastive loss to obtain the pre-trained multi-modal model.
7. The data processing method of claim 1, wherein, The method further comprises: training a to-be-trained generation model according to the audio features and the video features to obtain a pre-trained generation model; The pre-trained generation model is configured to generate a sound video according to a soundless video.
8. The data processing method of claim 1, wherein, The audio data and the video data are obtained by: In response to receiving a first retrieval request for indicating retrieval of a video related to a target sound video, the target sound video is obtained from the first retrieval request; The audio data and the video data are extracted from the target sound video, and the video data is soundless video data; The data processing method further comprises: According to the audio feature and the video feature, a retrieval result for the first retrieval request is obtained by searching in a preset video database.
9. The data processing method of claim 1, wherein, The audio data and the video data are obtained by: In response to receiving a second retrieval request for indicating retrieval of a video related to the audio data, the audio data is obtained from the second retrieval request, and the video data is obtained from a preset video database; The data processing method further comprises: A feature similarity between the audio feature and the video feature is determined; According to the feature similarity, a retrieval result for the second retrieval request is determined.
10. The data processing method of claim 1, wherein, The audio data and the video data are obtained by: In response to receiving a processing request for indicating audio-video synchronization of a to-be-processed sound video, the to-be-processed sound video is obtained from the processing request; The audio data and the video data are extracted from the to-be-processed sound video, and the video data is soundless video data; The data processing method further comprises: A feature similarity between the audio feature and the video feature is determined; According to the feature similarity, the to-be-processed sound video is processed by audio-video synchronization.
11. The data processing method of claim 1, wherein, The audio data and the video data are obtained by: In response to receiving an identification request for indicating content identification of a to-be-identified sound video, the to-be-identified sound video is obtained from the identification request; The audio data and the video data are extracted from the to-be-identified sound video, and the video data is soundless video data; The data processing method further comprises: According to the audio feature and the video feature, an identification result for the identification request is determined.
12. A data processing apparatus, characterized by The data processing apparatus is configured to perform the data processing method of any one of claims 1-11.
13. An electronic device, comprising: Comprises: A processor; A memory for storing processor-executable instructions; The processor is configured to execute the executable instructions to implement the data processing method of any one of claims 1-11.
14. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the data processing method of any one of claims 1-11.
15. A computer program product, characterised in that, The computer program is executed by the processor to implement the data processing method of any one of claims 1-11.