Media data matching and model training method and device, equipment and storage medium
By extracting and calculating semantic feature information of different time segments of media data, and using the matching model to determine the most suitable media data matching, the problem of poor matching effect in the prior art is solved and higher matching accuracy is achieved.
Patent Information
- Application Number
- CN202410016405.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
Existing media data matching methods cannot accurately match the most suitable music clips to the video, resulting in poor matching effects.
By obtaining semantic feature information of different time segments of the first media data and the second media data to be matched, the matching score between the two is calculated using the matching model to determine the most suitable target media data.
The matching accuracy of media data of different modalities is improved and the matching effect of media data across modalities is improved.
Smart Images

Figure CN120256675A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a method, apparatus, device, and storage medium for media data matching and model training. Background Art
[0002] With the rapid development of Artificial Intelligence (AI) technology, AI technology is used in various fields, such as media data matching. In the application scenario of media data matching, music can be matched to a video uploaded by a user, or music can be matched to a video generated by artificial intelligence, or tags can be added to the video based on the music information in the video.
[0003] The current media data matching method cannot match the most suitable music segment to the video, thus resulting in poor media data matching effect. Summary of the Invention
[0004] The present application provides a method, apparatus, device, and storage medium for media data matching and model training, which can achieve accurate matching of music and video, and thus improve the media data matching effect.
[0005] In a first aspect, the present application provides a media data matching method, including:
[0006] Obtain first media data and M second media data to be matched with the first media data, where if the first media data is music data, the second media data is video data, and if the first media data is video data, the second media data is music data, and M is a positive integer;
[0007] For the j-th second media data among the M second media data, extract first media semantic feature information of R different time segments of the first media data and second media semantic feature information of K different time segments of the j-th second media data through the matching model, where R and K are both positive integers, and j is a positive integer less than or equal to M;
[0008] Based on the first media semantic feature information of the R different time segments and the second media semantic feature information of the K different time segments, determine a matching score between the first media data and the j-th second media data;
[0009] Based on the matching scores between each second media data in the M second media data and the first media data, determine a target second media data that matches the first media data from the M second media data.
[0010] Second aspect, the present application provides a method for training a matching model, including:
[0011] Obtain N pairs of training samples, each pair of training samples including a music sample and a video sample, where N is a positive integer;
[0012] For the i-th pair of training samples among the N pairs of training samples, through the matching model, extract the music semantic feature information of P different time segments of the i-th music sample in the i-th pair of training samples, and the video semantic feature information of Q different time segments of the i-th video sample in the i-th pair of training samples, where P and Q are both positive integers, and i is a positive integer less than or equal to N;
[0013] Based on the music semantic feature information of the P different time segments and the video semantic feature information of the Q different time segments, determine the matching score between the i-th music sample and the i-th video sample;
[0014] Based on the matching scores between the music samples and the video samples of each pair of training samples among the N pairs of training samples, determine the model loss of the matching model, and based on the model loss, train the matching model.
[0015] Third aspect, the present application provides a media data matching device, including:
[0016] An acquisition unit, configured to acquire first media data and M pieces of second media data to be matched with the first media data, where if the first media data is music data, then the second media data is video data, and if the first media data is video data, then the second media data is music data, and M is a positive integer;
[0017] A feature extraction unit, configured to, for the j-th piece of second media data among the M pieces of second media data, through the matching model, extract the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th piece of second media data, where R and K are both positive integers, and j is a positive integer less than or equal to M;
[0018] A matching unit, configured to determine the matching score between the first media data and the j-th piece of second media data based on the first media semantic feature information of the R different time segments and the second media semantic feature information of the K different time segments;
[0019] A determination unit for determining, from the M second media data, target second media data that matches the first media data based on the matching scores between each of the M second media data and the first media data.
[0020] In some embodiments, the matching model includes a first media encoding module and a second media encoding module; a feature extraction unit specifically configured to extract first media semantic feature information of R different time segments of the first media data through the first media encoding module; and extract second media semantic feature information of K different time segments of the j-th second media data through the second media encoding module.
[0021] In some embodiments, when the first media data is music data, the feature extraction unit is specifically configured to extract MFCC feature information of the first media data; slice the MFCC feature information once every first time interval to obtain MFCC feature slice information of R different time segments of the first media data; and encode the MFCC feature slice information of the R different time segments through the first media encoding module to obtain first media semantic feature information of the R different time segments of the first media data.
[0022] In some embodiments, the first media encoding module includes a first Transformer unit, and the feature extraction unit is specifically configured to encode the MFCC feature slice information of the R different time segments through the first Transformer unit to obtain music semantic feature information of each time segment among the R different time segments of the first media data.
[0023] In some embodiments, when the second media data is video data, the feature extraction unit is specifically configured to select one video frame every second time interval from the video frames included in the j-th second media data to obtain K video frames; and extract features from the K video frames through the second media encoding module to obtain second media semantic feature information of the j-th second media data in K different time segments.
[0024] In some embodiments, the second media encoding module includes a second Transformer unit, and the feature extraction unit is specifically configured to extract features from the K video frames through the second Transformer unit to obtain video semantic feature information of each video frame among the K video frames, as second media semantic feature information of the j-th second media data in K different time segments.
[0025] In some embodiments, the matching unit is specifically configured to determine the semantic similarity between the first media semantic feature information of each of the R different time segments and the second media semantic feature information of each of the K different time segments as an element in the time matching matrix between the first media data and the j-th second media data, and obtain the time matching matrix between the first media data and the j-th second media data; based on the time matching matrix, determine the matching score between the first media data and the j-th second media data.
[0026] In some embodiments, the matching unit is specifically configured to perform bipartite graph optimal matching processing on the time matching matrix between the first media data and the j-th second media data to obtain the matching score between the first media data and the j-th second media data.
[0027] In some embodiments, the determining unit is further configured to obtain the label information of the target second media data;
[0028] Determine the label information of the target second media data as the label information of the second media data corresponding to the first media data.
[0029] Fourthly, the present application provides a training device for a matching model, including:
[0030] An obtaining unit, configured to obtain N training sample pairs, each training sample pair including a music sample and a video sample, where N is a positive integer;
[0031] A feature extraction unit, configured to, for the i-th training sample pair among the N training sample pairs, extract the music semantic feature information of P different time segments of the i-th music sample in the i-th training sample pair and the video semantic feature information of Q different time segments of the i-th video sample in the i-th training sample pair through the matching model, where P and Q are both positive integers, and i is a positive integer less than or equal to N;
[0032] A matching unit, configured to determine the matching score between the i-th music sample and the i-th video sample based on the music semantic feature information of the P different time segments and the video semantic feature information of the Q different time segments;
[0033] A training unit, configured to determine the model loss of the matching model based on the matching scores between the music samples and the video samples of each training sample pair among the N training sample pairs, and train the matching model based on the model loss.
[0034] In some embodiments, the matching model includes a music encoding module and a video encoding module; a feature extraction unit, specifically configured to extract music semantic feature information of P different time segments of the i-th music sample through the music encoding module; and extract video semantic feature information of Q different time segments of the i-th video sample through the video encoding module.
[0035] In some embodiments, the feature extraction unit is specifically configured to extract Mel Frequency Cepstral Coefficient (MFCC) feature information of the i-th music sample; slice the MFCC feature information once every first time interval to obtain MFCC feature slice information of P different time segments of the i-th music sample; and encode the MFCC feature slice information of the P different time segments through the music encoding module to obtain music semantic feature information of the P different time segments of the i-th music sample.
[0036] In some embodiments, the feature extraction unit is specifically configured to select one video frame every second time interval from the video frames included in the i-th video sample to obtain Q video frames; and extract features from the Q video frames through the video encoding module to obtain video semantic feature information of Q different time segments of the i-th video sample.
[0037] In some embodiments, the matching unit is specifically configured to determine the semantic similarity between the music semantic feature information of each time segment among the P different time segments and the video semantic feature information of each time segment among the Q different time segments as an element in the time matching matrix between the i-th music sample and the i-th video sample, to obtain the time matching matrix between the i-th music sample and the i-th video sample; and determine the matching score between the i-th music sample and the i-th video sample based on the time matching matrix between the i-th music sample and the i-th video sample.
[0038] In some embodiments, the i-th training sample pair further includes the label information of the i-th music sample and the label information of the i-th video sample; the training unit is specifically configured to determine the label similarity between the i-th music sample and the i-th video sample based on the label information of the i-th music sample and the label information of the i-th video sample; and determine the model loss of the matching model based on the matching scores and similarities between the music samples and video samples of each training sample pair among the N training sample pairs.
[0039] In some embodiments, the training unit is specifically configured to determine the label vector of the i-th music sample based on the label information of the i-th music sample and the preset label information; determine the label vector of the i-th video sample based on the label information of the i-th video sample and the preset label information; and determine the similarity between the label vector of the i-th music sample and the label vector of the i-th video sample as the label similarity between the i-th music sample and the i-th video sample.
[0040] In a fifth aspect, a computing device is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the methods in any implementation manner of the first aspect or the second aspect above.
[0041] In a sixth aspect, a chip is provided for implementing the methods in any aspect or its implementation manners of the first aspect above. Specifically, the chip includes: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the methods in any aspect or its implementation manners of the first aspect or the second aspect above.
[0042] In a seventh aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program causes a computer to execute the methods in any aspect or its implementation manners of the first aspect or the second aspect above.
[0043] In an eighth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions cause a computer to execute the methods in any aspect or its implementation manners of the first aspect or the second aspect above.
[0044] In a ninth aspect, a computer program is provided, which when running on a computer, causes the computer to execute the methods in any aspect of the first aspect or the implementation manners of the second aspect above.
[0045] In summary, the present application obtains first media data and M second media data to be matched with the first media data. Among them, if the first media data is music data, the second media data is video data; if the first media data is video data, the second media data is music data. For the j-th second media data among the M second media data, through a matching model, R first media semantic feature information of different time segments of the first media data and K second media semantic feature information of different time segments of the j-th second media data are extracted. Then, based on the R first media semantic feature information of different time segments and the K second media semantic feature information of different time segments, the matching score between the first media data and the j-th second media data is determined. Finally, based on the matching scores between each of the M second media data and the first media data, the target second media data that matches the first media data is determined from the M second media data. As can be seen from the above, when the present application embodiment matches the first media data and the second media data, the semantic feature information of the first media data and the second media data in different time periods is extracted, and then based on the semantic feature information of different time periods, the matching score between the first media data and the second media data is determined. In this way, by explicitly representing the features of cross-modal media data in different time periods, the most suitable second media data can be selected in the time dimension to match the first media data, thereby improving the matching accuracy of media data of different modalities and enhancing the matching effect of cross-modal media data. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 Schematic diagram for matching between different media data;
[0048] Figure 2 Schematic diagram of an implementation environment involved in the embodiment of the present application;
[0049] Figure 3 Schematic flowchart of a matching model training method provided by an embodiment of the present application;
[0050] Figures 4A to 4D Schematic diagram for obtaining training sample pairs;
[0051] Figure 5 Schematic diagram of a structure of a matching model proposed by an embodiment of the present application;
[0052] Figure 6 A schematic diagram of a music encoding module;
[0053] Figure 7 A schematic diagram of a video encoding module;
[0054] Figure 8 Another schematic structural diagram of the matching model proposed in the embodiment of the present application;
[0055] Figure 9 A schematic diagram of a matching module;
[0056] Figure 10 A schematic flowchart of a media data matching method provided in an embodiment of the present application;
[0057] Figures 11A to 11C Schematic diagrams of several application scenarios involved in the embodiment of the present application;
[0058] Figure 12 Another schematic structural diagram of the matching model proposed in the embodiment of the present application;
[0059] Figure 13 A schematic block diagram of a media data matching device provided in an embodiment of the present application;
[0060] Figure 14 A schematic block diagram of a training device for a matching model provided in an embodiment of the present application;
[0061] Figure 15 A schematic block diagram of a computing device provided in the embodiment of the present application. Detailed implementation manners
[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0063] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the description of this application, unless otherwise specified, "a plurality" means two or more than two.
[0064] The technical solution proposed in this application can be applied to technical fields such as artificial intelligence and media creation, and is used to improve the matching accuracy of media data in different modalities and enhance the matching effect of cross-modal media data.
[0065] The following introduces the related concepts involved in the embodiments of this application.
[0066] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, and is a theory, method, technology and application system that can perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0067] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called the large model or the basic model, and can be widely applied to downstream tasks in major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, music processing technology, natural language processing technology, and machine learning / deep learning.
[0068] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.
[0069] Representation: Data used to represent a content, usually a high-dimensional vector or a matrix.
[0070] AI generated content (AIGC) can refer to both the generated content and the technology for producing content.
[0071] Mel Frequency Cepstrum Coefficient (MFCC) feature: Mel frequency is proposed based on the auditory characteristics of the human ear, and it has a non-linear correspondence with the Hz frequency. MFCC is the Hz spectrum feature calculated using this relationship between them. MFCC can be used as the basic feature of sound, containing statistical values at different frequencies of the sound.
[0072] Transformer neural network: A model that uses the attention mechanism to process features. It usually consists of an encoder and a decoder, or can also consist of only a single encoder or decoder.
[0073] Cosine distance: The cosine distance between two vectors can be obtained by calculating the cosine value of the angle between them. The larger the cosine distance, the higher their similarity, and vice versa.
[0074] Bipartite graph: A graph that can be divided into two parts, where there are no edges connecting each part to each other, can be called a bipartite graph.
[0075] Hungarian algorithm: The input of this algorithm is a matrix, and the output is the optimal matching relationship of the bipartite graph corresponding to this matrix.
[0076] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0077] In the embodiments of the present application, artificial intelligence technology is applied to the matching of media data.
[0078] In one example, as Figure 1 shown, in the process of matching media data, it includes the matching of same-modal media data and cross-modal media data. In the process of matching same-modal media data, for example, in a music library, find music 2 that matches a piece of music 1 with unknown tags, and then use the tags of the matching music 2 as the tags of music 1. Another example is to find video 2 that matches a video 1 with unknown tags in a video library, and then use the tags of video 2 as the tags of this video 1. In the process of matching cross-modal media data, it is possible to find a matching video for a video in a music library, or find a matching music for a music in a video library.
[0079] The embodiments of the present application mainly relate to the matching of cross-modal media data, such as matching appropriate video data for music data, or finding appropriate music data for video data, so as to improve the matching accuracy of media data in different modalities and enhance the matching effect of cross-modal media data.
[0080] Existing methods for matching cross-modal media data do not explicitly express the time dimension. However, since both videos and music are continuous processes, the lack of expression of the time dimension will lead to incorrect matching. That is, instead of matching the most appropriate music segment to the video, an incorrect music segment and video segment are selected for matching, thus resulting in a poor matching effect of cross-modal media data.
[0081] To solve the above technical problems, in the matching process of cross-modal media data in the embodiments of the present application, an explicit representation of the time dimension is made. Specifically, the first media data and M second media data to be matched with the first media data are obtained. Among them, if the first media data is music data, the second media data is video data; if the first media data is video data, the second media data is music data. For the j-th second media data among the M second media data, through the matching model, the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th second media data are extracted. Then, based on the first media semantic feature information of R different time segments and the second media semantic feature information of K different time segments, the matching score between the first media data and the j-th second media data is determined. Finally, based on the matching scores between each of the M second media data and the first media data, the target second media data that matches the first media data is determined from the M second media data. As can be seen from the above, when the embodiments of the present application match the first media data and the second media data, the semantic feature information of the first media data and the second media data in different time periods is extracted, and then based on the semantic feature information of different time periods, the matching score between the first media data and the second media data is determined. In this way, by explicitly representing the features of different time periods of cross-modal media data, the most suitable second media data can be selected in the time dimension to match the first media data, thereby improving the matching accuracy of media data of different modalities and enhancing the matching effect of cross-modal media data.
[0082] The implementation environment of the embodiments of the present application will be introduced below.
[0083] Figure 2 FIG. is a schematic diagram of an implementation environment involved in the embodiments of the present application, including a terminal device 101 and a computing device 102.
[0084] As Figure 2 shown, the computing device 102 in the embodiments of the present application includes a matching model. Among them, the terminal device 101 obtains N training sample pairs, each training sample pair includes a music sample and a video sample, N is a positive integer, and sends the N training sample pairs to the computing device 102. The computing device 102 uses the N training sample pairs and adopts the model training method provided by the embodiments of the present application to train the matching model.
[0085] Exemplarily, the computing device 102 obtains N pairs of training samples, where each pair of training samples includes a music sample and a video sample, and N is a positive integer. For the i-th pair of training samples among the N pairs of training samples, through the matching model, P different time segments of music semantic feature information of the i-th music sample in the i-th pair of training samples and Q different time segments of video semantic feature information of the i-th video sample in the i-th pair of training samples are extracted, where P and Q are both positive integers, and i is a positive integer less than or equal to N. Based on the music semantic feature information of P different time segments and the video semantic feature information of Q different time segments, the matching score between the i-th music sample and the i-th video sample is determined. Based on the matching scores between the music samples and the video samples of each pair of training samples among the N pairs of training samples, the model loss of the matching model is determined, and based on the model loss, the matching model is trained. That is to say, in the embodiments of the present application, when matching the first media data and the second media data, the semantic feature information of the first media data and the second media data in different time periods is extracted, and then based on the semantic feature information of different time periods, the matching score between the first media data and the second media data is determined. In this way, by explicitly representing the features of cross-modal media data in different time periods, the most suitable second media data can be selected in the time dimension to match the first media data, thereby improving the matching accuracy of media data of different modalities and enhancing the matching effect of cross-modal media data.
[0086] In some embodiments, as Figure 1 shown, this application scenario may further include a database 103, and the database 103 includes historical data, such as historical media data. In the embodiments of the present application, the terminal device 101 is communicatively connected to the database 103 and can write data into the database 103, and the computing device 102 is also communicatively connected to the database 102 and can read data from the database 103. In one example, during the model training process of the embodiments of the present application, when the computing device 102 trains the model, it obtains historical data from the database 103 as training samples. Then, the computing device 102 uses the training samples to train the matching model to obtain a trained matching model. Optionally, the computing device 102 can save the trained matching model in the computing device 102. Optionally, the computing device 102 can send the trained matching model to the terminal device 101 for saving.
[0087] In the embodiments of the present application, the media data matching method can be implemented by the terminal device 101, or by the computing device 102, or by a system composed of the terminal device 101 and the computing device 102.
[0088] In some embodiments, when the media data matching method of the embodiments of the present application is executed by the terminal device 101, the terminal device 101 obtains the trained matching model from the computing device 102. In this way, when performing media data matching, the terminal device 101 obtains the first media data and the M second media data to be matched with the first media data. For the j-th second media data among the M second media data, through the matching model, the terminal device 101 extracts the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th second media data. Then, the terminal device 101 determines the matching score between the first media data and the j-th second media data based on the first media semantic feature information of R different time segments and the second media semantic feature information of K different time segments. Finally, the terminal device 101 determines the target second media data that matches the first media data from the M second media data based on the matching scores between each of the M second media data and the first media data.
[0089] In some embodiments, when the media data matching method of the embodiments of the present application is executed by the computing device 102, the terminal device 101 sends the obtained first media data and the M second media data to be matched with the first media data to the computing device 102. For the j-th second media data among the M second media data, the computing device 102 uses the matching model stored in itself to extract the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th second media data. Then, the computing device 102 determines the matching score between the first media data and the j-th second media data based on the first media semantic feature information of R different time segments and the second media semantic feature information of K different time segments. Finally, the computing device 102 determines the target second media data that matches the first media data from the M second media data based on the matching scores between each of the M second media data and the first media data.
[0090] The embodiments of the present application do not limit the specific type of the terminal device 101. In some embodiments, the terminal device 101 may include but is not limited to: mobile phones, computers, intelligent music interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, wearable smart devices, medical devices, and so on. The device is often configured with a display device, and the display device may also be a display, a display screen, a touch screen, and so on. The touch screen may also be a touch panel, a touch screen panel, and so on.
[0091] In some embodiments, the computing device 102 is a terminal device with data processing capabilities, such as a mobile phone, a computer, a smart music interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a wearable smart device, a medical device, and so on.
[0092] In some embodiments, the computing device 102 is a server. There may be one or more such servers. When there are multiple servers, at least two servers are used to provide different services, and / or at least two servers are used to provide the same service, such as providing the same service in a load balancing manner. The embodiments of the present application do not limit this. Among them, the above-mentioned server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server 102 may also become a node of the blockchain.
[0093] In the embodiments of the present application, the terminal device 101 and the computing device 102 may be directly or indirectly connected through wired communication or wireless communication. The present application does not limit this here.
[0094] It should be noted that the implementation environment of the embodiments of the present application includes but is not limited to Figure 2 as shown.
[0095] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0096] First, the training process of the matching model proposed in the embodiments of the present application will be introduced.
[0097] Figure 3 is a schematic flowchart of a method for training a matching model provided in an embodiment of the present application. The execution subject of the embodiments of the present application is a device with the function of training a model, such as a model training device. In some embodiments, the model training device may be Figure 2 the computing device in Figure 2 or the terminal device in Figure 2 or a system composed of a computing device and a terminal device in
[0098] For the sake of convenience of description, the embodiments of the present application will be described by taking the execution subject as a computing device as an example. Figure 3 The training process of the matching model will be introduced below in combination with
[0099] As Figure 3 shown, the matching model training process in the embodiments of this application includes:
[0100] S101. Obtain N training sample pairs.
[0101] Each training sample pair includes a music sample and a video sample, and N is a positive integer.
[0102] In the embodiments of this application, the matching model is mainly used for cross-modal media data matching. For example, it is used to find matching video data for music data, or find matching music data for video data. Therefore, when training the matching model, multiple training sample pairs are obtained, and each training sample pair includes a music sample and a video sample.
[0103] Both the music sample and the video sample included in each training sample pair include label information. That is to say, each training sample pair includes a music sample and its label information, and a video sample and its label information.
[0104] In some embodiments, the above N training sample pairs are all positive samples, that is, the music sample and the video sample included in each of the N training sample pairs have a matching relationship.
[0105] In some embodiments, the above N training sample pairs include positive samples and negative samples, that is, the music sample and the video sample included in some of the N training sample pairs have a matching relationship, and the music sample and the video sample included in some of the N training sample pairs do not have a matching relationship.
[0106] In some related technologies, annotators can adopt the Figures 4A to 4D shown annotation method to annotate a large number of music and video data pairs with matching relationships as training sample pairs. As Figure 4A shown, after watching the video, the annotator manually searches for matching music and establishes the corresponding relationship between the video and the music as a training sample pair. And / or, as Figure 4B shown, after listening to the music, the annotator manually searches for matching video and establishes the corresponding relationship between the video and the music as a training sample pair. And / or, as Figure 4C shown, the annotator directly uses the videos with music on the Internet and takes them and the corresponding matching relationships as training sample pairs. And / or, as Figure 4D shown, the annotator uses the videos with music on the Internet, and after secondary confirmation, selects a part of them to establish a matching relationship as a training sample pair.
[0107] However, Figure 4A finding suitable music by watching the video, orFigure 4B In the prior art, when finding a suitable video by listening to music, it is necessary for humans to find a suitable matching object from a vast amount of content, which requires a very high level of professionalism and proficiency of the annotators and has low annotation efficiency. Figure 4C In the prior art, using the matching relationship between videos and music on the Internet as training data pairs faces the problem of large data noise, which increases the difficulty of model training. Figure 4D In the prior art, when annotators confirm network data, there are still problems such as difficulty in unifying annotation standards, low efficiency, and large data noise.
[0108] To solve the above technical problems, in the embodiments of the present application, instead of directly establishing a training sample pair that matches music and videos, "content tags" are used as an intermediate medium to align music data and video data. The annotation of content tags is more objective and the data is easier to obtain, thus simultaneously solving the problems of low annotation efficiency, difficulty in unifying annotation, and large data noise. That is to say, in the embodiments of the present application, video data is tagged separately, and at the same time, music data is tagged separately. In this way, using tags as a bridge to construct video data and music data with a matching relationship is more efficient than directly annotating video data and music data with a matching relationship. This is because when directly annotating the matching relationship between music data and video data, only one pair of music data and video data with a matching relationship can be constructed in one annotation, while in the embodiments of the present application, by separately tagging music data and video data, video data and music data with similar tags can be formed into a matching relationship based on the tags, which can improve the construction speed and accuracy of training sample pairs, and thus improve the training speed and workload of the model.
[0109] That is to say, in the embodiments of the present application, by separately tagging music samples and video samples, and then based on the tag information of music samples and the tag information of video data, N training sample pairs are constructed, which can greatly reduce the difficulty of obtaining training data, thereby reducing the training complexity of the model and improving the training speed of the model. In addition, in the embodiments of the present application, by separately tagging music samples and video samples instead of directly annotating whether music samples and video samples match, on the one hand, the reusability of data is increased. Any set of video samples and music samples can use the tag information to calculate the similarity, while the method of directly annotating whether they match only generates a small number of positive samples. On the other hand, separately tagging music samples and video samples is a relatively objective task, which is easier for annotators to learn and master, while annotating whether music samples and video samples match is highly subjective and prone to introducing noise of inconsistent annotation in the data, thus increasing the difficulty of model training.
[0110] The embodiments of the present application do not limit the specific method of tagging video data.
[0111] In a possible implementation, the previously labeled video data can be reused to obtain the label information of the video data.
[0112] In a possible implementation, an off-the-shelf video tagging system can be used to tag the video data on the network, and the label information of the video data can be obtained.
[0113] In a possible implementation, the video data can be manually tagged to obtain the label information of the video data.
[0114] The embodiments of the present application do not limit the specific method of tagging music data.
[0115] In a possible implementation, the music data can be manually tagged to obtain the music data of the music data.
[0116] In a possible implementation, an existing text processing model can be used to perform text analysis on the lyrics of the music data to obtain the label information of the music data.
[0117] In the embodiments of the present application, after obtaining the label information of different music data and the label information of different video data through the above methods, N training sample pairs can be constructed based on the label information of the music data and the label information of the video data.
[0118] In some embodiments, if at least one positive training sample pair is included in the N training sample pairs of the embodiments of the present application, that is, one music data and video data included in each positive training sample pair of the at least one positive training sample pair have a matching relationship. At this time, the computing device can obtain the at least one positive training sample pair at least through the following manner:
[0119] In an example, the terminal device can determine the similarity between the label information of different music data and the label information of different video data based on the label information of different music data and the label information of different video data, and then determine the music data and video data with a similarity greater than a preset value as a pair of music data and video data with a matching relationship, as a positive training sample pair. In this way, the terminal device can obtain at least one positive training sample pair and send the obtained at least one positive training sample pair to the computing device.
[0120] In one example, the terminal device may send different music data with tag information and video data with tag information to the computing device. In this way, the computing device may determine the similarity between the tag information of different music data and the tag information of different video data based on the tag information of different music data and the tag information of different video data, and then determine the music data and video data with a similarity greater than a preset value as a pair of music data and video data with a matching relationship, as a positive training sample pair, and thus at least one positive training sample pair may be obtained.
[0121] In some embodiments, if at least one negative training sample pair is included in the N training sample pairs of the embodiments of the present application, that is, one music data and video data included in each negative training sample pair of the at least one negative training sample pair do not have a matching relationship. At this time, the computing device may obtain the at least one negative training sample pair at least through the following manner:
[0122] In one example, the terminal device may determine the similarity between the tag information of different music data and the tag information of different video data based on the tag information of different music data and the tag information of different video data, and then determine the music data and video data with a similarity less than a preset value as a pair of music data and video data without a matching relationship, as a negative training sample pair. In this way, the terminal device may obtain at least one negative training sample pair and send the obtained at least one negative training sample pair to the computing device.
[0123] In one example, the terminal device may send different music data with tag information and video data with tag information to the computing device. In this way, the computing device may determine the similarity between the tag information of different music data and the tag information of different video data based on the tag information of different music data and the tag information of different video data, and then determine the music data and video data with a similarity less than a preset value as a pair of music data and video data without a matching relationship, as a negative training sample pair, and thus at least one negative training sample pair may be obtained.
[0124] After the computing device obtains N training sample pairs, it executes the steps of S102 as follows.
[0125] S102: For the i-th training sample pair among the N training sample pairs, extract the music semantic feature information of P different time segments of the i-th music sample in the i-th training sample pair and the video semantic feature information of Q different time segments of the i-th video sample in the i-th training sample pair through a matching model.
[0126] Wherein, both P and Q are positive integers, and i is a positive integer less than or equal to N.
[0127] In the embodiments of the present application, after the computing device obtains N training sample pairs, it uses these N training sample pairs to train the matching model. Specifically, the process of the computing device using each of the N training sample pairs to train the matching model is basically the same. That is to say, the process of the matching model processing each of the N training sample pairs is basically the same. For the convenience of description, the i-th training sample pair among the N training sample pairs is taken as an example for illustration.
[0128] Specifically, the computing device processes each of the N training sample pairs through the matching model. For example, for the i-th training sample pair, the i-th training sample pair includes a music sample and a video sample. For the convenience of description, the music sample included in the i-th training sample pair is denoted as the i-th music sample, and the video sample included in the i-th training sample pair is denoted as the i-th video sample. The computing device inputs the i-th music sample and the i-th video sample into the matching model, and the matching model first extracts the music semantic feature information of P different time segments of the i-th music sample, and extracts the video semantic feature information of Q different time segments of the i-th video sample.
[0129] In the embodiments of the present application, since the i-th music sample is music data of a certain duration, and the i-th video sample is also video data of a certain duration. In order to improve the judgment accuracy of the matching relationship between the i-th music sample and the i-th video sample, in the embodiments of the present application, by extracting the music semantic feature information of different time segments of the i-th music sample, and extracting the video semantic feature information of different time segments of the i-th video sample, in this way, by refining the matching granularity and explicitly representing the semantic feature information of the music data and video data of different time segments, the matching accuracy of media data of different modalities can be improved, and the matching effect of cross-modal media data can be enhanced.
[0130] The following introduces the specific manner in which the computing device extracts the music semantic feature information of P different time segments of the i-th music sample and the video semantic feature information of Q different time segments of the i-th video sample through the matching model.
[0131] In some embodiments, the matching model includes a feature extraction module, and the feature extraction module can extract the music semantic feature information of the music data of P different time segments in the i-th music sample, and the video semantic feature information of the video data of Q different time segments in the i-th video sample.
[0132] In some embodiments, such as Figure 5As shown, the matching model of the embodiment of the present application includes a music encoding module and a video encoding module. The music encoding module is used to extract the music semantic feature information of different time segments of the input music sample. The video encoding module is used to extract the video semantic feature information of different time segments of the input video sample.
[0133] Based on the network structure of the matching model shown above, step S102 includes the following steps S102-A and S102-B: Figure 5 Based on the network structure of the matching model shown above, step S102 includes the following steps S102-A and S102-B:
[0134] S102-A: Extract the music semantic feature information of P different time segments of the i-th music sample through the music encoding module;
[0135] S102-B: Extract the video semantic feature information of Q different time segments of the i-th video sample through the video encoding module.
[0136] In this implementation, as Figure 5 shown, the computing device inputs the i-th music sample and the i-th video sample into the matching model. In this way, through the music encoding module in the matching model, the i-th music sample is processed to extract the music semantic feature information of P different time segments of the i-th music sample. Through the video encoding module in the matching model, the i-th video sample is processed to extract the video semantic feature information of Q different time segments of the i-th video sample.
[0137] Next, the specific process of the computing device extracting the music semantic feature information of P different time segments of the i-th music sample through the music encoding module will be introduced.
[0138] The embodiment of the present application does not limit the specific manner in which the computing device extracts the music semantic feature information of P different time segments of the i-th music sample through the music encoding module.
[0139] In some embodiments, the computing device divides the i-th music sample into P music segments in chronological order. For example, if the duration of the i-th music sample is 10 minutes, then the music data of each minute in the i-th music sample can be divided into a music segment. Thus, the 10-minute long i-th music sample can be divided into 10 music segments of 1 minute in length. Then, through the music encoding module, the music semantic feature information of each of these 10 music segments is extracted, and thus 10 music semantic feature information of different time segments can be obtained. At this time, P = 10. Exemplarily, the above division durations can be the same, such as all 1 minute, or different. For example, the music data of 1 minute in the i-th music sample can be divided into a music segment, or the music data of 2 minutes can be divided into a music segment.
[0140] In some embodiments, the above S102-A includes the following steps of S102-A1 to S102-A3:
[0141] S102-A1. Extract the MFCC feature information of the i-th music sample;
[0142] S102-A2. Slice the MFCC feature information once every first time interval to obtain the MFCC feature slice information of P different time segments of the i-th music sample;
[0143] S102-A3. Through the music encoding module, perform encoding processing on the MFCC feature slice information of P different time segments to obtain the music semantic feature information of P different time segments of the i-th music sample.
[0144] In this implementation manner, when the computing device extracts the music semantic feature information of P different time segments of the i-th music sample through the music encoding module, it first extracts the MFCC feature information of the i-th music sample. For example, pre-emphasize the i-th music sample to enhance the high-frequency information in the i-th music sample. Then, frame the pre-emphasized i-th music sample. Next, multiply the music signal of each frame by a window function. Finally, perform Fourier transform and other processing on the windowed music data to obtain the MFCC feature information of the i-th music sample. The specific process can refer to the relevant records of MFCC and will not be elaborated here.
[0145] After the computing device extracts the MFCC feature information of the i-th music sample, it slices the MFCC feature information of the i-th music sample according to the time information. The computing device slices the MFCC feature information once every first time interval preset, and obtains the MFCC feature slice information of P different time segments of the i-th music sample. As Figure 6 shown, the computing device slices the MFCC feature information of the i-th music sample once every first time interval, that is, summarizes the MFCC feature information within each time segment, so as to obtain the MFCC feature slice information of 5 (of course, it can also be other values) different time segments of the i-th music sample.
[0146] The present application embodiment does not limit the specific value of the first time interval. For example, the first time interval is 5 seconds, that is, the computing device slices the MFCC feature information of the i-th music sample once every 5 seconds, and can obtain the MFCC feature slice information of P different time segments of the i-th music sample.
[0147] As Figure 6As shown, the music encoding module of the embodiment of the present application includes a first Transformer unit, which is used to extract the semantic feature information of the input MFCC feature slice information. In this way, the computing device obtains the MFCC feature slice information of P different time segments of the i-th music sample, inputs the MFCC feature slice information of P different time segments of the i-th music sample into the first Transformer unit for semantic feature extraction, and obtains the semantic feature information corresponding to each time segment of the MFCC feature slice information of P different time segments. Furthermore, the semantic feature information corresponding to each of the P different time segments of the MFCC feature slice information obtained is determined as the music semantic feature information of P different time segments of the i-th music sample.
[0148] In some embodiments, the above first Transformer unit can also be replaced by other unit modules with semantic feature extraction functions, and the embodiments of the present application do not limit this.
[0149] The embodiments of the present application do not limit the specific size of each music semantic feature information among the music semantic feature information of P different time segments of the i-th music sample. In one example, each music semantic feature information among the music semantic feature information of P different time segments of the i-th music sample is 1024-dimensional feature information.
[0150] Next, in S102-B, the computing device extracts the video semantic feature information of Q different time segments of the i-th video sample through the video encoding module.
[0151] The embodiments of the present application do not limit the specific manner in which the computing device extracts the video semantic feature information of Q different time segments of the i-th video sample through the video encoding module.
[0152] In some embodiments, the computing device divides the i-th video sample into Q music segments in chronological order. For example, if the duration of the i-th video sample is 10 minutes, then the video data of each minute in the i-th video sample can be divided into a video segment. Furthermore, the 10-minute long i-th video sample can be divided into 10 video segments each with a length of 1 minute. Then, through the video encoding module, the video semantic feature information of each of these 10 video segments is extracted, and thus 10 video semantic feature information of different time segments can be obtained. At this time, Q = 10. Exemplarily, the above division durations can be the same, for example, all 1 minute, or different. For example, the video data of 1 minute in the i-th video sample can be divided into a video segment, or the video data of 2 minutes can be divided into a video segment.
[0153] In some embodiments, the above S102-B includes the following steps of S102-B1 to S102-B2:
[0154] S102-B1. Select one video frame every second time interval from the video frames included in the i-th video sample to obtain Q video frames;
[0155] S102-B2. Extract feature information of the Q video frames through a video encoding module to obtain video semantic feature information of Q different time segments of the i-th video sample.
[0156] In this implementation manner, when the computing device extracts the video semantic feature information of Q different time segments of the i-th video sample through the video encoding module, as Figure 7 shown, first extract the Q video frames of the i-th video sample.
[0157] For example, the computing device selects one video frame every second time interval from the video frames included in the i-th video sample to obtain Q video frames. By way of example, as Figure 7 shown, assume that the i-th video sample includes 1000 video frames and the second time interval is 5 seconds. In this way, the computing device first selects the 1st video frame of the i-th video sample, and then, after an interval of 5 seconds, selects one video frame from the 1000 video frames included in the i-th video sample, so that 5 video frames can be selected.
[0158] Next, the computing device inputs the selected Q video frames into the video encoding module, and the video encoding module extracts feature information of these Q video frames to obtain the semantic feature information of each of these Q video frames. Since these Q video frames correspond one-to-one to Q different time segments, the semantic feature information of these Q video frames is determined as the video semantic feature information of Q different time segments of the i-th video sample.
[0159] In one example, as Figure 7 shown, the video encoding module of the embodiment of the present application includes a second Transformer unit, and the second Transformer unit is used to extract the semantic feature information of the input video frames. In this way, after the computing device obtains the Q video frames of the i-th video sample, it inputs these Q video frames into the second Transformer unit for semantic feature extraction to obtain the semantic feature information of each of the Q video frames, and further determines the semantic feature information of each of the obtained Q video frames as the video semantic feature information of Q different time segments of the i-th video sample.
[0160] In some embodiments, the above-mentioned second Transformer unit can also be replaced by other unit modules with semantic feature extraction functions, and the embodiments of the present application do not limit this.
[0161] In the above, the computing device determines, in the i-th training sample pair among the N training sample pairs, the music semantic feature information of P different time segments of the i-th music sample and the video semantic feature information of Q different time segments of the i-th video sample.
[0162] Referring to the above method, the computing device can determine, in the music sample and the video sample included in each of the N training sample pairs, the music semantic feature information of different time segments of the music sample and the video semantic feature information of different time segments of the video sample. Then, the computing device performs the following step S103.
[0163] S103: Based on the music semantic feature information of P different time segments and the video semantic feature information of Q different time segments, determine the matching score between the i-th music sample and the i-th video sample.
[0164] Based on the above steps, after the computing device determines, in the music sample and the video sample included in each of the N training sample pairs, the music semantic feature information of different time segments of the music sample and the video semantic feature information of different time segments of the video sample, based on the music semantic feature information of different time segments of the music sample and the video semantic feature information of different time segments of the video sample included in each training sample pair, the computing device determines the matching score between the music sample and the video sample included in each training sample pair.
[0165] In the embodiments of the present application, the specific manner of determining the matching score between the music sample and the video sample included in each of the N training sample pairs is the same. In the embodiments of the present application, the i-th training sample among the N training samples is continued as an example for illustration.
[0166] In the embodiments of the present application, based on the above steps, the computing device determines, in the i-th music sample and the i-th video sample of the i-th training sample pair, the music semantic feature information of P different time segments of the i-th music sample and the video semantic feature information of Q different time segments of the i-th video sample. Then, the computing device determines the matching score between the i-th music sample and the i-th video sample based on the music semantic feature information of P different time segments and the video semantic feature information of Q different time segments.
[0167] The embodiments of the present application do not limit the specific manner in which the computing device determines the matching score between the i-th music sample and the i-th video sample based on the music semantic feature information of P different time segments of the i-th music sample and the video semantic feature information of Q different time segments of the i-th video sample.
[0168] In some possible implementation manners, the computing device forms a feature matrix 1 from the music semantic feature information of P different time segments of the i-th music sample, and forms a feature matrix 2 from the video semantic feature information of Q different time segments of the i-th video sample. Then, the similarity between the feature matrix 1 and the feature matrix 2 is determined, and further, the similarity between the feature distance 1 and the feature matrix 2 is determined as the matching score between the i-th music sample and the i-th video sample in the i-th training sample pair.
[0169] In some possible implementation manners, the computing device can determine the matching score between the i-th music sample and the i-th video sample through the following steps S103-A and S103-B:
[0170] S103-A: Determine the semantic similarity between the music semantic feature information of each time segment among the P different time segments and the video semantic feature information of each time segment among the Q different time segments, and use it as an element in the time matching matrix between the i-th music sample and the i-th video sample, to obtain the time matching matrix between the i-th music sample and the i-th video sample;
[0171] S103-B: Based on the time matching matrix between the i-th music sample and the i-th video sample, determine the matching score between the i-th music sample and the i-th video sample.
[0172] In this implementation manner, the computing device determines the similarity between the music semantic feature information of each different time segment in the i-th music sample and the video semantic feature information of each different time segment in the i-th video sample. Further, based on the similarity between the music semantic feature information of each different time segment in the i-th music sample and the video semantic feature information of each different time segment in the i-th video sample, the matching score between the i-th music sample and the i-th video sample is determined.
[0173] Specifically, the computing device determines the music semantic feature information of each of the P different time segments, and respectively determines the semantic similarity between the music semantic feature information of each of the P different time segments and the video semantic feature information of each of the Q different time segments, as an element in the time matching matrix between the i-th music sample and the i-th video sample, so as to obtain a time matching matrix of size P×Q. For example, the similarity between the music semantic feature information of the k-th time segment in the i-th music sample and the video semantic feature information of the j-th time segment in the i-th video sample is used as the element value at the k, j position in the time matching matrix. Referring to this method, a time matching matrix of size P×Q can be obtained.
[0174] Next, the computing device determines the matching score between the i-th music sample and the i-th video sample based on the time matching matrix between the i-th music sample and the i-th video sample.
[0175] The embodiments of the present application do not limit the specific manner in which the computing device determines the matching score between the i-th music sample and the i-th video sample based on the time matching matrix between the i-th music sample and the i-th video sample.
[0176] In a possible implementation manner, the computing device determines the sum value obtained by adding each element value in the time matching matrix between the i-th music sample and the i-th video sample as the matching score between the i-th music sample and the i-th video sample.
[0177] In a possible implementation manner, the computing device determines the average value obtained by adding each element value in the time matching matrix between the i-th music sample and the i-th video sample as the matching score between the i-th music sample and the i-th video sample.
[0178] In some embodiments, as Figure 8 shown, the matching model of the embodiments of the present application includes a matching module. The computing device extracts the music semantic feature information of P different time segments of the i-th music sample through a music encoding module, and extracts the video semantic feature information of Q different time segments of the i-th video sample. Next, through the matching module, the music semantic feature information of P different time segments of the i-th music sample and the video semantic feature information of Q different time segments of the i-th video sample are subjected to matching processing to obtain the matching score between the i-th music sample and the i-th video sample.
[0179] In a possible implementation, the computing device, through a matching module, performs bipartite graph optimal matching on the music semantic feature information of P different time segments of the i-th music sample and the video semantic feature information of Q different time segments of the i-th video sample, and obtains the matching score between the i-th music sample and the i-th video sample through the Hungarian algorithm.
[0180] Exemplarily, as Figure 9 shown, the computing device inputs the music semantic feature information of P different time segments of the i-th music sample and the video semantic feature information of Q different time segments of the i-th video sample into the matching module. The matching module determines the semantic similarity between the music semantic feature information of each time segment among the P different time segments and the video semantic feature information of each time segment among the Q different time segments as an element in the time matching matrix between the i-th music sample and the i-th video sample, so as to obtain a time matching matrix of size P×Q. Then, through the bipartite graph optimal matching algorithm, the time matching matrix between the i-th music sample and the i-th video sample is processed to obtain the matching score between the i-th music sample and the i-th video sample.
[0181] The above specifically introduces the process of the computing device determining the matching score between the i-th music sample and the i-th video sample in the i-th training sample pair. Referring to the above method, the computing device can determine the matching scores between the music samples and the video samples in each of the N training sample pairs. Then, the computing device executes the following step S104.
[0182] S104: Based on the matching scores between the music samples and the video samples in each of the N training sample pairs, determine the model loss of the matching model, and train the matching model based on the model loss.
[0183] After the computing device determines the matching scores between the music samples and the video samples in each of the N training sample pairs based on the above steps, it determines the model loss of the matching model based on the matching scores between the music samples and the video samples in each of the N training sample pairs.
[0184] The embodiments of the present application do not limit the specific manner in which the computing device determines the model loss of the matching model based on the matching scores between the music samples and the video samples in each of the N training sample pairs.
[0185] In some embodiments, the computing device determines the negative of the sum of the matching scores between the music samples and the video samples in each of the N training sample pairs as the model loss of the matching model.
[0186] In some embodiments, as described above, each of the N training sample pairs in the embodiments of the present application includes a music sample and the label information of the music sample, and a video sample and the label information of the video sample. At this time, the above S104 includes the following steps of S104-A and S104-B:
[0187] S104-A. Based on the label information of the i-th music sample and the label information of the i-th video sample, determine the label similarity between the i-th music sample and the i-th video sample;
[0188] S104-B. Based on the matching scores and similarities between the music samples and the video samples in each of the N training sample pairs, determine the model loss of the matching model.
[0189] In the embodiments of the present application, when calculating the device determines the model loss of the matching model, it not only determines the matching scores between the music samples and the video samples included in each of the N training sample pairs based on the above steps, but also determines the similarities between the label information of the music samples and the label information of the video samples included in each of the N training sample pairs.
[0190] In the embodiments of the present application, the specific ways for the calculating device to determine the similarities between the label information of the music samples and the label information of the video samples included in each of the N training sample pairs are basically the same. For the convenience of description, here, the similarity between the label information of the i-th music sample and the label information of the i-th video sample included in the i-th training sample pair among the N training sample pairs is taken as an example for illustration.
[0191] The embodiments of the present application do not limit the specific ways for the calculating device to determine the similarity between the label information of the i-th music sample and the label information of the i-th video sample.
[0192] In a possible implementation manner, the calculating device can determine the embedding representation of the label information of the i-th music sample and the embedding representation of the label information of the i-th video sample by looking up a dictionary, and then determine the similarity between the embedding representation of the label information of the i-th music sample and the embedding representation of the label information of the i-th video sample. For example, the cosine similarity between the embedding representation of the label information of the i-th music sample and the embedding representation of the label information of the i-th video sample is determined as the similarity between the label information of the i-th music sample and the label information of the i-th video sample.
[0193] In a possible implementation, the computing device determines the label vector of the i-th music sample based on the label information of the i-th music sample and the preset label information; determines the label vector of the i-th video sample based on the label information of the i-th video sample and the preset label information; and determines the similarity between the label vector of the i-th music sample and the label vector of the i-th video sample as the label similarity between the i-th music sample and the i-th video sample.
[0194] For example, assume that the preset label information includes 5 labels. Compare the labels included in the label information of the i-th music sample with these 5 labels. If the label at the corresponding position in the label information of the i-th music sample is the same as the label at the corresponding position among these 5 labels, then set the vector value at the corresponding position in the label vector of the i-th music sample to 1, otherwise set it to 0. For example, the 5 labels included in the preset label information are: movie, emotion, workplace, office, inspiration. If the label information of the i-th music sample includes: emotion, entertainment, at this time, the label vector of the i-th music sample can be determined as: 0, 1, 0, 0, 0. Similarly, compare the labels included in the label information of the i-th video sample with these 5 labels. If the label at the corresponding position in the label information of the i-th video sample is the same as the label at the corresponding position among these 5 labels, then set the vector value at the corresponding position in the label vector of the i-th video sample to 1, otherwise set it to 0. For example, the label information of the i-th video sample includes: movie, emotion, workplace, at this time, the label vector of the i-th video sample can be determined as: 1, 1, 1, 0, 0.
[0195] Next, determine the similarity between the label vector of the i-th music sample and the label vector of the i-th video sample. For example, determine the similarity between the label vector 0, 1, 0, 0, 0 of the i-th music sample and the label vector 1, 1, 1, 0, 0 of the i-th video sample, such as the distance similarity. Further obtain the similarity between the label vector of the i-th music sample and the label vector of the i-th video sample.
[0196] Referring to the above method, the computing device can determine the similarity between the label information of the music sample and the label information of the video sample included in each training sample pair among the N training sample pairs.
[0197] Finally, the computing device determines the model loss of the matching model based on the matching score and similarity between the music sample and the video sample of each training sample pair among the N training sample pairs. For example, the computing device determines the Euclidean distance between the matching score and the similarity between the music sample and the video sample of each training sample pair among the N training samples, and determines the loss of the matching model based on this Euclidean distance.
[0198] Next, the computing device calculates the loss of the matching model and updates the parameters in the matching model. For example, using this loss and adopting the process of backpropagation by the chain rule, the gradient of each network parameter can be obtained in sequence. After that, the gradient descent method can be used to adjust the parameter values of the network. Then, it is determined whether the updated matching model meets the training end condition. Exemplarily, the model training end condition includes whether the number of training times of the model reaches a preset number, or whether the loss of the model reaches a preset loss. If the matching model after parameter update does not meet the training end condition, then a new batch of training sample pairs is selected, and the above method is sampled to perform iterative training on the matching model, repeating the above steps until the model training end condition is reached, and finally the trained matching model is obtained.
[0199] The training method of the matching model provided by the embodiments of the present application obtains N training sample pairs, and each training sample pair includes a music sample and a video sample; for the i-th training sample pair among the N training sample pairs, through the matching model, P different time segment music semantic feature information of the i-th music sample in the i-th training sample pair and Q different time segment video semantic feature information of the i-th video sample in the i-th training sample pair are extracted; based on the P different time segment music semantic feature information and the Q different time segment video semantic feature information, the matching score between the i-th music sample and the i-th video sample is determined; based on the matching scores between the music samples and video samples of each training sample pair among the N training sample pairs, the model loss of the matching model is determined, and the matching model is trained based on the model loss. When training the matching model in the embodiments of the present application, the semantic feature information of the music samples and video samples included in each training sample pair among the N training sample pairs is extracted in different time segments, and then based on the semantic feature information in different time periods, the model loss is determined. In this way, by explicitly representing the features of cross-modal media data in different time segments, the matching model can learn the music data and video data in different time segments, thereby improving the matching accuracy of media data in different modalities and enhancing the matching effect of cross-modal media data.
[0200] As described above in conjunction with Figures 3 to 9 , the embodiments of the model training method of the present application are described in detail. Next, the media data matching method provided by the embodiments of the present application is introduced.
[0201] Figure 10 It is a schematic flowchart of the media data matching method provided by an embodiment of the present application. The execution subject of the embodiment of the present application is a device with a matching function, such as a media data matching device. In some embodiments, the media data matching device can be Figure 2 the computing device in Figure 2 or the terminal device inFigure 2 A system composed of a computing device and a terminal device. For the convenience of description, in the embodiments of the present application, the execution subject is taken as an example of a computing device for illustration. The computing device and the computing device for model training described above may be the same device or different devices. The embodiments of the present application do not limit this.
[0202] As Figure 10 shown, the method of the embodiments of the present application includes the following steps:
[0203] S201. Obtain first media data and M second media data to be matched with the first media data.
[0204] Among them, if the first media data is music data, the second media data is video data; if the first media data is video data, the second media data is music data, and M is a positive integer.
[0205] The media data matching method of the embodiments of the present application is used to match media data of different modalities. The media data matching method of the embodiments of the present application can be applied to at least the following scenarios:
[0206] Scenario 1: Use music information to label and classify videos by retrieval.
[0207] As Figure 11A shown, for video data 1 including music data, in order to classify and label the video data 1, the music data included in the video data 1 is used as the first media data, and in the video database, video data matching the music data is searched for. Since the video data included in the video database all includes label information, the label information of the video data 2 matching the music data in the video database can be determined as the label information of the video data 1 where the music data is located. Or, the classification of the video data 2 is determined as the classification of the video data 1. At this time, the video data included in the video database can be understood as M second media data.
[0208] Scenario 2: Recommend appropriate music data for video data uploaded by users.
[0209] Specifically, as Figure 11BAs shown, to recommend appropriate music data for the video data uploaded by the user, the video data uploaded by the user can be determined as the first media data at this time. The computing device can select one music data from the multiple music data included in the music library as the music data matching the video data. Or, select several candidate music data matching the video data from the multiple music data included in the music library and display them to the user, so that the user can select one or several of these selected music data as the background music for the video data, etc. At this time, the multiple music data included in the music library are determined as M second media data.
[0210] Scenario 3, in the AIGC-related scenario, match appropriate music data to the video data created by artificial intelligence.
[0211] As Figure 11C shown, in the AIGC-related scenario, query the music data matching the video data in the music data included in the music library for the video data created by AI. At this time, the video data created by this AI can be understood as the first media data, and the multiple music data included in the music library are determined as M second media data.
[0212] As can be seen from the above, in the embodiments of the present application, matching video data can be found for music data, and matching music data can also be found for video data.
[0213] S202. For the j-th second media data among the M second media data, through the matching model, extract the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th second media data.
[0214] Among them, both R and K are positive integers, and j is a positive integer less than or equal to M.
[0215] In the embodiments of the present application, the specific manner for the computing device to determine the matching score between each second media data among the M second media data and the first media data through the matching model is basically the same. For the sake of convenience, the j-th second media data among the M second media data is taken as an example for description here.
[0216] Specifically, the computing device inputs the first media data and the j-th second media data into the matching model through the matching model. The matching model first extracts the first media semantic feature information of R different time segments of the first media data and extracts the second media semantic feature information of K different time segments of the j-th second media data.
[0217] In the embodiments of the present application, in order to improve the judgment accuracy of the matching relationship between the first media data and the j-th second media data, in the embodiments of the present application, by extracting the first media semantic feature information of different time segments of the first media data, and extracting the second media semantic feature information of different time segments of the j-th second media data, in this way, by refining the matching granularity and explicitly representing the semantic feature information of the music data and video data of different time segments, the matching accuracy of media data of different modalities can be improved, and the matching effect of cross-modal media data can be enhanced.
[0218] The following introduces the specific manner in which the computing device extracts the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th second media data through the matching model.
[0219] In some embodiments, the matching model includes a feature extraction module, which can extract the first media semantic feature information of the music data of R different time segments in the first media data, and the second media semantic feature information of the video data of K different time segments in the j-th second media data.
[0220] In some embodiments, as Figure 12 shown, the matching model of the embodiments of the present application includes a first media encoding module and a second media encoding module, where the first media encoding module is used to extract the first media semantic feature information of different time segments of the input music sample. The second media encoding module is used to extract the second media semantic feature information of different time segments of the input video sample.
[0221] In Figure 5 Based on the network structure of the shown matching model, the above S202 includes the following steps of S202-A and S202-B:
[0222] S202-A: Through the first media encoding module, extract the first media semantic feature information of R different time segments of the first media data;
[0223] S202-B: Through the second media encoding module, extract the second media semantic feature information of K different time segments of the j-th second media data.
[0224] In this implementation manner, as Figure 12As shown in the figure, the computing device inputs the first media data and the j-th second media data into the matching model. In this way, through the first media encoding module in the matching model, the first media data is processed to extract the first media semantic feature information of R different time segments of the first media data. Through the second media encoding module in the matching model, the j-th second media data is processed to extract the second media semantic feature information of K different time segments of the j-th second media data. It should be noted that when the first media data is music data and the second media data is video data, the above first media encoding module is a music encoding module, and the second media module is a video encoding module. When the first media data is video data and the second media data is music data, the above first media encoding module is a video encoding module, and the second media module is a music encoding module.
[0225] Next, the specific process of the computing device extracting the first media semantic feature information of R different time segments of the first media data through the first media encoding module will be introduced.
[0226] The embodiments of the present application do not limit the specific manner in which the computing device extracts the first media semantic feature information of R different time segments of the first media data through the first media encoding module.
[0227] In some embodiments, when the first media data is music data, the computing device divides the first media data into R music segments in chronological order. For example, if the duration of the first media data is 10 minutes, the music data of each minute in the first media data can be divided into a music segment. Then, the first media data with a length of 10 minutes can be divided into 10 music segments with a length of 1 minute. Next, through the first media encoding module, the first media semantic feature information of each of these 10 music segments is extracted, and thus the first media semantic feature information of 10 different time segments can be obtained. At this time, R = 10. Exemplarily, the above division durations can be the same, such as 1 minute each, or different. For example, the music data of 1 minute in the first media data can be divided into a music segment, or the music data of 2 minutes can be divided into a music segment.
[0228] In some embodiments, when the first media data is music data, the above S202-A includes the following steps of S202-A1 to S202-A3:
[0229] S202-A1. Extract the MFCC feature information of the first media data;
[0230] S202-A2. Slice the MFCC feature information once every first time interval to obtain the MFCC feature slice information of R different time segments of the first media data;
[0231] S202-A3. Through the first media encoding module, encode the MFCC feature slice information of R different time segments to obtain the first media semantic feature information of R different time segments of the first media data.
[0232] In this implementation manner, when the computing device extracts the first media semantic feature information of R different time segments of the first media data through the first media encoding module, it first extracts the MFCC feature information of the first media data. Then, according to the time information, slice the MFCC feature information of the first media data. The computing device slices the MFCC feature information every first time interval according to the preset first time interval, and obtains the MFCC feature slice information of R different time segments of the first media data. As Figure 6 shown, the computing device slices the MFCC feature information of the first media data every first time interval, that is, summarizes the MFCC feature information within each time segment, so as to obtain the MFCC feature slice information of 5 (of course, it can also be other values) different time segments of the first media data.
[0233] The present application embodiment does not limit the specific value of the first time interval. For example, the first time interval is 5 seconds, that is, the computing device slices the MFCC feature information of the first media data every 5 seconds, and can obtain the MFCC feature slice information of R different time segments of the first media data.
[0234] As Figure 6 shown, the first media encoding module of the present application embodiment includes a first Transformer unit, and this first Transformer unit is used to extract the semantic feature information of the input MFCC feature slice information. In this way, the computing device obtains the MFCC feature slice information of R different time segments of the first media data, inputs the MFCC feature slice information of R different time segments of the first media data into this first Transformer unit for semantic feature extraction, obtains the semantic feature information corresponding to each time segment of the MFCC feature slice information of R different time segments, and further determines the semantic feature information corresponding to the MFCC feature slice information of R different time segments obtained as the first media semantic feature information of R different time segments of the first media data.
[0235] The embodiments of the present application do not limit the specific size of each piece of first media semantic feature information in the R different time segments of the first media data. In one example, each piece of first media semantic feature information in the R different time segments of the first media data is 1024-dimensional feature information.
[0236] Next, in S202-B, the computing device extracts the second media semantic feature information of K different time segments of the j-th second media data through the second media encoding module.
[0237] The embodiments of the present application do not limit the specific manner in which the computing device extracts the second media semantic feature information of K different time segments of the j-th second media data through the second media encoding module.
[0238] In some embodiments, when the second media data is video data, the computing device divides the j-th second media data into K video segments in chronological order. For example, if the duration of the j-th second media data is 10 minutes, then each minute of the video data in the j-th second media data can be divided into a video segment. Thus, the 10-minute long j-th second media data can be divided into 10 video segments each with a length of 1 minute. Then, through the second media encoding module, the second media semantic feature information of each of these 10 video segments is extracted, and thus the second media semantic feature information of 10 different time segments can be obtained. At this time, K = 10. Exemplarily, the above division durations can be the same, such as 1 minute each, or different. For example, 1 minute of the video data in the j-th second media data can be divided into a video segment, or 2 minutes of the video data can be divided into a video segment.
[0239] In some embodiments, when the second media data is video data, the above S202-B includes the following steps of S202-B1 to S202-B2:
[0240] S202-B1: Select one video frame every second time interval from the video frames included in the j-th second media data to obtain K video frames;
[0241] S202-B2: Perform feature extraction on the K video frames through the second media encoding module to obtain the second media semantic feature information of K different time segments of the j-th second media data.
[0242] In this implementation manner, when the computing device extracts the second media semantic feature information of K different time segments of the j-th second media data through the second media encoding module, as Figure 7 shown, first, the K video frames of the j-th second media data are extracted.
[0243] For example, the computing device selects one video frame from the video frames included in the j-th second media data at every second time interval, and obtains K video frames. For example, as Figure 7 shown, assume that the j-th second media data includes 1000 video frames and the second time interval is 5 seconds. In this way, the computing device first selects the 1st video frame of the j-th second media data, and then, at an interval of 5 seconds, selects one video frame from the 1000 video frames included in the j-th second media data, and in this way, 5 video frames can be selected.
[0244] Next, the computing device inputs the selected K video frames into the second media encoding module, and the second media encoding module extracts features from these K video frames to obtain the semantic feature information of each of these K video frames. Since these K video frames correspond one by one to K different time segments, the semantic feature information of these K video frames is determined as the second media semantic feature information of the K different time segments of the j-th second media data.
[0245] In one example, as Figure 7 shown, the second media encoding module of the embodiment of the present application includes a second Transformer unit, and the second Transformer unit is used to extract the semantic feature information of the input video frames. In this way, after the computing device obtains the K video frames of the j-th second media data, it inputs these K video frames into the second Transformer unit for semantic feature extraction to obtain the semantic feature information of each of the K video frames, and further determines the semantic feature information of each of the obtained K video frames as the second media semantic feature information of the K different time segments of the j-th second media data.
[0246] The embodiment of the present application does not limit the specific size of each second media semantic feature information among the second media semantic feature information of the K different time segments of the second media data. In one example, each second media semantic feature information among the second media semantic feature information of the K different time segments of the second media data is 1024-dimensional feature information.
[0247] As described above, the computing device determines the music semantic feature information of the R different time segments of the first media data, and the video semantic feature information of the K different time segments of the j-th second media data.
[0248] Referring to the above method, the computing device can determine the second media semantic feature information of the different time segments of each of the M second media data. Next, the computing device executes the following step S203.
[0249] S203. Determine a matching score between the first media data and the j-th second media data based on the first media semantic feature information of R different time segments and the second media semantic feature information of K different time segments.
[0250] Based on the above steps, the computing device determines the second media semantic feature information of different time segments of each second media data in the M second media data pairs, and the first media semantic feature information of different time segments of the first media data. Then, based on the second media semantic feature information of each second media data in the M second media data at different time segments and the first media semantic feature information of the first media data at R different time segments, the matching score is determined.
[0251] In the embodiments of the present application, the specific manner of determining the matching score between each second media data in the M second media data and the first media data is the same. In the embodiments of the present application, the j-th second media data in the M second media data is continued as an example for illustration.
[0252] In the embodiments of the present application, the computing device determines the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th second media data based on the above steps. Then, the computing device determines the matching score between the first media data and the j-th second media data based on the first media semantic feature information of R different time segments and the second media semantic feature information of K different time segments.
[0253] The embodiments of the present application do not limit the specific manner in which the computing device determines the matching score between the first media data and the j-th second media data based on the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th second media data.
[0254] In some possible implementation manners, the computing device forms a feature matrix 1 with the first media semantic feature information of R different time segments of the first media data, and forms a feature matrix 2 with the second media semantic feature information of K different time segments of the j-th second media data. Then, the similarity between the feature matrix 1 and the feature matrix 2 is determined, and further, the similarity between the feature distance 1 and the feature matrix 2 is determined as the matching score between the first media data and the j-th second media data.
[0255] In some possible implementation manners, the computing device can determine the matching score between the first media data and the j-th second media data through the following steps S203-A and S203-B:
[0256] S203-A. Determine the semantic similarity between the first media semantic feature information of each of the R different time segments and the second media semantic feature information of each of the K different time segments, and use it as an element in the time matching matrix between the first media data and the j-th second media data, to obtain the time matching matrix between the first media data and the j-th second media data;
[0257] S203-B. Based on the time matching matrix between the first media data and the j-th second media data, determine the matching score between the first media data and the j-th second media data.
[0258] In this implementation manner, the computing device determines the similarity between the first media semantic feature information of each different time segment in the first media data and the second media semantic feature information of each different time segment in the j-th second media data. Further, based on the similarity between the first media semantic feature information of each different time segment in the first media data and the second media semantic feature information of each different time segment in the j-th second media data, the matching score between the first media data and the j-th second media data is determined.
[0259] Specifically, the computing device determines the semantic similarity between the first media semantic feature information of each of the R different time segments and the second media semantic feature information of each of the K different time segments, and uses it as an element in the time matching matrix between the first media data and the j-th second media data, so as to obtain a time matching matrix with a size of R×K. For example, use the similarity between the first media semantic feature information of the k-th time segment in the first media data and the second media semantic feature information of the j-th time segment in the j-th second media data as the element value at the k, j position in the time matching matrix. Referring to this method, a time matching matrix with a size of R×K can be obtained.
[0260] Next, the computing device determines the matching score between the first media data and the j-th second media data based on the time matching matrix between the first media data and the j-th second media data.
[0261] The embodiments of the present application do not limit the specific manner of determining the matching score between the first media data of the computing device and the j-th second media data based on the time matching matrix between the first media data and the j-th second media data.
[0262] In a possible implementation manner, the computing device determines the sum value obtained by adding each element value in the time matching matrix between the first media data and the j-th second media data as the matching score between the first media data and the j-th second media data.
[0263] In a possible implementation, the computing device determines the average value of the sum of each element value in the time matching matrix between the first media data and the j-th second media data as the matching score between the first media data and the j-th second media data.
[0264] In some embodiments, as Figure 8 shown, the matching model of the embodiments of the present application includes a matching module. The computing device extracts the first media semantic feature information of R different time segments of the first media data through the first media encoding module, and extracts the second media semantic feature information of K different time segments of the second media data. Then, through the matching module, the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the second media data are subjected to a matching process to obtain the matching score between the first media data and the j-th second media data.
[0265] In a possible implementation, the computing device performs a bipartite graph optimal matching process on the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the second media data through the matching module, and obtains the matching score between the first media data and the j-th second media data through the Hungarian algorithm.
[0266] Exemplarily, as Figure 9 shown, the computing device inputs the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the second media data into the matching module. The matching module determines the semantic similarity between the first media semantic feature information of each time segment among the R different time segments and the second media semantic feature information of each time segment among the K different time segments as an element in the time matching matrix between the first media data and the j-th second media data, so as to obtain a time matching matrix with a size of R×K. Then, through the bipartite graph optimal matching algorithm, the time matching matrix between the first media data and the j-th second media data is processed to obtain the matching score between the first media data and the j-th second media data.
[0267] The above specifically introduces the process of the computing device determining the matching score between the first media data and the j-th second media data. Referring to the above method, the computing device can determine the matching score between each of the M second media data and the first media data.
[0268] S204. Based on the matching scores between each of the M second media data and the first media data, determine the target second media data that matches the first media data from the M second media data.
[0269] For example, determine the second media data with the highest matching score between the M second media data and the first media data as the target second media data that matches the first media data.
[0270] For another example, determine several second media data with the highest matching scores between the M second media data and the first media data as several target second media data that match the first media data.
[0271] In some embodiments, the computing device may further obtain the label information of the target second media data; determine the label information of the target second media data as the label information of the second media data corresponding to the first media data. For example, in the above scenario 1, the computing device may use the above method of the embodiments of the present application to find the matching video data 2 for the music data included in the video data 1 in the video database, and then determine the label information of the video data 2 as the label data of the video data 1.
[0272] The media data matching method provided by the embodiments of the present application obtains the first media data and M second media data to be matched with the first media data. Among them, if the first media data is music data, the second media data is video data; if the first media data is video data, the second media data is music data. For the j-th second media data among the M second media data, through a matching model, R first media semantic feature information of different time segments of the first media data and K second media semantic feature information of different time segments of the j-th second media data are extracted. Then, based on the R first media semantic feature information of different time segments and the K second media semantic feature information of different time segments, the matching score between the first media data and the j-th second media data is determined. Finally, based on the matching scores between each of the M second media data and the first media data, the target second media data that matches the first media data is determined from the M second media data. As can be seen from the above, when the embodiments of the present application match the first media data and the second media data, the semantic feature information of the first media data and the second media data in different time periods is extracted, and then based on the semantic feature information of different time periods, the matching score between the first media data and the second media data is determined. In this way, by explicitly representing the features of cross-modal media data in different time periods, the most suitable second media data can be selected in the time dimension to match the first media data, thereby improving the matching accuracy of media data of different modalities and enhancing the matching effect of cross-modal media data.
[0273] As described above in conjunction with Figures 3 to 12 , the embodiments of the model training and media data matching methods of the present application are described in detail. Below in conjunction with Figures 13 to 14 , the device embodiments of the present application are described in detail.
[0274] Figure 13 FIG. is a schematic block diagram of a media data matching device provided by an embodiment of the present application. The device 10 can be applied to a computing device.
[0275] As Figure 13 shown, the media data matching device 10 includes:
[0276] An obtaining unit 11, configured to obtain the first media data and M second media data to be matched with the first media data, where if the first media data is music data, the second media data is video data; if the first media data is video data, the second media data is music data, and M is a positive integer;
[0277] A feature extraction unit 12, configured to, for the j-th second media data among the M second media data, extract first media semantic feature information of R different time segments of the first media data and second media semantic feature information of K different time segments of the j-th second media data through the matching model, where both R and K are positive integers, and j is a positive integer less than or equal to M;
[0278] A matching unit 13, configured to determine a matching score between the first media data and the j-th second media data based on the first media semantic feature information of the R different time segments and the second media semantic feature information of the K different time segments;
[0279] A determination unit 14, configured to determine a target second media data that matches the first media data from the M second media data based on the matching scores between each second media data in the M second media data and the first media data.
[0280] In some embodiments, the matching model includes a first media encoding module and a second media encoding module; the feature extraction unit 12 is specifically configured to extract first media semantic feature information of R different time segments of the first media data through the first media encoding module; and extract second media semantic feature information of K different time segments of the j-th second media data through the second media encoding module.
[0281] In some embodiments, when the first media data is music data, the feature extraction unit 12 is specifically configured to extract MFCC feature information of the first media data; slice the MFCC feature information once every first time interval to obtain MFCC feature slice information of R different time segments of the first media data; and encode the MFCC feature slice information of the R different time segments through the first media encoding module to obtain first media semantic feature information of the R different time segments of the first media data.
[0282] In some embodiments, the first media encoding module includes a first Transformer unit, and the feature extraction unit 12 is specifically configured to encode the MFCC feature slice information of the R different time segments through the first Transformer unit to obtain music semantic feature information of each time segment among the R different time segments of the first media data.
[0283] In some embodiments, when the second media data is video data, the feature extraction unit 12 is specifically configured to select one video frame from the video frames included in the j-th second media data at every second time interval to obtain K video frames; perform feature extraction on the K video frames through the second media encoding module to obtain second media semantic feature information of the j-th second media data in K different time segments.
[0284] In some embodiments, the second media encoding module includes a second Transformer unit. The feature extraction unit 12 is specifically configured to perform feature extraction on the K video frames through the second Transformer unit to obtain video semantic feature information of each of the K video frames, and use it as the second media semantic feature information of the j-th second media data in K different time segments.
[0285] In some embodiments, the matching unit 13 is specifically configured to determine the semantic similarity between the first media semantic feature information of each time segment among the R different time segments and the second media semantic feature information of each time segment among the K different time segments, and use it as an element in the time matching matrix between the first media data and the j-th second media data, so as to obtain the time matching matrix between the first media data and the j-th second media data; based on the time matching matrix, determine the matching score between the first media data and the j-th second media data.
[0286] In some embodiments, the matching unit 13 is specifically configured to perform bipartite graph optimal matching processing on the time matching matrix between the first media data and the j-th second media data to obtain the matching score between the first media data and the j-th second media data.
[0287] In some embodiments, the determination unit 14 is further configured to obtain the label information of the target second media data;
[0288] Determine the label information of the target second media data as the label information of the second media data corresponding to the first media data.
[0289] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 13 The device shown can execute the embodiments of the above-mentioned media data matching method, and the foregoing and other operations and / or functions of each module in the device respectively implement the corresponding method embodiments of the computing device. For the sake of brevity, it will not be elaborated here.
[0290] Figure 14 Figure 14 is a schematic block diagram of a training device for a matching model provided by an embodiment of the present application. The device 20 can be applied to a computing device.
[0291] As Figure 14 shown, the training device 20 of the matching model includes:
[0292] An obtaining unit 21, configured to obtain N training sample pairs, each training sample pair including a music sample and a video sample, where N is a positive integer;
[0293] A feature extraction unit 22, configured to, for the i-th training sample pair among the N training sample pairs, extract, through the matching model, music semantic feature information of P different time segments of the i-th music sample in the i-th training sample pair, and video semantic feature information of Q different time segments of the i-th video sample in the i-th training sample pair, where P and Q are both positive integers, and i is a positive integer less than or equal to N;
[0294] A matching unit 23, configured to determine a matching score between the i-th music sample and the i-th video sample based on the music semantic feature information of the P different time segments and the video semantic feature information of the Q different time segments;
[0295] A training unit 24, configured to determine a model loss of the matching model based on the matching scores between the music samples and the video samples of each training sample pair among the N training sample pairs, and train the matching model based on the model loss.
[0296] In some embodiments, the matching model includes a music encoding module and a video encoding module; the feature extraction unit 22 is specifically configured to extract, through the music encoding module, music semantic feature information of P different time segments of the i-th music sample; and extract, through the video encoding module, video semantic feature information of Q different time segments of the i-th video sample.
[0297] In some embodiments, the feature extraction unit 22 is specifically configured to extract Mel Frequency Cepstral Coefficient (MFCC) feature information of the i-th music sample; slice the MFCC feature information once every first time interval to obtain MFCC feature slice information of P different time segments of the i-th music sample; and encode the MFCC feature slice information of the P different time segments through the music encoding module to obtain music semantic feature information of P different time segments of the i-th music sample.
[0298] In some embodiments, the feature extraction unit 22 is specifically configured to select one video frame from the video frames included in the i-th video sample at every second time interval, so as to obtain Q video frames; and perform feature extraction on the Q video frames through the video encoding module to obtain video semantic feature information of Q different time segments of the i-th video sample.
[0299] In some embodiments, the matching unit 23 is specifically configured to determine the semantic similarity between the music semantic feature information of each time segment among the P different time segments and the video semantic feature information of each time segment among the Q different time segments, and use it as an element in the time matching matrix between the i-th music sample and the i-th video sample, so as to obtain the time matching matrix between the i-th music sample and the i-th video sample; and determine the matching score between the i-th music sample and the i-th video sample based on the time matching matrix between the i-th music sample and the i-th video sample.
[0300] In some embodiments, the i-th training sample pair further includes the label information of the i-th music sample and the label information of the i-th video sample; the training unit 24 is specifically configured to determine the label similarity between the i-th music sample and the i-th video sample based on the label information of the i-th music sample and the label information of the i-th video sample; and determine the model loss of the matching model based on the matching scores and similarities between the music samples and the video samples of each training sample pair among the N training sample pairs.
[0301] In some embodiments, the training unit 24 is specifically configured to determine the label vector of the i-th music sample based on the label information of the i-th music sample and the preset label information; determine the label vector of the i-th video sample based on the label information of the i-th video sample and the preset label information; and determine the similarity between the label vector of the i-th music sample and the label vector of the i-th video sample as the label similarity between the i-th music sample and the i-th video sample.
[0302] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, it will not be elaborated here. Specifically, Figure 14 The illustrated apparatus can execute the embodiments of the above model training method, and the foregoing and other operations and / or functions of each module in the apparatus respectively implement the corresponding method embodiments of the computing device. For the sake of brevity, it will not be elaborated here.
[0303] In the above, the device according to the embodiments of the present application has been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.
[0304] Figure 15 is a schematic block diagram of a computing device provided by an embodiment of the present application, Figure 15 and the computing device can be used to execute the above model training method and / or media data matching method.
[0305] As Figure 15 shown, the computing device 30 may include:
[0306] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiments of the present application.
[0307] For example, the processor 32 can be used to execute the steps in the above method according to the instructions in the computer program 33.
[0308] In some embodiments of the present application, the processor 32 may include, but is not limited to:
[0309] A general-purpose processor, a digital signal processor (Digital Signal Processor, DSR), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.
[0310] In some embodiments of the present application, the memory 31 includes, but is not limited to:
[0311] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synch Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0312] In some embodiments of the present application, the computer program 33 may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method for recording a page provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the computing device.
[0313] As Figure 15 shown, the computing device 30 may further include:
[0314] A transceiver 34, which may be connected to the processor 32 or the memory 31.
[0315] Among them, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas may be one or more.
[0316] It should be understood that the various components in the computing device 30 are connected through a bus system. Among them, the bus system includes, in addition to the data bus, a power bus, a control bus, and a status signal bus.
[0317] According to one aspect of the present application, there is provided a computer storage medium having a computer program stored thereon. When the computer program is executed by a computer, the computer is enabled to execute the methods of the above method embodiments. Or rather, the embodiments of the present application further provide a computer program product containing instructions. When the instructions are executed by a computer, the computer is enabled to execute the methods of the above method embodiments.
[0318] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the methods of the above method embodiments.
[0319] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, a data center, etc. that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0320] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0321] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or module can be in an electrical, mechanical, or other form.
[0322] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0323] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A media data matching method, characterized in that, Including: Obtain first media data and M second media data to be matched with the first media data, where if the first media data is music data, then the second media data is video data, and if the first media data is video data, then the second media data is music data, and M is a positive integer; For the j-th second media data among the M second media data, through the matching model, extract first media semantic feature information of R different time segments of the first media data and second media semantic feature information of K different time segments of the j-th second media data, where both R and K are positive integers, and j is a positive integer less than or equal to M; Based on the first media semantic feature information of the R different time segments and the second media semantic feature information of the K different time segments, determine the matching score between the first media data and the j-th second media data; Based on the matching scores between each of the M second media data and the first media data, determine the target second media data that matches the first media data from the M second media data.
2. The method according to claim 1, characterized in that, The matching model includes a first media encoding module and a second media encoding module; The step of, through the matching model, extracting first media semantic feature information of R different time segments of the first media data and second media semantic feature information of K different time segments of the j-th second media data includes: Through the first media encoding module, extract first media semantic feature information of R different time segments of the first media data; Through the second media encoding module, extract second media semantic feature information of K different time segments of the j-th second media data.
3. The method according to claim 2, characterized in that, If the first media data is music data, then the step of, through the first media encoding module, extracting first media semantic feature information of R different time segments of the first media data includes: Extract MFCC feature information of the first media data; Slice the MFCC feature information once every first time interval to obtain MFCC feature slice information of R different time segments of the first media data; Through the first media encoding module, perform encoding processing on the MFCC feature slice information of the R different time segments to obtain first media semantic feature information of R different time segments of the first media data.
4. The method according to claim 3, characterized in that, The first media encoding module includes a first Transformer unit, and the step of, through the first media encoding module, performing encoding processing on the MFCC feature slice information of the R different time segments to obtain first media semantic feature information of R different time segments of the first media data includes: Through the first Transformer unit, perform encoding processing on the MFCC feature slice information of the R different time segments to obtain music semantic feature information of each time segment among the R different time segments of the first media data.
5. The method according to claim 2, wherein When the second media data is video data, the second media semantic feature information of the j-th second media data in K different time segments is extracted through the second media encoding module, including: Select a video frame every second time interval from the video frames included in the j-th second media data to obtain K video frames; Perform feature extraction on the K video frames through the second media encoding module to obtain the second media semantic feature information of the j-th second media data in K different time segments.
6. The method according to claim 5, wherein The second media encoding module includes a second Transformer unit. The step of performing feature extraction on the K video frames through the second media encoding module to obtain the second media semantic feature information of the j-th second media data in K different time segments includes: Perform feature extraction on the K video frames through the second Transformer unit to obtain the video semantic feature information of each video frame in the K video frames, which is used as the second media semantic feature information of the j-th second media data in K different time segments.
7. The method according to claim 2, wherein The step of determining the matching score between the first media data and the j-th second media data based on the first media semantic feature information in R different time segments and the second media semantic feature information in K different time segments includes: Determine the semantic similarity between the first media semantic feature information of each time segment in the R different time segments and the second media semantic feature information of each time segment in the K different time segments, and use it as an element in the time matching matrix between the first media data and the j-th second media data, so as to obtain the time matching matrix between the first media data and the j-th second media data; Based on the time matching matrix, determine the matching score between the first media data and the j-th second media data.
8. The method according to claim 7, wherein The step of determining the matching score between the first media data and the j-th second media data based on the time matching matrix includes: Perform bipartite graph optimal matching processing on the time matching matrix between the first media data and the j-th second media data to obtain the matching score between the first media data and the j-th second media data.
9. The method according to claim 1, characterized in that, The method further includes: Obtain the label information of the target second media data; Determine the label information of the target second media data as the label information of the second media data corresponding to the first media data.
10. A training method for a matching model, characterized in that It includes: Obtain N training sample pairs, each training sample pair includes a music sample and a video sample, and N is a positive integer; For the i-th training sample pair in the N training sample pairs, through the matching model, extract the music semantic feature information of P different time segments of the i-th music sample in the i-th training sample pair, and the video semantic feature information of Q different time segments of the i-th video sample in the i-th training sample pair, where P and Q are both positive integers, and i is a positive integer less than or equal to N; Determine a matching score between the \(i\)-th music sample and the \(i\)-th video sample based on the music semantic feature information of the \(P\) different time segments and the video semantic feature information of the \(Q\) different time segments; Determine a model loss of the matching model based on the matching scores between the music samples and the video samples in each of the \(N\) training sample pairs, and train the matching model based on the model loss.
11. The method according to claim 10, wherein The matching model includes a music encoding module and a video encoding module; Extracting, by the matching model, music semantic feature information of \(P\) different time segments of the \(i\)-th music sample in the \(i\)-th training sample pair, and video semantic feature information of \(Q\) different time segments of the \(i\)-th video sample in the \(i\)-th training sample pair includes: Extract music semantic feature information of \(P\) different time segments of the \(i\)-th music sample through the music encoding module; Extract video semantic feature information of \(Q\) different time segments of the \(i\)-th video sample through the video encoding module.
12. The method according to claim 11, wherein The extracting, by the music encoding module, music semantic feature information of \(P\) different time segments of the \(i\)-th music sample includes: Extract the Mel Frequency Cepstral Coefficient (MFCC) feature information of the \(i\)-th music sample; Slice the MFCC feature information once every first time interval to obtain MFCC feature slice information of \(P\) different time segments of the \(i\)-th music sample; Encode the MFCC feature slice information of the \(P\) different time segments through the music encoding module to obtain music semantic feature information of \(P\) different time segments of the \(i\)-th music sample.
13. The method according to claim 11, wherein The extracting, by the video encoding module, video semantic feature information of \(Q\) different time segments of the \(i\)-th video sample includes: Select one video frame every second time interval from the video frames included in the \(i\)-th video sample to obtain \(Q\) video frames; Extract features from the \(Q\) video frames through the video encoding module to obtain video semantic feature information of \(Q\) different time segments of the \(i\)-th video sample.
14. The method according to claim 10, characterized in that, The determining a matching score between the \(i\)-th music sample and the \(i\)-th video sample based on the music semantic feature information of the \(P\) different time segments and the video semantic feature information of the \(Q\) different time segments includes: Determine the semantic similarity between the music semantic feature information of each of the \(P\) different time segments and the video semantic feature information of each of the \(Q\) different time segments as an element in the time matching matrix between the \(i\)-th music sample and the \(i\)-th video sample, to obtain the time matching matrix between the \(i\)-th music sample and the \(i\)-th video sample; Determine the matching score between the \(i\)-th music sample and the \(i\)-th video sample based on the time matching matrix between the \(i\)-th music sample and the \(i\)-th video sample.
15. The method according to any one of claims 10-14, characterized in that, The i-th training sample pair further includes the label information of the i-th music sample and the label information of the i-th video sample; Determining the model loss of the matching model based on the matching scores between the music samples and the video samples in each of the N training sample pairs includes: Based on the label information of the i-th music sample and the label information of the i-th video sample, determining the label similarity between the i-th music sample and the i-th video sample; Based on the matching scores and similarities between the music samples and the video samples in each of the N training sample pairs, determining the model loss of the matching model.
16. The method according to claim 15, wherein The determining the label similarity between the i-th music sample and the i-th video sample based on the label information of the i-th music sample and the label information of the i-th video sample includes: Based on the label information of the i-th music sample and the preset label information, determining the label vector of the i-th music sample; Based on the label information of the i-th video sample and the preset label information, determining the label vector of the i-th video sample; Determining the similarity between the label vector of the i-th music sample and the label vector of the i-th video sample as the label similarity between the i-th music sample and the i-th video sample.
17. A media data matching device, characterized in that, Includes: An acquisition unit, configured to acquire first media data and M pieces of second media data to be matched with the first media data, where if the first media data is music data, the second media data is video data, and if the first media data is video data, the second media data is music data, and M is a positive integer; A feature extraction unit, configured to, for the j-th piece of second media data among the M pieces of second media data, extract the first media semantic feature information of R different time segments of the first media data and the second media semantic feature information of K different time segments of the j-th piece of second media data through the matching model, where R and K are both positive integers, and j is a positive integer less than or equal to M; A matching unit, configured to determine the matching score between the first media data and the j-th piece of second media data based on the first media semantic feature information of the R different time segments and the second media semantic feature information of the K different time segments; A determination unit, configured to determine a target second media data that matches the first media data from the M pieces of second media data based on the matching scores between each piece of second media data in the M pieces of second media data and the first media data.
18. A training device for a matching model, characterized in that, Includes: An acquisition unit, configured to acquire N training sample pairs, each training sample pair including a music sample and a video sample, where N is a positive integer; A feature extraction unit, for the i-th training sample pair among the N training sample pairs, through the matching model, extract the music semantic feature information of P different time segments of the i-th music sample in the i-th training sample pair, and the video semantic feature information of Q different time segments of the i-th video sample in the i-th training sample pair, where both P and Q are positive integers, and i is a positive integer less than or equal to N; A matching unit, for determining the matching score between the i-th music sample and the i-th video sample based on the music semantic feature information of the P different time segments and the video semantic feature information of the Q different time segments; A training unit, for determining the model loss of the matching model based on the matching scores between the music samples and the video samples of each training sample pair among the N training sample pairs, and training the matching model based on the model loss.
19. A computer device, comprising a processor and a memory; The memory is used for storing a computer program; The processor is used for executing the computer program to implement the method according to any one of claims 1 to 9 or 10 to 16 above.
20. A computer-readable storage medium, characterized in that, For storing a computer program; The computer program causes the computer to execute the method according to any one of claims 1 to 9 or 10 to 16 above.
Citation Information
Cited By
Audio-visual video semantic analysis method and system suitable for cross-media information retrieval
CN122240879A
Audiovisual video semantic parsing method and system applicable to cross-media information retrieval
CN122240879B