Audio Data Processing Method, Apparatus, Device, and Storage Medium
By fusing the characteristic information of the target object, video data and candidate audio data, and using multiple audio recognition models for identification, the problem of low accuracy in user selection of audio data is solved, and more efficient and accurate audio data recommendation is achieved.
Patent Information
- Application Number
- CN202111017197.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-08-31
AI Technical Summary
In the prior art, when users select audio data for video data, their lack of expertise leads to a low accuracy in the selection of audio data.
By obtaining the object feature information of the target object, the video feature information of the target video data, and the audio feature information of the candidate audio data, the audio feature information of the candidate audio data is fused, and multiple audio recognition models are used for audio recognition, and the most suitable audio data is recommended.
Improve the accuracy and efficiency of audio data recommendation, avoid single model identification bias, and the recommended audio data is more robust, more accurate and more reliable.
Smart Images

Figure CN115734024B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine learning in artificial intelligence, and particularly relates to an audio data processing method, apparatus, device, and storage medium. Background Art
[0002] With the development of Internet technology, people can record and publish video data (such as short videos) anytime and anywhere, and can also watch video data published by others. Usually, when a user publishes video data, the user needs to select an audio data (such as background music) that matches the theme of the video data from the local terminal, and then use the audio data to score the video data. The audio data can be used to strengthen the theme of the video data, which is beneficial for users to more intuitively understand the theme of the video data, and can enhance the interest and rhythm of the video data. Currently, usually the audio data that matches the video data is selected manually. However, if the user does not have professional knowledge related to audio data, it is difficult to select appropriate audio data, resulting in a relatively low accuracy of the selected audio data. Summary of the Invention
[0003] The technical problem to be solved by the embodiments of the present application is to provide an audio data processing method, apparatus, device, and storage medium, which can effectively improve the accuracy of recommended audio data.
[0004] An embodiment of the present application provides an audio data processing method, including:
[0005] Obtaining object feature information of a target object, video feature information of target video data belonging to the target object, and audio feature information of at least two candidate audio data associated with the target video data;
[0006] Respectively fusing the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain audio fusion feature information of the at least two candidate audio data;
[0007] Performing audio recognition on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the at least two candidate audio data by using at least two target audio recognition models to obtain target audio data for scoring the target video data; the target audio data belongs to the at least two candidate audio data;
[0008] Recommending the target audio data to the target object.
[0009] An embodiment of the present application provides an audio data processing method, including:
[0010] Obtain the object feature information of the sample object, the sample video data belonging to the sample object, the sample audio data for scoring the sample video data, and the labeled audio matching degree of the sample audio data; the labeled audio matching degree is used to reflect the matching degree between the sample audio data and the sample object and the sample video data;
[0011] Extract video signs from the sample video data to obtain the video feature information of the sample video data, and extract audio features from the sample audio data to obtain the audio feature information of the sample audio data;
[0012] Fuse the audio feature information of the sample audio data with the video feature information of the sample video data and the object feature information of the sample object to obtain the audio fusion feature information of the sample audio data;
[0013] According to the labeled audio matching degree, the audio feature information of the sample audio data, and the audio fusion feature information of the sample audio data, adjust at least two candidate audio recognition models respectively to obtain at least two target audio recognition models.
[0014] One aspect of the embodiments of the present application provides an audio data processing device, including:
[0015] An acquisition module, configured to acquire the object feature information of the target object, the video feature information of the target video data belonging to the target object, and the audio feature information of at least two candidate audio data associated with the target video data;
[0016] A fusion module, configured to fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object respectively to obtain the audio fusion feature information of the at least two candidate audio data;
[0017] An identification module, configured to perform audio identification on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the at least two candidate audio data respectively by using at least two target audio recognition models to obtain target audio data for scoring the target video data; the target audio data belongs to the at least two candidate audio data;
[0018] A recommendation module, configured to recommend the target audio data to the target object.
[0019] One aspect of the embodiments of the present application provides an audio data processing device, including:
[0020] An acquisition module, configured to acquire object feature information of a sample object, sample video data belonging to the sample object, sample audio data for scoring the sample video data, and an annotation audio matching degree of the sample audio data; the annotation audio matching degree is used to reflect the matching degree between the sample audio data and the sample object and the sample video data;
[0021] An extraction module, configured to perform video feature extraction on the sample video data to obtain video feature information of the sample video data, and perform audio feature extraction on the sample audio data to obtain audio feature information of the sample audio data;
[0022] A fusion module, configured to fuse the audio feature information of the sample audio data, the video feature information of the sample video data, and the object feature information of the sample object to obtain audio fusion feature information of the sample audio data;
[0023] An adjustment module, configured to adjust at least two candidate audio recognition models respectively according to the annotation audio matching degree, the audio feature information of the sample audio data, and the audio fusion feature information of the sample audio data, to obtain at least two target audio recognition models.
[0024] On the one hand, the present application provides a computer device, including: a processor and a memory;
[0025] Wherein, the above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call the above-mentioned computer program to execute the following steps:
[0026] Acquire object feature information of a target object, video feature information of target video data belonging to the target object, and audio feature information of at least two candidate audio data associated with the target video data;
[0027] Respectively fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain audio fusion feature information of the at least two candidate audio data;
[0028] Use at least two target audio recognition models to perform audio recognition on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the at least two candidate audio data respectively, to obtain target audio data for scoring the target video data; the target audio data belongs to the at least two candidate audio data;
[0029] Recommend the target audio data to the target object.
[0030] Among them, the above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call the above-mentioned computer program to execute the following steps:
[0031] Obtain the object feature information of the sample object, the sample video data belonging to the sample object, the sample audio data for scoring the sample video data, and the labeled audio matching degree of the sample audio data; the labeled audio matching degree is used to reflect the matching degree between the sample audio data and the sample object and the sample video data;
[0032] Extract video signs from the sample video data to obtain the video feature information of the sample video data, and extract audio features from the sample audio data to obtain the audio feature information of the sample audio data;
[0033] Fuse the audio feature information of the sample audio data with the video feature information of the sample video data and the object feature information of the sample object to obtain the audio fusion feature information of the sample audio data;
[0034] Adjust at least two candidate audio recognition models respectively according to the labeled audio matching degree, the audio feature information of the sample audio data and the audio fusion feature information of the sample audio data to obtain at least two target audio recognition models.
[0035] An embodiment of the present application provides a computer-readable storage medium on the one hand. The above-mentioned computer-readable storage medium stores a computer program. The above-mentioned computer program includes program instructions. When the above-mentioned program instructions are executed by a processor, the steps of the above-mentioned method are executed.
[0036] An embodiment of the present application provides a computer program product on the one hand, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the steps of the above-mentioned method are implemented.
[0037] In this application, by fusing the audio feature information of at least two candidate audio data with the object feature information of the target object and the video feature information of the target video data, the audio fusion feature information of at least two candidate audio data is obtained. That is, by fusing multi-modal feature information, it is beneficial to provide more information for the recommended audio data and improve the accuracy of the recommended audio data. Further, by using at least two target audio recognition models to respectively recognize the audio fusion feature information of at least two candidate audio data and the audio feature information of at least two candidate audio data, the target audio data for scoring the target video data is obtained, and this target audio data is recommended to the target object; that is, by comprehensively considering the audio recognition results of multi-modal audio recognition models, the audio data is automatically recommended to the target object, which can improve the efficiency of the recommended audio data; at the same time, the advantages of different audio recognition models are fully utilized, which can effectively avoid the problem that a single model produces deviations and leads to a relatively low accuracy of the recommended audio data, and can make the recommended audio data more robust, more accurate, and more credible. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0039] Figure 1 It is a schematic diagram of the architecture of an audio data processing system provided by the present application;
[0040] Figure 2 It is a schematic diagram of a scenario for recommending audio data based on a single modality provided by the present application;
[0041] Figure 3 It is a schematic diagram of a scenario for recommending audio data based on multi-modal provided by the present application;
[0042] Figure 4 It is a flowchart of an audio data processing method provided by the present application;
[0043] Figure 5 It is a flowchart of an audio data processing method provided by the present application;
[0044] Figure 6 It is a schematic diagram of a scenario for obtaining the object feature information of the target object provided by the present application;
[0045] Figure 7 It is a schematic diagram of a scenario for obtaining the audio feature information of candidate audio data provided by the present application;
[0046] Figure 8 It is a schematic diagram of a scenario for obtaining video feature information of target video data provided by this application;
[0047] Figure 9 It is a schematic diagram of a scenario for obtaining audio fusion feature information of candidate audio data provided by this application;
[0048] Figure 10 It is a flowchart of a method for processing audio data provided by this application;
[0049] Figure 11 It is a schematic structural diagram of an audio data processing device provided by an embodiment of this application;
[0050] Figure 12 It is a schematic structural diagram of an audio data processing device provided by an embodiment of this application;
[0051] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of this application. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0053] This application mainly relates to machine learning technology in artificial intelligence. The so-called artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0054] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0055] Among them, the above-mentioned Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in studying how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0056] To facilitate a clearer understanding of this application, first, an audio data processing system for implementing the audio data processing method of this application is introduced, as Figure 1 shown, the audio data processing system includes as Figure 1 shown, the audio data processing system includes a server and a terminal.
[0057] Among them, the terminal can refer to a user-facing device, and the terminal may include a multimedia application platform (i.e., a multimedia application program) for playing multimedia data (such as audio and video data); here, the multimedia application platform can refer to a multimedia website platform (such as a forum, Tieba), a social application platform, a shopping application platform, a content interaction platform (such as an audio and video playback application platform), and so on. The server can refer to a device for providing multimedia backend services, and specifically can be used to identify the audio data for scoring the video data and recommend the audio data to the user.
[0058] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of at least two physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be an intelligent vehicle terminal, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a screen speaker, a smart watch, a smart TV, etc., but is not limited thereto. Each terminal and the server can be directly or indirectly connected through wired or wireless communication methods. At the same time, the number of terminals and servers can be one or at least two, and this application does not make any restrictions here.
[0059] Based on the above audio data processing system, the audio data recommendation method in this application can be implemented. The audio data recommendation method includes a single-modal based audio data recommendation method and a multi-modal based audio data recommendation method. AsFigure 2 As shown, the single-modal based audio data recommendation method refers to using an audio recognition model to analyze the audio feature information of at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object, so as to obtain the target audio data for scoring the target video data. Specifically, as Figure 2 shown, the single-modal based audio data recommendation method includes the training process of the candidate audio recognition model and the process of using the target audio recognition model to identify the target audio data. As Figure 2 shown, the training process of the candidate audio recognition model includes the following steps 1-2:
[0060] 1. The server obtains training samples for training the candidate audio recognition model. The candidate audio recognition model refers to a to-be-trained model for obtaining the music score for video data. That is to say, the candidate audio recognition model is a recognition model with relatively low audio data recognition accuracy. The candidate audio recognition model can be a classifier, and the classifier can be one of a machine learning model, a deep learning model, and a graph network model, etc. Machine learning models include SVM (Support Vector Machine), FM (Factorization Machines), XGBoost (eXtreme Gradient Boosting), deep learning models include DNN (Deep Neural Networks), W&D (Wide&Deep Learning for Recommender System), etc., and graph network models include DeepWalk, GraphSAGE, GCN (Graph Convolutional Network), etc. In order to improve the audio data recognition accuracy of the candidate audio recognition model, first, the server can obtain the object feature information belonging to the sample object, the sample video data belonging to the sample object, and the sample audio data for scoring the sample video data from the terminal. The sample object can be a user who has published video data on the multimedia application platform. The object feature information of the sample object includes the age, gender, hobbies, etc. of the sample object. The sample video data can be the video data published by the sample object on the multimedia application platform within a historical time period. The sample video data can be the video data shot by the sample object, or the sample video data can be the video data downloaded from the Internet by the sample object and edited; the sample audio data refers to the audio data (such as background music, voice data for poetry recitation) for scoring the sample video data when the sample object publishes the sample video data. As Figure 2Among them, User 1 publishes video data 1 on the multimedia application platform. The background music of this video data 1 is Music 1,..., User N publishes video data N on the multimedia application platform. The background music of this video data N is Music N. The video data 1 to video data N and audio data 1 to audio data N can be filtered. The filtered video data is used as sample video data, and the filtered audio data is determined as sample audio data. The filtering process includes copyright filtering, quality filtering, and tonality filtering. Copyright filtering means filtering out audio data and video data without copyright. Quality filtering can mean filtering out video data and audio data with relatively low quality (such as clarity). Tonality filtering can mean filtering out audio data whose melody does not meet the conditions, such as audio data with too much noise.
[0061] It should be noted that the object feature information in this solution can refer to user portrait data, which is obtained after obtaining user authorization. The audio data in this solution can refer to audio data authorized by the creator of the audio data. The video data in this solution can refer to original video data or video data authorized by the creator. Further, the server can extract video features from the sample video data to obtain the video feature information (i.e., video portrait) of the sample video data; the video feature information of the sample video data is used to reflect the theme information, scene, color information, quality information, etc. of the sample video data. Similarly, the server can extract audio features from the sample audio data to obtain the audio feature information (i.e., audio portrait) of the sample audio data. The audio feature information of the sample audio data is used to reflect the lyric information, score information, object feature information of the creator of the sample audio data, etc. Then, obtain the labeled audio matching degree of the sample audio data, which can be used to reflect the matching degree between the sample audio data and the sample object and sample video data. The video feature information of the sample video data, the audio feature information of the sample audio data, the object feature information of the sample object, and the labeled audio matching degree are determined as training samples for training the candidate audio recognition model.
[0062] 2. The server uses the training samples to train the candidate audio recognition model to obtain the target audio recognition model. The server uses the candidate audio recognition model to perform audio prediction on the video feature information of the sample video data, the audio feature information of the sample audio data, and the object feature information of the sample object to obtain the predicted audio matching degree of the sample audio data; adjust the candidate audio recognition model according to the predicted audio matching degree and the labeled audio matching degree to obtain the adjusted candidate audio recognition model, and determine the adjusted audio recognition model as the target audio recognition model.
[0063] Such asFigure 2 As shown, the process of using the target audio recognition model to recognize the target audio data includes the following steps 3-5:
[0064] 3. The server obtains the object feature information of the target object, the target video data belonging to the target object, and at least two candidate audio data associated with the target video data. The target object refers to the user who needs to publish video data to the multimedia application platform (such as Figure 2 user W in), and the target video data may refer to the video data to be published to the multimedia application platform (such as Figure 2 video W in). The target video data may be obtained by the target object through shooting, or the target video data may be obtained by the target object by editing the video data downloaded from the Internet. The at least two candidate audio data may refer to the audio data that matches the attribute information such as the theme information and scene of the target video data, and the at least two candidate audio data refer to the audio data that the target object has the right to use.
[0065] 4. The server obtains the video feature information of the target video data and the audio feature information of at least two candidate audio data. The server can perform video feature extraction on the target video data to obtain the video feature information of the target video data; the video feature information of the target video data is used to reflect the theme information, scene, color information, quality information, etc. of the target video data. Similarly, the server can perform audio feature extraction on each candidate audio data respectively to obtain the audio feature information of each candidate audio data, and the audio feature information of the candidate audio data is used to reflect the lyric information, musical score information, object feature information of the creator of the candidate audio data, etc.
[0066] 5. The server can use the target audio recognition model to recognize the target video data. The server can use the target audio recognition model to perform audio recognition on the audio feature information of the at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object to obtain the target video data for scoring the target video data, and recommend the target video data to the target object.
[0067] In practice, it is found that the audio recommendation result of the single-modal based audio data recommendation method completely depends on the knowledge accumulation of this one candidate audio recognition model. If there is a deviation in the knowledge accumulation process (i.e., the training process) of the candidate audio recognition model, it will lead to a relatively low accuracy of the recommended audio data. Based on this, the present application proposes a multi-modal based audio data recommendation method, such as Figure 3As shown, the multi-modal based audio data recommendation method refers to analyzing the audio feature information of at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object using at least two audio recognition models to obtain the target audio data for scoring the target video data. For example, Figure 3 As shown, the following improvements are made to the multi-modal based audio data recommendation method compared to the single-modal based audio data recommendation method:
[0068] a. The improvements made to step 2 in the training process of the above audio recognition models include: 1. The server can obtain at least two candidate audio recognition models, and the at least two candidate audio recognition models can include at least two of machine learning models, deep learning models, and graph network models, etc. The network attributes of each candidate audio recognition model are different, and the network attributes include at least one of network structure, network parameters, network algorithms, etc. Since the network attributes of each candidate audio recognition model are inconsistent, the feature processing capabilities of each candidate audio recognition model are inconsistent. For example, the candidate audio recognition model based on FM is good at mining the correlation relationships between feature information, and the candidate audio recognition model based on XGBoost is good at mining key split points (such as key feature points in video data), etc. 2. Implement multi-modal feature fusion: fuse the audio feature information of the sample audio data with the video feature information of the sample video data and the object feature information of the sample object to obtain the fused audio feature information of the sample audio data; the fused feature information of the sample audio data is used to reflect the preference of the sample object for the sample audio data, the correlation relationship between the sample audio data and the sample video data, etc. 3. Implement multi-modal training: use the fused audio feature information of the sample audio data and the audio feature information of the sample audio data to train at least two candidate audio recognition models respectively to obtain at least two target audio recognition models.
[0069] b. The improvements made to step 5 in the process of using the target audio recognition model to identify the target audio data include: 1. Multi-modal feature fusion: fuse the audio feature information of at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object respectively to obtain the fused audio feature information of at least two candidate audio data; the fused feature information of the candidate audio data is used to reflect the preference of the target object for the candidate audio data, the correlation relationship between the candidate audio data and the target video data, etc. 2. Recommendation decision fusion: use at least two target audio recognition models to identify the fused feature information of at least two candidate audio data and the audio feature information of at least two candidate audio data respectively to obtain the target audio data for scoring the target video data; that is, comprehensively consider the audio recognition results of each target audio recognition model and recommend the target video data to the target object to achieve recommendation decision fusion.
[0070] In summary, in the multi-modal based audio data recommendation method, by using the fused audio feature information of the sample audio data and the audio feature information of the sample audio data to train at least two candidate audio recognition models, at least two target audio recognition models are obtained, which can avoid the problem that a single audio recognition model has biases in the knowledge accumulation process, resulting in a relatively low accuracy of the recommended audio data. By fusing the audio feature information of at least two candidate audio data with the object feature information of the target object and the video feature information of the target video data, it is beneficial for the target audio recognition model to explore the implicit relationships between various feature information, and further improve the accuracy of the recommended audio data. By comprehensively considering the recognition results of at least two target audio data models and recommending audio data, the problem that over-reliance on a single audio recognition model leads to a relatively low accuracy of the recommended audio data can be avoided, and the accuracy of the recommended audio data can be improved.
[0071] It should be noted that the modality in this application refers to: the source or form of each type of information can be called a modality. For example, a person has tactile, auditory, visual, and olfactory senses; the media of information include speech, video, text, etc. A variety of sensors, such as radar, infrared, accelerometers, etc. Each of the above can be called a modality. Therefore, the multi-modal features of this application can include at least two of video feature information, audio feature information, and object feature information; the multi-modal audio recognition model in this application is also called multi-modal machine learning. Multi-modal machine learning (abbreviated as MMML) aims to achieve the ability to process and understand multi-modal information through machine learning methods, and establish a model that can process and associate information from multiple modes. It is a vibrant multi-disciplinary field with extraordinary potential.
[0072] Furthermore, please refer to Figure 4 , which is a schematic flowchart of an audio data processing method provided by an embodiment of this application. As Figure 4 shown, this method can be executed by a computer device, and the computer device can refer to the Figure 1 terminal, or the computer device can refer to the Figure 1 server, or the computer device includes the Figure 1 terminal and the server in Figure 1 , that is, this method can be jointly executed by the
[0073] S101. Obtain the object feature information of the target object, the video feature information of the target video data belonging to the above target object, and the audio feature information of at least two candidate audio data associated with the above target video data.
[0074] In this application, when a user needs to publish video data on a multimedia application platform, the user can be referred to as the target object, and the video data to be published can be referred to as the target video data. To select a suitable background music for the target video, the computer device can obtain the object feature information of the target object, the target video data belonging to the target object, and at least two candidate audio data associated with the target video data. Further, video feature extraction is performed on the target video data to obtain the video feature information of the target video data, and audio feature extraction is respectively performed on at least two candidate audio data to obtain the audio feature information of the at least two candidate audio data.
[0075] Among them, the object feature information of the target object may refer to at least one of the basic portrait feature information, multimedia portrait feature information, and portrait association feature information of the target object. The basic portrait feature information is used to reflect basic information such as the age and gender of the target object; the multimedia portrait feature information is used to reflect the multimedia preferences of the target object, such as favorite movies, poems, music, favorite singers, etc.; the portrait association feature information is used to reflect the association relationship between the basic portrait feature information and the multimedia portrait feature information. For example, the portrait association feature is used to reflect that the user group aged between [18, 25] years old likes singer A more. The video feature information of the target video data is used to reflect the theme information, scene, color information, quality information, etc. of the target video data. The audio feature information of the candidate audio data can be used to reflect the lyric information of the candidate audio data, the object feature information of the creator, the musical score information, etc. The at least two candidate audio data may refer to audio data that matches the theme information, scene, etc. of the target video data, or the at least two audio data may refer to the audio data played by the target object within a historical time period (such as the recent week, the recent month), or the at least two audio data may refer to the audio data created by the target object, or the at least two audio data may be currently popular music, such as audio data with a current playback volume greater than the playback volume threshold. It should be noted that the audio data involved in this application may refer to music, voice data for poem recitation, voice data for story narration, etc.
[0076] S102. Respectively fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain the audio fusion feature information of the at least two candidate audio data.
[0077] In this application, the computer device may fuse the audio feature information of the first candidate audio data among the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain the audio fusion feature information of the first candidate audio data; similarly, fuse the audio feature information of the second candidate audio data among the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain the audio fusion feature information of the second candidate audio data. By analogy, the audio fusion feature information of each candidate audio data among the at least two candidate audio data can be obtained.
[0078] It should be noted that the fusion implementation methods here include direct fusion and processing fusion. Direct fusion may refer to directly merging two or more feature information into one fusion feature information. For example, assuming that the audio feature information of the candidate audio data is (1, 2, 3) and the video feature information of the target video data is (4, 5, 6), directly merge the audio feature information of the candidate audio data with the video feature information of the target video data to obtain the audio fusion feature information (1, 2, 3, 4, 5, 6) of the candidate audio data. Direct fusion may also refer to merging the feature parameters with an associated relationship among two or more feature information into one fusion feature information. For example, assuming that the audio feature parameter 2 in the audio feature information of the candidate audio data has an associated relationship with the video feature parameter 5 in the video feature information of the target video data, then merge the feature parameters with an associated relationship in the audio feature information of the candidate audio data and the video feature information of the target video data to obtain the audio fusion feature information (2, 5) of the candidate audio data. Processing fusion means performing averaging processing or extracting the maximum value, etc. on two or more feature information to obtain one fusion feature information. For example, assuming that the audio feature information of the candidate audio data is (1, 2, 3) and the video feature information of the target video data is (4, 5, 6), perform averaging processing on the audio feature information of the candidate audio data and the video feature information of the target video data to obtain the audio fusion feature information (2.5, 3.5, 4.5) of the candidate audio data. Or, perform maximum value extraction processing on the audio feature information of the candidate audio data and the video feature information of the target video data to obtain the audio fusion feature information (4, 5, 6) of the candidate audio data.
[0079] S103. Use at least two target audio recognition models to perform audio recognition on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the at least two candidate audio data respectively to obtain target audio data for scoring the target video data; the target audio data belongs to the at least two candidate audio data.
[0080] In this application, the computer device can use at least two target audio recognition models to respectively perform audio recognition on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the at least two candidate audio data, obtain at least two audio recognition results, and determine, according to the at least two audio recognition results, target audio data for scoring the above-mentioned target video data from the at least two candidate audio data; by fusing the audio recognition results of multiple target audio recognition models to determine the target audio data, the advantages of different audio recognition models are fully utilized, which can effectively avoid the problem that a single model generates deviations, resulting in a relatively low accuracy of the recommended audio data, and can improve the robustness, accuracy, and credibility of the recommended audio data.
[0081] It should be noted that the audio recognition method includes non-discriminative recognition or discriminative recognition. Non-discriminative recognition means that the feature information processed by each target audio recognition model is the same. For example, assuming that the at least two target audio recognition models include a first target audio recognition model and a second target audio recognition model, the computer device can use the first target audio recognition model to perform audio recognition on the audio fusion feature information to obtain a first audio recognition result. Then, use the first target audio recognition model to perform audio recognition on the audio feature information of the at least two candidate audio data to obtain a second audio recognition result. Similarly, use the second target audio recognition model to perform audio recognition on the audio fusion feature information to obtain a third audio recognition result. Then, use the second target audio recognition model to perform audio recognition on the audio feature information of the at least two candidate audio data to obtain a fourth audio recognition result. Further, determine the target audio data for scoring the above-mentioned target video data according to the first audio recognition result, the second audio recognition result, the third audio recognition result, and the fourth audio recognition result; here, the first audio recognition result and the third audio recognition result are used to reflect the audio joint matching degree between each candidate audio data and the target object and the target video data. The audio joint matching degree is specifically used to reflect the preference degree of the target object for the candidate audio data and the matching degree between the candidate audio data and the target video data. The second audio recognition result and the third audio recognition result are used to reflect the applicability (i.e., audio self-matching degree) of each candidate audio data for scoring.
[0082] Similarly, differential recognition means that the feature information processed by each target audio recognition model is different. For example, a computer device may use a first target audio recognition model to perform audio recognition on the audio fusion feature information of the at least two candidate audio data to obtain a fifth audio recognition result, and then use a second target audio recognition model to perform audio recognition on the audio feature information of the at least two candidate audio data to obtain a sixth audio recognition result. Further, based on the fifth audio recognition result and the sixth audio recognition result, the target audio data for scoring the target video data is determined; here, the fifth audio recognition result is used to reflect the audio joint matching degree between each candidate audio data and the target object and the target video data, and the sixth audio recognition result is used to reflect the suitability of each candidate audio data for scoring (i.e., the audio self-matching degree).
[0083] S104. Recommend the target audio data to the above target object.
[0084] In this application, the number of the target audio data may be one or more. When the number of the target audio data is one, the computer device may display the target audio data on the release interface of the target video data, and in response to a selection request for the target audio data, use the target audio data to score the target video data. When the number of the target audio data is multiple, the computer device may display each target audio data on the release interface of the target video data in turn according to the total matching degree of each target audio data (here, the total matching degree may be determined according to the above audio joint matching degree and audio self-matching degree). For example, each target audio data may be displayed on the release interface of the target video data simultaneously in the order of the total matching degree of each target audio data from large to small, or each target audio data may be scrolled and displayed on the release interface of the target video data in the order of the total matching degree of each target audio data from large to small. For example, the target audio data with the total matching degree ranked 1-10 is displayed on the release interface of the target video data at the first time, and the target audio data with the total matching degree ranked 11-20 is displayed at the second time. Then, in response to a selection operation for any one of the multiple target audio data, the selected target audio data may be used to score the target video data. Through the audio recognition model, audio data can be automatically recommended to the target object, improving the accuracy of the recommended audio data and the efficiency of the recommended audio data without manual participation.
[0085] Optionally, each target audio recognition model can be trained based on the sample video data belonging to the sample object, the sample audio data used for scoring the sample video data, the object feature information of the sample object, and the labeled audio matching degree. The labeled audio data can be determined according to the object behavior data regarding the sample video data. The object behavior data includes at least one of the like count, follow count, share count, favorite count, and click count of the audience users for the sample video data. That is to say, the object behavior data reflects to a certain extent the preference degree of the audience users for the sample video data and the sample audio data. The target audio recognition model trained in this way has the ability to recommend audio data to the creator (creator of the video data) based on the preference of the audience users for the video data and the audio data. In summary, the audio recognition results output by each target audio recognition model can not only reflect the preference degree of the target object (creator of the target video) for the candidate audio data, the matching degree between the candidate audio data and the target video data, and the applicability of the candidate audio data for scoring, but also reflect to a certain extent the preference degree of the audience users for the candidate audio data. Therefore, by synthesizing the audio recognition results of each target audio recognition model to recommend the target audio data, the preference of the audience users for the multimedia (i.e., audio data and video data) can be transmitted to the creator, effectively breaking down the barrier between the creator and the audience users, expanding the creative ideas of the creator, and at the same time, more works that are liked by both the audience users and the creator will be produced under the guidance of the recommendation.
[0086] In this application, by fusing the audio feature information of at least two candidate audio data with the object feature information of the target object and the video feature information of the target video data, the audio fusion feature information of at least two candidate audio data is obtained. That is, by fusing the multi-modal feature information, it is beneficial to provide more information for recommending audio data and improve the accuracy of the recommended audio data. Further, by using at least two target audio recognition models to respectively recognize the audio fusion feature information of at least two candidate audio data and the audio feature information of at least two candidate audio data, the target audio data used for scoring the target video data is obtained, and the target audio data is recommended to the target object. That is, by comprehensively considering the audio recognition results of the multi-modal audio recognition models, the audio data is automatically recommended to the target object, which can improve the efficiency of the recommended audio data. At the same time, the advantages of different audio recognition models are fully utilized, which can effectively avoid the problem that a single model produces deviations and leads to a relatively low accuracy of the recommended audio data, and can make the recommended audio data more robust, accurate, and credible.
[0087] Further, please refer to Figure 5 , which is a schematic flowchart of a method for processing audio data provided by an embodiment of this application. AsFigure 5 As shown, this method can be executed by a computer device, which can refer to Figure 1 the terminal in Figure 1 the server in Figure 1 the terminal and server in Figure 1 That is, this method can be jointly executed by the terminal and server in . The audio data processing method may include the following steps S201 to S206:
[0088] S201. Obtain the object feature information of the target object, the video feature information of the target video data belonging to the above target object, and the audio feature information of at least two candidate audio data associated with the above target video data.
[0089] Optionally, obtaining the object feature information of the above target object in step S201 may include the following steps s11 to s13:
[0090] s11. Obtain the basic portrait feature information and multimedia portrait feature information of the above target object.
[0091] s12. Perform portrait association recognition on the basic portrait feature information and the multimedia portrait feature information of the above target object to obtain portrait association feature information.
[0092] s13. Determine the basic portrait feature information, the multimedia portrait feature information, and the portrait association feature information of the above target object as the object feature information of the above target object.
[0093] In steps s11 to s13, as Figure 6 shown, the computer device can obtain the basic portrait feature information and multimedia portrait feature information of the target object; the basic portrait feature information includes basic attribute features such as age and gender, and the multimedia portrait feature information includes singers, movie actors, favorite movies, songs, etc. that the target object likes. Further, the portrait association recognition model can perform portrait association recognition on the basic portrait feature information and the multimedia portrait feature information of the above target object to obtain portrait association feature information, which is used to reflect the implicit association relationship between the basic portrait feature information and the multimedia portrait feature information of the target object. As Figure 6Among them, the association recognition model can be a deep neural network, which is composed of multiple neural network layers. The output result of the previous neural network layer is input to the next neural network layer for processing in a forward propagation manner between different neural network layers. Through this deep neural network, the implicit relationship between the basic portrait feature information and the multimedia portrait feature information of the target object can be mined, and the expression ability of the feature information can be improved. Then, the basic portrait feature information, the multimedia portrait feature information, and the portrait association feature information of the target object are determined as the object feature information of the target object; by mining the implicit relationship between the basic portrait feature information and the multimedia portrait feature information of the target object, rich information is provided for the recommended audio data, and the accuracy of the recommended audio data is improved.
[0094] Optionally, the obtaining of the audio feature information of at least two candidate audio data associated with the target video data in step S201 may include the following steps s21 to s25:
[0095] s21. Obtain at least two candidate audio data associated with the target video data.
[0096] s22. Determine the object feature information of the creators of the at least two candidate audio data.
[0097] s23. Extract the lyric feature of the at least two candidate audio data to obtain the lyric feature information of the at least two candidate audio data.
[0098] s24. Extract the score feature of the at least two candidate audio data to obtain the score feature information of the at least two candidate audio data.
[0099] s25. Integrate the object feature information of the creators, the lyric feature information of the at least two candidate audio data, and the score feature information of the at least two candidate audio data to obtain the audio feature information of the at least two candidate audio data.
[0100] In steps s21 to s25, such as Figure 7As shown, when the candidate audio data is music, the computer device can obtain video attributes of the target video data, such as theme information and scene information (e.g., shooting scene), and obtain at least two candidate audio data associated with the target video data according to the video attributes. Then, obtain the corresponding creator information (i.e., singer information) of the creators of the at least two candidate audio data. The creator information includes basic portrait feature information and multimedia portrait feature information of the creator. Use an association recognition model (such as a deep neural network) to perform association recognition on the basic portrait feature information and multimedia portrait feature information of the creator to obtain the portrait association feature information of the creator, and determine the portrait association feature information, basic portrait feature information, and multimedia portrait feature information of the creator as the object feature information of the creator. The object feature information of the creator can be called the song meta-information vector. Next, the above at least two candidate audio data can be text-converted to obtain the text information of the at least two candidate audio data, perform word segmentation on the text information to obtain multiple segmented words, and use a word statistics method such as TF-IDF (Term Frequency–Inverse Document Frequency) or WordRank to extract the backbone entity words of each candidate audio data from the multiple segmented words. The backbone entity words can refer to the keywords of the candidate audio data, that is, the words that reflect the theme of the candidate audio data. Use a word vector conversion model such as WordVec or Bert to convert the backbone entity words of the candidate audio data into lyric vectors, and the lyric vectors can be called lyric feature information. Next, the above at least two candidate audio data can be pre-emphasized, framed, etc. to obtain the score feature information of the above at least two candidate audio data, and fuse the object feature information of the above creator, the lyric feature information of the above at least two candidate audio data, and the score feature information of the above at least two candidate audio data to obtain the audio feature information of the above at least two candidate audio data.
[0101] Optionally, step s44 above may include steps s31 to s33 as follows:
[0102] s31. Perform frame segmentation on the candidate audio data Yi in the above at least two candidate audio data to obtain at least two frames of audio data belonging to the candidate audio data Yi; i is a positive integer less than or equal to N, and N is the number of candidate audio data in the above at least two candidate audio data.
[0103] s32. Perform frequency domain transformation on at least two frames of audio data belonging to the candidate audio data Yi to obtain the frequency domain information of the candidate audio data Yi.
[0104] s33. Extract the musical score features from the frequency domain information of the above candidate audio data Yi to obtain the musical score feature information of the at least two candidate audio data.
[0105] In steps s31 to s33, as Figure 7 shown, the computer device can perform pre-emphasis processing on each candidate audio data, and the processed candidate audio data; the role of pre-emphasis processing is to eliminate the effects caused by the vocal cords and lips during the sound production process, to compensate for the high-frequency part of the speech signal suppressed by the pronunciation system; and to highlight the high-frequency formants. Then, frame each of the processed candidate audio data according to the frame parameters to obtain at least two frames of audio data for each candidate audio data; the frame parameters can include the frame length and the frame shift, for example, the frame length can be 20 to 40 ms, and the frame shift can be 10 ms. Next, windowing processing can be performed on each frame of audio data to make the attenuation at both ends of each frame of audio data close to zero; perform frequency domain transformation on the windowed audio data of each frame to obtain the frequency domain information of each candidate audio data, and this frequency domain information is used to reflect the frequency and amplitude of the candidate audio data. Then, musical score features can be extracted from the frequency domain information of each candidate audio data to obtain the musical score feature information of each candidate audio data, and the musical score feature information is used to reflect parameters such as the frequency and energy of the candidate audio data. By performing pre-emphasis, frequency domain transformation and other processing on the candidate audio data to obtain the musical score feature information of the candidate audio data, the complexity of obtaining the musical score feature information of the candidate audio data can be reduced, and the significance of the musical score feature information of the candidate audio data can be improved.
[0106] Optionally, the above step s33 may include steps s41 to s43 as follows:
[0107] s41. Determine the energy information of the candidate audio data Yi according to the frequency domain information of the candidate audio data Yi.
[0108] s42. Perform filtering processing on the energy information of the candidate audio data Yi to obtain the filtered energy information.
[0109] s43. Determine the filtered energy information as the musical score feature information of the at least two candidate audio data.
[0110] In steps s41 to s43, as Figure 7As shown, the computer device can determine the energy information of the candidate audio data Yi based on the frequency domain information of the candidate audio data Yi. Since the frequency range of the sound that the human ear can perceive is limited, that is, the audio corresponding to the frequency that the human ear cannot perceive is called noise. Therefore, a filter can be generated according to the auditory characteristics of the human ear, and the filter is used to filter the energy information of the candidate audio data Yi to obtain the filtered energy information. Further, the filtered energy information is determined as the musical score feature information of the at least two candidate audio data. By filtering the energy information of the candidate audio data, the problem that the accuracy of the obtained musical score feature information is not high due to noise interference can be effectively avoided, and the subsequent processing of invalid noise can be avoided, which can save processing resources.
[0111] Optionally, obtaining the video feature information of the target video data belonging to the target object in step S201 may include the following steps s51 to s54:
[0112] s51. Obtain the target video data belonging to the target object.
[0113] s52. Extract at least two key video frames of the target video data.
[0114] s53. Perform video feature extraction on the at least two key video frames to obtain the video feature information of the at least two key video frames.
[0115] s54. Perform fusion on the video feature information of the at least two key video frames to obtain the video feature information of the target video data.
[0116] In steps s51 to s54, as Figure 8 shown, the computer device can obtain the target video data belonging to the target object, extract at least two key frames (i.e., representative frames) of the target video data. The key frame may refer to an audio data frame in the target video data that can reflect the theme information of the target video data. Further, a video feature extraction network is used to perform video feature extraction on the at least two key frames to obtain the video feature information of the at least two key video frames; as Figure 8Among them, the video feature extraction network may refer to a convolutional neural network (CNN), which is composed of multiple convolutional layers and pooling layers. Convolutional layer: The input of each node in the convolutional layer is only a small piece of the previous layer of the neural network (usually with a size of 3*3 or 5*5). The convolutional layer attempts to analyze each small piece in the neural network more deeply to obtain features with a higher degree of abstraction. Pooling layer: The pooling layer will not change the depth of the three-dimensional matrix, but it can reduce the size of the matrix, further reducing the number of nodes in the last fully connected layer, thereby achieving the purpose of reducing the parameters in the entire neural network. Therefore, through the convolutional neural network, more in-depth and less redundant video feature information of the video data can be extracted. Then, the video feature information of at least two of the above key video frames can be fused to obtain the video feature information of the above target video data. Extracting video features from key video frames is beneficial to mining the implicit video feature information in the target video and reducing the redundancy of the video feature information.
[0117] S202. Respectively fuse the audio feature information of at least two of the above candidate audio data with the video feature information of the above target video data and the object feature information of the above target object to obtain the audio fusion feature information of at least two of the above candidate audio data.
[0118] Optionally, step S202 may include the following steps s61 to s63:
[0119] s61. Fuse the audio feature information of at least two of the above candidate audio data with the video feature information of the above target video data to obtain the first fusion feature information, and fuse the audio feature information of at least two of the above candidate audio data with the object feature information of the above target object to obtain the second fusion feature information.
[0120] s62. Fuse the audio feature information of at least two of the above candidate audio data, the video feature information of the above target video data, and the object feature information of the above target object to obtain the third fusion feature information.
[0121] s63. Determine the first fusion feature information, the second fusion feature information, and the third fusion feature information as the audio fusion feature information of at least two of the above candidate audio data.
[0122] In steps S61 - S63, the computer device may use the direct fusion method or the processing fusion method to fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data to obtain the first fusion feature information, and use the direct fusion method or the processing fusion method to fuse the audio feature information of the at least two candidate audio data with the object feature information of the target object to obtain the second fusion feature information. Similarly, use the direct fusion method or the processing fusion method to fuse the audio feature information of the at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object to obtain the third fusion feature information. The first fusion feature information, the second fusion feature information, and the third fusion feature information may be determined as the audio fusion feature information of the at least two candidate audio data; or, the first fusion feature information and the second fusion feature information may be determined as the audio fusion feature information of the at least two candidate audio data; or, the first fusion feature information and the third fusion feature information may be determined as the audio fusion feature information of the at least two candidate audio data, or the second fusion feature information and the third fusion feature information may be determined as the audio fusion feature information of the at least two candidate audio data.
[0123] Optionally, when selecting the method of extracting correlation parameters (i.e., the direct fusion method) to fuse the feature information, step S61 may include the following steps S71 - S74:
[0124] S71. Obtain the first video feature parameter and the first audio feature parameter with a correlation relationship; the first video feature parameter belongs to the video feature information of the target video data, and the first audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0125] S72. Generate the first fusion feature information according to the first video feature parameter and the first audio feature parameter;
[0126] S73. Obtain the first object feature parameter and the second audio feature parameter with a correlation relationship; the first object feature parameter belongs to the object feature information of the target object, and the second audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0127] S74. Generate the second fusion feature information according to the first object feature parameter and the second audio feature parameter.
[0128] In steps s71 to s74, the computer device may obtain a first video feature parameter and a first audio feature parameter having an associated relationship. The first video feature parameter and the first audio feature information having an associated relationship may refer to video feature parameters and audio feature parameters that have a positive effect on the recommended audio data. The first fusion feature information may be generated based on the first video feature parameter and the first audio feature parameter. Similarly, the first object feature parameter and the second audio feature parameter having an associated relationship are obtained; the first object feature parameter and the second audio feature parameter having an associated relationship may refer to object feature parameters and audio feature parameters that have a positive effect on the recommended audio data; then, the second fusion feature information is generated based on the first object feature parameter and the second audio feature parameter. By extracting video feature parameters and audio feature parameters having an associated relationship from the video feature information and the audio feature information, it is beneficial to mine the implicit information and implicit relationship in the video feature information and the audio feature information, greatly reducing the reliance on manual work and improving the accuracy of the recommended audio data.
[0129] Optionally, when the method of extracting associated parameters is selected to fuse the feature information, step s62 may include the following steps s75 to s76:
[0130] s75. Obtain a second object feature parameter, a second video feature parameter and a third audio feature parameter that have an associated relationship; the second object feature parameter belongs to the object feature information of the target object, the second video feature parameter belongs to the video feature information of the target video data, and the third audio feature parameter belongs to the audio feature information of the at least two candidate audio data.
[0131] s76. Generate third fusion feature information based on the second object feature parameter, the second video feature information and the third audio feature parameter.
[0132] In steps s75 to s76, the computer device can obtain the second object feature parameter, the second video feature parameter and the third audio feature parameter with an associated relationship; the second object feature parameter, the second video feature parameter and the third audio feature parameter with an associated relationship refer to the object feature parameter, the audio feature parameter and the video feature parameter that have a positive effect on the recommended audio data; then, the third fusion feature information can be generated according to the second object feature parameter, the second video feature information and the third audio feature parameter. By extracting the video feature parameter, the audio feature parameter and the object feature parameter with an associated relationship from the video feature information, the audio feature information and the object feature information, it is beneficial to mine the implicit information and implicit relationship in the video feature information, the audio feature information and the object feature information, greatly reducing the dependence on manual work and improving the accuracy of the recommended audio data.
[0133] S203. Respectively determine a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data and a second target audio recognition model that matches the audio feature information of the at least two candidate audio data from the at least two target audio recognition models described above.
[0134] In this application, the computer device can use different target audio recognition models to process different characteristic information. Specifically, the computer device can select the target audio recognition model in a random selection method or according to the feature processing ability selection method. For example, when the computer device uses the random selection method, randomly select a target audio recognition model from the at least two target audio recognition models as the first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data, and randomly select a target audio recognition model from the remaining target audio recognition models as the second target audio recognition model that matches the audio feature information of the at least two candidate audio data.
[0135] Optionally, when selecting the target audio recognition model according to the feature processing ability selection method, step S203 may include the following steps: obtain the feature processing ability information of the at least two target audio recognition models; respectively determine a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data and a second target audio recognition model that matches the audio feature information of the target audio data from the at least two target audio recognition models according to the feature processing ability information.
[0136] The computer device can obtain the feature processing capability information of each of at least two target audio recognition models. The feature processing capability information is used to reflect the feature information that the target audio recognition model is good at processing. Then, according to the feature processing capability information, the first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data and the second target audio recognition model that matches the audio feature information of the target audio data can be determined from the at least two target audio recognition models respectively. By selecting the target audio recognition model for processing feature information according to the feature processing capability information, it is beneficial to improve the accuracy of feature information processing. For example, the target audio recognition model based on FM is good at mining the correlation relationships between feature information, and the target audio recognition model based on XGBoost is good at mining key split points, etc.; therefore, the target audio recognition model based on FM can be determined as the target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data, so as to be able to mine the implicit relationships between audio feature information, video feature information, and object feature information; the target audio recognition model based on XGBoost can be determined as the target audio recognition model that matches the audio feature information of the at least two candidate audio data, so as to be able to mine the key audio feature information in the audio feature information (i.e., the implicit information in the audio feature information).
[0137] S204. Use the above first target audio recognition model to perform audio joint relationship recognition on the audio fusion feature information of the at least two candidate audio data to obtain an audio joint matching degree; use the above second target audio recognition model to perform audio autocorrelation recognition on the audio feature information of the target audio data to obtain an audio autocorrelation matching degree.
[0138] In this application, the computer device can use the first target audio recognition model to perform audio joint relationship recognition on the audio fusion feature information of the at least two candidate audio data, and obtain an audio joint matching degree, which is used to reflect the association relationship between the candidate audio data, the target object, and the target video data. Specifically, the audio joint matching degree is used to reflect the degree of preference of the target object for the candidate audio data and the matching degree between the candidate audio data and the target video data. Further, the second target audio recognition model can be used to perform audio autocorrelation recognition on the audio feature information of the target audio data to obtain an audio autocorrelation matching degree, which is used to reflect the suitability of the candidate audio data for background music. Performing audio joint relationship recognition on the audio fusion feature information through the first target audio recognition model is beneficial to exploring the implicit relationships among audio feature information, video feature information, and object feature information; performing audio autocorrelation recognition on the audio feature information through the second target audio recognition model is beneficial to exploring the implicit information within the audio feature information. By exploring the deeper information in the feature information, it is beneficial to improve the accuracy of the recommended audio data.
[0139] S205. Select target audio data for performing background music on the target video data from the at least two candidate audio data according to the above audio joint matching degree and the above audio autocorrelation matching degree.
[0140] In this application, the computer device can determine the total matching degree of each candidate audio data according to the audio joint matching degree and the audio autocorrelation matching degree, and select target video data for performing background music on the target video data from the at least two candidate audio data according to the total matching degree of each candidate audio data. By comprehensively considering the audio recognition results of the multi-modal audio recognition model to determine the target audio data, it can effectively avoid the problem of low accuracy of the recommended audio data caused by the deviation of a single model, and can make the recommended audio data more robust, accurate, and credible.
[0141] Optionally, step S205 may include the following steps s81 to s82:
[0142] s81. Perform a summation process on the above audio joint matching degree and the above audio autocorrelation matching degree to obtain the total matching degree.
[0143] s82. Determine the candidate audio data with a total matching degree greater than the matching degree threshold among the at least two candidate audio data as the target audio data for performing background music on the target video data.
[0144] In steps S81 - S82, the computer device can accumulate the above - mentioned audio joint matching degree and the above - mentioned audio self - matching degree to obtain the total matching degree; alternatively, it can perform a weighted summation process on the above - mentioned audio joint matching degree and the above - mentioned audio self - matching degree to obtain the total matching degree. Further, the computer device can determine the candidate audio data with a total matching degree greater than the matching degree threshold among the above - mentioned at least two candidate audio data as the target audio data for scoring the above - mentioned target video data. By performing a summation process on the audio recognition results of the multi - modal audio recognition model to determine the target video data, it can effectively avoid the problem that a single model produces deviations, resulting in a relatively low accuracy of the recommended audio data, and can improve the robustness, accuracy, and credibility of the recommended audio data.
[0145] Optionally, when the computer device obtains the total matching degree by performing a weighted summation process on the above - mentioned audio joint matching degree and the above - mentioned audio self - matching degree, step S81 may include: obtaining the recognition weight of the first target audio recognition model and the recognition weight of the second target audio recognition model; performing a weighted process on the above - mentioned audio joint matching degree using the recognition weight of the first target audio recognition model to obtain the weighted audio joint matching degree; performing a weighted process on the above - mentioned audio self - matching degree using the recognition weight of the second target audio recognition model to obtain the weighted audio self - matching degree; performing a summation process on the weighted audio joint matching degree and the weighted audio self - matching degree to obtain the total matching degree.
[0146] The computer device can obtain the recognition weight of the first target audio recognition model and the recognition weight of the second target audio recognition model; the recognition weight of the first target audio recognition model and the recognition weight of the second target audio recognition model can be determined according to the audio recognition accuracy of the corresponding target audio recognition model; or the recognition weight of the first target audio recognition model and the recognition weight of the second target audio recognition model can be set according to the application scenario; for example, the target video data is obtained by editing a certain movie, and the creator of the target video data hopes that the target video data can get more clicks. Therefore, the recognition weight of the first target audio recognition model can be set higher than the recognition weight of the second target audio recognition model, which is beneficial to highlighting the audio joint matching degree and is beneficial to recommending audio data that the public likes. For example, the computer device can use the following formula (1) to calculate the total matching degree of each candidate audio data:
[0147]
[0148] Among them, in formula (1), P j represents the total matching degree of the j - th candidate audio data, that is, P jThe final inference score given for the entire multi-modal target audio recognition model (i.e., at least two target audio recognition models), Q ji During the process of recognizing the j-th candidate audio data, it is the audio recognition result output by the i-th target audio recognition model. This target audio recognition model can be a classifier, i.e., Q ji It is the score given by a single classifier to the candidate audio data, w i It is the recognition weight of the i-th target audio recognition model, and N is the number of models in at least two target audio recognition models.
[0149] S206. Recommend the above-mentioned target audio data to the above-mentioned target object.
[0150] For example, as Figure 9 shown, the computer device analyzes at least two candidate audio data, target video data, and target objects to obtain the feature information of modalities 1 to N. The feature information included in the feature information of modalities 1 to N is different. For example, modality 1 is video feature information, modality 2 is the first audio feature information, and the first audio feature information includes lyric feature information, musical score feature information, and singer information,..., modality N is the j-th audio feature information, and the j-th audio feature information includes lyric feature information and singer feature information. Further, feature parameters can be extracted from the multi-modal feature information. As Figure 9 shown, in it, the video features in modality 1 are fused with each audio feature parameter in modality 2 to obtain the audio fusion feature information 1. The lyric feature information and singer features are extracted from modality 2 as the audio feature information 2,... There are N target audio recognition models in the computer device, which are models 1 to n respectively. The audio fusion feature information 1 can be audio-recognized by using model 1 to obtain the audio recognition result 1 (i.e., the matching degree), the audio feature information 1 can be audio-recognized by using model 2 to obtain the audio recognition result 2,..., and the audio feature information in modality N can be audio-recognized by using model n to obtain the audio recognition result n. Then, the audio recognition results 1 to n can be fused (summed) to obtain the total matching degree of each candidate audio data. The candidate audio data with the top 10 total matching degrees can be selected from at least two candidate audio data as the target audio data for scoring the target audio data, and this target audio data is recommended to the target object.
[0151] In this application, by comprehensively considering the audio recognition results of the multi-modal audio recognition model, audio data is automatically recommended to the target object, which can improve the efficiency of recommending audio data; at the same time, the advantages of different audio recognition models are fully utilized, which can effectively avoid the problem that a single model generates deviations and leads to a relatively low accuracy of recommending audio data, and can make the recommended audio data more robust, accurate, and credible.
[0152] Further, please refer to Figure 10 , which is a schematic flowchart of an audio data processing method provided by an embodiment of the present application. As Figure 10 shown, this method can be executed by a computer device, and the computer device can refer to Figure 1 the terminal in Figure 1 , or the computer device can refer to Figure 1 the server in Figure 1 , or the computer device includes Figure 1 the terminal and the server in Figure 1 , that is, this method can be jointly executed by Figure 1 the terminal and the server in
[0153] S301. Obtain the object feature information of the sample object, the sample video data belonging to the sample object, the sample audio data for scoring the sample video data, and the labeled audio matching degree of the sample audio data; the labeled audio matching degree is used to reflect the matching degree between the sample audio data and the sample object and the sample video data.
[0154] In the present application, the sample object refers to a user who has published video data on a multimedia application platform. The published video data is called sample video data, and the audio data used to score the sample video data is called sample audio data. The computer device can obtain the object feature information of the sample object, the sample video data belonging to the sample object, the sample audio data for scoring the sample video data, and the labeled audio matching degree of the sample audio data from the multimedia application platform. The labeled audio matching degree is used to reflect the matching degree between the sample audio data and the sample object and the sample video data. The labeled audio matching degree can be obtained by multiple professional users labeling the sample audio data; or the labeled audio matching degree can be obtained according to the object behavior data of the sample video data, and the object behavior data includes at least one of the like volume, attention volume, forwarding volume, collection volume, and click volume of the sample video data.
[0155] It should be noted that the sample video data in this application can refer to short videos or non - short videos. A short video can refer to video data with a playing duration less than a duration threshold. At the same time, the sample data can be obtained by screening candidate video data according to the attribute information of the candidate video data. For example, the attribute information can refer to clarity, duration, whether it is original, etc. That is, the sample data can be original video data, video data with relatively high clarity, etc. Optionally, the above - mentioned labeled audio matching degree is obtained according to the object behavior data of the sample video data. Specifically, a computer device can obtain the object behavior data about the sample video data and determine the labeled audio matching degree of the sample audio data according to the object behavior data about the sample video data.
[0156] A computer device can obtain the object behavior data about the sample video data from a multimedia application platform. The object behavior data includes at least one of the like count, follow count, share count, favorite count, and click count of the sample video data. Determine the labeled audio matching degree of the sample audio data according to the object behavior data about the sample video data. For example, the labeled audio matching degree increases as at least one of the like count, follow count, share count, favorite count, and click count of the sample video data increases. In particular, if the sample video data has been clicked, followed, shared, etc., then the sample video data is regarded as a positive sample. Conversely, if the sample video data has not been clicked, followed, shared, etc., then the sample video data is regarded as a negative sample.
[0157] This labeled matching degree can not only reflect the matching degree between the above - mentioned sample audio data and the above - mentioned sample object and the above - mentioned sample video data, but also be used to reflect the preferences of viewer users for the sample video data and the sample video data. Therefore, determining the labeled audio matching degree of the sample audio data according to the object behavior data about the sample video data has the following beneficial effects: 1. By training a candidate audio recognition model according to this labeled audio matching degree, the trained target audio recognition model has the ability to recommend audio data to creators (creators of video data) based on the preferences of viewer users for video data and audio data. That is, it can transmit the multimedia preferences of viewer users to creators, effectively breaking down the barriers between creators and viewer users, expanding the creative ideas of creators, and at the same time, more works that are liked by both viewer users and creators will be produced under the guidance of the recommendation. 2. There is no need for manual annotation of sample audio data, which can avoid problems such as missed annotation and mis - annotation in manual annotation; it can improve the accuracy of the labeled audio matching degree and the efficiency of obtaining the labeled audio matching degree.
[0158] S302. Extract video signs from the above sample video data to obtain the video feature information of the above sample video data, and extract audio features from the above sample audio data to obtain the audio feature information of the above sample audio data.
[0159] S303. Integrate the audio feature information of the above sample audio data with the video feature information of the above sample video data and the object feature information of the above sample object to obtain the audio integration feature information of the above sample audio data.
[0160] In this application, the computer device can integrate the audio feature information of the above sample audio data with the video feature information of the above sample video data and the object feature information of the above sample object in a direct integration manner or a processing integration manner to obtain the audio integration feature information of the above sample audio data.
[0161] S304. Adjust at least two candidate audio recognition models respectively according to the above-mentioned labeled audio matching degree, the audio feature information of the above sample audio data, and the audio integration feature information of the above sample audio data to obtain the above at least two target audio recognition models.
[0162] In this application, using the above-mentioned labeled audio matching degree, the audio feature information of the above sample audio data, and the audio integration feature information of the above sample audio data as training data, iteratively train at least two candidate audio recognition models respectively to obtain the above at least two target audio recognition models. By training the candidate audio recognition models, the accuracy of the recommended audio data can be improved.
[0163] Optionally, the above step S304 may include the following steps s91 to s92:
[0164] s91. Respectively use the above at least two candidate audio recognition models to perform audio matching prediction on the audio feature information of the above sample audio data and the audio integration feature information of the above sample audio data to obtain the predicted audio matching degree.
[0165] s92. Adjust at least two candidate audio recognition models respectively according to the above predicted audio matching degree and the above labeled audio matching degree to obtain the above at least two target audio recognition models.
[0166] In steps S91 - S92, the training methods for candidate audio recognition models include undifferentiated training and differentiated training. Undifferentiated training means training each candidate audio recognition model with the same feature information. For example, if at least two candidate audio recognition models include a first candidate audio recognition model and a second candidate audio recognition model, the first candidate audio recognition model can be used to predict the audio joint relationship of the audio fusion feature information of the sample audio data to obtain a first prediction result, and the first candidate audio recognition model can be used to identify the audio autocorrelation relationship of the audio feature information of the sample audio data to obtain a second prediction result. The prediction audio matching degree of the first candidate audio recognition model is determined based on the first prediction result and the second prediction result. The first candidate audio recognition model is adjusted according to the labeled audio matching degree and the prediction audio matching degree of the first candidate audio recognition model to obtain a first target audio recognition model. Similarly, the second candidate audio recognition model can be used to predict the audio joint relationship of the audio fusion feature information of the sample audio data to obtain a third prediction result, and the second candidate audio recognition model can be used to identify the audio autocorrelation relationship of the audio feature information of the sample audio data to obtain a fourth prediction result. The prediction audio matching degree of the second candidate audio recognition model is determined based on the third prediction result and the fourth prediction result; the second candidate audio recognition model is adjusted according to the labeled audio matching degree and the prediction audio matching degree of the second candidate audio recognition model to obtain a second target audio recognition model.
[0167] Similarly, differentiated training trains each candidate audio recognition model with different feature information. For example, training can be performed on the candidate audio recognition models according to their feature processing capabilities. For example, the first candidate audio recognition model is good at processing fused audio feature information, and the second candidate audio recognition model is good at processing audio feature information; therefore, the first candidate audio recognition model can be used to predict the audio joint relationship of the audio fusion feature information of the sample audio data to obtain the prediction audio matching degree of the first candidate audio recognition model, and the first candidate audio recognition model is adjusted according to the labeled audio matching degree and the prediction audio matching degree of the first candidate audio recognition model to obtain a first target audio recognition model. The second candidate audio recognition model is used to identify the audio autocorrelation relationship of the audio feature information of the sample audio data to obtain the prediction audio matching degree of the second candidate audio recognition model, and the second candidate audio recognition model is adjusted according to the labeled audio matching degree and the prediction audio matching degree of the second candidate audio recognition model to obtain a second target audio recognition model.
[0168] It should be noted that when the target audio recognition model is obtained through an undifferentiated training method, the above audio recognition method is an undifferentiated recognition method; when the target audio recognition model is obtained through a differentiated training method, the above audio recognition method is a differentiated recognition method.
[0169] Optionally, step s92 may include: respectively determining the prediction errors of the at least two candidate audio recognition models according to the prediction audio matching degree and the labeled audio matching degree; if the prediction errors are not in a convergent state, adjusting the at least two candidate audio recognition models respectively according to the prediction errors to obtain the at least two target audio recognition models.
[0170] If the difference between the prediction audio matching degree and the labeled audio matching degree is relatively small, it indicates that the audio recognition accuracy of the candidate audio recognition model is relatively high (i.e., the prediction error is relatively low); if the difference between the prediction audio matching degree and the labeled audio matching degree is relatively large, it indicates that the audio recognition accuracy of the candidate audio recognition model is relatively low (i.e., the prediction error is relatively high). Therefore, the computer device can respectively determine the prediction errors of the at least two candidate audio recognition models according to the prediction audio matching degree and the labeled audio matching degree; if the prediction errors are in a convergent state, it indicates that the audio recognition accuracy of the candidate audio recognition model is relatively high, so the candidate audio recognition model can be used as the target audio recognition model. If the prediction errors are not in a convergent state, it indicates that the audio recognition accuracy of the candidate audio recognition model is relatively low, then the at least two candidate audio recognition models are adjusted respectively according to the prediction errors to obtain the at least two target audio recognition models.
[0171] In this application, by using the fused audio feature information of the sample audio data and the audio feature information of the sample audio data to train at least two candidate audio recognition models to obtain at least two target audio recognition models, it is possible to avoid the problem that a single audio recognition model has biases in the knowledge accumulation process, resulting in relatively low accuracy of the recommended audio data.
[0172] Please refer to Figure 11 , which is a schematic structural diagram of an audio data processing device provided by an embodiment of this application. The above audio data processing device may be a computer program (including program code) running in a computer device. For example, the audio data processing device is an application software; the device can be used to execute the corresponding steps in the method provided by the embodiment of this application. As Figure 11 shown, the audio data processing device may include: an acquisition module 111, a fusion module 112, an identification module 113, and a recommendation module 114.
[0173] An acquisition module, configured to acquire object feature information of a target object, video feature information of target video data belonging to the target object, and audio feature information of at least two candidate audio data associated with the target video data;
[0174] A fusion module, configured to respectively fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain audio fusion feature information of the at least two candidate audio data;
[0175] An identification module, configured to respectively perform audio identification on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the at least two candidate audio data by using at least two target audio identification models to obtain target audio data for scoring the target video data; the target audio data belongs to the at least two candidate audio data;
[0176] A recommendation module, configured to recommend the target audio data to the target object.
[0177] Optionally, the fusion module respectively fuses the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain audio fusion feature information of the at least two candidate audio data, including:
[0178] Fusing the audio feature information of the at least two candidate audio data with the video feature information of the target video data to obtain first fusion feature information, and fusing the audio feature information of the at least two candidate audio data with the object feature information of the target object to obtain second fusion feature information;
[0179] Fusing the audio feature information of the at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object to obtain third fusion feature information;
[0180] Determining the first fusion feature information, the second fusion feature information, and the third fusion feature information as the audio fusion feature information of the at least two candidate audio data.
[0181] Optionally, the fusion module fuses the audio feature information of the at least two candidate audio data with the video feature information of the target video data to obtain first fusion feature information, and fuses the audio feature information of the at least two candidate audio data with the object feature information of the target object to obtain second fusion feature information, including:
[0182] Obtain a first video feature parameter and a first audio feature parameter with an associated relationship; the first video feature parameter belongs to the video feature information of the target video data, and the first audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0183] Generate first fusion feature information according to the first video feature parameter and the first audio feature parameter;
[0184] Obtain a first object feature parameter and a second audio feature parameter with an associated relationship; the first object feature parameter belongs to the object feature information of the target object, and the second audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0185] Generate second fusion feature information according to the first object feature parameter and the second audio feature parameter.
[0186] Optionally, the fusion module fuses the audio feature information of the at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object to obtain third fusion feature information, including:
[0187] Obtain a second object feature parameter, a second video feature parameter, and a third audio feature parameter with an associated relationship; the second object feature parameter belongs to the object feature information of the target object, the second video feature parameter belongs to the video feature information of the target video data, and the third audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0188] Generate third fusion feature information according to the second object feature parameter, the second video feature information, and the third audio feature parameter.
[0189] Optionally, the recognition module respectively performs audio recognition on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the target audio data by using the at least two target audio recognition models to obtain target audio data for scoring the target video data, including:
[0190] Respectively determine a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data and a second target audio recognition model that matches the audio feature information of the at least two candidate audio data from the at least two target audio recognition models;
[0191] Use the first target audio recognition model to perform audio joint relationship recognition on the audio fusion feature information of the at least two candidate audio data to obtain an audio joint matching degree; use the second target audio recognition model to perform audio autocorrelation recognition on the audio feature information of the target audio data to obtain an audio autocorrelation matching degree;
[0192] Select target audio data for scoring the target video data from the at least two candidate audio data according to the audio joint matching degree and the audio autocorrelation matching degree.
[0193] Optionally, the recognition module respectively determines a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data and a second target audio recognition model that matches the audio feature information of the at least two candidate audio data from the at least two target audio recognition models, including:
[0194] Obtain the feature processing ability information of the at least two target audio recognition models;
[0195] According to the feature processing ability information, respectively determine a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data and a second target audio recognition model that matches the audio feature information of the target audio data from the at least two target audio recognition models.
[0196] Optionally, the recognition module selects target audio data for scoring the target video data from the at least two candidate audio data according to the audio joint matching degree and the audio autocorrelation matching degree, including:
[0197] Perform a summation process on the audio joint matching degree and the audio autocorrelation matching degree to obtain a total matching degree;
[0198] Determine the candidate audio data with a total matching degree greater than the matching degree threshold among the at least two candidate audio data as the target audio data for scoring the target video data.
[0199] Optionally, the recognition module performs a summation process on the audio joint matching degree and the audio autocorrelation matching degree to obtain a total matching degree, including:
[0200] Obtain the recognition weight of the first target audio recognition model and the recognition weight of the second target audio recognition model;
[0201] Perform a weighted process on the audio joint matching degree using the recognition weight of the first target audio recognition model to obtain a weighted audio joint matching degree;
[0202] Weight the audio self - matching degree using the recognition weights of the second target audio recognition model to obtain the weighted audio self - matching degree;
[0203] Sum the weighted audio joint matching degree and the weighted audio self - matching degree to obtain the total matching degree.
[0204] Optionally, the obtaining module obtains the object feature information of the target object, including:
[0205] Obtain the basic portrait feature information and multimedia portrait feature information of the target object;
[0206] Perform portrait association recognition on the basic portrait feature information and the multimedia portrait feature information of the target object to obtain portrait association feature information;
[0207] Determine the basic portrait feature information, the multimedia portrait feature information, and the portrait association feature information of the target object as the object feature information of the target object.
[0208] Optionally, the obtaining module obtains the audio feature information of at least two candidate audio data associated with the target video data, including:
[0209] Obtain at least two candidate audio data associated with the target video data;
[0210] Determine the object feature information of the creators of the at least two candidate audio data;
[0211] Extract the lyric feature information of the at least two candidate audio data to obtain the lyric feature information of the at least two candidate audio data;
[0212] Extract the score feature information of the at least two candidate audio data to obtain the score feature information of the at least two candidate audio data;
[0213] Fuse the object feature information of the creators, the lyric feature information of the at least two candidate audio data, and the score feature information of the at least two candidate audio data to obtain the audio feature information of the at least two candidate audio data.
[0214] Optionally, when the obtaining module extracts the score feature information of the at least two candidate audio data to obtain the score feature information of the at least two candidate audio data, it includes:
[0215] Perform frame segmentation on the candidate audio data Yi among the at least two candidate audio data to obtain at least two frames of audio data belonging to the candidate audio data Yi; i is a positive integer less than or equal to N, and N is the number of candidate audio data among the at least two candidate audio data;
[0216] Perform frequency-domain transformation on at least two frames of audio data belonging to the candidate audio data Yi to obtain the frequency-domain information of the candidate audio data Yi;
[0217] Extract musical score features from the frequency-domain information of the candidate audio data Yi to obtain the musical score feature information of the at least two candidate audio data.
[0218] Optionally, the obtaining module extracts musical score features from the frequency-domain information of the candidate audio data Yi to obtain the musical score feature information of the at least two candidate audio data, including:
[0219] Determine the energy information of the candidate audio data Yi according to the frequency-domain information of the candidate audio data Yi;
[0220] Perform filtering processing on the energy information of the candidate audio data Yi to obtain the filtered energy information;
[0221] Determine the filtered energy information as the musical score feature information of the at least two candidate audio data.
[0222] Optionally, the obtaining module obtains the video feature information of the target video data belonging to the target object, including:
[0223] Obtain the target video data belonging to the target object;
[0224] Extract at least two key video frames of the target video data;
[0225] Perform video feature extraction on the at least two key video frames to obtain the video feature information of the at least two key video frames;
[0226] Fuse the video feature information of the at least two key video frames to obtain the video feature information of the target video data.
[0227] According to an embodiment of the present application, Figure 4 The steps involved in the audio data processing method shown can be Figure 11 executed by each module in the audio data processing device shown. For example, Figure 4 The step S101 shown in Figure 11 can be executed by the obtaining module 111 in Figure 4 The step S102 shown in Figure 11executed by the fusion module 112 in Figure 4 The step S103 shown in Figure 11 can be executed by the recognition block 113 in Figure 4 The step S104 shown in Figure 11 can be executed by the recommendation module 114 in
[0228] According to an embodiment of the present application, Figure 11 Each module in the audio data processing device shown in
[0229] can be separately or wholly combined into one or several units to form, or a certain one (or some) of the units can be further split into at least two smaller sub-units in terms of function, and the same operations can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module can also be realized by at least two units, or the functions of at least two modules can be realized by one unit. In other embodiments of the present application, the audio data processing device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of at least two units.
[0229] According to an embodiment of the present application, it is possible to construct the audio data processing device shown in Figure 4 by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method shown in Figure 11 on a general computer device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), a read-only storage medium (ROM), etc., and to implement the audio data processing method of the embodiments of the present application. The above computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0230] In this application, by fusing the audio feature information of at least two candidate audio data with the object feature information of the target object and the video feature information of the target video data, the audio fusion feature information of at least two candidate audio data is obtained. That is, by fusing multi-modal feature information, it is beneficial to provide more information for the recommended audio data and improve the accuracy of the recommended audio data. Further, by using at least two target audio recognition models to respectively recognize the audio fusion feature information of at least two candidate audio data and the audio feature information of at least two candidate audio data, the target audio data for scoring the target video data is obtained, and the target audio data is recommended to the target object; that is, by comprehensively considering the audio recognition results of multi-modal audio recognition models, the audio data is automatically recommended to the target object, which can improve the efficiency of the recommended audio data; at the same time, the advantages of different audio recognition models are fully utilized, which can effectively avoid the problem that a single model produces deviations and leads to a relatively low accuracy of the recommended audio data, and can make the recommended audio data more robust, more accurate and more credible.
[0231] Please refer to Figure 12 , which is a schematic structural diagram of an audio data processing device provided by an embodiment of this application. The above audio data processing device can be a computer program (including program code) running in a computer device. For example, the audio data processing device is an application software; the device can be used to execute the corresponding steps in the method provided by the embodiment of this application. As Figure 12 shown, the audio data processing device can include: an acquisition module 121, an extraction module 122, a fusion module 123, and an adjustment module 124.
[0232] The acquisition module is used to acquire the object feature information of the sample object, the sample video data belonging to the sample object, the sample audio data for scoring the sample video data, and the labeled audio matching degree of the sample audio data; the labeled audio matching degree is used to reflect the matching degree between the sample audio data and the sample object and the sample video data;
[0233] The extraction module is used to extract video signs from the sample video data to obtain the video feature information of the sample video data, and extract audio features from the sample audio data to obtain the audio feature information of the sample audio data;
[0234] The fusion module is used to fuse the audio feature information of the sample audio data with the video feature information of the sample video data and the object feature information of the sample object to obtain the audio fusion feature information of the sample audio data;
[0235] An adjustment module, configured to adjust at least two candidate audio recognition models respectively according to the labeled audio matching degree, the audio feature information of the sample audio data, and the audio fusion feature information of the sample audio data, to obtain at least two target audio recognition models.
[0236] Optionally, the obtaining module obtains the labeled audio matching degree of the sample audio data, including:
[0237] Obtain object behavior data regarding the sample video data;
[0238] Determine the labeled audio matching degree of the sample audio data according to the object behavior data regarding the sample video data.
[0239] Optionally, the adjustment module adjusts at least two candidate audio recognition models respectively according to the labeled audio matching degree, the audio feature information of the sample audio data, and the audio fusion feature information of the sample audio data, to obtain the at least two target audio recognition models, including:
[0240] Respectively use the at least two candidate audio recognition models to perform audio matching prediction on the audio feature information of the sample audio data and the audio fusion feature information of the sample audio data, to obtain predicted audio matching degrees;
[0241] Adjust at least two candidate audio recognition models respectively according to the predicted audio matching degrees and the labeled audio matching degree, to obtain the at least two target audio recognition models.
[0242] According to an embodiment of the present application, Figure 10 the steps involved in the audio data processing method shown can be Figure 12 executed by each module in the audio data processing device shown. For example, Figure 10 the step S301 shown in Figure 12 can be executed by the obtaining module 121 in Figure 10 the step S302 shown in Figure 12 can be executed by the extraction module 122 in Figure 10 the step S303 shown in Figure 12 can be executed by the fusion module 123 in Figure 10 the step S304 shown in Figure 12 can be executed by the adjustment module 124 in
[0243] According to an embodiment of the present application, Figure 12Each module in the audio data processing device shown can be separately or all combined into one or several units to form, or a certain one (or some) of the units can be further split into at least two smaller sub-units in terms of function, and the same operations can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module can also be realized by at least two units, or the functions of at least two modules can be realized by one unit. In other embodiments of the present application, the audio data processing device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of at least two units.
[0244] According to an embodiment of the present application, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding method shown in Figure 10 on a general-purpose computer device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct the audio data processing device shown in Figure 12 and to implement the audio data processing method of the embodiments of the present application. The above computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above computing device through the computer-readable recording medium and run therein.
[0245] In the present application, by using the fused audio feature information of the sample audio data and the audio feature information of the sample audio data to train at least two candidate audio recognition models, at least two target audio recognition models are obtained, which can avoid the problem that a single audio recognition model has deviations in the knowledge accumulation process, resulting in a relatively low accuracy of the recommended audio data.
[0246] Please refer to Figure 13 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 13 shown, the above computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 may further include: an object interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection communication between these components. Among them, the object interface 1003 may include a display screen (Display), a keyboard (Keyboard). Optionally, the object interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as W I -F IInterface). The memory 1005 can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 can also be at least one storage device located far from the aforementioned processor 1001. As Figure 13 shown, the memory 1005, as a computer-readable storage medium, can include an operating system, a network communication module, an object interface module, and a device control application program.
[0247] In Figure 13 the computer device 1000 shown, the network interface 1004 can provide network communication functions; the object interface 1003 is mainly used to provide an input interface for objects; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0248] Obtain the object feature information of the target object, the video feature information of the target video data belonging to the target object, and the audio feature information of at least two candidate audio data associated with the target video data;
[0249] Respectively fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain the audio fusion feature information of the at least two candidate audio data;
[0250] Use at least two target audio recognition models to respectively perform audio recognition on the audio fusion feature information of the at least two candidate audio data and the audio feature information of the at least two candidate audio data to obtain target audio data for scoring the target video data; the target audio data belongs to the at least two candidate audio data;
[0251] Recommend the target audio data to the target object.
[0252] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0253] Fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data to obtain first fusion feature information, and fuse the audio feature information of the at least two candidate audio data with the object feature information of the target object to obtain second fusion feature information;
[0254] Fuse the audio feature information of the at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object to obtain third fusion feature information;
[0255] Determine the first fusion feature information, the second fusion feature information, and the third fusion feature information as the audio fusion feature information of the at least two candidate audio data.
[0256] Optionally, the processor 1001 may be used to call a device control application program stored in the memory 1005 to implement:
[0257] Obtain a first video feature parameter and a first audio feature parameter having an association relationship; the first video feature parameter belongs to the video feature information of the target video data, and the first audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0258] Generate first fusion feature information according to the first video feature parameter and the first audio feature parameter;
[0259] Obtain a first object feature parameter and a second audio feature parameter having an association relationship; the first object feature parameter belongs to the object feature information of the target object, and the second audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0260] Generate second fusion feature information according to the first object feature parameter and the second audio feature parameter.
[0261] Optionally, the processor 1001 may be used to call a device control application program stored in the memory 1005 to implement:
[0262] Obtain a second object feature parameter, a second video feature parameter, and a third audio feature parameter having an association relationship; the second object feature parameter belongs to the object feature information of the target object, the second video feature parameter belongs to the video feature information of the target video data, and the third audio feature parameter belongs to the audio feature information of the at least two candidate audio data;
[0263] Generate third fusion feature information according to the second object feature parameter, the second video feature information, and the third audio feature parameter.
[0264] Optionally, the processor 1001 may be used to call a device control application program stored in the memory 1005 to implement:
[0265] Respectively determine a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data, and a second target audio recognition model that matches the audio feature information of the at least two candidate audio data from the at least two target audio recognition models;
[0266] Use the first target audio recognition model to perform audio joint relationship recognition on the audio fusion feature information of the at least two candidate audio data to obtain an audio joint matching degree; use the second target audio recognition model to perform audio autocorrelation recognition on the audio feature information of the target audio data to obtain an audio autocorrelation matching degree;
[0267] According to the audio joint matching degree and the audio autocorrelation matching degree, select target audio data from the at least two candidate audio data for scoring the target video data.
[0268] Optionally, the processor 1001 may be used to call a device control application program stored in the memory 1005 to implement:
[0269] Obtain the feature processing ability information of the at least two target audio recognition models;
[0270] According to the feature processing ability information, respectively determine a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data and a second target audio recognition model that matches the audio feature information of the target audio data from the at least two target audio recognition models.
[0271] Optionally, the processor 1001 may be used to call a device control application program stored in the memory 1005 to implement:
[0272] Perform a summation process on the audio joint matching degree and the audio autocorrelation matching degree to obtain a total matching degree;
[0273] Determine the candidate audio data with a total matching degree greater than the matching degree threshold among the at least two candidate audio data as the target audio data for scoring the target video data.
[0274] Optionally, the processor 1001 may be used to call a device control application program stored in the memory 1005 to implement:
[0275] Obtain the recognition weight of the first target audio recognition model and the recognition weight of the second target audio recognition model;
[0276] Perform weighted processing on the audio joint matching degree using the recognition weight of the first target audio recognition model to obtain a weighted audio joint matching degree;
[0277] Perform weighted processing on the audio autocorrelation matching degree using the recognition weight of the second target audio recognition model to obtain a weighted audio autocorrelation matching degree;
[0278] Sum the jointly matched degree of the weighted audio and the self-matched degree of the weighted audio to obtain the total matching degree.
[0279] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0280] Obtain the basic portrait feature information and multimedia portrait feature information of the target object;
[0281] Perform portrait association recognition on the basic portrait feature information and the multimedia portrait feature information of the target object to obtain portrait association feature information;
[0282] Determine the basic portrait feature information, the multimedia portrait feature information, and the portrait association feature information of the target object as the object feature information of the target object.
[0283] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0284] Obtain at least two candidate audio data associated with the target video data;
[0285] Determine the object feature information of the creators of the at least two candidate audio data;
[0286] Extract the lyric feature information of the at least two candidate audio data to obtain the lyric feature information of the at least two candidate audio data;
[0287] Extract the music score feature information of the at least two candidate audio data to obtain the music score feature information of the at least two candidate audio data;
[0288] Fuse the object feature information of the creators, the lyric feature information of the at least two candidate audio data, and the music score feature information of the at least two candidate audio data to obtain the audio feature information of the at least two candidate audio data.
[0289] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0290] Perform frame splitting on the candidate audio data Yi among the at least two candidate audio data to obtain at least two frames of audio data belonging to the candidate audio data Yi; i is a positive integer less than or equal to N, and N is the number of candidate audio data among the at least two candidate audio data;
[0291] Perform a frequency-domain transformation on at least two frames of audio data belonging to the candidate audio data Yi to obtain the frequency-domain information of the candidate audio data Yi;
[0292] Extract musical score features from the frequency-domain information of the candidate audio data Yi to obtain the musical score feature information of the at least two candidate audio data.
[0293] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0294] Determine the energy information of the candidate audio data Yi according to the frequency-domain information of the candidate audio data Yi;
[0295] Perform filtering processing on the energy information of the candidate audio data Yi to obtain the filtered energy information;
[0296] Determine the filtered energy information as the musical score feature information of the at least two candidate audio data.
[0297] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0298] Obtain the target video data belonging to the target object;
[0299] Extract at least two key video frames from the target video data;
[0300] Perform video feature extraction on the at least two key video frames to obtain the video feature information of the at least two key video frames;
[0301] Fuse the video feature information of the at least two key video frames to obtain the video feature information of the target video data.
[0302] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0303] Obtain the object feature information of the sample object, the sample video data belonging to the sample object, the sample audio data used for scoring the sample video data, and the labeled audio matching degree of the sample audio data; the labeled audio matching degree is used to reflect the matching degree between the sample audio data and the sample object and the sample video data;
[0304] Perform video feature extraction on the sample video data to obtain the video feature information of the sample video data, and perform audio feature extraction on the sample audio data to obtain the audio feature information of the sample audio data;
[0305] Fuse the audio feature information of the sample audio data with the video feature information of the sample video data and the object feature information of the sample object to obtain the audio fusion feature information of the sample audio data;
[0306] Adjust at least two candidate audio recognition models respectively according to the labeled audio matching degree, the audio feature information of the sample audio data, and the audio fusion feature information of the sample audio data to obtain the at least two target audio recognition models.
[0307] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0308] Obtain the object behavior data regarding the sample video data;
[0309] Determine the labeled audio matching degree of the sample audio data according to the object behavior data regarding the sample video data.
[0310] Optionally, the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0311] Perform audio matching prediction on the audio feature information of the sample audio data and the audio fusion feature information of the sample audio data respectively by using the at least two candidate audio recognition models to obtain the predicted audio matching degree;
[0312] Adjust at least two candidate audio recognition models respectively according to the predicted audio matching degree and the labeled audio matching degree to obtain the at least two target audio recognition models.
[0313] In this application, by fusing the audio feature information of at least two candidate audio data with the object feature information of the target object and the video feature information of the target video data, the audio fusion feature information of at least two candidate audio data is obtained. That is, by fusing multi-modal feature information, it is beneficial to provide more information for the recommended audio data and improve the accuracy of the recommended audio data. Further, by using at least two target audio recognition models to identify the audio fusion feature information of at least two candidate audio data and the audio feature information of at least two candidate audio data respectively, the target audio data for scoring the target video data is obtained, and this target audio data is recommended to the target object; that is, by comprehensively considering the audio recognition results of multi-modal audio recognition models, audio data is automatically recommended to the target object, which can improve the efficiency of the recommended audio data; at the same time, the advantages of different audio recognition models are fully utilized, which can effectively avoid the problem that a single model produces deviations and leads to a relatively low accuracy of the recommended audio data, and can make the recommended audio data more robust, more accurate, and more credible.
[0314] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the foregoing Figure 4 and the foregoing Figure 10 description of the above audio data processing method in the corresponding embodiments, and can also execute the foregoing Figure 11 and Figure 12 description of the above audio data processing device in the corresponding embodiments, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0315] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the foregoing-mentioned audio data processing device, and the computer program includes program instructions. When the above-mentioned processor executes the above-mentioned program instructions, it can execute the foregoing Figure 4 and Figure 10 description of the above audio data processing method in the corresponding embodiments. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.
[0316] As an example, the above program instructions can be deployed to be executed on a computer device, or be deployed on at least two computer devices located at one location, or, on at least two computer devices distributed at least two locations and interconnected through a communication network. The at least two computer devices distributed at least two locations and interconnected through a communication network can form a blockchain network.
[0317] The above computer-readable storage medium can be the data processing device provided in any of the foregoing embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0318] An embodiment of the present application further provides a computer program product, including a computer program / instructions, and the computer program / instructions are executed by a processor as described above Figure 4 and Figure 10 the descriptions of the above audio data processing method in the corresponding embodiments. Therefore, they will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer program product involved in the present application, please refer to the descriptions of the method embodiments of the present application.
[0319] The terms "first", "second", etc. in the specification, claims and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products or equipment.
[0320] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0321] The method and related device provided by the embodiments of the present application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in the processFigure 1 One process or multiple processes and / or structural schematic Figure 1 The functions specified in one box or multiple boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 One process or multiple processes and / or structural schematic steps for the functions specified in one box or multiple boxes.
[0322] What is disclosed above is only the preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. An audio data processing method, characterized in that, Including: Obtaining the object feature information of the target object, the video feature information of the target video data belonging to the target object, and the audio feature information of at least two candidate audio data associated with the target video data; Respectively fusing the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain the audio fusion feature information of the at least two candidate audio data; Respectively determining, from at least two target audio recognition models, a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data, and a second target audio recognition model that matches the audio feature information of the at least two candidate audio data; the network attributes of the at least two target audio recognition models are different from each other; Using the first target audio recognition model to perform audio joint relationship recognition on the audio fusion feature information of the at least two candidate audio data to obtain an audio joint matching degree; Using the second target audio recognition model to perform audio autocorrelation recognition on the audio feature information of the at least two candidate audio data to obtain an audio autocorrelation matching degree; According to the audio joint matching degree and the audio autocorrelation matching degree, selecting target audio data for scoring the target video data from the at least two candidate audio data; the target audio data belongs to the at least two candidate audio data; Recommending the target audio data to the target object.
2. The method according to claim 1, wherein The step of respectively fusing the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object to obtain the audio fusion feature information of the at least two candidate audio data includes: Fusing the audio feature information of the at least two candidate audio data with the video feature information of the target video data to obtain first fusion feature information, and fusing the audio feature information of the at least two candidate audio data with the object feature information of the target object to obtain second fusion feature information; Fusing the audio feature information of the at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object to obtain third fusion feature information; Determining the first fusion feature information, the second fusion feature information, and the third fusion feature information as the audio fusion feature information of the at least two candidate audio data.
3. The method according to claim 2, wherein The step of fusing the audio feature information of the at least two candidate audio data with the video feature information of the target video data to obtain first fusion feature information, and fusing the audio feature information of the at least two candidate audio data with the object feature information of the target object to obtain second fusion feature information includes: Obtaining a first video feature parameter and a first audio feature parameter having an association relationship; the first video feature parameter belongs to the video feature information of the target video data, and the first audio feature parameter belongs to the audio feature information of the at least two candidate audio data; Generate first fusion feature information based on the first video feature parameter and the first audio feature parameter; Obtain a first object feature parameter and a second audio feature parameter with an associated relationship; the first object feature parameter belongs to the object feature information of the target object, and the second audio feature parameter belongs to the audio feature information of the at least two candidate audio data; Generate second fusion feature information based on the first object feature parameter and the second audio feature parameter.
4. The method according to claim 2, characterized in that The fusing the audio feature information of the at least two candidate audio data, the video feature information of the target video data, and the object feature information of the target object to obtain third fusion feature information includes: Obtain a second object feature parameter, a second video feature parameter, and a third audio feature parameter with an associated relationship; the second object feature parameter belongs to the object feature information of the target object, the second video feature parameter belongs to the video feature information of the target video data, and the third audio feature parameter belongs to the audio feature information of the at least two candidate audio data; Generate third fusion feature information based on the second object feature parameter, the second video feature information, and the third audio feature parameter.
5. The method according to claim 1, wherein The determining, from the at least two target audio recognition models, a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data, and a second target audio recognition model that matches the audio feature information of the at least two candidate audio data includes: Obtain the feature processing ability information of the at least two target audio recognition models; Based on the feature processing ability information, determine, from the at least two target audio recognition models, a first target audio recognition model that matches the audio fusion feature information of the at least two candidate audio data, and a second target audio recognition model that matches the audio feature information of the target audio data.
6. The method according to claim 1, wherein The selecting, from the at least two candidate audio data, target audio data for scoring the target video data according to the audio joint matching degree and the audio self-matching degree includes: Perform a summation process on the audio joint matching degree and the audio self-matching degree to obtain a total matching degree; Determine, as the target audio data for scoring the target video data, the candidate audio data among the at least two candidate audio data whose total matching degree is greater than the matching degree threshold.
7. The method according to claim 6, wherein The performing a summation process on the audio joint matching degree and the audio self-matching degree to obtain a total matching degree includes: Obtain the recognition weight of the first target audio recognition model and the recognition weight of the second target audio recognition model; Perform a weighting process on the audio joint matching degree using the recognition weight of the first target audio recognition model to obtain a weighted audio joint matching degree; Perform a weighting process on the audio self-matching degree using the recognition weight of the second target audio recognition model to obtain a weighted audio self-matching degree; Sum the jointly matched degree of the weighted processed audio and the self-matched degree of the weighted processed audio to obtain the total matching degree.
8. The method according to claim 1, wherein The obtaining of the object feature information of the target object includes: Obtain the basic portrait feature information and the multimedia portrait feature information of the target object; Perform portrait association recognition on the basic portrait feature information and the multimedia portrait feature information of the target object to obtain portrait association feature information; Determine the basic portrait feature information, the multimedia portrait feature information, and the portrait association feature information of the target object as the object feature information of the target object.
9. The method according to claim 1, characterized in that, The obtaining of the audio feature information of at least two candidate audio data associated with the target video data includes: Obtain at least two candidate audio data associated with the target video data; Determine the object feature information of the creators of the at least two candidate audio data; Extract the lyric feature information of the at least two candidate audio data to obtain the lyric feature information of the at least two candidate audio data; Extract the musical score feature information of the at least two candidate audio data to obtain the musical score feature information of the at least two candidate audio data; Fuse the object feature information of the creators, the lyric feature information of the at least two candidate audio data, and the musical score feature information of the at least two candidate audio data to obtain the audio feature information of the at least two candidate audio data.
10. The method according to claim 9, characterized in that The extracting of the musical score feature information of the at least two candidate audio data to obtain the musical score feature information of the at least two candidate audio data includes: Perform frame splitting on the candidate audio data Yi in the at least two candidate audio data to obtain at least two frames of audio data belonging to the candidate audio data Yi; i is a positive integer less than or equal to N, and N is the number of candidate audio data in the at least two candidate audio data; Perform frequency domain transformation on the at least two frames of audio data belonging to the candidate audio data Yi to obtain the frequency domain information of the candidate audio data Yi; Extract the musical score feature information from the frequency domain information of the candidate audio data Yi to obtain the musical score feature information of the at least two candidate audio data.
11. The method according to claim 10, wherein The extracting of the musical score feature information from the frequency domain information of the candidate audio data Yi to obtain the musical score feature information of the at least two candidate audio data includes: Determine the energy information of the candidate audio data Yi according to the frequency domain information of the candidate audio data Yi; Perform filtering processing on the energy information of the candidate audio data Yi to obtain the filtered energy information; Determine the filtered energy information as the musical score feature information of the at least two candidate audio data.
12. The method according to claim 1, characterized in that, The obtaining of the video feature information of the target video data belonging to the target object includes: Obtain the target video data belonging to the target object; Extract at least two key video frames of the target video data; Extract video feature information from the at least two key video frames to obtain the video feature information of the at least two key video frames. Fuse the video feature information of the at least two key video frames to obtain the video feature information of the target video data.
13. An audio data processing method, characterized in that, Including: Obtain the object feature information of the sample object, the sample video data belonging to the sample object, the sample audio data for scoring the sample video data, and the labeled audio matching degree of the sample audio data; The labeled audio matching degree is used to reflect the matching degree between the sample audio data, the sample object, and the sample video data; Extract video signs from the sample video data to obtain the video feature information of the sample video data, and extract audio features from the sample audio data to obtain the audio feature information of the sample audio data; Fuse the audio feature information of the sample audio data, the video feature information of the sample video data, and the object feature information of the sample object to obtain the audio fusion feature information of the sample audio data; According to the labeled audio matching degree, the audio feature information of the sample audio data, and the audio fusion feature information of the sample audio data, adjust at least two candidate audio recognition models respectively to obtain the at least two target audio recognition models as described in any one of claims 1-12; the network attributes of the at least two candidate audio recognition models are different from each other, the at least two candidate audio recognition models include a first candidate audio recognition model and a second candidate audio recognition model, the first candidate audio recognition model is used to predict the audio joint relationship of the audio fusion feature information of the sample audio data, and the second candidate audio recognition model is used to identify the audio autocorrelation relationship of the audio feature information of the sample audio data.
14. The method according to claim 13, wherein The obtaining the labeled audio matching degree of the sample audio data includes: Obtain the object behavior data of the sample video data; Determine the labeled audio matching degree of the sample audio data according to the object behavior data of the sample video data.
15. The method according to claim 13, wherein According to the labeled audio matching degree, the audio feature information of the sample audio data, and the audio fusion feature information of the sample audio data, adjust at least two candidate audio recognition models respectively to obtain the at least two target audio recognition models, including: Respectively use the at least two candidate audio recognition models to perform audio matching prediction on the audio feature information of the sample audio data and the audio fusion feature information of the sample audio data to obtain the predicted audio matching degree; Adjust at least two candidate audio recognition models respectively according to the predicted audio matching degree and the labeled audio matching degree to obtain the at least two target audio recognition models.
16. An audio data processing device, characterized in that, Including: An obtaining module, configured to obtain the object feature information of the target object, the video feature information of the target video data belonging to the target object, and the audio feature information of at least two candidate audio data associated with the target video data; A fusion module, configured to fuse the audio feature information of the at least two candidate audio data with the video feature information of the target video data and the object feature information of the target object respectively, so as to obtain the audio fusion feature information of the at least two candidate audio data; An identification module, configured to respectively determine a first target audio identification model that matches the audio fusion feature information of the at least two candidate audio data and a second target audio identification model that matches the audio feature information of the at least two candidate audio data from at least two target audio identification models; perform audio joint relationship identification on the audio fusion feature information of the at least two candidate audio data by using the first target audio identification model to obtain an audio joint matching degree; perform audio autocorrelation identification on the audio feature information of the at least two candidate audio data by using the second target audio identification model to obtain an audio autocorrelation matching degree; select, according to the audio joint matching degree and the audio autocorrelation matching degree, target audio data for scoring the target video data from the at least two candidate audio data; the target audio data belongs to the at least two candidate audio data; A recommendation module, configured to recommend the target audio data to the target object.
17. A computer device, characterized in that, Comprising: A processor and a memory; The above-mentioned processor is connected to the memory; the memory is used to store program codes, and the processor is used to call the program codes to execute the method according to any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and the program instructions, when executed by a processor, cause the processor to execute the method according to any one of claims 1-15.
19. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-15 are implemented.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and storage medium
CN111918094A
Audio recommendation method and device, electronic equipment and computer storage medium
CN112380377A