A multimedia data processing method, device, and readable storage medium
By using the target audio recognition model to perform vocal separation and speech recognition on the original audio data, the misidentification problem caused by noise interference is solved, and the accuracy of audio data recognition and automated recognition capabilities are improved.
Patent Information
- Application Number
- CN202111361702.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-11-17
AI Technical Summary
During the audio data recognition process, noise interference causes misidentification of classification models, reducing the accuracy of audio data recognition.
By obtaining the target audio recognition model, including the target vocal separation model and the target speech recognition model, the original audio data is separated, the audio tracks associated with the object are separated, and the speech data is text-recognized through the speech recognition model to determine the audio type.
Reduces noise interference, improves the accuracy of audio data recognition, and enables automatic identification of audio types in multimedia files.
Smart Images

Figure CN114329041B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a method and apparatus for processing multimedia data and a readable storage medium. Background Art
[0002] Currently, in some service scenarios (for example, an audio data recognition scenario), if it is necessary to classify a certain audio data (for example, audio data A), the audio type of the audio data can be identified first, so as to provide the audio type of the audio data A when storing or using the audio data A by classification.
[0003] For example, in an existing audio classification solution, the complete audio data A can be directly input into an audio classification model, and then the audio classification model can perform overall recognition on the audio data A to obtain the audio type of the audio data A. However, when performing overall recognition on the audio data A, there are often some noises (for example, the harmony in the audio data A) interfering, which easily leads to the possibility of misrecognition by the audio classification model, thus reducing the accuracy of audio data recognition. Summary of the Invention
[0004] Embodiments of the present application provide a method and apparatus for processing multimedia data and a readable storage medium, which can improve the accuracy of audio data recognition.
[0005] On the one hand, an embodiment of the present application provides a method for processing multimedia data, including:
[0006] When obtaining the original audio data in a multimedia file, obtaining a target audio recognition model for audio processing of the original audio data; the target audio recognition model includes a target vocal separation model and a target speech recognition model;
[0007] Inputting the original audio data into the target vocal separation model, and the target vocal separation model performs vocal separation on the original audio data to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data;
[0008] Obtaining the speech data of the first object from the first type of audio track, inputting the speech data of the first object into the target speech recognition model, and the target speech recognition model performs text recognition on the speech data of the first object to obtain a text recognition result of the first object;
[0009] Determining the audio type of the original audio data based on the text recognition result, and storing the audio data associated with the second object in the second type of audio track.
[0010] On the one hand, an embodiment of the present application provides a method for processing multimedia data, including:
[0011] Obtain sample audio data for training an initial audio recognition model, and use the labeled audio type corresponding to the sample audio data as the sample type label of the sample audio data; the sample audio data is obtained from sample multimedia files; the initial audio recognition model includes an initial vocal separation model and an initial speech recognition model;
[0012] Input the sample audio data into the initial vocal separation model, and the initial vocal separation model separates the vocals from the sample audio data to obtain a first type of sample audio track associated with a first sample object in the sample audio data and a second type of sample audio track associated with a second sample object in the sample audio data;
[0013] Obtain the speech data of the first sample object from the first type of sample audio track, input the speech data of the first sample object into the initial speech recognition model, and the initial speech recognition model performs text recognition on the speech data of the first sample object. Based on the obtained text recognition result of the first sample object, determine the predicted audio type of the sample audio data, and use the predicted audio type as the predicted type label;
[0014] Iteratively train the initial audio recognition model based on the predicted type label and the sample type label to obtain a target audio recognition model for audio processing of the original audio data in the multimedia file.
[0015] An embodiment of the present application provides a multimedia data processing device on the one hand, including:
[0016] An acquisition module, configured to obtain a target audio recognition model for audio processing of the original audio data when the original audio data in the multimedia file is obtained; the target audio recognition model includes a target vocal separation model and a target speech recognition model;
[0017] A separation module, configured to input the original audio data into the target vocal separation model, and the target vocal separation model separates the vocals from the original audio data to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data;
[0018] An identification module, configured to obtain the speech data of the first object from the first type of audio track, input the speech data of the first object into the target speech recognition model, and the target speech recognition model performs text recognition on the speech data of the first object to obtain the text recognition result of the first object;
[0019] A first determination module, configured to determine the audio type of the original audio data based on the text recognition result and store the audio data associated with the second object in the second type of audio track.
[0020] Among them, the above-mentioned target vocal separation model includes a first segmentation network for extracting speech features corresponding to the original audio data and a second segmentation network for extracting audio features corresponding to the original audio data;
[0021] The above-mentioned separation module includes:
[0022] A spectrum acquisition unit for acquiring the original track amplitude spectrum corresponding to the original audio data;
[0023] A first feature extraction unit for inputting the original track amplitude spectrum into the first segmentation network, and outputting speech features associated with the original track amplitude spectrum by the first segmentation network;
[0024] A second feature extraction unit for inputting the original track amplitude spectrum into the second segmentation network, and outputting audio features associated with the original track amplitude spectrum by the second segmentation network;
[0025] A track separation unit for obtaining a first type of track associated with the first object in the original audio data and a second type of track associated with the second object in the original audio data based on the speech features, audio features, and the original track amplitude spectrum.
[0026] Among them, the above-mentioned first segmentation network and second segmentation network are symmetric networks with the same network structure; the symmetric network includes U convolutional layers and U upsampling layers; the p-th convolutional layer among the U convolutional layers is used to obtain convolutional features associated with the original track amplitude spectrum; p is a positive integer less than or equal to U; the q-th upsampling layer among the U upsampling layers is used to splice the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer when obtaining the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer; the (q - 1)-th upsampling layer is the previous upsampling layer of the q-th upsampling layer; q is a positive integer less than or equal to U, and q = U - p + 1.
[0027] Among them, the above-mentioned track separation unit includes:
[0028] A feature fusion sub-unit for fusing the speech features and audio features to obtain target fusion features;
[0029] A first track acquisition sub-unit for generating a first mask associated with the speech features based on the target fusion features and the speech features, generating a first track amplitude spectrum based on the first mask and the original track amplitude spectrum, and performing inverse spectrum transformation on the first track amplitude spectrum to obtain a first type of track associated with the first object in the original audio data;
[0030] A second audio track obtaining subunit, configured to generate a second mask associated with the audio feature based on the target fusion feature and the audio feature, generate a second audio track amplitude spectrum based on the second mask and the original audio track amplitude spectrum, and perform an inverse spectral transform on the second audio track amplitude spectrum to obtain a second type of audio track associated with a second object in the original audio data.
[0031] Wherein, if the multimedia file is a video file, the first object includes a role object and a first background music object in the video file;
[0032] The above-mentioned first audio track obtaining subunit is specifically configured to perform an inverse spectral transform on the first audio track amplitude spectrum to obtain a first mixed speech audio track associated with the first object; the first mixed speech audio track carries object speech data associated with the role object and first background music speech data associated with the first background music object; perform voiceprint feature recognition on the object speech data and the first background music speech data carried in the first mixed speech audio track, use the recognized voiceprint feature of the role object as the first voiceprint feature, and use the recognized voiceprint feature of the first background music object as the second voiceprint feature; based on the first voiceprint feature and the second voiceprint feature, perform voice segmentation on the object speech data and the first background music speech data in the first mixed speech audio track to obtain object speech data corresponding to the first voiceprint feature and first background music speech data corresponding to the second voiceprint feature; use the object speech data corresponding to the first voiceprint feature and the first background music speech data corresponding to the second voiceprint feature as the first type of audio track associated with the first object.
[0033] Wherein, if the multimedia file is an audio file, the first object includes a second background music object in the audio file;
[0034] The above-mentioned first audio track obtaining subunit is specifically configured to perform an inverse spectral transform on the first audio track amplitude spectrum to obtain a second mixed speech audio track associated with the first object; the second mixed speech audio track carries second background music speech data associated with the second background music object; use the second background music speech data obtained in the second mixed speech audio track as the first type of audio track associated with the first object.
[0035] Wherein, if the multimedia file is a video file, the first type of audio track includes object speech data associated with a role object in the video file, and the second type of audio track includes audio data associated with a background object in the video file; the background object includes a third background music object and an accompaniment object;
[0036] The above-mentioned apparatus further includes:
[0037] A separation and update module, configured to input audio data associated with a background object in a second type of audio track into a target vocal separation model, perform vocal separation on the audio data associated with the background object through the target vocal separation model to obtain third background music voice data associated with a third background music object and accompaniment audio data associated with an accompaniment object; add the separated third background music voice data to a first type of audio track containing object voice data to obtain a first type of updated audio track, and use the separated accompaniment audio data as a second type of updated audio track.
[0038] Among them, the above recognition module includes:
[0039] A third feature extraction unit, configured to obtain a to-be-processed voice sequence corresponding to the voice data of a first object included in a first type of audio track, input the to-be-processed voice sequence into an encoding network in a target speech recognition model, extract voice sequence features of the to-be-processed voice sequence by the encoding network, and use the extracted voice sequence features as target voice sequence features corresponding to the first object;
[0040] A vector conversion unit, configured to obtain a first decoding result output by a decoding network in the target speech recognition model at the i-th moment, input the first decoding result into a vector conversion network in the target speech recognition model, and convert the first decoding result into a target word vector by the vector conversion network; i is a positive integer;
[0041] A decoding output unit, configured to obtain a second decoding result output by the decoding network at the (i + 1)-th moment based on the target voice sequence features, the target word vector, and the decoding network in the target speech recognition model;
[0042] A result determination unit, configured to determine a text recognition result of the first object based on the first decoding result and the second decoding result.
[0043] Among them, the target speech recognition model includes an encoding network, and the encoding network in the target speech recognition model is a bidirectional long short-term memory network; the bidirectional long short-term memory network includes a forward long short-term memory network and a backward long short-term memory network; the forward long short-term memory network includes memory network B j and memory network B j+1 , memory network B j+1 is the next memory network of memory network B j ; the backward long short-term memory network includes memory network C j+1 and memory network C j ; memory network C j+1 is the previous memory network of memory network C j ; j is a positive integer less than or equal to M; the number of memory networks in both the forward long short-term memory network and the backward long short-term memory network is M;
[0044] The third feature extraction unit comprises:
[0045] The forward feature extraction subunit is used to obtain the memory network B in the forward long short-term memory network. j The associated positive history hidden feature h j-1 , the speech sequence to be processed and the forward historical hidden feature h j-1 Input memory network B j , by memory network B j At the jth moment, the positive target hidden feature h is extracted j , the positive target hidden feature h j And the speech sequence to be processed is input into the memory network B j+1 , by memory network B j+1 At the j+1th moment, the positive target hidden feature h is extracted j+1 ;
[0046] The reverse feature extraction subunit is used to obtain the memory network C in the reverse long short-term memory network. j+1 The associated reverse history hidden feature k j+1 , the speech sequence to be processed and the reverse history hidden feature k j+1 Input memory network C j+1 , by the memory network C j+1 At the j+1th moment, the reverse target hidden feature k is extracted j , the reverse target hidden feature k j And the speech sequence to be processed is input into the memory network C j , by the memory network C j At the jth moment, the reverse target hidden feature k is extracted j-1 ;
[0047] Feature concatenation subunit, used to combine memory network B j The positive target hidden feature h extracted at the jth moment j With memory network C j The reverse target hidden feature k extracted at the jth moment j-1 Perform feature splicing to obtain the first splicing feature and store the memory network B j+1 The positive target hidden feature h extracted at the j+1th time j+1 With memory network C j+1 The reverse target hidden feature k extracted at the j+1th time j Perform feature splicing to obtain a second splicing feature;
[0048] The feature determination subunit is used to determine, based on the first splicing feature and the second splicing feature, a target speech sequence feature corresponding to the first object extracted from the speech sequence to be processed.
[0049] Among them, the above decoding output unit includes:
[0050] A weight acquisition sub-unit, configured to generate an initial weight coefficient based on the target speech sequence feature and the target word vector, and perform normalization processing on the initial weight coefficient to obtain a target weight coefficient;
[0051] A vector generation sub-unit, configured to generate a semantic coding vector based on the target weight coefficient and the target speech sequence feature;
[0052] A vector splicing sub-unit, configured to splice the target word vector and the semantic coding vector to obtain a target splicing vector;
[0053] A decoding output sub-unit, configured to input the target splicing vector into a decoding network in the target speech recognition model, and the decoding network outputs a second decoding result at the (i + 1)-th moment.
[0054] Among them, the decoding network in the target speech recognition model is a unidirectional long short-term memory network; the unidirectional long short-term memory network includes a memory network D i and a memory network D i+1 , the memory network D i+1 is the next memory network of the memory network D i ; i is a positive integer less than or equal to N; the number of memory networks in the unidirectional long short-term memory network is N;
[0055] The above decoding output sub-unit is specifically configured to obtain a unidirectional target hidden feature s i extracted by the memory network D i in the unidirectional long short-term memory network at the i-th moment; the unidirectional target hidden feature s i is obtained based on the unidirectional historical hidden feature s i associated with the memory network D i-1 ; input the unidirectional target hidden feature s i and the target splicing vector into the memory network D i+1 , and the memory network D i+1 extracts a unidirectional target hidden feature s i+1 at the (i + 1)-th moment, and based on the target splicing vector and the unidirectional target hidden feature s i+1 , obtains a second decoding result at the (i + 1)-th moment.
[0056] Among them, the first object includes a background music object, and the text recognition result includes the target recognition result of the background music object; the second object includes an accompaniment object;
[0057] The above first determination module is specifically configured to determine that the audio type of the original audio data is a pure music type if the target recognition result in the text recognition result is a null value; store the accompaniment audio data associated with the accompaniment object in the second type of audio track.
[0058] Wherein, the above device further includes:
[0059] A second determination module, configured to determine that the audio type of the original audio data is a non-pure music type if the target recognition result in the text recognition result is a non-null value; use the decoding result associated with the text recognition result as the text information associated with the first object; associate and store the speech data and text information associated with the first object in the first type of audio track.
[0060] An embodiment of the present application provides a multimedia data processing device on the one hand, including:
[0061] A sample acquisition module, configured to acquire sample audio data for training an initial audio recognition model, and use the labeled audio type corresponding to the sample audio data as the sample type label of the sample audio data; the sample audio data is acquired from a sample multimedia file; the initial audio recognition model includes an initial vocal separation model and an initial speech recognition model;
[0062] An audio track separation module, configured to input the sample audio data into the initial vocal separation model, and the initial vocal separation model performs vocal separation on the sample audio data to obtain a first type of sample audio track associated with the first sample object in the sample audio data and a second type of sample audio track associated with the second sample object in the sample audio data;
[0063] A text recognition module, configured to acquire the speech data of the first sample object from the first type of sample audio track, input the speech data of the first sample object into the initial speech recognition model, the initial speech recognition model performs text recognition on the speech data of the first sample object, determines the predicted audio type of the sample audio data based on the obtained text recognition result of the first sample object, and uses the predicted audio type as the predicted type label;
[0064] A model training module, configured to perform iterative training on the initial audio recognition model based on the predicted type label and the sample type label to obtain a target audio recognition model for performing audio processing on the original audio data in the multimedia file.
[0065] An embodiment of the present application provides a computer device on the one hand, including: a processor and a memory;
[0066] The processor is connected to the memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the computer device executes the method provided by the embodiment of the present application.
[0067] In one aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiment of the present application.
[0068] In one aspect, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the method provided by the embodiment of the present application.
[0069] When the computer device involved in the embodiment of the present application obtains the original audio data in the multimedia file, by introducing a target audio recognition model, the target vocal separation model in the target audio recognition model can perform vocal separation on the original audio data, so as to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data. Further, after obtaining the speech data of the first object from the first type of audio track, the target speech recognition model in the target audio recognition model can perform text recognition on the speech data of the first object, so as to obtain the text recognition result of the first object. Furthermore, the audio type of the original audio data can be determined based on the text recognition result, and the audio data associated with the second object in the second type of audio track can be stored. It can be seen that in the process of intelligently identifying the audio type of the original audio data, the embodiment of the present application can first separate the individual first type of audio track and the second type of audio track from the original audio data, and then perform text recognition on the separated first type of audio track, thereby reducing the interference of the audio data (such as accompaniment audio data) in the second type of audio track on the target speech recognition model, and further improving the accuracy of audio data recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0071] Figure 1 It is a schematic diagram of a system architecture provided by an embodiment of the present application;
[0072] Figure 2 It is a schematic diagram of a scenario for audio data recognition provided by an embodiment of the present application;
[0073] Figure 3 It is a schematic flowchart of a multimedia data processing method provided by an embodiment of the present application;
[0074] Figure 4 It is a schematic structural diagram of a symmetric network provided by an embodiment of the present application;
[0075] Figure 5 It is a schematic diagram of a scenario for vocal separation provided by an embodiment of the present application;
[0076] Figure 6 It is a schematic diagram of a scenario for a target speech recognition model provided by an embodiment of the present application;
[0077] Figure 7 It is a schematic diagram of a scenario for audio classification provided by an embodiment of the present application;
[0078] Figure 8 It is a schematic flowchart of a multimedia data processing method provided by an embodiment of the present application;
[0079] Figure 9 It is a schematic structural diagram of a multimedia data processing device provided by an embodiment of the present application;
[0080] Figure 10 It is a schematic structural diagram of a multimedia data processing device provided by an embodiment of the present application;
[0081] Figure 11 It is a schematic structural diagram of a multimedia data processing device provided by an embodiment of the present application;
[0082] Figure 12 It is a schematic structural diagram of a computer device provided by an embodiment of the present application;
[0083] Figure 13 It is a schematic structural diagram of a multimedia data processing system provided by an embodiment of the present application. Detailed implementation manners
[0084] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0085] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0086] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0087] The key technologies of speech processing technology (Speech Technology) include automatic speech recognition technology (Automatic Speech Recognition, ASR), speech synthesis technology, and voiceprint recognition technology. Among them, automatic speech recognition technology is also called speech recognition technology, and its goal is to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. Different from speaker recognition and speaker verification, the latter attempts to identify or verify the speaker who emits the speech rather than the lexical content contained therein. In the embodiments of the present application, automatic speech recognition technology can be used for text recognition of speech data.
[0088] The solution provided by the embodiments of this application belongs to machine learning (ML) in the field of artificial intelligence. It can be understood that machine learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. In the embodiments of this application, the target audio recognition model (including the target vocal separation model and the target speech recognition model) is an AI model based on machine learning technology and can be used for audio processing of raw audio data.
[0089] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a system architecture provided by the embodiments of this application. As Figure 1 shown, the system architecture may include a business server 100 and a user terminal cluster. Among them, the user terminal cluster may include one or more user terminals, and the number of user terminals in the user terminal cluster will not be limited here. As Figure 1 shown, the multiple user terminals in the user terminal cluster may specifically include: user terminal 200a, user terminal 200b, user terminal 200c,..., user terminal 200n. Among them, there may be a communication connection between the user terminal clusters. For example, there is a communication connection between user terminal 200a and user terminal 200b, and there is a communication connection between user terminal 200a and user terminal 200c. At the same time, any user terminal in the user terminal cluster may have a communication connection with the business server 100, so that each user terminal in the user terminal cluster can perform data interaction with the business server 100 through this communication connection. For example, there is a communication connection between user terminal 200a and the business server 100. Among them, the above communication connection does not limit the connection method and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, or can also be connected through other methods, which are not limited in this application.
[0090] It should be understood that each user terminal in the user terminal cluster as Figure 1 shown may be installed with an application client. When the application client runs on each user terminal, it can be respectively connected to the above Figure 1Data interaction is carried out between the business servers 100 shown. Among them, the application client can be an application client with functions such as displaying text, images, audio, and video data information, such as a short video application, a video application, a live broadcast application, a music application, a social application, an instant messaging application, a game application, a shopping application, a novel application, a payment application, a browser, etc. Among them, the application client can be an independent client or an embedded sub-client integrated in a certain client (such as a social client, a video client, etc.), which is not limited here. Taking the short video application as an example, the business server 100 can be a collection of multiple servers including the background server, data processing server, etc. corresponding to the short video application. Therefore, each user terminal can perform data transmission with the business server 100 through the application client corresponding to the short video application. For example, each user terminal can upload the short videos produced by it to the business server 100 through the application client of the short video application, and then the business server 100 can send these short videos to other user terminals. Among them, short videos have the characteristics of short duration, fast spread, low production threshold, strong participation, etc., and are one of the important dissemination methods of content entrepreneurship and social media platforms. In addition, during the process of making short videos, the business server 100 can also recommend appropriate audio data (such as background music) for each user terminal to enrich the content of the short videos. For example, it can recommend audio data of pure music type or non-pure music type.
[0091] It should be understood that in order to obtain the type information (i.e., audio type) of a to-be-processed audio data, an audio data recognition method is provided in an embodiment of the present application, that is, the to-be-processed audio data is intelligently processed through a trained audio recognition model. For the convenience of subsequent understanding and description, the to-be-processed audio data can be collectively referred to as the original audio data in the embodiment of the present application, and the audio recognition model used to perform audio processing on the original audio data is called the target audio recognition model. Among them, the original audio data can be audio data obtained from a multimedia file, and may include voices, singing, musical instrument sounds, noises, etc. Here, the multimedia file refers to a file carrying audio data, including video files (such as short videos, TV dramas, movies, music videos (MVs), animations, etc.) that carry both image data and audio data, and audio files mainly composed of audio data (such as songs / music, audiobooks, radio dramas, radio programs, etc.). These multimedia files can be files from business platforms on the network (such as video platforms, music platforms, etc.), or locally produced files (such as files composed of audio-visual data collected by a camera, such as short videos recorded through a short video application), or files uploaded or shared by a certain user (such as user X).
[0092] It should be noted that different multimedia files can be encapsulated into different file formats. For example, video files can be encapsulated into file formats such as MKV (Matroska Video File), AVI (Audio Video Interleaved), MP4 (an abbreviation of MPEG-4 (Moving Picture Experts Group 4)), etc., and audio files can be encapsulated into file formats such as MP3 (Moving Picture Experts Group Audio Layer III), OGG (OGGVobis (oggVorbis)), AAC (Moving Picture Experts Group 4), etc. The specific file format used for encapsulation is not limited in the embodiments of this application.
[0093] It can be understood that the method provided in the embodiments of this application can be executed by a computer device, and the computer device includes but is not limited to a user terminal (for example, Figure 1 any one of the user terminals in the user terminal cluster shown) or a business server (for example, Figure 1 the business server 100 shown). Among them, the business server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The user terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a handheld computer, a wearable device (such as a smart watch, a smart bracelet, etc.), a smart computer, a smart vehicle, etc., which are smart terminals that can run the above applications. Among them, the user terminal and the business server can be directly or indirectly connected by wired or wireless means, and this is not limited in the embodiments of this application.
[0094] Taking the business server 100 as an example, it should be understood that when the business server 100 obtains a certain multimedia file (for example, multimedia file A), it can first obtain the original audio data (for example, audio data B) in the multimedia file. For example, a certain user terminal (for example, user terminal 200a) in the user terminal cluster can first upload the multimedia file A to the business server 100, and then the business server 100 performs audio extraction processing (including decompression processing, decoding processing, etc.) on the multimedia file A to obtain the audio data B in the multimedia file A. Or, the user terminal (for example, user terminal 200a) can first obtain the original audio data in the multimedia file and then upload the original audio data to the business server 100.
[0095] Furthermore, after obtaining the original audio data, the service server 100 can obtain a target audio recognition model integrated with a target vocal separation model and a target speech recognition model. Then, the obtained original audio data (e.g., audio data B) can be input into the target vocal separation model, and the target vocal separation model can perform vocal separation on the original audio data, so as to extract a first type of audio track (e.g., audio track C1) associated with the first object and a second type of audio track (e.g., audio track C2) associated with the second object from the original audio data. Here, both the first object and the second object belong to the objects with the ability to produce sound in the original audio data (or multimedia file). That is to say, the original audio data can be composed of audio data associated with all the objects with the ability to produce sound in the original audio data. In the embodiments of the present application, the objects with the ability to produce sound can be classified into role objects, background music objects, accompaniment objects, etc. according to the object type. Among them, the role object refers to the role that speaks (e.g., says lines, including forms such as dialogue, monologue, and narration) in the multimedia file; the background music object refers to the singer who sings the lyrics in the multimedia file; and the accompaniment object refers to the object that produces accompaniment in the multimedia file. Generally speaking, it can be other audio data in the original audio data except for human voices, including but not limited to objects that hum without lyrics, musical instruments that play music (such as piano, violin, flute, etc.), and in some business scenarios, it can also include sounding objects in the real / virtual environment (such as thunder, rain, wind, stream, and even interfering objects that emit noise, etc.). In the embodiments of the present application, both the role object and the background music object belong to the first object, and the accompaniment object belongs to the second object. In addition, the background music object and the accompaniment object can be collectively referred to as background objects.
[0096] Based on this, the above-mentioned first type of audio track can include speech data associated with the role object / background music object, such as dialogue between roles, singing vocals, etc.; the second type of audio track can include audio data associated with the accompaniment object, such as the sound of a piano being played. For the convenience of distinction, in the embodiments of the present application, the speech data associated with the role object can be collectively referred to as object speech data, the speech data associated with the background music object can be collectively referred to as background music speech data, and the audio data associated with the accompaniment object can be collectively referred to as accompaniment audio data. In addition, subsequently, the target speech recognition model can perform text recognition on the object speech data and the background music speech data to obtain corresponding text information. For the convenience of distinction, in the embodiments of the present application, the text information corresponding to the object speech data can be called background music text information (e.g., lyrics), and the text information corresponding to the background music speech data can be called object text information (e.g., lines).
[0097] Further, the service server 100 can obtain the voice data of the first object (e.g., voice data D) from the above-mentioned first type of audio track (e.g., audio track C1), and input the voice data of the first object into the target voice recognition model. Then, the target voice recognition model can perform text recognition on the voice data of the first object, thereby obtaining the text recognition result of the first object. Finally, based on the text recognition result, the audio type of the original audio data can be determined. In addition, the audio data (e.g., audio data E) associated with the second object in the above-mentioned second type of audio track (e.g., audio track C2) can be stored.
[0098] It can be understood that in the embodiments of the present application, the audio type can include a pure music type and a non-pure music type. Among them, the pure music type is used to represent that the music (or song) contained in the original audio data is pure music, that is, there is no soundtrack text information (such as lyrics) in the original audio data, and the non-pure music type is used to represent that the music contained in the original audio data is non-pure music, that is, there is soundtrack text information in the original audio data.
[0099] Optionally, it can be understood that Figure 1 The shown system architecture may include multiple service servers. A user terminal can be connected to one service server, and each service server can obtain the multimedia files uploaded by the user terminal connected to it, so as to load the target audio recognition model to perform audio processing on the original audio data in the multimedia files.
[0100] Optionally, it can be understood that the user terminal can also perform audio processing on the original audio data obtained from the multimedia file by loading the trained target audio recognition model.
[0101] Among them, it can be understood that the specific business scenarios applicable to the above system architecture can include: audio classification scenarios, audio recommendation scenarios, audio search scenarios, audio extraction scenarios, audio production scenarios, video production scenarios, etc. Here, the specific business scenarios will not be listed one by one.
[0102] For example, in the audio classification scenario, a computer device (e.g., the above-mentioned service server 100) can identify the audio data (e.g., music F1) uploaded by a certain user (e.g., user X) through an application client on a user terminal (e.g., the above-mentioned user terminal 200a), so as to obtain the audio type of the music F1, and can use the obtained audio type as the type label of the music F1. Subsequently, the music F1 with the type label added can be stored in the music database, that is, the music F1 is classified, labeled, and stored to promote the label construction of the music database. For example, if the audio type of the music F1 is pure music type, the type label of the music F1 can be "pure music"; conversely, if the audio type of the music F1 is non-pure music type, the type label of the music F1 can be "non-pure music". It can be understood that the type label of the music F1 and the style label of the music F1 (e.g., soothing, lively, fresh, healing, etc.) can jointly serve as the music label of the music F1.
[0103] For another example, in the audio recommendation scenario, a computer device (e.g., the above-mentioned service server 100) can recommend at least one associated music to the user X based on the user profile of the user X, where at least one of the associated music here can have the same music label (including type label) as the music that the user X likes (or has listened to). For example, by analyzing the user profile of the user X, it can be determined that the music label that the user X likes (or has listened to) is "pure music", then the computer device can obtain one or more associated music with the music label of "pure music" (e.g., music F2 and music F3), and then the computer device can push the music F2 and music F3 to the application client corresponding to the user X. It should be understood that the number of music labels that the user X likes can also be multiple (e.g., 2), and the embodiments of the present application do not limit the number of music labels here.
[0104] For another example, in the audio search scenario, similar to the audio recommendation scenario, when the user X conducts a search, the computer device (e.g., the above-mentioned service server 100) can identify the music label corresponding to the content searched by the user X and recommend at least one associated music to the user X. Among them, at least one of the associated music here can have the same music label (including type label) as the content searched by the user X.
[0105] For another example, in the scenario of audio extraction, a computer device (e.g., the above-mentioned service server 100) can obtain a certain multimedia file uploaded by user X (e.g., video F4), and separate a first type of audio track (e.g., audio track F41) and a second type of audio track (e.g., audio track F42) from the audio data of the multimedia file. Furthermore, it can select the required audio track (e.g., audio track F42) for storage. For example, if user X hopes to extract the accompaniment contained in video F4, the computer device can obtain the corresponding accompaniment (i.e., audio track F42) through the target vocal separation model in the target audio recognition model. The obtained accompaniment can be stored locally, shared with other users, used to expand the music database, or used in other multimedia files. For another example, if user X hopes to extract the character dialogue contained in video F4 (i.e., the voice data associated with the character object in audio track F41), the computer device can obtain the corresponding voice data through the target vocal separation model. The extracted voice data can be used in other multimedia files, or the object text information corresponding to the voice data can be automatically generated through the target speech recognition model in the target audio recognition model.
[0106] For another example, in the scenario of audio production, a computer device (e.g., the above-mentioned service server 100) can obtain a certain audio data to be processed uploaded by user X (e.g., music F5), and separate a first type of audio track (e.g., audio track F51) and a second type of audio track (e.g., audio track F52) from the audio data to be processed. Furthermore, it can select the corresponding audio track for audio track update to obtain updated audio data. For example, user X can delete a certain audio track (e.g., audio track F51); or user X can add a certain audio track (e.g., audio track F53); or user X can adjust a certain audio track (e.g., audio track F52). For example, adjust the audio track of the piano part in audio track F52 (i.e., the piano sound), for example, adjust the speed of the piano sound, such as from the original fast speed to a slow speed.
[0107] For another example, in the scenario of video production, a computer device (e.g., the above-mentioned service server 100) can obtain a certain video uploaded by user X (e.g., video F6), and separate a first type of audio track (e.g., audio track F61) and a second type of audio track (e.g., audio track F62) from the audio data contained in the video. At the same time, in combination with the above-described audio recommendation scenario, the computer device can also recommend at least one associated music (e.g., music F7) to user X based on the user profile of user X. Furthermore, user X can select from at least one associated music, and the selected associated music can be used to replace audio track F62 in video F6 to obtain a video with updated accompaniment.
[0108] It is understandable that in the specific embodiments of the present application, data related to user portraits is involved. When the embodiments in the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0109] For ease of understanding, please also refer to Figure 2 , Figure 2 which is a schematic diagram of a scenario for audio data recognition provided by an embodiment of the present application. Among them, as Figure 2 shown, the computer device 20A can be any one of the business server 100 or the user terminals in the user terminal cluster corresponding to the above Figure 1 corresponding embodiment (for example, the user terminal 200a). As Figure 2 shown, the user terminal 20B can be any one of the user terminals in the user terminal cluster corresponding to the above Figure 1 corresponding embodiment (for example, the user terminal 200b), and user A has a binding relationship with the user terminal 20B.
[0110] As Figure 2 shown, the multimedia file 201A is the multimedia file to be processed. After the computer device 20A obtains the multimedia file 201A, it can perform operations such as unpacking and decoding on the multimedia file 201A, so as to obtain the original audio data (for example, the audio data 202A) in the multimedia file 201A. Among them, the multimedia file 201A can be a video file or an audio file, which is not limited here. The multimedia file 201A can be a file stored in the local cache of the computer device 20A, a file obtained from a certain business platform (for example, the business platform B), or a file uploaded by a certain user (for example, the file uploaded by user A through the user terminal 20B). The source of the multimedia file 201A in the embodiments of the present application is not limited.
[0111] Furthermore, the computer device 20A can obtain a pre-trained audio recognition model 203A (i.e., the target audio recognition model), as Figure 2As shown, the audio recognition model 203A may include a vocal separation model M1 (i.e., the target vocal separation model) and a speech recognition model M2 (i.e., the target speech recognition model). Furthermore, the above audio data 202A can be input into the vocal separation model M1. By using the vocal separation model M1 to perform vocal separation on the audio data 202A, a track 204A and a track 205A can be obtained. Among them, the track 204A is the first type of track associated with the first object, and the track 205A is the second type of track associated with the second object. Furthermore, the track 204A can be input into the speech recognition model M2. By using the speech recognition model M2 to perform text recognition on the speech data of the first object carried in the track 204A, the corresponding text recognition result can be obtained. Finally, based on this text recognition result, the audio type of the above audio data 202A can be determined.
[0112] It can be understood that the above text recognition result may include two types of recognition results, namely, the recognition result for the background music object (which can also be called the target recognition result) and the recognition result for the character object (which can also be called the character recognition result).
[0113] Optionally, if in Figure 2 the shown text recognition result, the recognition result for the background music object is a null value, it means that there is no background music text information (i.e., the text information associated with the background music object) in the audio data 202A. Therefore, it can be determined that the audio type of the audio data 202A is the pure music type; conversely, optionally, if in Figure 2 the shown text recognition result, the recognition result for the background music object is a non-null value, it means that there is background music text information in the audio data 202A. Therefore, it can be determined that the audio type of the audio data 202A is a non-pure music type.
[0114] Similarly, optionally, if in Figure 2 the shown text recognition result, the recognition result for the character object is a null value, it means that there is no object text information (i.e., the text information associated with the character object, such as the lines of the character object) in the audio data 202A. Therefore, it can be determined that there is no character object in the audio data 202A; conversely, optionally, if in Figure 2 the shown text recognition result, the recognition result for the character object is a non-null value, it means that there is object text information in the audio data 202A. Therefore, it can be determined that there is a character object in the audio data 202A.
[0115] It should be understood that when the first object (including the background music object and the character object) does not exist in the audio data 202A, the audio track 204A is an empty audio track, that is, there is no voice data of the first object in this audio track 204A. Therefore, the final text recognition result (including the target recognition result and the character recognition result) is a null value. Similarly, when the second object (including the accompaniment object) does not exist in the audio data 202A, the audio track 205A is an empty audio track, that is, there is no audio data of the second object in this audio track 205A.
[0116] It can be understood that if the above text recognition result is a non-null value, that is, at least one of the target recognition result and the character recognition result is a non-null value, the text information associated with this text recognition result output by the speech recognition model M2 can be obtained (for example, Figure 2 the text information 206A shown). It should be understood that the text information 206A here can include at least one of the background music text information and the object text information.
[0117] Among them, the computer device 20A can use a multimedia database with a large amount of multimedia data (which can include video data and audio data) to train a deep neural network to obtain the above audio recognition model 203A. The specific training process can be referred to the subsequent Figure 8 corresponding embodiments. It should be noted that the vocal separation model M1 and the speech recognition model M2 can be built and trained separately, or they can also be built and trained jointly. The embodiments of the present application do not limit this.
[0118] It can be understood that the obtained audio track 205A can contain audio data associated with the accompaniment object (that is, the accompaniment audio data). Therefore, the computer device 20A can store the accompaniment audio data in the audio track 205A. For example, it can be stored in a music database 207A associated with the computer device 20A, or it can also be stored in the local cache of the computer device 20A. Among them, the obtained accompaniment audio data can also be applied to many business scenarios. For example, it can be pushed as associated music to the corresponding user (for example, user A), and user A can play the associated music through the user terminal 20B; another example is that user A can use the user terminal 20B to use the associated music as the background music of a multimedia file being edited (such as video, music, slides, documents, games, etc.).
[0119] It can be understood that the computer device 20A can also associatively store the above automatically generated text information 206A and the audio track 204A. Optionally, the object text information in the text information 206A can be added to the multimedia file 201A as a line, and the background music text information in the text information 206A can be added to the multimedia file 201A as lyrics. Optionally, the computer device 20A can send the multimedia file with the added text information to the user terminal 20B for playback.
[0120] Among them, the computer device 20A obtains the target audio recognition model by training the initial audio recognition model, and performs audio processing on the original audio data in the multimedia file through the target audio recognition model to obtain the audio type of the original audio data and the specific implementation manners of storing and using the relevant data. For details, please refer to the description in the following Figures 3 - 8 corresponding embodiments.
[0121] Please refer to Figure 3 , Figure 3 FIG. is a schematic flowchart of a multimedia data processing method provided by an embodiment of the present application. It can be understood that the method provided by the embodiment of the present application can be executed by a computer device, and the computer device herein includes but is not limited to a user terminal or a service server running a target audio recognition model. For ease of understanding, the embodiment of the present application takes the computer device as a user terminal as an example to illustrate the specific process of audio processing on the trained target audio recognition model in the user terminal. As Figure 3 shown, the method may at least include the following steps S101 to S104:
[0122] Step S101, when obtaining the original audio data in the multimedia file, obtain a target audio recognition model for performing audio processing on the original audio data; the target audio recognition model includes a target vocal separation model and a target speech recognition model;
[0123] Specifically, after obtaining the multimedia file, the computer device can perform audio extraction processing on the multimedia file to obtain the original audio data carried in the multimedia file. For example, the computer device can perform decompression processing based on the file format adopted when the multimedia file is encapsulated to obtain the decompressed multimedia data, and then can perform decoding processing on the decompressed multimedia data to obtain the decoded multimedia data, and then can use the audio data carried in the decoded multimedia data as the original audio data. It can be understood that the original audio data can include one or more types of audio data, such as speech, singing, musical instrument sounds, noise, etc. The specific content of the original audio data will not be limited here.
[0124] Further, the computer device can obtain a target audio recognition model to perform audio processing on the above-mentioned original audio data. Here, the target audio recognition model can be an audio processing system composed of a target vocal separation model and a target speech recognition model. That is to say, the target vocal separation model and the target speech recognition model can be used as two functional modules in this audio processing system. In the embodiments of the present application, the target vocal separation model is used to separate vocals from the original audio data, and the target speech recognition model is used to perform text recognition on the input speech data (such as the speech data of the first object). Among them, the target audio recognition model is obtained by the computer device training the initial audio recognition model. The specific training process can be referred to the subsequent Figure 8 corresponding embodiments.
[0125] Step S102: Input the original audio data into the target vocal separation model, and the target vocal separation model separates the vocals from the original audio data to obtain a first type of audio track associated with the first object in the original audio data and a second type of audio track associated with the second object in the original audio data;
[0126] It should be understood that, optionally, the target vocal separation model in the embodiments of the present application may include one or more (for example, 2) segmentation networks to extract corresponding features from the original audio data through the segmentation networks. Among them, the multiple segmentation networks may specifically include a first segmentation network and a second segmentation network. The first segmentation network can be used to extract the speech features corresponding to the original audio data, and the second segmentation network can be used to extract the audio features corresponding to the original audio data.
[0127] Specifically, the computer device can first obtain the original audio track amplitude spectrum corresponding to the original audio data, and then can input the original audio track amplitude spectrum into the first segmentation network, and the first segmentation network outputs the speech features associated with the original audio track amplitude spectrum; at the same time, the original audio track amplitude spectrum can be input into the second segmentation network, and the second segmentation network outputs the audio features associated with the original audio track amplitude spectrum. Further, based on the extracted speech features, audio features, and the above-mentioned original audio track amplitude spectrum, a first type of audio track associated with the first object in the original audio data and a second type of audio track associated with the second object in the original audio data can be obtained.
[0128] It should be understood that the first type of audio track in the embodiments of the present application can also be referred to as a vocal track (Vocals), which is the speech data of the first object obtained by separating the source of the entire mixed audio signal (i.e., the original audio data). This first type of audio track can be separately written into a wav file (i.e., a waveform audio file, WaveForm, which is a standard digital audio file). For example, when the original audio data contains a song, the data stored in the wav file is the vocal part of the song, including singing, rapping, and other vocals.
[0129] It should be understood that the second type of audio track in the embodiments of the present application can also be referred to as a background music track (Bgm), that is, the audio track of the background music part, which is the audio data of the second object obtained after separating the source of the entire mixed audio signal. This second type of audio track can be separately written into a wav file, and the data stored in the wav file is other audio tracks except the vocal track. For example, when the original audio data is a song, the second type of audio track obtained by source separation contains the background music audio data of the song.
[0130] Among them, the original audio data is an audio signal in the time domain. For the convenience of analysis, the computer device can transform the original audio data into an original audio track spectrum through spectrum transformation (for example, Fourier transform), that is, observe and analyze it in the frequency domain. Further, the phase of the original audio track spectrum can be eliminated, so as to obtain the spectrum amplitude spectrum after eliminating the phase characteristics, that is, the original audio track amplitude spectrum.
[0131] Among them, the target vocal separation model in the embodiments of the present application can be a source separation model. It should be understood that in a whole piece of audio data (such as the original audio data), there may be various audio signals mixed in, so the whole piece of audio data may be generated by mixing various audio signals. Source separation is to separate such a mixed audio signal through signal processing or other algorithms, extract the audio signal sequence of a specified type from the mixed audio signal, and finally generate a separate audio file (such as a file for storing audio tracks). In the embodiments of the present application, source separation can also be referred to as vocal separation.
[0132] Optionally, the above first segmentation network and second segmentation network can be symmetric networks with the same network structure. The symmetric network can include U convolutional layers and U upsampling layers. Among them, the p-th convolutional layer in the U convolutional layers can be used to obtain convolutional features associated with the amplitude spectrum of the original audio track, where p is a positive integer less than or equal to U. Among them, the q-th upsampling layer in the U upsampling layers can be used to concatenate the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer when the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer are obtained, and finally corresponding speech features or audio features can be output. Among them, the (q - 1)-th upsampling layer is the previous upsampling layer of the q-th upsampling layer, where q is a positive integer less than or equal to U, and q = U - p + 1.
[0133] It can be understood that the above symmetric network refers to a network with a symmetric structure. For example, U-Net (U-shaped network). Here, U-Net is one of the algorithms for semantic segmentation using a fully convolutional network (FCN). It uses a U-shaped symmetric structure including a contracting path and an expanding path, which to a certain extent affects the design of several subsequent segmentation networks. Among them, a convolutional neural network (CNN) is a feedforward neural network. Its artificial neurons can respond to surrounding units within a certain coverage range and have excellent performance in large-scale image processing. A convolutional neural network can be composed of one or more convolutional layers and a fully connected layer at the top (corresponding to a classic neural network), and can also include associated weights and pooling layers.
[0134] For ease of understanding, please also refer to Figure 4 , Figure 4 which is a schematic structural diagram of a symmetric network provided by an embodiment of the present application. As Figure 4 shown, the symmetric network is a symmetric network based on U-Net, which adopts a U-shaped symmetric structure. Among them, the symmetric network includes U (for example, 4) convolutional layers and U (for example, 4) upsampling layers. Taking U = 4 as an example for illustration, as Figure 4As shown, the symmetric network adopts an encoding-decoding structure. The left side belongs to the encoding part, which consists of 4 convolutional layers; the right side belongs to the decoding part, which consists of 4 upsampling layers. Among them, the first convolutional layer to the fourth convolutional layer in the above 4 convolutional layers are: convolutional layer 401a, convolutional layer 402a, convolutional layer 403a, and convolutional layer 404a. The first upsampling layer to the fourth upsampling layer in the above 4 upsampling layers are: upsampling layer 401c, upsampling layer 402c, upsampling layer 403c, and upsampling layer 404c. In addition, a convolutional connection layer 40b is used to connect the encoding part and the decoding part.
[0135] As Figure 4 shown, the convolutional layer 401a, convolutional layer 402a, convolutional layer 403a, and convolutional layer 404a can each include two 3×3 convolutional networks, and a ReLU function (Rectified Linear Unit, linear rectifier function, also known as the modified linear unit) or other function can be used as the activation function after each convolutional network. In addition, a 2×2 maxpooling layer can be connected after each convolutional layer for downsampling. Similarly, the upsampling layer 401c, upsampling layer 402c, and upsampling layer 403c can each include two 3×3 convolutional networks, and a ReLU function or other function can be used as the activation function after each convolutional network. In addition, a 2×2 up-conv can be connected after these three upsampling layers for upsampling. The upsampling layer 404c can include two 3×3 convolutional networks and a 1×1 convolutional network. As Figure 4 shown, the convolutional connection layer 40b can also include two 3×3 convolutional networks, and a ReLU function or other function can be used as the activation function after each convolutional network. In addition, a 2×2 up-conv can be connected after the convolutional connection layer 40b.
[0136] It should be understood that Figure 4Each of the upsampling layers shown can obtain the features output by the previous layer of the network and the convolutional features output by the corresponding convolutional layer for feature concatenation, and the obtained target concatenated features are used as input features and input to the upsampling layer. Among them, when performing feature concatenation, in order to keep the feature sizes on both sides consistent, feature cropping can be performed, and the embodiments of the present application do not limit the size (dimension) of the features output by each network. For example, for the upsampling layer 401c, when obtaining the convolutional features of the convolutional connection layer 40b, the convolutional features of the convolutional connection layer 40b and the convolutional features of the convolutional layer 404a can be concatenated to obtain the target concatenated feature Y1. Then, the target concatenated feature Y1 can be calculated through the upsampling layer 401c and the connected transposed convolutional layer, so as to obtain the upsampled features of the upsampling layer 401c; for the upsampling layer 402c, when obtaining the upsampled features of the upsampling layer 401c, the upsampled features of the upsampling layer 401c and the convolutional features of the convolutional layer 403a can be concatenated to obtain the target concatenated feature Y2. Then, the target concatenated feature Y2 can be calculated through the upsampling layer 402c and the connected transposed convolutional layer, so as to obtain the upsampled features of the upsampling layer 402c; for the upsampling layer 403c, the upsampled features of the upsampling layer 402c and the convolutional features of the convolutional layer 402a can be concatenated to obtain the target concatenated feature Y3. Then, the target concatenated feature Y3 can be calculated through the upsampling layer 403c and the connected transposed convolutional layer, so as to obtain the upsampled features of the upsampling layer 403c; for the upsampling layer 404c, the upsampled features of the upsampling layer 403c and the convolutional features of the convolutional layer 401a can be concatenated to obtain the target concatenated feature Y4. Then, the target concatenated feature Y4 can be calculated through the upsampling layer 404c, so as to obtain the finally output features (i.e., voice features or audio features).
[0137] It can be understood that in order to retain some important feature information in the previous downsampling process, the feature maps (such as convolutional features) of each convolutional layer in Unet will be concatenated (i.e., concatenate) to the upsampling layer at the corresponding position. Therefore, each layer of feature maps can be effectively used in subsequent calculations, that is, skip-connection is achieved. In this way, compared with some other network structures (such as FCN), Unet avoids directly performing supervision and loss calculation in high-level feature maps, but combines the features in low-level feature maps, so that the finally obtained feature maps contain both high-level features and many low-level features. Therefore, feature fusion in different dimensions is achieved, and the utilization rate of feature information in each layer of the network can be improved, and then the separation accuracy of the target vocal separation model can be improved.
[0138] It can be understood that when implementing the above symmetric network, one can either implement the symmetric network from scratch, initialize the weights, and then train the model; or borrow the convolutional layer structures of some existing networks (such as vgg (a convolutional network) in resnet (i.e., residual neural network)) and the corresponding pre-trained weight files, and then add the subsequent upsampling layer to train the model. It can be understood that in the model training of deep learning, if existing weight files can be used, the training speed can be greatly accelerated. The embodiments of this application do not limit the specific manner of implementing the symmetric network.
[0139] Among them, the specific process of obtaining the first type of audio track and the second type of audio track based on the speech feature, audio feature, and the original audio track amplitude spectrum can be as follows: The computer device first fuses the speech feature and the audio feature to obtain a target fusion feature. Then, based on the target fusion feature and the speech feature, a first mask associated with the speech feature can be generated. Further, based on the first mask and the original audio track amplitude spectrum, a first audio track amplitude spectrum can be generated. Subsequently, an inverse spectral transformation can be performed on the first audio track amplitude spectrum to obtain the first type of audio track associated with the first object in the original audio data. Similarly, the computer device can generate a second mask associated with the audio feature based on the target fusion feature and the audio feature. Further, based on the second mask and the original audio track amplitude spectrum, a second audio track amplitude spectrum can be generated. Subsequently, an inverse spectral transformation can be performed on the second audio track amplitude spectrum to obtain the second type of audio track associated with the second object in the original audio data.
[0140] For ease of understanding, please also refer to Figure 5 , Figure 5 which is a schematic diagram of a scenario for vocal separation provided by the embodiments of this application. Figure 5 The shown vocal separation scenario can be implemented in the above target vocal separation model. As Figure 5As shown, the computer device can perform a spectral transformation on the audio data 51a (i.e., the original audio data) to obtain the corresponding track spectrum 52a (i.e., the original track spectrum), and then the track amplitude spectrum 53a (i.e., the original track amplitude spectrum) can be calculated from the track spectrum 52a. Further, the track amplitude spectrum 53a can be input into the segmentation network 54a (i.e., the first segmentation network) and the segmentation network 55a (i.e., the second segmentation network) respectively. Among them, the segmentation network 54a is used to process the features associated with the first object, and the segmentation network 55a is used to process the features associated with the second object. Therefore, through the segmentation network 54a, the speech feature 56a corresponding to the track amplitude spectrum 53a can be extracted, and through the segmentation network 55a, the audio feature 57a corresponding to the track amplitude spectrum 53a can be extracted. Further, the speech feature 56a and the audio feature 57a extracted by these two segmentation networks can be feature fused. For example, the elements at the corresponding positions of the speech feature 56a and the audio feature 57a are linearly added to obtain the corresponding fusion feature (i.e., the target fusion feature).
[0141] Further, based on the target fusion feature, the speech feature 56a, and the audio feature 57a, mask calculations can be performed to calculate the mask associated with the speech feature 56a and the mask associated with the audio feature 57a respectively. For example, the elements at the corresponding positions of the target fusion feature and the speech feature 56a can be proportionally calculated (such as multiplied), and the calculated weighted matrix can be used as the mask associated with the speech feature 56a (i.e., the first mask). Similarly, the elements at the corresponding positions of the target fusion feature and the audio feature 57a are proportionally calculated (such as multiplied), and the calculated weighted matrix can be used as the mask associated with the audio feature 57a (i.e., the second mask). Further, the calculated first mask can be proportionally calculated (such as multiplied) with the elements at the corresponding positions of the track amplitude spectrum 53a, so as to obtain the track amplitude spectrum 58a (i.e., the first track amplitude spectrum, for example, the track amplitude spectrum 5A). Subsequently, an inverse spectral transformation can be performed on the track amplitude spectrum 58a to obtain the track 59a (i.e., the first type of track, for example, the track 5B). Similarly, the calculated second mask can be proportionally calculated (such as multiplied) with the elements at the corresponding positions of the track amplitude spectrum 53a to obtain the track amplitude spectrum 510a (i.e., the second track amplitude spectrum, for example, the track amplitude spectrum 5C). Subsequently, an inverse spectral transformation can be performed on the track amplitude spectrum 510a to obtain the track 511a (i.e., the second type of track, for example, the track 5D).
[0142] For example, when Figure 5When the audio data 51a shown is a song, through the vocal separation process described above, the track amplitude spectrum 5A and the track 5B corresponding to the track amplitude spectrum 5A, as well as the track amplitude spectrum 5C and the track 5D corresponding to the track amplitude spectrum 5C can be obtained. Among them, the track 5B is the track of the vocal singing part separated from the song, and the track 5D is the track of the accompaniment part separated from the song. For the convenience of observation and analysis, both the track amplitude spectrum 5A and the track amplitude spectrum 5C here can be displayed in the form of a spectrogram. It can be understood that the horizontal axis (i.e., the x-axis) of the spectrogram represents time, the vertical axis (i.e., the y-axis) represents frequency, and the depth (i.e., the z-axis) represents amplitude. Both the track 5B and the track 5D can be displayed in the form of an audio graph (or waveform graph). It can be understood that the horizontal axis (i.e., the x-axis) of the audio graph represents time, and the vertical axis (i.e., the y-axis) represents amplitude. Other forms can also be used for display, and the embodiments of the present application do not limit this.
[0143] It can be understood that for different types of multimedia files, the object types of the first objects they contain may be different. Therefore, the data types of the first type of tracks separated may also be different.
[0144] Optionally, if the multimedia file is a video file, the first object may include a role object in the video file (for the convenience of subsequent distinction, the role object in this scenario can be called the first role object) and a first background music object. At this time, the specific process of obtaining the first type of track in the video file can be as follows: The computer device performs an inverse spectrum transformation on the first track amplitude spectrum to obtain a first mixed speech track associated with the first object. Among them, the first mixed speech track carries object speech data associated with the role object (for the convenience of subsequent distinction, the object speech data in this scenario can be called the first object speech data, for example, character dialogue) and first background music speech data associated with the first background music object (for example, singing). Further, perform voiceprint feature recognition on the object speech data and the first background music speech data carried in the first mixed speech track, so that the voiceprint feature of the recognized role object can be used as the first voiceprint feature, and the voiceprint feature of the recognized first background music object can be used as the second voiceprint feature. Further, based on the first voiceprint feature and the second voiceprint feature, perform voice segmentation on the object speech data and the first background music speech data in the first mixed speech track to obtain the object speech data corresponding to the first voiceprint feature and the first background music speech data corresponding to the second voiceprint feature. Finally, the object speech data corresponding to the first voiceprint feature and the first background music speech data corresponding to the second voiceprint feature can be used as the first type of track associated with the first object.
[0145] Optionally, if the multimedia file is an audio file (e.g., a song), the first object may include a second background music object in the audio file. In this case, the specific process of obtaining the first type of audio track in the audio file may be as follows: The computer device performs an inverse spectral transform on the first audio track amplitude spectrum to obtain a second mixed speech track associated with the first object, where the second mixed speech track carries second background music speech data (e.g., singing voice) associated with the second background music object. Then, the second background music speech data obtained from the second mixed speech track can be used as the first type of audio track associated with the first object. Optionally, the first type of audio track can also be obtained by using voiceprint feature recognition. For example, voiceprint feature recognition can be performed on the second background music speech data carried in the second mixed speech track, and the voiceprint feature of the recognized second background music object can be used as the third voiceprint feature. Then, based on the third voiceprint feature, the second background music speech data corresponding to the third voiceprint feature can be extracted from the second mixed speech track. Finally, the second background music speech data corresponding to the third voiceprint feature can be used as the first type of audio track associated with the first object.
[0146] It can be understood that optionally, when both a second background music object and a second character object (e.g., the host of a radio program) exist in the above audio file (e.g., a radio program), similar to the video file, in this case, the first type of audio track in the audio file can also be obtained by using voiceprint feature recognition. The specific process may be as follows: The computer device performs an inverse spectral transform on the first audio track amplitude spectrum to obtain a third mixed speech track associated with the first object, where the third mixed speech track carries second object speech data (e.g., the speaking voice of the radio host) associated with the second character object and second background music speech data (e.g., singing voice) associated with the second background music object. Further, voiceprint feature recognition is performed on the second object speech data and the second background music speech data carried in the third mixed speech track, so that the voiceprint feature of the recognized second character object can be used as the fourth voiceprint feature, and the voiceprint feature of the recognized second background music object can be used as the fifth voiceprint feature. Further, based on the fourth voiceprint feature and the fifth voiceprint feature, voice segmentation is performed on the second object speech data and the second background music speech data in the third mixed speech track to obtain the second object speech data corresponding to the fourth voiceprint feature and the second background music speech data corresponding to the fifth voiceprint feature. Finally, the second object speech data corresponding to the fourth voiceprint feature and the second background music speech data corresponding to the fifth voiceprint feature can be used as the first type of audio track associated with the first object.
[0147] It can be understood that in some alternative embodiments, the first type of audio track and the second type of audio track obtained by the first vocal separation through the target vocal separation model may not be pure enough, which will inevitably affect the accuracy of subsequent text recognition. For example, if the multimedia file is a video file containing a character object (or, it can also be an audio file containing a character object, such as a radio program), the separated first type of audio track may contain object voice data associated with the character object in the video file, and the second type of audio track may contain audio data associated with the background object in the video file. That is to say, the first vocal separation through the target vocal separation model may only separate the human voice of the character object speaking and the background music / song, which obviously does not meet the expectations. Here, the background object includes the third background music object and the accompaniment object. In this scenario, since the second type of audio track obtained at this time still contains singing (i.e., the object voice data associated with the third background music object), the singing and accompaniment in it can be separated by performing vocal separation on the second type of audio track again, so as to improve the accuracy of subsequent text recognition. The specific process can be as follows: The computer device inputs the audio data associated with the background object in the second type of audio track into the target vocal separation model, and the target vocal separation model performs vocal separation on the audio data associated with the background object to obtain the third background music voice data associated with the third background music object and the accompaniment audio data associated with the accompaniment object. The vocal separation process here is similar to the process of separating the first type of audio track and the second type of audio track above, and will not be elaborated here. Further, the separated third background music voice data can be added to the first type of audio track containing the object voice data to obtain the first type of updated audio track, and the separated accompaniment audio data is used as the second type of updated audio track. Therefore, subsequently, the first type of updated audio track can replace the original first type of audio track and be input into the target speech recognition model for text recognition, and the audio data (i.e., the accompaniment audio data) associated with the accompaniment object in the second type of updated audio track can also be stored. Optionally, the above second vocal separation can also adopt the method of voiceprint feature recognition to obtain the first type of updated audio track and the second type of updated audio track, which will not be elaborated here.
[0148] Step S103, obtain the voice data of the first object from the first type of audio track, input the voice data of the first object into the target speech recognition model, and the target speech recognition model performs text recognition on the voice data of the first object to obtain the text recognition result of the first object;
[0149] Specifically, the computer device can obtain the to-be-processed speech sequence corresponding to the speech data of the first object included in the first type of audio track. For example, the first type of audio track can be frame-processed, that is, the first type of audio track is divided into multiple frames of speech data in chronological order, so as to obtain the to-be-processed speech sequence composed of multiple frames of speech data. The embodiments of the present application will not limit the specific number of multiple frames of speech data. Optionally, the to-be-processed speech sequence here can be represented in the form of a mel spectrogram (log-melspectrogram).
[0150] For ease of understanding, please also refer to Figure 6 , Figure 6 which is a schematic diagram of a scenario of a target speech recognition model provided by an embodiment of the present application. As Figure 6 shown, by performing frame processing on the audio track 601a (i.e., the first type of audio track), a speech sequence 602a (i.e., the to-be-processed speech sequence) can be obtained. The speech sequence 602a can include M (M is a positive integer) frames of speech data. The M frames of speech data can specifically include speech data x 1 , speech data x 2 , speech data x 3 , speech data x 4 , speech data x 5 , …, speech data x M .
[0151] Furthermore, the computer device can input the above-mentioned to-be-processed speech sequence into the target speech recognition model for text recognition. Optionally, in the embodiments of the present application, the target speech recognition model can be an ASR model based on the attention mechanism. The attention mechanism is a method for solving problems proposed by imitating human attention. Simply put, it is to quickly screen out high-value information from a large amount of information. It is mainly used to solve the problem that it is difficult to obtain a reasonable final vector representation when the input sequence of the LSTM (Long-Short Term Memory) / RNN (Recurrent Neural Network) model is long. The approach is to retain the intermediate results of the LSTM, learn from them with a new model, and associate them with the output, so as to achieve the purpose of information screening.
[0152] The target speech recognition model may include three parts, namely an encoding network (encoder), a decoding network (decoder), and an embedding network (embedding). The specific process of text recognition may be as follows: The computer device inputs the speech sequence to be processed into the encoding network in the target speech recognition model. The encoding network extracts the speech sequence features of the speech sequence to be processed and uses the extracted speech sequence features as the target speech sequence features corresponding to the first object. Further, obtain the first decoding result output by the decoding network in the target speech recognition model at the i-th moment, and input the first decoding result into the embedding network in the target speech recognition model. The embedding network converts the first decoding result into a target word vector, where i is a positive integer. Further, based on the target speech sequence features, the target word vector, and the decoding network in the target speech recognition model, the second decoding result output by the decoding network at the (i + 1)-th moment can be obtained. Finally, based on the first decoding result and the second decoding result, the text recognition result of the first object can be determined.
[0153] Among them, optionally, the encoding network included in the target speech recognition model may specifically be a bidirectional long short-term memory network (for example, Bi-directional Long Short Term Memory network, abbreviated as Bi-LSTM network). The bidirectional long short-term memory network may specifically include a forward long short-term memory network and a backward long short-term memory network. Among them, as Figure 6 shown, the forward long short-term memory network can be used to calculate the forward target hidden features at each moment of the speech sequence 602a along the first feature calculation direction from left to right. For example, the forward hidden features extracted by the computer device at any two adjacent moments (for example, the j-th moment and the (j + 1)-th moment) may include: the forward target hidden feature h j at the j-th moment and the forward target hidden feature h j+1 at the (j + 1)-th moment. Among them, it should be understood that since the forward long short-term memory network is essentially a recurrent neural network, the forward hidden feature (for example, the forward target hidden feature h j ) extracted by the computer device at the j-th moment can essentially be used as the input feature for the next moment (i.e., the (j + 1)-th moment). Therefore, the forward hidden feature (for example, the forward target hidden feature h j ) extracted by the computer device at the j-th moment is essentially determined jointly by the forward historical hidden feature h j-1 and the speech sequence to be processed (for example, the speech sequence 602a). Similarly, the forward hidden feature (for example, the forward target hidden feature h j+1)Essentially, it is jointly determined by the forward target hidden feature h extracted at the previous moment (i.e., the j-th moment) of the (j + 1)-th moment j and the speech sequence 602a. It should be understood that the forward long short-term memory network may include M memory networks, and any two adjacent memory networks among the M memory networks may include memory network B j and memory network B j+1 , and in the above first feature calculation direction, memory network B j+1 is the next memory network of memory network B j . The memory network used by the computer device at the j-th moment may be memory network B j , and similarly, the memory network used by the computer device at the (j + 1)-th moment may be memory network B j+1 . Where j is a positive integer less than or equal to M
[0154] Similarly, as Figure 6 shown, the backward long short-term memory network can be used to calculate the backward target hidden feature of the speech sequence 602a at each moment along the second feature calculation direction from right to left. For example, the backward hidden features extracted by the computer device at any two adjacent moments (e.g., the j-th moment and the (j + 1)-th moment) may include: the backward target hidden feature k j at the (j + 1)-th moment and the backward target hidden feature k j-1 at the j-th moment. It should be understood that since the backward long short-term memory network is essentially a recurrent neural network, the backward hidden feature (e.g., the backward target hidden feature k j-1 ) extracted by the computer device at the j-th moment is essentially jointly determined by the backward target hidden feature k j extracted at the next moment (i.e., the (j + 1)-th moment) of the j-th moment and the speech sequence 602a. Similarly, the backward hidden feature (e.g., the backward target hidden feature k j ) extracted by the computer device at the (j + 1)-th moment is essentially jointly determined by the backward historical hidden feature k j+1 extracted at the next moment (i.e., the (j + 2)-th moment) of the (j + 1)-th moment and the speech sequence 602a. It should be understood that the backward long short-term memory network may also include M memory networks. In the backward long short-term memory network, any two adjacent memory networks among the M memory networks may include memory network C j and memory network C j+1 , and in the above second feature calculation direction, memory network C j+1 is the previous memory network of memory network C j . The memory network used by the computer device at the j-th moment may be memory network Cj Similarly, the memory network used by the computer device at the (j + 1)-th moment can be memory network C j+1 .
[0155] It can be seen that when the computer device obtains the forward historical hidden feature h j associated with the memory network B in the forward long short-term memory network j-1 , it can input the speech sequence to be processed (for example, Figure 6 the speech sequence 602a shown) and the forward historical hidden feature h j-1 into the memory network B j . The memory network B j extracts the forward target hidden feature h j at the j-th moment. Further, the forward target hidden feature h j and the speech sequence to be processed can be input into the memory network B j+1 (i.e., the next memory network of the memory network B j ), and the memory network B j+1 extracts the forward target hidden feature h j+1 at the (j + 1)-th moment.
[0156] Similarly, when the computer device obtains the backward historical hidden feature k j+1 associated with the memory network C in the backward long short-term memory network j+1 , it can input the speech sequence to be processed (for example, Figure 6 the speech sequence 602a shown) and the backward historical hidden feature k j+1 into the memory network C j+1 . The memory network C j+1 extracts the backward target hidden feature k j at the (j + 1)-th moment. Further, the backward target hidden feature k j and the speech sequence to be processed can be input into the memory network C j , and the memory network C j extracts the backward target hidden feature k j-1 at the j-th moment.
[0157] Furthermore, as shown in Figure 6 , the computer device can combine the forward hidden feature (for example, the above-mentioned forward target hidden feature h j ) extracted by the memory network B at the j-th moment and the backward hidden feature (for example, the above-mentioned backward target hidden feature k j ) extracted by the memory network C at the j-th moment j with the memory network C j-1Perform feature splicing to obtain the spliced feature at the j-th moment. It should be understood that in the embodiments of the present application, the spliced feature obtained at the j-th moment can be collectively referred to as the first spliced feature. Similarly, the computer device can use the memory network B j+1 The forward hidden feature extracted at the (j + 1)-th moment (e.g., the above-mentioned forward target hidden feature h j+1 ) and the memory network C j+1 The backward hidden feature extracted at the (j + 1)-th moment (e.g., the above-mentioned backward target hidden feature k j ) are subjected to feature splicing to obtain the spliced feature at the (j + 1)-th moment. It should be understood that in the embodiments of the present application, the spliced feature obtained at the (j + 1)-th moment can be collectively referred to as the second spliced feature.
[0158] As Figure 6 shown, the computer device can determine the target speech sequence feature corresponding to the first object extracted from the speech sequence to be processed based on the first spliced feature and the second spliced feature. It can be understood that at this time, the target speech sequence feature carries certain physical information corresponding to the speech data of the first object. For example, as Figure 6 shown, the computer device can input the speech sequence 602a into the encoding network 603a, and the encoding network 603a can extract the speech sequence feature 604a from the speech sequence 602a. The speech sequence feature 604a can specifically include feature R 1 , feature R 2 , feature R 3 , feature R 4 , feature R 5 , …, feature R M .
[0159] As Figure 6 shown, the computer device can collectively refer to the decoding result y i output by the decoding network at the i-th moment (e.g., the decoding result 605a) as the first decoding result, and collectively refer to the decoding result y i+1 output by the decoding network at the (i + 1)-th moment (e.g., the decoding result 613a) as the second decoding result. The decoding result here is the text output by the decoding network. It can be understood that the decoding network can use the target speech sequence feature output by the encoding network and the decoding result (e.g., the first decoding result) output by the encoding network at the previous moment to jointly generate the decoding result at the current moment (e.g., the second decoding result).
[0160] Among them, the vector conversion network (e.g., Figure 6The vector conversion network 606a) shown can convert the input data into word vectors (embedding vectors) to provide embedding information at the i-th moment for the decoding network. The vector conversion network may include an embedding matrix, and the computer device can pre-train the embedding matrix using a transcript. For example, the word2vec open-source model can be used to increase the spatial distance of all embeddings in the transcript and increase the divergence of the embedding matrix. As Figure 6 shown, inputting the decoding result 605a into the vector conversion network 606a can convert the decoding result 605a into the corresponding word vector 607a (i.e., the target word vector). It can be understood that at the initial moment (i.e., when i = 1), a decoding result of the previous moment (i.e., the 0-th moment) can be assumed as the first decoding result at this time.
[0161] Furthermore, the computer device can generate initial weight coefficients based on the target speech sequence features (e.g., Figure 6 the speech sequence feature 604a shown) and the target word vector (e.g., Figure 6 the word vector 607a shown), and then the initial weight coefficients can be normalized (e.g., using the softmax function), and the obtained target weight coefficients (e.g., Figure 6 the weight coefficient 608a shown). Subsequently, based on the target weight coefficients and the target speech sequence features, for example, a weighted sum of the target weight coefficients and the target speech sequence features is performed to obtain a semantic encoding vector (i.e., the context vector, e.g., Figure 6 the vector 609a shown). The process described here is the attention mechanism.
[0162] Furthermore, the computer device can concatenate the target word vector and the semantic encoding vector to obtain the current target concatenated vector, and then the target concatenated vector can be input into the decoding network in the target speech recognition model, and the decoding network outputs the second decoding result at the i + 1-th moment.
[0163] Among them, the decoding network in the embodiments of the present application can adopt a unidirectional long short-term memory network (e.g., LSTM) because in the process of text recognition of speech data, the associated speech and text strictly follow the time order. The unidirectional long short-term memory network can be N memory networks, and the N memory networks can specifically include memory network D i and memory network D i+1 , and memory network D i+1 is memory network D iThe next memory network. It can be understood that the network structure of this unidirectional long short-term memory network is similar to the network structure of the forward long short-term memory network in the encoding network 603a shown in Figure 6 . Therefore, this unidirectional long short-term memory network can also be used to calculate the unidirectional hidden features of the above-mentioned target splicing vector at each moment along the first feature calculation direction from left to right. For example, the unidirectional hidden features extracted by the computer device at any two adjacent moments (for example, the i-th moment and the i+1-th moment) may include: the unidirectional target hidden feature s i at the i-th moment and the unidirectional target hidden feature s i+1 at the i+1-th moment. Therefore, the unidirectional hidden feature (for example, the unidirectional target hidden feature s i ) extracted by the computer device at the i-th moment is essentially jointly determined by the unidirectional historical hidden feature s i-1 and the splicing vector of the previous moment. Similarly, the unidirectional hidden feature (for example, the unidirectional target hidden feature s i+1 ) extracted by the computer device at the i+1-th moment is essentially jointly determined by the unidirectional target hidden feature s i and the target splicing vector.
[0164] It can be seen that when the computer device obtains the unidirectional target hidden feature s i extracted by the memory network D i in the unidirectional long short-term memory network at the i-th moment, it can input the unidirectional target hidden feature s i and the target splicing vector into the memory network D i+1 . The memory network D i+1 extracts the unidirectional target hidden feature s i+1 at the i+1-th moment, and can obtain the second decoding result at the i+1-th moment based on the target splicing vector and the unidirectional target hidden feature s i+1 .
[0165] As Figure 6 shown, the unidirectional hidden feature 612a can be obtained based on the unidirectional hidden feature 611a. After merging the vector 609a and the word vector 607a and inputting them into the decoding network 610a, the decoding network 610a can output the decoding result 613a at the i+1-th moment. Finally, the text recognition result can be determined based on the decoding result 613a and the decoding result 605a. It can be understood that the decoding results output by the decoding network at all moments can be used as the text information associated with the first object.
[0166] Step S104, determine the audio type of the original audio data based on the text recognition result, and store the audio data associated with the second object in the second type of audio track.
[0167] It should be understood that the text recognition result in the embodiment of the present application may include the target recognition result of the background music object and the role recognition result of the role object.
[0168] Optionally, if the target recognition result in the text recognition result is a null value, it can be determined that the audio type of the original audio data is pure music type; conversely, optionally, if the target recognition result in the text recognition result is a non-null value, it can be determined that the audio type of the original audio data is non-pure music type.
[0169] Optionally, if the role recognition result in the text recognition result is a null value, it can be determined that there is no object voice data in the original audio data; conversely, optionally, if the role recognition result in the text recognition result is a non-null value, it can be determined that there is object voice data in the original audio data.
[0170] It can be understood that the computer device can store the audio data associated with the second object in the second type of audio track. For example, store the accompaniment audio data associated with the accompaniment object in the second type of audio track, that is, store the separated accompaniment in the original audio data.
[0171] In addition, the computer device can also use the decoding result associated with the text recognition result as the text information associated with the first object, and can associate and store the voice data associated with the first object in the first type of audio track and this text information. For example, when the target recognition result is a non-null value, the decoding result associated with the target recognition result can be used as the background music text information (such as lyrics); and for another example, when the role recognition result is a non-null value, the decoding result associated with the role recognition result can be used as the object text information (such as lines). Therefore, the method provided in the embodiment of the present application can also be applied to the business scenario of intelligently generating lines / lyrics.
[0172] Please also refer to Figure 7 , Figure 7 which is a schematic diagram of a scenario for audio classification provided by the embodiment of the present application. As Figure 7 shown, assuming that for a song 7A, the finally obtained text recognition result is text recognition result 701a, and assuming that this text recognition result 701a is a non-null value, the corresponding text information 702a (such as "How are you") can be obtained, and at the same time, it can be determined that the song 7A is non-pure music; assuming that this text recognition result 701a is a null value, it can be seen that the corresponding text information 703a is also a null value, so it can be determined that the song 7A is pure music.
[0173] It can be understood that the above-mentioned target vocal separation model and target speech recognition model are both model solutions that are fully applicable to other actual business scenarios and have a certain degree of universality for commercial implementation. It should be understood that in addition to the network described in the above embodiments, the target vocal separation model and target speech recognition model can also be built using other networks. The embodiments of the present application do not limit the network used for the models. For example, an ASR model built based on transformer can be used to replace the above-mentioned attention-based ASR model.
[0174] As can be seen from the above, the present application proposes an audio data recognition method based on a target vocal separation model and a target speech recognition model, which can automatically and efficiently recognize all audio data, thereby getting rid of the problems of slow speed and low efficiency caused by manual annotation and manual recognition in the prior art. In the embodiments of the present application, the target vocal separation model can be used to separate the two types of audio tracks, namely, the human voice and the accompaniment, in the original audio data, and then the target speech recognition model is used to perform speech recognition on the separated human voice track (i.e., the first type of audio track), and based on the information obtained from the speech recognition, the audio type of the original audio data is determined (i.e., whether the song in the original audio data is pure music). At the same time, the embodiments of the present application can reduce the interference of the audio data in the second type of audio track (for example, the accompaniment audio data) on the target speech recognition model, and since the target speech recognition model has a certain degree of noise resistance, it can reduce the interference of impure separation and also reduce the impact of occasional human voice screams, etc. Therefore, the accuracy of audio data recognition can be improved. In addition, the target audio recognition model can also obtain the accompaniment track (i.e., the second type of audio track) of the song and the corresponding text information, for example, the background music text information (in the case of non-pure music). When the original audio data contains object speech data, the object text information corresponding to the object speech data can also be obtained. That is to say, the target audio recognition model belongs to a multi-task system.
[0175] Further, please refer to Figure 8 , Figure 8 which is a schematic flowchart of a multimedia data processing method provided by the embodiments of the present application. This method can be executed by the above-mentioned computer device. Among them, this method can at least include the following steps:
[0176] Step S201, obtain sample audio data for training the initial audio recognition model, and use the labeled audio type corresponding to the sample audio data as the sample type label of the sample audio data; the sample audio data is obtained from the sample multimedia file; the initial audio recognition model includes an initial vocal separation model and an initial speech recognition model;
[0177] Specifically, after obtaining the sample multimedia file, the computer device can obtain the initial audio recognition model and simultaneously perform audio extraction processing on the sample multimedia file to obtain sample audio data. The specific process can refer to step S101 in the corresponding embodiment above, which will not be elaborated here. In addition, for subsequent calculation of the loss function, the computer device can also obtain the labeled audio type corresponding to the sample audio data and use the labeled audio type as the sample type label of the sample audio data. Figure 3 In step S202, input the sample audio data into the initial vocal separation model, and the initial vocal separation model performs vocal separation on the sample audio data to obtain a first type of sample audio track associated with the first sample object in the sample audio data and a second type of sample audio track associated with the second sample object in the sample audio data.
[0178] It should be understood that the initial vocal separation model may include a first initial segmentation network and a second initial segmentation network. The computer device can first obtain the sample audio track amplitude spectrum corresponding to the sample audio data, and then input the sample audio track amplitude spectrum into the first initial segmentation network and the second initial segmentation network respectively, so as to obtain the sample speech features and sample audio features associated with the sample audio track amplitude spectrum. Furthermore, based on the sample speech features, sample audio features, and sample audio track amplitude spectrum, a first type of sample audio track associated with the first sample object in the sample audio data and a second type of sample audio track associated with the second sample object in the sample audio data can be obtained. The specific implementation manner of this step can refer to step S102 in the corresponding embodiment above, which will not be elaborated here.
[0179] In step S203, obtain the speech data of the first sample object from the first type of sample audio track, input the speech data of the first sample object into the initial speech recognition model, and the initial speech recognition model performs text recognition on the speech data of the first sample object. Based on the obtained text recognition result of the first sample object, determine the predicted audio type of the sample audio data, and use the predicted audio type as the predicted type label. Figure 3 Specifically, the computer device can obtain the sample speech sequence corresponding to the speech data of the first sample object included in the first type of sample audio track. Furthermore, the sample speech sequence can be input into the initial speech recognition model, and the initial speech recognition model can obtain the text recognition result of the first sample object, and can determine the predicted audio type of the sample audio data based on the text recognition result of the first sample object. Subsequently, the predicted audio type can be used as the predicted type label. The specific implementation manner of this step can refer to steps S103 - S104 in the corresponding embodiment above, which will not be elaborated here.
[0180]
[0181] Figure 3
[0182] Step S204: Iteratively train the initial audio recognition model based on the prediction type label and the sample type label to obtain a target audio recognition model for audio processing of the original audio data in the multimedia file.
[0183] Specifically, the computer device can generate a target loss function based on the prediction type label and the sample type label, and then can correct the model parameters in the initial audio recognition model based on the target loss function. Through multiple iterative trainings, finally, a target audio recognition model for audio processing of the original audio data in the multimedia file can be obtained.
[0184] It should be understood that the training method adopted in the embodiments of the present application is to jointly train the initial vocal separation model and the initial speech recognition model. Optionally, the initial vocal separation model and the initial speech recognition model can also be separately trained. The embodiments of the present application do not limit the adopted training method.
[0185] As can be seen from the above, through training the initial vocal separation model and the initial speech recognition model, the embodiments of the present application can obtain an audio processing system jointly composed of a target vocal separation model and a target speech recognition model, that is, a target audio recognition model. The target vocal separation model and the target speech recognition model in the embodiments of the present application can be separately built and trained, and both of these models are fully applicable to other landing business scenarios. For example, the target vocal separation model can be used in business scenarios that require vocal separation (for example, extracting the speech data of the human voice part from audio files), and the target speech recognition model can be used in business scenarios that require text recognition of speech data (for example, intelligently recognizing the speech data in video files or the speech data input by certain users and automatically generating corresponding text information). In addition, the target audio recognition model provided by the embodiments of the present application is applicable to all audio data, that is, it can automatically and efficiently recognize all audio data. For example, it can also recognize pure music with pure human voice humming, thus getting rid of the problems of slow speed and low efficiency caused by manual annotation and manual recognition in the prior art. Therefore, this target audio recognition model has better versatility.
[0186] Please refer to Figure 9 , which is a schematic structural diagram of a multimedia data processing device provided by the embodiments of the present application. The multimedia data processing device 1 can be a computer program (including program code) running on a computer device. For example, the multimedia data processing device 1 is an application software; this device can be used to execute the corresponding steps in the multimedia data processing method provided by the embodiments of the present application. As Figure 9 shown, the multimedia data processing 1 can include: an acquisition module 11, a separation module 12, an identification module 13, and a first determination module 14;
[0187] An obtaining module 11, configured to obtain a target audio recognition model for performing audio processing on the original audio data when the original audio data in the multimedia file is obtained; the target audio recognition model includes a target vocal separation model and a target speech recognition model;
[0188] A separation module 12, configured to input the original audio data into the target vocal separation model, and the target vocal separation model performs vocal separation on the original audio data to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data;
[0189] An identification module 13, configured to obtain the speech data of the first object from the first type of audio track, input the speech data of the first object into the target speech recognition model, and the target speech recognition model performs text recognition on the speech data of the first object to obtain the text recognition result of the first object;
[0190] A first determination module 14, configured to determine the audio type of the original audio data based on the text recognition result, and store the audio data associated with the second object in the second type of audio track.
[0191] Among them, the specific functional implementation manners of the obtaining module 11, the separation module 12, the identification module 13, and the first determination module 14 can refer to steps S101 - S104 in the corresponding embodiments above, which will not be elaborated here. Figure 3 For details, please refer to
[0192] Please refer to Figure 10 , which is a schematic structural diagram of a multimedia data processing device provided by an embodiment of the present application. The multimedia data processing device 2 may be a computer program (including program code) running on a computer device. For example, the multimedia data processing device 2 is an application software; the device can be used to execute the corresponding steps in the multimedia data processing method provided by an embodiment of the present application. As Figure 10 shown, the multimedia data processing device 2 may include: an obtaining module 21, a separation module 22, an identification module 23, a first determination module 24, a separation update module 25, and a second determination module 26;
[0193] An obtaining module 21, configured to obtain a target audio recognition model for performing audio processing on the original audio data when the original audio data in the multimedia file is obtained; the target audio recognition model includes a target vocal separation model and a target speech recognition model;
[0194] A separation module 22 for inputting the original audio data into a target vocal separation model, which separates the vocals from the original audio data to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data;
[0195] Wherein, the above target vocal separation model includes a first segmentation network for extracting the speech features corresponding to the original audio data and a second segmentation network for extracting the audio features corresponding to the original audio data;
[0196] The separation module 22 may include: a spectrum acquisition unit 221, a first feature extraction unit 222, a second feature extraction unit 223, and an audio track separation unit 224;
[0197] The spectrum acquisition unit 221 is used to acquire the original audio track amplitude spectrum corresponding to the original audio data;
[0198] The first feature extraction unit 222 is used to input the original audio track amplitude spectrum into the first segmentation network, and the first segmentation network outputs the speech features associated with the original audio track amplitude spectrum;
[0199] The second feature extraction unit 223 is used to input the original audio track amplitude spectrum into the second segmentation network, and the second segmentation network outputs the audio features associated with the original audio track amplitude spectrum;
[0200] Wherein, the above first segmentation network and the second segmentation network are symmetric networks with the same network structure; the symmetric network includes U convolutional layers and U upsampling layers; the p-th convolutional layer in the U convolutional layers is used to obtain the convolutional features associated with the original audio track amplitude spectrum; p is a positive integer less than or equal to U; the q-th upsampling layer in the U upsampling layers is used to splice the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer when obtaining the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer; the (q - 1)-th upsampling layer is the previous upsampling layer of the q-th upsampling layer; q is a positive integer less than or equal to U, and q = U - p + 1;
[0201] The audio track separation unit 224 is used to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data based on the speech features, audio features, and the original audio track amplitude spectrum;
[0202] Wherein, the audio track separation unit 224 may include: a feature fusion subunit 2241, a first audio track acquisition subunit 2242, and a second audio track acquisition subunit 2243;
[0203] A feature fusion subunit 2241, configured to perform feature fusion on the speech feature and the audio feature to obtain a target fusion feature;
[0204] A first audio track obtaining subunit 2242, configured to generate a first mask associated with the speech feature based on the target fusion feature and the speech feature, generate a first audio track amplitude spectrum based on the first mask and the original audio track amplitude spectrum, and perform inverse spectral transformation on the first audio track amplitude spectrum to obtain a first type of audio track associated with a first object in the original audio data;
[0205] Wherein, if the multimedia file is a video file, the first object includes a role object and a first background music object in the video file;
[0206] Specifically, the first audio track obtaining subunit 2242 is configured to perform inverse spectral transformation on the first audio track amplitude spectrum to obtain a first mixed speech audio track associated with the first object; the first mixed speech audio track carries object speech data associated with the role object and first background music speech data associated with the first background music object; perform voiceprint feature recognition on the object speech data and the first background music speech data carried in the first mixed speech audio track, use the recognized voiceprint feature of the role object as the first voiceprint feature, and use the recognized voiceprint feature of the first background music object as the second voiceprint feature; based on the first voiceprint feature and the second voiceprint feature, perform speech segmentation on the object speech data and the first background music speech data in the first mixed speech audio track to obtain object speech data corresponding to the first voiceprint feature and first background music speech data corresponding to the second voiceprint feature; use the object speech data corresponding to the first voiceprint feature and the first background music speech data corresponding to the second voiceprint feature as the first type of audio track associated with the first object;
[0207] Wherein, if the multimedia file is an audio file, the first object includes a second background music object in the audio file;
[0208] Specifically, the first audio track obtaining subunit 2242 is configured to perform inverse spectral transformation on the first audio track amplitude spectrum to obtain a second mixed speech audio track associated with the first object; the second mixed speech audio track carries second background music speech data associated with the second background music object; use the second background music speech data obtained in the second mixed speech audio track as the first type of audio track associated with the first object;
[0209] A second audio track obtaining subunit 2243, configured to generate a second mask associated with the audio feature based on the target fusion feature and the audio feature, generate a second audio track amplitude spectrum based on the second mask and the original audio track amplitude spectrum, and perform inverse spectral transformation on the second audio track amplitude spectrum to obtain a second type of audio track associated with a second object in the original audio data;
[0210] Among them, the specific implementation manners of the feature fusion subunit 2241, the first audio track acquisition subunit 2242, and the second audio track acquisition subunit 2243 can be referred to the description of step S102 in the corresponding embodiments above, and will not be elaborated here. Figure 3 The description of step S102 in the corresponding embodiments will not be elaborated here.
[0211] Among them, the specific implementation manners of the spectrum acquisition unit 221, the first feature extraction unit 222, the second feature extraction unit 223, and the audio track separation unit 224 can be referred to the description of step S102 in the corresponding embodiments above, and will not be elaborated here. Figure 3 The description of step S102 in the corresponding embodiments will not be elaborated here.
[0212] The recognition module 23 is configured to obtain the speech data of the first object from the first type of audio track, input the speech data of the first object into the target speech recognition model, and the target speech recognition model performs text recognition on the speech data of the first object to obtain the text recognition result of the first object.
[0213] Among them, the recognition module 23 may include: a third feature extraction unit 231, a vector conversion unit 232, a decoding output unit 233, and a result determination unit 234.
[0214] The third feature extraction unit 231 is configured to obtain the to-be-processed speech sequence corresponding to the speech data of the first object included in the first type of audio track, input the to-be-processed speech sequence into the encoding network in the target speech recognition model, and the encoding network extracts the speech sequence features of the to-be-processed speech sequence, and uses the extracted speech sequence features as the target speech sequence features corresponding to the first object.
[0215] Among them, the target speech recognition model includes an encoding network, and the encoding network in the target speech recognition model is a bidirectional long short-term memory network; the bidirectional long short-term memory network includes a forward long short-term memory network and a backward long short-term memory network; the forward long short-term memory network includes memory network B j and memory network B j+1 , memory network B j+1 is the next memory network of memory network B j ; the backward long short-term memory network includes memory network C j+1 and memory network C j ; memory network C j+1 is the previous memory network of memory network C j ; j is a positive integer less than or equal to M; the number of memory networks in the forward long short-term memory network and the backward long short-term memory network is both M.
[0216] The third feature extraction unit 231 may include: a forward feature extraction subunit 2311, a reverse feature extraction subunit 2312, a feature splicing subunit 2313, and a feature determination subunit 2314;
[0217] The forward feature extraction subunit 2311 is configured to obtain a forward historical hidden feature h j associated with the memory network B in the forward long short-term memory network j-1 , input the speech sequence to be processed and the forward historical hidden feature h j-1 into the memory network B j , and extract, by the memory network B j , a forward target hidden feature h at the j-th moment j , input the forward target hidden feature h j and the speech sequence to be processed into the memory network B j+1 , and extract, by the memory network B j+1 , a forward target hidden feature h at the (j + 1)-th moment j+1 ;
[0218] The reverse feature extraction subunit 2312 is configured to obtain a reverse historical hidden feature k j+1 associated with the memory network C in the reverse long short-term memory network j+1 , input the speech sequence to be processed and the reverse historical hidden feature k j+1 into the memory network C j+1 , and extract, by the memory network C j+1 , a reverse target hidden feature k at the (j + 1)-th moment j , input the reverse target hidden feature k j and the speech sequence to be processed into the memory network C j , and extract, by the memory network C j , a reverse target hidden feature k at the j-th moment j-1 ;
[0219] The feature splicing subunit 2313 is configured to splice the forward target hidden feature h j extracted by the memory network B at the j-th moment j with the reverse target hidden feature k j extracted by the memory network C at the j-th moment j-1 to obtain a first spliced feature, and splice the forward target hidden feature h j+1 extracted by the memory network B at the (j + 1)-th moment j+1 with the reverse target hidden feature k j+1 extracted by the memory network C at the (j + 1)-th moment j to obtain a second spliced feature;
[0220] A feature determination subunit 2314, configured to determine, based on a first splicing feature and a second splicing feature, a target speech sequence feature corresponding to a first object extracted from a speech sequence to be processed.
[0221] Wherein, the specific implementation manners of the forward feature extraction subunit 2311, the reverse feature extraction subunit 2312, the feature splicing subunit 2313, and the feature determination subunit 2314 may refer to the description of step S103 in the corresponding embodiment above, and will not be elaborated here. Figure 3 The description will not be continued here.
[0222] A vector conversion unit 232, configured to obtain a first decoding result output by a decoding network in a target speech recognition model at the i-th moment, input the first decoding result into a vector conversion network in the target speech recognition model, and convert the first decoding result into a target word vector by the vector conversion network; i is a positive integer;
[0223] A decoding output unit 233, configured to obtain a second decoding result output by the decoding network at the (i + 1)-th moment based on the target speech sequence feature, the target word vector, and the decoding network in the target speech recognition model;
[0224] Wherein, the decoding output unit 233 may include: a weight acquisition subunit 2331, a vector generation subunit 2332, a vector splicing subunit 2333, and a decoding output subunit 2334;
[0225] The weight acquisition subunit 2331 is configured to generate an initial weight coefficient based on the target speech sequence feature and the target word vector, and perform normalization processing on the initial weight coefficient to obtain a target weight coefficient;
[0226] The vector generation subunit 2332 is configured to generate a semantic coding vector based on the target weight coefficient and the target speech sequence feature;
[0227] The vector splicing subunit 2333 is configured to perform vector splicing on the target word vector and the semantic coding vector to obtain a target spliced vector;
[0228] The decoding output subunit 2334 is configured to input the target spliced vector into the decoding network in the target speech recognition model, and the decoding network outputs a second decoding result at the (i + 1)-th moment;
[0229] Wherein, the decoding network in the target speech recognition model is a unidirectional long short-term memory network; the unidirectional long short-term memory network includes a memory network D i and a memory network D i+1 , the memory network D i+1 is a memory network D ithe next memory network; i is a positive integer less than or equal to N; the number of memory networks in the unidirectional long short-term memory network is N;
[0230] The decoding output subunit 2334 is specifically configured to obtain the unidirectional target hidden feature s i extracted from the memory network D in the unidirectional long short-term memory network at the i-th moment i ; the unidirectional target hidden feature s i is obtained based on the unidirectional historical hidden feature s i associated with the memory network D; the unidirectional target hidden feature s i-1 and the target splicing vector are input into the memory network D i , and the memory network D i+1 extracts the unidirectional target hidden feature s at the (i + 1)-th moment i+1 , and based on the target splicing vector and the unidirectional target hidden feature s i+1 , obtains the second decoding result at the (i + 1)-th moment; i+1
[0231] Among them, the specific implementation manners of the weight acquisition subunit 2331, the vector generation subunit 2332, the vector splicing subunit 2333, and the decoding output subunit 2334 can refer to the description of step S103 in the corresponding embodiment above, and will not be elaborated here. Figure 3
[0232] The result determination unit 234 is configured to determine the text recognition result of the first object based on the first decoding result and the second decoding result;
[0233] Among them, the specific implementation manners of the third feature extraction unit 231, the vector conversion unit 232, the decoding output unit 233, and the result determination unit 234 can refer to the description of step S103 in the corresponding embodiment above, and will not be elaborated here. Figure 3
[0234] The first determination module 24 is configured to determine the audio type of the original audio data based on the text recognition result, and store the audio data associated with the second object in the second type of audio track;
[0235] Among them, the first object includes a background music object, and the text recognition result includes the target recognition result of the background music object; the second object includes an accompaniment object;
[0236] The first determination module 24 is specifically configured to determine that the audio type of the original audio data is a pure music type if the target recognition result in the text recognition result is a null value; store the accompaniment audio data associated with the accompaniment object in the second type of audio track.
[0237] Among them, if the multimedia file is a video file, the first type of audio track includes object voice data associated with the character objects in the video file, and the second type of audio track includes audio data associated with the background objects in the video file; the background objects include the third background music object and the accompaniment object;
[0238] The separation and update module 25 is configured to input the audio data associated with the background objects in the second type of audio track into the target vocal separation model, perform vocal separation on the audio data associated with the background objects through the target vocal separation model, and obtain the third background music voice data associated with the third background music object and the accompaniment audio data associated with the accompaniment object; add the separated third background music voice data to the first type of audio track including the object voice data to obtain the first type of updated audio track, and use the separated accompaniment audio data as the second type of updated audio track;
[0239] The second determination module 26 is configured to, if the target recognition result in the text recognition result is a non-empty value, determine that the audio type of the original audio data is a non-pure music type; use the decoding result associated with the text recognition result as the text information associated with the first object; and perform associated storage on the voice data and the text information associated with the first object in the first type of audio track.
[0240] Among them, the specific implementation manners of the acquisition module 21, the separation module 22, the recognition module 23, the first determination module 24, the separation and update module 25, and the second determination module 26 can refer to the descriptions of steps S101 - S104 in the corresponding embodiments above, and will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Figure 3 Please refer to
[0241] which is a schematic structural diagram of a multimedia data processing device provided in an embodiment of the present application. The multimedia data processing device 3 may be a computer program (including program codes) running on a computer device. For example, the multimedia data processing device 3 is an application software; the device may be used to execute the corresponding steps in the multimedia data processing method provided in the embodiment of the present application. As Figure 11 shown, the multimedia data processing 3 may include: a sample acquisition module 31, an audio track separation module 32, a text recognition module 33, and a model training module 34; Figure 11
[0242] The sample acquisition module 31 is configured to acquire sample audio data for training an initial audio recognition model, and use the labeled audio type corresponding to the sample audio data as the sample type label of the sample audio data; the sample audio data is acquired from a sample multimedia file; the initial audio recognition model includes an initial vocal separation model and an initial speech recognition model;
[0243] The audio track separation module 32 is configured to input the sample audio data into the initial vocal separation model. The initial vocal separation model performs vocal separation on the sample audio data to obtain a first type of sample audio track associated with the first sample object in the sample audio data and a second type of sample audio track associated with the second sample object in the sample audio data;
[0244] The text recognition module 33 is configured to obtain the speech data of the first sample object from the first type of sample audio track, input the speech data of the first sample object into the initial speech recognition model. The initial speech recognition model performs text recognition on the speech data of the first sample object, determines the predicted audio type of the sample audio data based on the obtained text recognition result of the first sample object, and uses the predicted audio type as the predicted type label;
[0245] The model training module 34 is configured to iteratively train the initial audio recognition model based on the predicted type label and the sample type label to obtain a target audio recognition model for processing the original audio data in the multimedia file.
[0246] Among them, the specific implementation manners of the sample acquisition module 31, the audio track separation module 32, the text recognition module 33, and the model training module 34 can refer to the descriptions of steps S201 - S204 in the corresponding embodiments above, and will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Figure 8 Please refer to
[0247] which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 12 shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to implement connection communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1004 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally be at least one storage device located far from the aforementioned processor 1001. As Figure 12 shown, Figure 12As shown in the figure, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0248] In the computer device 1000 as shown in Figure 12 the figure, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to execute the multimedia data processing method described in any one of the foregoing Figure 3 , Figure 8 corresponding embodiments, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0249] In addition, it should be pointed out here that: The embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the foregoing multimedia data processing device 1, multimedia data processing device 2, or multimedia data processing device 3. The computer program includes program instructions. When the processor executes the program instructions, it can execute the multimedia data processing method described in any one of the foregoing Figure 3 , Figure 9 corresponding embodiments, so it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiment of the present application.
[0250] The above computer-readable storage medium may be an internal storage unit of the multimedia data processing device provided in any of the foregoing embodiments or the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.
[0251] In addition, it should be noted here that: The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in any of the foregoing Figure 3 , Figure 8 corresponding embodiments. In addition, the description of the beneficial effects of the same method will not be repeated. For the technical details not disclosed in the computer program product or the computer program embodiments involved in the present application, please refer to the description of the method embodiments of the present application.
[0252] Please refer to Figure 13 , Figure 13 , which is a schematic structural diagram of a multimedia data processing system provided by an embodiment of the present application. The multimedia data processing system 4 may include a multimedia data processing device 1a and a multimedia data processing device 2a. Among them, the multimedia data processing device 1a may be the multimedia data processing device 1 in the corresponding embodiment above Figure 9 , or may be the multimedia data processing device 2 in the corresponding embodiment above Figure 10 . It can be understood that the multimedia data processing device 1a may be integrated into the audio recognition model 203A in the corresponding embodiment above Figure 2 , so it will not be repeated here. Among them, the multimedia data processing device 2a may be the multimedia data processing device 3 in the corresponding embodiment above Figure 11 , so it will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For the technical details not disclosed in the multimedia data processing system embodiments involved in the present application, please refer to the description of the method embodiments of the present application.
[0253] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products or equipment.
[0254] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0255] The above-disclosed are only the preferred embodiments of this application, and of course, the scope of rights of this application cannot be limited thereby. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. A multimedia data processing method, characterized in that, it includes: when obtaining the original audio data in a multimedia file, obtaining a target audio recognition model for performing audio processing on the original audio data; the target audio recognition model includes a target vocal separation model and a target speech recognition model; inputting the original audio data into the target vocal separation model, and separating vocals from the original audio data by the target vocal separation model to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data; obtaining the speech data of the first object from the first type of audio track, inputting the speech data of the first object into the target speech recognition model, and performing text recognition on the speech data of the first object by the target speech recognition model to obtain the text recognition result of the first object; determining the audio type of the original audio data based on the text recognition result, and storing the audio data associated with the second object in the second type of audio track; wherein, if the multimedia file is a video file, the first type of audio track contains object speech data associated with a role object in the video file, and the second type of audio track contains audio data associated with a background object in the video file; the background object includes a third background music object and an accompaniment object; the third background music object refers to the singer who sings the lyrics in the video file, and the accompaniment object refers to the object that generates accompaniment in the video file, and the accompaniment generated by the accompaniment object refers to other audio data in the original audio data except for human voices; the method further includes: inputting the audio data associated with the background object in the second type of audio track into the target vocal separation model, and separating vocals from the audio data associated with the background object by the target vocal separation model to obtain third background music speech data associated with the third background music object and accompaniment audio data associated with the accompaniment object; adding the separated third background music speech data to the first type of audio track containing the object speech data to obtain a first type of updated audio track, and using the separated accompaniment audio data as a second type of updated audio track.
2. The method according to claim 1, characterized in that, the target vocal separation model includes a first segmentation network for extracting the speech features corresponding to the original audio data and a second segmentation network for extracting the audio features corresponding to the original audio data; the inputting the original audio data into the target vocal separation model, and separating vocals from the original audio data by the target vocal separation model to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data includes: obtaining the original audio track amplitude spectrum corresponding to the original audio data; Input the original audio track amplitude spectrum into the first segmentation network, and output the speech features associated with the original audio track amplitude spectrum by the first segmentation network; Input the original audio track amplitude spectrum into the second segmentation network, and output the audio features associated with the original audio track amplitude spectrum by the second segmentation network; Based on the speech features, the audio features, and the original audio track amplitude spectrum, obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data.
3. The method according to claim 2, wherein, the first segmentation network and the second segmentation network are symmetric networks with the same network structure; the symmetric network includes U convolutional layers and U upsampling layers; the p-th convolutional layer in the U convolutional layers is used to obtain the convolutional features associated with the original audio track amplitude spectrum; p is a positive integer less than or equal to U; the q-th upsampling layer in the U upsampling layers is used to splice the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer when obtaining the convolutional features of the p-th convolutional layer and the upsampling features of the (q - 1)-th upsampling layer; the (q - 1)-th upsampling layer is the previous upsampling layer of the q-th upsampling layer; q is a positive integer less than or equal to U, and q = U - p + 1.
4. The method according to claim 2, wherein, the obtaining a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data based on the speech features, the audio features, and the original audio track amplitude spectrum includes: Fuse the speech features and the audio features to obtain target fusion features; Based on the target fusion features and the speech features, generate a first mask associated with the speech features, based on the first mask and the original audio track amplitude spectrum, generate a first audio track amplitude spectrum, and perform inverse spectrum transformation on the first audio track amplitude spectrum to obtain a first type of audio track associated with a first object in the original audio data; Based on the target fusion features and the audio features, generate a second mask associated with the audio features, based on the second mask and the original audio track amplitude spectrum, generate a second audio track amplitude spectrum, and perform inverse spectrum transformation on the second audio track amplitude spectrum to obtain a second type of audio track associated with a second object in the original audio data.
5. The method according to claim 4, wherein, if the multimedia file is a video file, the first object includes a role object and a first background music object in the video file; the performing inverse spectrum transformation on the first audio track amplitude spectrum to obtain a first type of audio track associated with a first object in the original audio data includes: Perform an inverse spectrum transformation on the amplitude spectrum of the first audio track to obtain a first mixed speech audio track associated with the first object; the first mixed speech audio track carries object speech data associated with the character object and first background music speech data associated with the first background music object; Perform voiceprint feature recognition on the object speech data and the first background music speech data carried in the first mixed speech audio track, use the recognized voiceprint feature of the character object as the first voiceprint feature, and use the recognized voiceprint feature of the first background music object as the second voiceprint feature; Based on the first voiceprint feature and the second voiceprint feature, perform voice segmentation on the object speech data and the first background music speech data in the first mixed speech audio track to obtain the object speech data corresponding to the first voiceprint feature and the first background music speech data corresponding to the second voiceprint feature; Use the object speech data corresponding to the first voiceprint feature and the first background music speech data corresponding to the second voiceprint feature as the first type of audio track associated with the first object.
6. The method according to claim 4, wherein, if the multimedia file is an audio file, the first object includes a second background music object in the audio file; The performing an inverse spectrum transformation on the amplitude spectrum of the first audio track to obtain a first type of audio track associated with the first object in the original audio data includes: Perform an inverse spectrum transformation on the amplitude spectrum of the first audio track to obtain a second mixed speech audio track associated with the first object; the second mixed speech audio track carries second background music speech data associated with the second background music object; Use the second background music speech data obtained in the second mixed speech audio track as the first type of audio track associated with the first object.
7. The method according to claim 1, wherein, The obtaining the speech data of the first object from the first type of audio track, inputting the speech data of the first object into the target speech recognition model, and having the target speech recognition model perform text recognition on the speech data of the first object to obtain the text recognition result of the first object includes: Obtain a to-be-processed speech sequence corresponding to the speech data of the first object included in the first type of audio track, input the to-be-processed speech sequence into the encoding network in the target speech recognition model, have the encoding network extract the speech sequence feature of the to-be-processed speech sequence, and use the extracted speech sequence feature as the target speech sequence feature corresponding to the first object; Obtain a first decoding result output by the decoding network in the target speech recognition model at the i-th moment, input the first decoding result into the vector conversion network in the target speech recognition model, and have the vector conversion network convert the first decoding result into a target word vector; i is a positive integer; Based on the target speech sequence feature, the target word vector, and the decoding network in the target speech recognition model, obtain a second decoding result output by the decoding network at the (i + 1)-th moment; Determine the text recognition result of the first object based on the first decoding result and the second decoding result.
8. The method according to claim 7, wherein, The target speech recognition model includes an encoding network, and the encoding network in the target speech recognition model is a bidirectional long short-term memory network; the bidirectional long short-term memory network includes a forward long short-term memory network and a backward long short-term memory network; the forward long short-term memory network includes memory network B j and memory network B j+1 , and the memory network B j+1 is the next memory network of the memory network B j ; the backward long short-term memory network includes memory network C j+1 and memory network C j ; the memory network C j+1 is the previous memory network of the memory network C j ; the j is a positive integer less than or equal to M; the number of memory networks in both the forward long short-term memory network and the backward long short-term memory network is M; the step of inputting the to-be-processed speech sequence into the encoding network of the target speech recognition model, extracting the speech sequence feature of the to-be-processed speech sequence by the encoding network, and using the extracted speech sequence feature as the target speech sequence feature corresponding to the first object includes: Obtain the forward historical hidden feature h associated with the memory network B in the forward long short-term memory network j , and input the speech sequence to be processed and the forward historical hidden feature h j-1 into the memory network B j-1 . The memory network B j extracts the forward target hidden feature h at the j-th moment j . Then input the forward target hidden feature h j and the speech sequence to be processed into the memory network B j . The memory network B j+1 extracts the forward target hidden feature h at the (j + 1)-th moment j+1 ; j+1 Obtain the reverse historical hidden feature k associated with the memory network C in the reverse long short-term memory network j+1 , and input the speech sequence to be processed and the reverse historical hidden feature k j+1 into the memory network C j+1 . The memory network C j+1 extracts the reverse target hidden feature k at the (j + 1)-th moment j+1 . Input the reverse target hidden feature k j and the speech sequence to be processed into the memory network C j . The memory network C j extracts the reverse target hidden feature k at the j-th moment j ; j-1 The memory network B j The forward target hidden feature h obtained by extraction at the j-th moment j and the memory network C j The backward target hidden feature k obtained by extraction at the j-th moment j-1 are subjected to feature splicing to obtain a first spliced feature. The memory network B j+1 The forward target hidden feature h obtained by extraction at the (j + 1)-th moment j+1 and the memory network C j+1 The backward target hidden feature k obtained by extraction at the (j + 1)-th moment j are subjected to feature splicing to obtain a second spliced feature; Determine the target speech sequence feature corresponding to the first object extracted from the to-be-processed speech sequence based on the first splicing feature and the second splicing feature.
9. The method according to claim 7, wherein, the step of obtaining the second decoding result output by the decoding network at the (i + 1)-th moment based on the target speech sequence feature, the target word vector, and the decoding network in the target speech recognition model includes: Generate an initial weight coefficient based on the target speech sequence feature and the target word vector, and perform normalization processing on the initial weight coefficient to obtain a target weight coefficient; Generate a semantic encoding vector based on the target weight coefficient and the target speech sequence feature; Perform vector splicing on the target word vector and the semantic encoding vector to obtain a target splicing vector; Input the target splicing vector into the decoding network in the target speech recognition model, and the decoding network outputs the second decoding result at the (i + 1)-th moment.
10. The method according to claim 9, wherein, The decoding network in the target speech recognition model is a unidirectional long short-term memory network; the unidirectional long short-term memory network includes a memory network D i and a memory network D i+1 , the memory network D i+1 is the next memory network of the memory network D i ; i is a positive integer less than or equal to N; the number of memory networks in the unidirectional long short-term memory network is N; the step of inputting the target splicing vector into the decoding network in the target speech recognition model, and the decoding network outputs the second decoding result at the (i + 1)-th moment includes: Obtain the memory network D in the one-way long short-term memory network i The one-way target hidden feature s extracted at the i-th moment i ; The one-way target hidden feature s i is based on the memory network D i The associated one-way historical hidden feature s i-1 is obtained; Input the unidirectional target hidden feature s i and the target splicing vector into the memory network D i+1 , and the memory network D i+1 extracts the unidirectional target hidden feature s at the (i + 1)-th moment i+1 . Based on the target splicing vector and the unidirectional target hidden feature s i+1 , obtain the second decoding result at the (i + 1)-th moment 11. The method according to claim 1, wherein, the first object includes a background music object, and the text recognition result includes the target recognition result of the background music object; the second object includes an accompaniment object; the step of determining the audio type of the original audio data based on the text recognition result and storing the audio data associated with the second object in the second type of audio track includes: If the target recognition result in the text recognition result is a null value, determine that the audio type of the original audio data is a pure music type; Store the accompaniment audio data associated with the accompaniment object in the second type of audio track.
12. The method according to claim 11, wherein, further includes: If the target recognition result in the text recognition result is a non-null value, determine that the audio type of the original audio data is a non-pure music type; Use the decoding result associated with the text recognition result as the text information associated with the first object; Associatively store the speech data associated with the first object and the text information in the first type of audio track.
13. A multimedia data processing method, wherein, includes: Obtain sample audio data for training an initial audio recognition model, and use the labeled audio type corresponding to the sample audio data as the sample type label of the sample audio data; the sample audio data is obtained from sample multimedia files; the initial audio recognition model includes an initial vocal separation model and an initial speech recognition model; Input the sample audio data into the initial vocal separation model, and have the initial vocal separation model perform vocal separation on the sample audio data to obtain a first type of sample audio track associated with a first sample object in the sample audio data and a second type of sample audio track associated with a second sample object in the sample audio data; Obtain the speech data of the first sample object from the first type of sample audio track, input the speech data of the first sample object into the initial speech recognition model, have the initial speech recognition model perform text recognition on the speech data of the first sample object, determine the predicted audio type of the sample audio data based on the obtained text recognition result of the first sample object, and use the predicted audio type as the predicted type label; Iteratively train the initial audio recognition model based on the predicted type label and the sample type label to obtain the target audio recognition model in the method according to any one of claims 1-12.
14. A multimedia data processing device, characterized in that, it includes: an acquisition module, configured to obtain a target audio recognition model for audio processing of the original audio data when the original audio data in the multimedia file is obtained; the target audio recognition model includes a target vocal separation model and a target speech recognition model; a separation module, configured to input the original audio data into the target vocal separation model, and have the target vocal separation model perform vocal separation on the original audio data to obtain a first type of audio track associated with a first object in the original audio data and a second type of audio track associated with a second object in the original audio data; an identification module, configured to obtain the speech data of the first object from the first type of audio track, input the speech data of the first object into the target speech recognition model, and have the target speech recognition model perform text recognition on the speech data of the first object to obtain the text recognition result of the first object; a first determination module, configured to determine the audio type of the original audio data based on the text recognition result and store the audio data associated with the second object in the second type of audio track; Wherein, if the multimedia file is a video file, the first type of audio track includes object voice data associated with the character object in the video file, and the second type of audio track includes audio data associated with the background object in the video file; the background object includes a third background music object and an accompaniment object; the third background music object refers to the singer who sings the lyrics in the video file, and the accompaniment object refers to the object that generates the accompaniment in the video file, and the accompaniment generated by the accompaniment object refers to other audio data in the original audio data except for the human voice; The device further includes: A separation and update module, configured to input the audio data associated with the background object in the second type of audio track into the target vocal separation model, and perform vocal separation on the audio data associated with the background object through the target vocal separation model to obtain third background music voice data associated with the third background music object and accompaniment audio data associated with the accompaniment object; The separation and update module is further configured to add the separated third background music voice data to the first type of audio track including the object voice data to obtain a first type of updated audio track, and use the separated accompaniment audio data as the second type of updated audio track.
15. A multimedia data processing device, Characterized in that, It includes: A sample acquisition module, configured to acquire sample audio data for training an initial audio recognition model, and use the labeled audio type corresponding to the sample audio data as the sample type label of the sample audio data; the sample audio data is acquired from a sample multimedia file; the initial audio recognition model includes an initial vocal separation model and an initial speech recognition model; An audio track separation module, configured to input the sample audio data into the initial vocal separation model, and perform vocal separation on the sample audio data by the initial vocal separation model to obtain a first type of sample audio track associated with a first sample object in the sample audio data and a second type of sample audio track associated with a second sample object in the sample audio data; A text recognition module, configured to obtain the voice data of the first sample object from the first type of sample audio track, input the voice data of the first sample object into the initial speech recognition model, perform text recognition on the voice data of the first sample object by the initial speech recognition model, determine the predicted audio type of the sample audio data based on the obtained text recognition result of the first sample object, and use the predicted audio type as the predicted type label; A model training module, configured to perform iterative training on the initial audio recognition model based on the predicted type label and the sample type label to obtain the target audio recognition model in the method according to any one of claims 1-12.
16. A computer device, Characterized in that, It includes: A processor and a memory; The processor is connected to the memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1-13.
17. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-13.
18. A computer program product, characterized in that the computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and the computer instructions are adapted to be read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-13.
Citation Information
Patent Citations
Language recognition method and device, model training method and device, and facility
CN110853618A
Video processing method and device, computer equipment and storage medium
CN113573136A