Singing object recognition method and device, electronic device and storage medium

Through the combined model of hollow convolutional network and convolutional classification network, the feature extraction and recognition of audio data in the metaverse is solved, and the problem of difficulty in identifying singing objects in the metaverse is achieved and higher recognition accuracy is achieved.

CN115238122BActive Publication Date: 2025-05-09PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210906248.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-05-09
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

在元宇宙中,随着演唱人员数量增加,现有的识别方法难以准确识别演唱对象的身份。

Method used

The combination model of hollow convolution network and convolution classification network is used to extract and identify the target audio data, expand the receptive field through the hollow convolution network, and combine the convolution classification network to perform feature fusion and prediction to obtain the target identity tag.

Benefits of technology

It improves the accuracy of the recognition of singing objects and can effectively solve the problem of identity recognition difficulties caused by the increase in the number of singing objects in the metaverse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115238122B_ABST
    Figure CN115238122B_ABST
Patent Text Reader

Abstract

The present application provides a method and device, electronic device and storage medium for identifying a singing object, and belongs to the field of artificial intelligence technology. The method includes: obtaining target audio data of a target singing object; inputting the target audio data into a character recognition model; the character recognition model includes a dilated convolutional network and a convolutional classification network; extracting features of the target audio data through a dilated convolutional network to obtain audio timing features; activating the audio timing features through a dilated convolutional network to obtain an initial audio feature vector; fusing features of multiple initial audio feature vectors to obtain a fused audio feature vector; extracting features of the fused audio feature vector through a convolutional classification network to obtain a target audio feature vector; predicting the target audio feature vector through a convolutional classification network to obtain a target identity label of the target singing object. The present application can improve the recognition accuracy of singing objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a singing object recognition method and device, electronic equipment and storage medium. Background Art

[0002] With the development of metaverse technology, all aspects of daily life can be extended to a world that combines the real and the virtual through the metaverse. However, as the number of singers in the metaverse increases, commonly used recognition methods often find it difficult to accurately identify the identities of the singers. Therefore, how to improve the accuracy of singer recognition has become a technical problem that needs to be solved urgently. Summary of the invention

[0003] The main purpose of the embodiments of the present application is to propose a singing object recognition method and device, electronic device and storage medium, aiming to improve the recognition accuracy of the singing object.

[0004] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a singing object recognition method, the method comprising:

[0005] Acquire target audio data of a target singing object;

[0006] Inputting the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network;

[0007] Extracting features of the target audio data through the atrous convolutional network to obtain audio time series features;

[0008] Activating the audio time series features through the dilated convolutional network to obtain an initial audio feature vector;

[0009] Performing feature fusion on the multiple initial audio feature vectors to obtain a fused audio feature vector;

[0010] Performing feature extraction on the fused audio feature vector through the convolution classification network to obtain a target audio feature vector;

[0011] The target audio feature vector is predicted and processed by the convolution classification network to obtain a target identity label of the target singing object, wherein the target identity label is used to characterize the identity of the target singing object.

[0012] In some embodiments, the step of activating the audio time series features through the dilated convolutional network to obtain an initial audio feature vector includes:

[0013] Activate the audio time series feature through the first activation function of the dilated convolutional network to obtain a first activated feature vector;

[0014] Activate the audio time series feature through the second activation function of the dilated convolutional network to obtain a second activated feature vector;

[0015] Performing a dot multiplication process on the first activated feature vector and the second activated feature vector according to a preset weight parameter to obtain the initial audio feature vector.

[0016] In some embodiments, the convolution classification network includes a two-dimensional convolution layer and a pooling layer, and the step of extracting features from the fused audio feature vector through the convolution classification network to obtain a target audio feature vector includes:

[0017] Performing feature extraction on the fused audio feature vector through the two-dimensional convolutional layer to obtain an intermediate audio feature vector;

[0018] The intermediate audio feature vector is subjected to dimensionality reduction processing by the pooling layer to obtain the target audio feature vector.

[0019] In some embodiments, the convolution classification network includes a flattening layer and a fully connected layer, and the step of performing prediction processing on the target audio feature vector through the convolution classification network to obtain the target identity label of the target singing object includes:

[0020] The target audio feature vector is stretched by the flattening layer to obtain a variable-dimensional audio feature vector;

[0021] The variable-dimensional audio feature vector is subjected to label prediction processing through the fully connected layer to obtain a target identity label of the target singing object.

[0022] In some embodiments, the step of performing label prediction processing on the variable-dimensional audio feature vector through the fully connected layer to obtain the target identity label of the target singing object includes:

[0023] Performing label probability calculation on the variable-dimensional audio feature vector using the classification function of the fully connected layer and the preset identity label to obtain a label probability vector corresponding to each preset identity label;

[0024] Select the preset identity label corresponding to the label probability vector with the largest value to obtain the candidate identity label;

[0025] According to the candidate identity tags, a target identity tag is obtained.

[0026] In some embodiments, the step of performing feature fusion on the multiple initial audio feature vectors to obtain a fused audio feature vector includes:

[0027] Get the preset splicing order;

[0028] Performing vector addition on a plurality of the initial audio feature vectors according to the concatenation order to obtain the fused audio feature vector.

[0029] In some embodiments, the step of obtaining target audio data of a target singing object includes:

[0030] Acquire original audio data of the target singing object;

[0031] The original audio data is filtered and processed according to a preset audio length to obtain the target audio data.

[0032] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a singing object recognition device, the device comprising:

[0033] A data acquisition module, used to acquire target audio data of a target singing object;

[0034] An input module, used to input the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network;

[0035] A first feature extraction module, configured to extract features of the target audio data through the atrous convolutional network to obtain audio time series features;

[0036] An activation module, used for performing activation processing on the audio time series feature through the dilated convolutional network to obtain an initial audio feature vector;

[0037] A feature fusion module, used for performing feature fusion on the multiple initial audio feature vectors to obtain a fused audio feature vector;

[0038] A second feature extraction module is used to extract features from the fused audio feature vector through the convolution classification network to obtain a target audio feature vector;

[0039] The recognition module is used to predict and process the target audio feature vector through the convolution classification network to obtain the target identity label of the target singing object, wherein the target identity label is used to characterize the identity of the target singing object.

[0040] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory, and the program, when executed by the processor, implements the method described in the first aspect above.

[0041] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium used for computer-readable storage, and the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in the first aspect above.

[0042] The singing object recognition method, singing object recognition device, electronic device and storage medium proposed in the present application obtain the target audio data of the target singing object; input the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network; extract the features of the target audio data through the dilated convolutional network to obtain the audio time series features, and activate the audio time series features through the dilated convolutional network to obtain the initial audio feature vector, which can expand the receptive field of the target audio data through the dilated convolutional network to improve the recognition accuracy. Furthermore, feature fusion is performed on multiple initial audio feature vectors to obtain a fused audio feature vector; feature extraction is performed on the fused audio feature vector through the convolutional classification network to obtain a target audio feature vector, which can better capture the frequency domain spatial features of the fused audio feature vector to obtain the target audio feature vector. Finally, the target audio feature vector is predicted and processed through the convolutional classification network to obtain the target identity label of the target singing object, where the target identity label is used to characterize the identity of the target singing object, thereby conveniently determining the identity of the target singing object, which can effectively solve the problem of difficulty in identifying the identity of the singing object caused by the increase in the number of singing objects in the metaverse, and improve the recognition accuracy of the singing object. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of a singing object identification method provided in an embodiment of the present application;

[0044] Figure 2 yes Figure 1 Flow chart of step S101 in FIG.

[0045] Figure 3 yes Figure 1 Flow chart of step S104 in FIG.

[0046] Figure 4 yes Figure 1Flow chart of step S105 in FIG.

[0047] Figure 5 yes Figure 1 Flow chart of step S106 in FIG.

[0048] Figure 6 yes Figure 1 Flow chart of step S107 in FIG.

[0049] Figure 7 yes Figure 6 Flowchart of step S602 in FIG.

[0050] Figure 8 is a structural diagram of a singing object recognition device provided in an embodiment of the present application;

[0051] Fig. 9 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0055] First, some nouns involved in this application are analyzed:

[0056] Artificial intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a similar way to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing and expert systems. AI can simulate the information process of human consciousness and thinking. AI is also a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0057] Natural language processing (NLP): NLP uses computers to process, understand and apply human languages ​​(such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary subject between computer science and linguistics. It is often referred to as computational linguistics. Natural language processing includes grammatical analysis, semantic analysis, and text understanding. Natural language processing is often used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computing.

[0058] Information Extraction (NER): A text processing technology that extracts specified types of entity, relationship, event and other factual information from natural language text and forms structured data output. Information extraction is a technology that extracts specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and chapters. Text information is composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or a combination of these specific units. Extracting noun phrases, names, place names, etc. from text data are all text information extraction. Of course, the information extracted by text information extraction technology can be various types of information.

[0059] Metaverse: It is a virtual world that is linked and created by scientific and technological means, mapped and interacted with the real world, and has a digital living space with a new social system. The metaverse is essentially a virtualization and digitization process of the real world, which requires a lot of transformation of content production, economic system, user experience, and content in the physical world. However, the development of the metaverse is gradual. It is supported by shared infrastructure, standards, and protocols, and is finally formed by the continuous integration and evolution of many tools and platforms. It provides an immersive experience based on extended reality technology, generates a mirror image of the real world based on digital twin technology, and builds an economic system based on blockchain technology. It closely integrates the virtual world with the real world in the economic system, social system, and identity system, and allows each user to produce content and edit the world.

[0060] Dilated Convolution: Dilated convolution is also called dilated convolution or expanded convolution. Simply put, it is the process of adding some spaces (zeros) between the elements of the convolution kernel to expand the convolution kernel.

[0061] Receptive Field: It is the size of the area where the pixels on the feature map output by each layer of the convolutional neural network are mapped on the input image.

[0062] Fourier transform: It means that a function that satisfies certain conditions can be expressed as a linear combination of trigonometric functions (sine and / or cosine functions) or their integrals. In different research fields, Fourier transform has many different variants, such as continuous Fourier transform and discrete Fourier transform.

[0063] Mel-Frequency Cipstal Coefficients (MFCC): A set of key coefficients used to create a Mel-Cepstrum. From a fragment of a music signal, a set of cepstrum that is sufficient to represent the music signal can be obtained, and the Mel-Cepstrum coefficients are the cepstrum (i.e. the spectrum of the spectrum) derived from this cepstrum. Different from the general cepstrum, the biggest feature of the Mel-Cepstrum is that the frequency bands on the Mel-Cepstrum are evenly distributed on the Mel scale. In other words, compared to the generally seen, linear cepstrum representation method, such frequency bands are closer to the nonlinear human auditory system (Audio System). For example: In audio compression technology, Mel-Cepstrum is often used for processing.

[0064] Activation Function: It is a function that runs on the neurons of the artificial neural network and is responsible for mapping the input of the neuron to the output.

[0065] Hyperbolic tangent function: The hyperbolic tangent function is a type of hyperbolic function. In mathematical language, the hyperbolic tangent function is generally written as tanh, and can also be abbreviated as th. Like trigonometric functions, hyperbolic functions are also divided into six types: hyperbolic sine, hyperbolic cosine, hyperbolic tangent, hyperbolic cotangent, hyperbolic secant, and hyperbolic cosecant. The hyperbolic tangent function is one of them. Similar to the tangent function, the hyperbolic tangent function is equal to the ratio of the hyperbolic sine to the hyperbolic cosine, that is, tanh(x) = sinh(x) / cosh(x).

[0066] Dimensionality reduction: It is an operation that converts a single image into a data set in a high-dimensional space by making the single image data high-dimensional.

[0067] Decoder: It converts the previously generated fixed vector into an output sequence; the input sequence can be text, voice, image, or video; the output sequence can be text or image.

[0068] Pooling: It is essentially a kind of sampling. It selects a certain method to reduce the dimension and compress the input feature map to speed up the operation. The most common pooling process is Max Pooling.

[0069] Softmax function: The Softmax function is a normalized exponential function that can "compress" a K-dimensional vector z containing any real number into another K-dimensional real vector σ(z) so that the range of each element is between (0,1) and the sum of all elements is 1. This function is often used in multi-classification problems.

[0070] Residual connection: It allows the output of a previous layer to be used as the input of a subsequent layer, thus effectively creating a shortcut in the sequence network.

[0071] With the development of Metaverse technology, all aspects of daily life can be extended to a virtual and real world through Metaverse. As the number of virtual singers in Metaverse increases, the number of singers in Metaverse will continue to accumulate over time, far exceeding the number of singers in the real world. As the number of singers in Metaverse increases, commonly used recognition methods often have difficulty in accurately identifying the identities of singers. Therefore, how to improve the recognition accuracy of singers has become a technical problem that needs to be solved urgently.

[0072] Therefore, in response to the surge in the number of singing objects in the metaverse and the problem of identifying virtual singing objects and real singing objects, an embodiment of the present application provides a singing object identification method, a singing object identification device, an electronic device and a storage medium, aiming to improve the identification accuracy of singing objects.

[0073] The singing object identification method, singing object identification device, electronic device and storage medium provided in the embodiments of the present application are specifically explained through the following embodiments. First, the singing object identification method in the embodiments of the present application is described.

[0074] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0075] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0076] The singing object recognition method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The singing object recognition method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the singing object recognition method, etc., but is not limited to the above forms.

[0077] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0078] Figure 1 is an optional flow chart of the singing object recognition method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S107.

[0079] Step S101, obtaining target audio data of a target singing object;

[0080] Step S102, inputting the target audio data into a preset person recognition model; wherein the person recognition model includes a dilated convolutional network and a convolutional classification network;

[0081] Step S103, extracting features of the target audio data through a dilated convolutional network to obtain audio time series features;

[0082] Step S104, activating the audio time series features through a dilated convolutional network to obtain an initial audio feature vector;

[0083] Step S105, performing feature fusion on multiple initial audio feature vectors to obtain a fused audio feature vector;

[0084] Step S106, extracting features from the fused audio feature vector through a convolutional classification network to obtain a target audio feature vector;

[0085] Step S107, predicting the target audio feature vector through a convolution classification network to obtain a target identity label of the target singing object, wherein the target identity label is used to characterize the identity of the target singing object.

[0086] Steps S101 to S107 shown in the embodiment of the present application are as follows: obtaining target audio data of the target singing object; inputting the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network; extracting features of the target audio data through the dilated convolutional network to obtain audio timing features, and activating the audio timing features through the dilated convolutional network to obtain an initial audio feature vector, which can expand the receptive field of the target audio data through the dilated convolutional network to improve recognition accuracy. Furthermore, feature fusion is performed on multiple initial audio feature vectors to obtain a fused audio feature vector; feature extraction is performed on the fused audio feature vector through the convolutional classification network to obtain a target audio feature vector, which can better capture the frequency domain spatial features of the fused audio feature vector to obtain a target audio feature vector. Finally, the target audio feature vector is predicted and processed through the convolutional classification network to obtain the target identity label of the target singing object, where the target identity label is used to characterize the identity of the target singing object, thereby conveniently determining the identity of the target singing object, which can effectively solve the problem of difficulty in identifying the identity of the singing object caused by the increase in the number of singing objects in the metaverse, and improve the recognition accuracy of the singing object.

[0087] See also Figure 2 In some embodiments, step S101 may include but is not limited to steps S201 to S202:

[0088] Step S201, obtaining the original audio data of the target singing object;

[0089] Step S202, filtering and processing the original audio data according to a preset audio length to obtain target audio data.

[0090] In step S201 of some embodiments, a web crawler may be written to crawl data in a targeted manner after setting a data source, to obtain the original audio data of the target singing object, wherein the data source may be various types of network platforms, social media, or certain specific audio databases, etc., and the original audio data may be the music material, speech report, chat dialogue, etc. of the target singing object. The original audio data may also be obtained by other means, not limited thereto.

[0091] In step S202 of some embodiments, the preset audio length can be set according to actual business needs without restriction. For example, the preset audio length is 5 minutes. According to the preset audio length, the audio data in the original audio data with an audio length less than or equal to 5 minutes is screened and processed, and these original audio data that meet the requirements are used as target audio data.

[0092] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0093] In step S102 of some embodiments, the target audio data is input into a preset character recognition model; wherein the target audio data is an audio waveform file, and the character recognition model includes a dilated convolutional network and a convolutional classification network, wherein the dilated convolutional network is mainly used to perform dimension expansion on the input target audio data, and the convolutional classification network is mainly used to perform label prediction on the dimensionally expanded target audio data to obtain an identity label of a target singing object corresponding to the target audio data, and by inputting the audio waveform file as the target audio data into the preset character recognition model, completely end-to-end recognition of the singing object can be achieved, and the process of extracting spectral features of the target audio data is omitted, which can effectively reduce the influence of the prior knowledge of short-time Fourier transform on the target audio data during the frame processing, thereby improving the recognition accuracy.

[0094] In step S103 of some embodiments, feature extraction is performed on the target audio data through the atrous convolution layer of the atrous convolution network, the receptive field of the target audio data is expanded exponentially, and one-dimensional atrous convolution processing is performed on the target audio data to obtain audio timing features, wherein the receptive field mainly refers to the mapping area range of a pixel on the audio feature map of the input target audio data corresponding to the input space. In this way, the acceptable dimension of the target audio data can be greatly expanded.

[0095] See also Figure 3 In some embodiments, step S104 may include but is not limited to steps S301 to S303:

[0096] Step S301, activating the audio time series features through a first activation function of a dilated convolutional network to obtain a first activated feature vector;

[0097] Step S302, activating the audio time series features through the second activation function of the dilated convolutional network to obtain a second activated feature vector;

[0098] Step S303: Perform a dot multiplication process on the first activated feature vector and the second activated feature vector according to a preset weight parameter to obtain an initial audio feature vector.

[0099] In step S301 of some embodiments, the first activation function is the tanh function, which is the hyperbolic tangent function, equal to the hyperbolic cosine divided by the hyperbolic sine. The tanh function is an odd function. The element value of the input audio timing feature can be transformed to between -1 and 1 through the tanh function to obtain the first activation feature vector.

[0100] In step S302 of some embodiments, the second activation function is a Sigmoid function, and its function form is x is the audio time series feature. The audio time series feature is activated by the Sigmoid function, and the activated vector is mapped to [0, 1] to obtain the second activated feature vector.

[0101] In step S303 of some embodiments, the preset weight parameter can be set according to actual business needs without limitation. The process of performing dot multiplication on the first activated feature vector and the second activated feature vector according to the preset weight parameter to obtain the initial audio feature vector can be expressed as:

[0102] Z=tanh(W1,x)·sign(W2,x);

[0103] Among them, Z is the initial audio feature vector, W1 is the weight parameter of the first activation function, W2 is the weight parameter of the second activation function, and x is the input audio time series feature.

[0104] It should be noted that the above-mentioned dilated convolution network includes multiple dilated convolution modules, each of which includes a dilated convolution layer, a first activation function and a second activation function, and each dilated convolution module is connected by residual connection. For example, the dilated convolution network includes two dilated convolution modules, the input of the first dilated convolution module is the target audio data, and the input of the second dilated convolution module is the result of splicing the output of the previous dilated convolution module with the target audio data. In order to better expand the receptive field of the target audio data, the dilated convolution network in the embodiment of the present application adopts K layers of stacked dilated convolution modules, where K is equal to 9.

[0105] The atrous convolution network of the above-mentioned steps S103 and S104 can extract the timing features of the target audio data (waveform file) in the time domain, and the receptive field of the target audio data can be exponentially increased by stacking K layers of atrous convolution modules, so that the singing object recognition method of the present embodiment can support waveform input, simplify the recognition process, and improve the efficiency of singing object recognition.

[0106] See also Figure 4 In some embodiments, the convolution classification network includes a two-dimensional convolution layer and a pooling layer, and step S105 may include but is not limited to steps S401 to S402:

[0107] Step S401, obtaining a preset splicing order;

[0108] Step S402, performing vector addition on a plurality of initial audio feature vectors according to a concatenation order to obtain a fused audio feature vector.

[0109] In step S401 of some embodiments, the preset splicing order is the stacking order of the dilated convolution modules of the dilated convolution network described above.

[0110] In step S402 of some embodiments, the initial audio feature vectors output by each dilated convolution module are jump-connected in turn according to the stacking order of the dilated convolution modules. The jump connection method here can be vector addition or vector connection, etc. For example, vector addition is performed on multiple initial audio feature vectors according to the stacking order of the dilated convolution modules to obtain a fused audio feature vector.

[0111] See also Figure 5 In some embodiments, the convolution classification network includes a two-dimensional convolution layer and a pooling layer, and step S106 may include but is not limited to steps S501 to S502:

[0112] Step S501, extracting features from the fused audio feature vector through a two-dimensional convolutional layer to obtain an intermediate audio feature vector;

[0113] Step S502: Perform dimensionality reduction processing on the intermediate audio feature vector through a pooling layer to obtain a target audio feature vector.

[0114] In step S501 of some embodiments, feature extraction is performed on the fused audio feature vector through a two-dimensional convolution layer to capture the frequency domain spatial features of the fused audio feature vector and obtain an intermediate audio feature vector.

[0115] In step S502 of some embodiments, the intermediate audio feature vector is subjected to dimensionality reduction processing through a pooling layer so that the intermediate audio feature vector is in a preset feature dimensional space, thereby obtaining a target audio feature vector.

[0116] It should be noted that the above-mentioned two-dimensional convolutional layer and pooling layer may include one or more. When multiple two-dimensional convolutional layers and multiple pooling layers are involved, the two-dimensional convolutional layers and the pooling layers are alternately connected. For example, when 2 two-dimensional convolutional layers and 2 pooling layers are involved, the input of the first two-dimensional convolutional layer is the fused audio feature vector, the output of the first two-dimensional convolutional layer is used as the input of the first pooling layer, the output of the first pooling layer is used as the input of the second two-dimensional convolutional layer, the output of the second two-dimensional convolutional layer is used as the input of the second pooling layer, and so on. In the embodiment of the present application, 4 groups of two-dimensional convolutional layers and pooling layers may be included, that is, 4 two-dimensional convolutional layers and 4 pooling layers.

[0117] See also Figure 6 In some embodiments, the convolution classification network includes a flattening layer and a fully connected layer, and step S107 includes but is not limited to steps S601 to S602:

[0118] Step S601, stretching the target audio feature vector through a flattening layer to obtain a variable-dimensional audio feature vector;

[0119] Step S602, performing label prediction processing on the variable-dimensional audio feature vector through a fully connected layer to obtain a target identity label of the target singing object.

[0120] In step S601 of some embodiments, the target audio feature vector is stretched through a flattening layer to stretch the two-dimensional target audio feature vector into a one-dimensional feature vector to obtain a variable-dimensional audio feature vector.

[0121] In step S602 of some embodiments, the label probability of the variable-dimensional audio feature vector is calculated by the classification function of the fully connected layer to obtain a label probability vector corresponding to each preset identity label, and the preset identity label corresponding to the label probability vector with the largest value is selected as the target identity label of the target singing object. The identity of the target singing object is determined according to the target identity label, and the identity can represent whether the target singing object is a virtual character or a real person.

[0122] See also Figure 7 In some embodiments, step S602 may include but is not limited to steps S701 to S703:

[0123] Step S701, calculating the label probability of the variable-dimensional audio feature vector using the classification function of the fully connected layer and the preset identity label to obtain a label probability vector corresponding to each preset identity label;

[0124] Step S702, selecting a preset identity tag corresponding to the tag probability vector with the largest value to obtain a candidate identity tag;

[0125] Step S703: Obtain a target identity tag according to the candidate identity tags.

[0126] In step S701 of some embodiments, the classification function may be a function such as a softmax function, and the preset identity tags may be extracted from different data sources, for example, basic information of various characters, including identity information, personal information, and related audio and video data, etc., may be obtained from online media and social platforms. Taking the softmax function as an example, the probability distribution of the variable-dimensional audio feature vector on each preset identity tag may be created through the softmax function, and the probability distribution reflects the probability of the variable-dimensional audio feature vector belonging to each preset identity tag, thereby obtaining a tag probability vector corresponding to each preset identity tag.

[0127] In steps S702 and S703 of some embodiments, the size of the label probability vector can intuitively reflect the possibility that the variable-dimensional audio feature vector belongs to each preset identity tag. The larger the value of the label probability vector, the higher the degree of matching between the variable-dimensional audio feature vector and the corresponding preset identity tag, indicating that the variable-dimensional audio feature vector is more likely to come from the person corresponding to this preset identity tag. Therefore, the preset identity tag corresponding to the label probability vector with the largest value is selected to obtain one or more candidate identity tags, and then an identity tag is selected from the candidate identity tags as the target identity tag, thereby representing the identity of the target singing object through the target identity tag.

[0128] The above steps S701 to S703 can conveniently quantify the possibility that the variable-dimensional audio feature vector belongs to each preset identity label through the classification function to obtain the label probability vector, and then select the most appropriate preset identity label as the target identity label according to the size of the label probability vector, so as to confirm the identity of the target singer according to the target identity label, thereby improving the recognition accuracy of the singer.

[0129] The singing object recognition method of the embodiment of the present application obtains the target audio data of the target singing object; inputs the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network; extracts features of the target audio data through the dilated convolutional network to obtain audio time series features, and activates the audio time series features through the dilated convolutional network to obtain an initial audio feature vector, which can expand the receptive field of the target audio data through the dilated convolutional network to improve recognition accuracy. Furthermore, feature fusion is performed on multiple initial audio feature vectors to obtain a fused audio feature vector; feature extraction is performed on the fused audio feature vector through the convolutional classification network to obtain a target audio feature vector, which can better capture the frequency domain spatial features of the fused audio feature vector to obtain a target audio feature vector. Finally, the target audio feature vector is predicted and processed through the convolutional classification network, which can conveniently quantify the possibility that the target audio feature vector belongs to each preset identity label to obtain the label probability vector, and then select the most appropriate preset identity label as the target identity label according to the size of the label probability vector, so as to confirm the identity of the target singing object according to the target identity label, thereby conveniently determining the identity of the target singing object, which can effectively solve the problem of difficulty in identifying the singing object caused by the increase in the number of singing objects in the metaverse, and improve the recognition accuracy of the singing object.

[0130] See also Figure 8 The embodiment of the present application further provides a singing object recognition device, which can implement the above singing object recognition method, and the device includes:

[0131] Data acquisition module 801, used to acquire target audio data of a target singing object;

[0132] An input module 802 is used to input the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network;

[0133] A first feature extraction module 803 is used to extract features of target audio data through a dilated convolutional network to obtain audio time series features;

[0134] An activation module 804 is used to activate the audio time series features through a dilated convolutional network to obtain an initial audio feature vector;

[0135] A feature fusion module 805 is used to perform feature fusion on multiple initial audio feature vectors to obtain a fused audio feature vector;

[0136] A second feature extraction module 806 is used to extract features from the fused audio feature vector through a convolution classification network to obtain a target audio feature vector;

[0137] The identification module 807 is used to predict and process the target audio feature vector through a convolution classification network to obtain a target identity label of the target singing object, wherein the target identity label is used to characterize the identity of the target singing object.

[0138] In some embodiments, the data acquisition module 801 includes:

[0139] A data acquisition unit, used to acquire original audio data of a target singing object;

[0140] The screening unit is used to screen the original audio data according to a preset audio length to obtain target audio data.

[0141] In some embodiments, the activation module 804 includes:

[0142] A first activation unit, configured to activate the audio time series feature through a first activation function of the dilated convolutional network to obtain a first activated feature vector;

[0143] A second activation unit is used to activate the audio time series feature through a second activation function of the dilated convolutional network to obtain a second activated feature vector;

[0144] The dot product unit is used to perform dot product processing on the first activated feature vector and the second activated feature vector according to a preset weight parameter to obtain an initial audio feature vector.

[0145] In some embodiments, the feature fusion module 805 includes:

[0146] A sequence acquisition unit, used to acquire a preset splicing sequence;

[0147] The vector addition unit is used to perform vector addition on multiple initial audio feature vectors according to a splicing order to obtain a fused audio feature vector.

[0148] In some embodiments, the convolution classification network includes a two-dimensional convolution layer and a pooling layer, and the second feature extraction module 806 includes:

[0149] An extraction unit, used for performing feature extraction on the fused audio feature vector through a two-dimensional convolution layer to obtain an intermediate audio feature vector;

[0150] The dimension reduction unit is used to perform dimension reduction processing on the intermediate audio feature vector through a pooling layer to obtain a target audio feature vector.

[0151] In some embodiments, the convolution classification network includes a flattening layer and a fully connected layer, and the recognition module 807 includes:

[0152] A stretching unit, used for stretching the target audio feature vector through a flattening layer to obtain a variable-dimensional audio feature vector;

[0153] The prediction unit is used to perform label prediction processing on the variable-dimensional audio feature vector through a fully connected layer to obtain a target identity label of the target singing object.

[0154] In some embodiments, the prediction unit includes:

[0155] A probability calculation subunit, used to calculate the label probability of the variable-dimensional audio feature vector through the classification function of the fully connected layer and the preset identity label, and obtain a label probability vector corresponding to each preset identity label;

[0156] A selection subunit is used to select a preset identity label corresponding to the label probability vector with the largest value to obtain a candidate identity label;

[0157] The label determination subunit is used to obtain a target identity label according to the candidate identity labels.

[0158] The specific implementation of the singing object identification device is basically the same as the specific implementation of the above-mentioned singing object identification method, and will not be repeated here.

[0159] The embodiment of the present application also provides an electronic device, the electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory, and the program is executed by the processor to realize the above-mentioned singing object recognition method. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0160] See also Fig. 9 , Fig. 9 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0161] The processor 901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0162] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the singing object recognition method of the embodiment of this application;

[0163] Input / output interface 903, used to implement information input and output;

[0164] Communication interface 904, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0165] A bus 905 that transmits information between various components of the device (e.g., the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0166] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0167] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium used for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned singing object recognition method.

[0168] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0169] The singing object recognition method, singing object recognition device, electronic device and storage medium provided in the embodiments of the present application obtain the target audio data of the target singing object; input the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network; extract the features of the target audio data through the dilated convolutional network to obtain the audio time series features, and activate the audio time series features through the dilated convolutional network to obtain the initial audio feature vector, which can expand the receptive field of the target audio data through the dilated convolutional network to improve the recognition accuracy. Furthermore, feature fusion is performed on multiple initial audio feature vectors to obtain a fused audio feature vector; feature extraction is performed on the fused audio feature vector through the convolutional classification network to obtain a target audio feature vector, which can better capture the frequency domain spatial features of the fused audio feature vector to obtain the target audio feature vector. Finally, the target audio feature vector is predicted and processed through the convolutional classification network, which can conveniently quantify the possibility that the target audio feature vector belongs to each preset identity label to obtain the label probability vector, and then select the most appropriate preset identity label as the target identity label according to the size of the label probability vector, so as to confirm the identity of the target singing object according to the target identity label, thereby conveniently determining the identity of the target singing object, which can effectively solve the problem of difficulty in identifying the singing object caused by the increase in the number of singing objects in the metaverse, and improve the recognition accuracy of the singing object.

[0170] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0171] It can be understood by those skilled in the art that Figure 1-7 The technical solutions shown in the figure do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or a combination of certain steps, or different steps.

[0172] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0173] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0174] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0175] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0176] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0177] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0179] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0180] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A singing object recognition method, characterized in that: The method comprises: Acquire target audio data of a target singing object; Inputting the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network; Extracting features of the target audio data through the atrous convolutional network to obtain audio time series features; activating the audio time series features through the atrous convolutional network to obtain an initial audio feature vector; Performing feature fusion on the multiple initial audio feature vectors to obtain a fused audio feature vector; performing feature extraction on the fused audio feature vector through the convolution classification network to obtain a target audio feature vector; The target audio feature vector is predicted and processed by the convolution classification network to obtain a target identity label of the target singing object, wherein the target identity label is used to characterize the identity of the target singing object; The convolution classification network includes a two-dimensional convolution layer, a pooling layer, a flattening layer, and a fully connected layer. The convolution classification network is used to extract features from the fused audio feature vector to obtain a target audio feature vector, including: Performing feature extraction on the fused audio feature vector through the two-dimensional convolution layer to obtain an intermediate audio feature vector; performing dimensionality reduction processing on the intermediate audio feature vector through the pooling layer to obtain the target audio feature vector; The step of predicting the target audio feature vector by the convolution classification network to obtain the target identity label of the target singing object includes: The target audio feature vector is stretched through the flattening layer to obtain a variable-dimensional audio feature vector; the label probability of the variable-dimensional audio feature vector is calculated through the classification function of the fully connected layer and the preset identity label to obtain a label probability vector corresponding to each preset identity label; the preset identity label corresponding to the label probability vector with the largest value is selected to obtain a candidate identity label; and the target identity label is obtained based on the candidate identity label.

2. The singing object recognition method according to claim 1, characterized in that: The step of activating the audio time series features through the atrous convolutional network to obtain an initial audio feature vector comprises: Activate the audio time series feature through the first activation function of the dilated convolutional network to obtain a first activated feature vector; Activate the audio time series feature through the second activation function of the dilated convolutional network to obtain a second activated feature vector; The first activated feature vector and the second activated feature vector are subjected to a dot multiplication process according to a preset weight parameter to obtain the initial audio feature vector.

3. The singing object recognition method according to any one of claims 1 to 2, characterized in that: The step of fusing the features of the plurality of initial audio feature vectors to obtain a fused audio feature vector comprises: Get the preset splicing order; Performing vector addition on a plurality of the initial audio feature vectors according to the concatenation order to obtain the fused audio feature vector.

4. The singing object recognition method according to any one of claims 1 to 2, characterized in that: The step of obtaining target audio data of a target singing object comprises: Acquire original audio data of the target singing object; The original audio data is filtered and processed according to a preset audio length to obtain the target audio data.

5. A singing object recognition device, characterized in that: The device comprises: A data acquisition module, used to acquire target audio data of a target singing object; An input module, used to input the target audio data into a preset character recognition model; wherein the character recognition model includes a dilated convolutional network and a convolutional classification network; A first feature extraction module, used to extract features of the target audio data through the dilated convolutional network to obtain audio time series features; An activation module, used for performing activation processing on the audio time series feature through the dilated convolutional network to obtain an initial audio feature vector; A feature fusion module, used for performing feature fusion on the multiple initial audio feature vectors to obtain a fused audio feature vector; A second feature extraction module is used to extract features from the fused audio feature vector through the convolution classification network to obtain a target audio feature vector; An identification module, used for predicting and processing the target audio feature vector through the convolution classification network to obtain a target identity tag of the target singing object, wherein the target identity tag is used to characterize the identity of the target singing object; The convolution classification network includes a two-dimensional convolution layer, a pooling layer, a flattening layer, and a fully connected layer. The convolution classification network is used to extract features from the fused audio feature vector to obtain a target audio feature vector, including: Performing feature extraction on the fused audio feature vector through the two-dimensional convolution layer to obtain an intermediate audio feature vector; performing dimensionality reduction processing on the intermediate audio feature vector through the pooling layer to obtain the target audio feature vector; The step of predicting the target audio feature vector by the convolution classification network to obtain the target identity label of the target singing object includes: The target audio feature vector is stretched through the flattening layer to obtain a variable-dimensional audio feature vector; the label probability of the variable-dimensional audio feature vector is calculated through the classification function of the fully connected layer and the preset identity label to obtain a label probability vector corresponding to each preset identity label; the preset identity label corresponding to the label probability vector with the largest value is selected to obtain a candidate identity label; and the target identity label is obtained based on the candidate identity label.

6. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the singing object identification method as described in any one of claims 1 to 4 are realized.

7. A storage medium, the storage medium being a computer-readable storage medium, used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the singing object recognition method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Voiceprint recognition method and device, and terminal equipment

    CN113223536A

  • Audio source separation and audio dubbing

    WO2021239285A1