Audio stream identification method and device, electronic equipment and computer storage medium
By extracting the audio stream in electronic devices and inputting the trained scene recognition model for holographic scene recognition, the problem of insufficient accuracy in holographic scene type determination under holographic audio is solved, and higher accuracy of holographic audio effects and holographic scene type recognition is achieved.
Patent Information
- Application Number
- CN202311452191.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-06
AI Technical Summary
The accuracy of determining the holographic scene type under existing holographic audio is poor, especially because there are many types of audio streams that the application can play, resulting in insufficient accuracy of determining the holographic scene type based on the package name.
By obtaining the audio stream of the electronic device, feature extraction is performed, and these features are input into the trained scene recognition model for holographic scene recognition to determine the holographic scene type of the audio stream. At the same time, the scene recognition model is trained using the sample training set, and the trained model is sent to the electronic device.
The accuracy of determining the holographic scene type under holographic audio is improved, thereby improving the sound effect of holographic audio, making the holographic scene type recognition of the audio stream more accurate and suitable.
Smart Images

Figure CN119943091A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a technology for playing spatial audio in holographic audio, and more particularly to an audio stream recognition method, device, electronic device and computer storage medium. Background Art
[0002] At present, the mobile terminal sets the holographic scene type of the application through the package name (Package name) of the application where the audio stream type (Stream) is located; for example, when a music-type application (Application, APP) plays sound, the holographic scene type is determined according to the package name of the APP, and the spatial audio of the APP is played based on the determined holographic scene type.
[0003] However, when using the above method to determine the holographic scene type, since there are many types of audio streams that can be played by the application, determining the holographic scene type based on the software package name of the application results in poor accuracy of the result. It can be seen that the existing method of determining the holographic scene type under holographic audio has a technical problem of low accuracy. Summary of the invention
[0004] The embodiments of the present application provide an audio stream recognition method, device, electronic device and computer storage medium, which can improve the accuracy of determining the type of holographic scene under holographic audio.
[0005] The technical solution of this application is implemented as follows:
[0006] The present application provides an audio stream recognition method, which is applied in an electronic device and includes:
[0007] Acquiring an audio stream of the electronic device;
[0008] Extracting features from the audio stream to obtain features of the audio stream;
[0009] The features of the audio stream are input into the trained scene recognition model to perform holographic scene recognition, and the holographic scene type of the audio stream is obtained.
[0010] The embodiment of the present application provides a method for identifying an audio stream, which is applied to a cloud server and includes:
[0011] Acquire a sample training set; wherein the sample training set includes: a collected audio stream and a holographic scene type of the collected audio stream;
[0012] Using the sample training set to train the scene recognition model to obtain a trained scene recognition model;
[0013] The trained scene recognition model is sent to the electronic device.
[0014] The embodiment of the present application provides an audio stream recognition device, which is arranged in an electronic device and includes:
[0015] A first acquisition module, used to acquire an audio stream of the electronic device;
[0016] An extraction module, used to extract features from the audio stream to obtain features of the audio stream;
[0017] The recognition module is used to input the features of the audio stream into the trained scene recognition model to perform holographic scene recognition and obtain the holographic scene type of the audio stream.
[0018] The embodiment of the present application provides an audio stream recognition device, which is arranged in a cloud server and includes:
[0019] A second acquisition module is used to acquire a sample training set; wherein the sample training set includes: a collected audio stream and a holographic scene type of the collected audio stream;
[0020] A training module, used to train the scene recognition model using the sample training set to obtain a trained scene recognition model;
[0021] The sending module is used to send the trained scene recognition model to the electronic device.
[0022] An embodiment of the present application provides an electronic device, including:
[0023] A processor and a storage medium storing instructions executable by the processor, wherein the storage medium relies on the processor to perform operations through a communication bus, and when the instructions are executed by the processor, the audio stream recognition method described in one or more of the above embodiments is executed.
[0024] The present application provides a cloud server, including:
[0025] A processor and a storage medium storing instructions executable by the processor, wherein the storage medium relies on the processor to perform operations through a communication bus, and when the instructions are executed by the processor, the audio stream recognition method described in one or more of the above embodiments is executed.
[0026] An embodiment of the present application provides a computer storage medium storing executable instructions. When the executable instructions are executed by one or more processors, the processors execute the audio stream recognition method as described in one or more embodiments.
[0027] The embodiments of the present application provide an audio stream recognition method, device, electronic device and computer storage medium. The method is applied to an electronic device, comprising: acquiring an audio stream of the electronic device, performing feature extraction on the audio stream to obtain features of the audio stream, inputting the features of the audio stream into a trained scene recognition model for holographic scene recognition, and obtaining the holographic scene type of the audio stream; that is, in the embodiments of the present application, the holographic scene type of the audio stream can be obtained by inputting the audio stream of the electronic device into a trained scene recognition model for holographic scene recognition. In this way, the holographic scene type of the audio stream is identified using the trained scene recognition model, so that the determined holographic scene type is more suitable for holographic audio, thereby improving the accuracy of determining the holographic scene type under holographic audio, and further improving the sound effect of the holographic audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A schematic diagram of an optional audio stream recognition method provided in an embodiment of the present application Figure 1 ;
[0029] Figure 2 A schematic diagram of an optional audio stream recognition method provided in an embodiment of the present application Figure 2 ;
[0030] Figure 3 A flowchart of Example 1 of an optional audio stream recognition method provided in an embodiment of the present application;
[0031] Figure 4 A flowchart of Example 2 of an optional audio stream recognition method provided in an embodiment of the present application;
[0032] Figure 5 A schematic diagram of the structure of an optional audio stream recognition device provided in an embodiment of the present application Figure 1 ;
[0033] Figure 6 A schematic diagram of the structure of an optional audio stream recognition device provided in an embodiment of the present application Figure 2 ;
[0034] Figure 7 A schematic diagram of the structure of an optional electronic device provided in an embodiment of the present application;
[0035] Figure 8 A schematic diagram of the structure of an optional cloud server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0037] In the related art, in determining the holographic scene type of the audio stream, for example, a common music APP has both live broadcast, video, and music. For another example, a commonly used communication APP can watch videos on a social platform, watch videos on the communication interface with friends, listen to voice messages, make voice calls, and use navigation functions. At this time, it is often inappropriate to classify the holographic scene type according to the type of application.
[0038] Secondly, it is often inappropriate to judge the type of holographic scene based on the type of audio stream played by the application, because the type of audio stream is set by the application when it is played, and its accuracy is poor.
[0039] Finally, if the holographic scene type is identified by the software package name of the application, the identification of the holographic scene type can only be applied to known software packages. The limited list does not include unknown applications, and the accuracy is poor.
[0040] In addition, holographic audio (also known as holographic sound effect) is a technology that simulates the way of auditory perception, that is, 3D sound effect technology. It uses complex algorithms and processing methods to transform the sound from a single plane into three-dimensional, so that the audience can clearly perceive the position and distance of the sound in space; then, for the control of applications under holographic audio.
[0041] In order to solve the technical problem that the method of determining the type of holographic scene under holographic audio in the related art has low accuracy, the embodiment of the present application provides an audio stream recognition method, which is applied to an electronic device. Figure 1 A flowchart of an optional audio stream recognition method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the audio stream identification method may include:
[0042] S101: Acquire an audio stream of an electronic device;
[0043] In S101, the electronic device first obtains an audio stream of the electronic device, wherein the number of the audio streams may be one, two, or more than two, and this embodiment of the present application does not specifically limit this.
[0044] In addition, the audio stream of the above-mentioned electronic device can be one audio stream of an application, or two or more audio streams of an application, or audio streams of two applications, or audio streams of more than two applications. Here, the embodiment of the present application does not make any specific limitations on this.
[0045] Specifically, the application program writes the audio stream into the memory of the electronic device through the interface of AudioTrack, so that the electronic device obtains the audio stream of the electronic device.
[0046] S102: extracting features from the audio stream to obtain features of the audio stream;
[0047] After S101, the electronic device performs feature extraction on the audio stream, wherein, before the feature extraction, the audio stream may be sequentially processed by data cleaning, file format alignment, data enhancement, and spectrum conversion to obtain a processed audio stream. It should be noted that the above processing is not limited to this.
[0048] After the above processing, the processed audio stream is subjected to feature extraction by a front-end processing algorithm (FiliterBank, Fbank) or a Mel-frequency cepstral coefficient (Mel-frequency cepstral coefficients, MFCC) algorithm, so that the features of the audio stream can be obtained.
[0049] It should be noted that the above-mentioned feature extraction is performed on audio streams one by one. For example, when there is one audio stream, feature extraction is performed on the audio stream to obtain the features of this audio stream. When there are two or more audio streams, feature extraction is performed on each audio stream to obtain the features of each audio stream.
[0050] In this way, the characteristics of the audio stream can be obtained to determine the holographic scene type of the audio stream.
[0051] S103: Input the features of the audio stream into the trained scene recognition model to perform holographic scene recognition, and obtain the holographic scene type of the audio stream.
[0052] After determining the characteristics of the audio stream through S102, since the electronic device stores a trained scene recognition model, the electronic device can input the characteristics of the audio stream into the trained scene recognition model for holographic scene recognition, thereby obtaining the holographic scene type of the audio stream.
[0053] Among them, the above-mentioned trained scene recognition model can be obtained by training the electronic device based on a sample data set of the model, or it can be obtained by training the cloud server based on a sample data set of the model. After obtaining the trained scene recognition model, the trained scene recognition model is sent to the electronic device, so that the trained scene recognition model is stored in the electronic device; here, the embodiment of the present application does not make specific limitations on this.
[0054] Among them, when there is one audio stream, the characteristics of this audio stream are input into the trained scene recognition model to obtain the holographic scene type of this audio stream. When there are two or more audio streams, the characteristics of each audio stream are input into the trained scene recognition model to obtain the holographic scene type of each audio stream.
[0055] In addition, the above-mentioned holographic scene types may include: game background sound, music, voice, type I video, type II video, system calls, voice transmission based on IP (Voice over Internet Protocol, VoIP) type voice, ringtones, alarms, notifications, navigation, game voice, etc. Here, the embodiments of the present application do not make specific limitations on this.
[0056] In this way, the holographic scene type of the audio stream can be obtained. The trained scene recognition model can be used to determine the holographic scene type more accurately, and the audio effect under holographic audio is better.
[0057] Generally, in a holographic audio scenario, there are multiple audio streams that need to be played. In order to identify the holographic scene types of the multiple audio streams, in an optional embodiment, when the number of audio streams is at least two, S102 may include:
[0058] Extract features from each audio stream in the audio stream to obtain features of each audio stream;
[0059] Accordingly, S103 may include:
[0060] The features of each audio stream are input into the trained scene recognition model to perform holographic scene recognition respectively, and the holographic scene type of each audio stream is obtained.
[0061] It can be understood that when the number of audio streams is at least two, when the electronic device obtains at least two audio streams, the electronic device extracts features from each audio stream, thereby obtaining the features of each audio stream, and then inputs the features of each audio stream into the trained scene recognition model to perform holographic scene recognition respectively, thereby obtaining the holographic scene type of each audio stream.
[0062] That is to say, for two or more audio streams, feature extraction and scene recognition are performed separately to determine the holographic scene type of each audio stream, which helps to improve the accuracy of holographic scene type determination, so that electronic devices can know the holographic scene type of each audio stream to better utilize the holographic audio to play the spatial audio corresponding to the audio stream.
[0063] In order to improve the sound effect of the audio stream in the holographic audio, in an optional embodiment, the method may further include:
[0064] Based on the holographic scene type of the audio stream, play the spatial audio corresponding to the audio stream.
[0065] It can be understood that when the holographic scene type of the audio stream is obtained, the spatial audio corresponding to the audio stream can be played based on the determined holographic scene type, wherein the holographic scene type of the audio stream is mainly used to determine the spatial position of the audio stream, and then the electronic device plays the spatial audio based on the determined spatial position of the audio stream.
[0066] In this way, playing the spatial audio corresponding to the audio stream based on the holographic scene type of the audio stream can improve the spatial sound effect of the audio stream in the holographic scene.
[0067] When the number of audio streams is one, in order to improve the spatial sound effect of the audio stream under holographic audio, in an optional embodiment, based on the holographic scene type of the audio stream, playing the spatial audio corresponding to the audio stream may include:
[0068] Based on the holographic scene type of the audio stream, determine the spatial position corresponding to the holographic scene type;
[0069] Play the spatial audio of the audio stream based on the spatial position corresponding to the holographic scene type.
[0070] It can be understood that the electronic device first determines the spatial position corresponding to the holographic scene type based on the holographic scene type of the audio stream. Here, the electronic device can be pre-set with a correspondence between the scene type and the spatial position, and then determines the spatial position corresponding to the holographic scene type based on the correspondence, and plays the spatial audio of the audio stream at the determined spatial position corresponding to the holographic scene type.
[0071] It is also possible to directly determine the center position in the space where the holographic audio is located as the spatial position corresponding to the holographic scene type, and play the spatial audio of the audio stream at the center position; here, the embodiment of the present application does not make specific limitations on this.
[0072] In this way, the spatial audio of the audio stream can be played based on the holographic scene type of the audio stream, thereby improving the spatial sound effect under the holographic audio.
[0073] In addition, when the number of audio streams is at least two, in order to improve the spatial sound effect in the holographic scene, in an optional embodiment, when the number of audio streams is at least two, based on the holographic scene type of the audio stream, playing the spatial audio corresponding to the audio stream may include:
[0074] Determine the spatial position corresponding to each audio stream based on the priority of the holographic scene type of each audio stream in the audio stream;
[0075] Based on the spatial position corresponding to each audio stream, the spatial audio of each audio stream is played.
[0076] It can be understood that the electronic device determines the priority of the holographic scene type of each audio stream. Here, it should be noted that each holographic scene type corresponds to a priority. For example, the priorities are from high to low: game background sound, music, voice, first-class video, second-class video, system call, VoIP-class voice, ringtone, alarm, notification, navigation, and game voice.
[0077] After the electronic device determines the priority of the holographic scene type of each audio stream, since the correspondence between the priority and the spatial position can be pre-stored in the electronic device, the spatial position of each audio stream can be determined based on the correspondence, or the spatial position can be determined from high to low priority. Here, the embodiments of the present application do not make specific limitations on this.
[0078] After determining the spatial position corresponding to each audio stream, the electronic device can play the spatial audio of each audio stream at the spatial position corresponding to each audio stream.
[0079] In this way, the electronic device determines the spatial position corresponding to each audio stream based on the priority of the holographic scene type of each audio, and plays the spatial audio of each audio stream at the spatial position corresponding to each audio stream, which helps to improve the spatial sound effect under the holographic audio.
[0080] In order to improve the accuracy of holographic scene type recognition, in an optional embodiment, the above method may further include:
[0081] When the trained scene recognition model continuously outputs the same holographic scene type of the audio stream, and the number of continuous outputs reaches a preset number, the characteristics of the audio stream and the holographic scene type of the audio stream are sent to the cloud server.
[0082] It can be understood that after the electronic device obtains the audio stream, it will obtain the audio stream at preset time periods to identify the holographic scene type of the audio stream in the above manner. Then, in model recognition, when the trained scene recognition model continuously outputs the same holographic scene type of the audio stream, and the number of consecutive outputs reaches a preset number of times, it means that for the audio stream, the recognition results of the preset number of consecutive recognitions are the same. It can be considered that the recognition result has a high accuracy and can be used for model retraining. Therefore, here, the characteristics of the audio stream and the holographic scene type of the audio stream are sent to the cloud server.
[0083] Among them, the characteristics of the audio stream and the holographic scene type of the audio stream are used by the cloud server to train the locally trained scene recognition model, and obtain new parameters of the trained scene recognition model to update the locally trained scene recognition model; that is, the cloud server can use the characteristics of the audio stream and the holographic scene type of the audio stream to send to the cloud server to continue training the locally trained scene recognition model to obtain new parameters of the trained scene recognition model, and then update the locally trained scene recognition model. In addition, the new parameters can also be used by electronic devices to update their locally trained scene recognition models.
[0084] In this way, the parameters of the model are updated on both the cloud server side and the electronic device side, thereby improving the accuracy of the holographic scene type recognized by the model.
[0085] In order to obtain a trained scene recognition model, in an optional embodiment, the method may further include:
[0086] Get the trained scene recognition model from the cloud server.
[0087] It can be understood that in addition to the electronic device being able to train the trained scene recognition model, the trained scene recognition model can also be trained in the cloud service, so that the electronic device can obtain the trained scene recognition model from the cloud server, so that after the electronic device obtains the audio stream and extracts features thereof, it can perform holographic scene recognition on the audio stream, thereby obtaining the holographic scene type of the audio stream.
[0088] In this way, by training a scene recognition model through a cloud server, it is possible to reduce the energy consumption of electronic devices while improving the recognition accuracy of holographic scene types.
[0089] In addition, in order to update the parameters of the trained scene recognition model to improve the accuracy of holographic scene recognition, in an optional embodiment, the method may further include:
[0090] Get new parameters of the trained scene recognition model from the cloud server;
[0091] The parameters of the trained scene recognition model are updated to new parameters to obtain the trained scene recognition model again.
[0092] It can be understood that after obtaining the trained scene recognition model, the electronic device can obtain new parameters of the locally trained scene recognition model from the cloud server, and then update the parameters of the trained scene recognition model to the new parameters, and then obtain the trained scene recognition model again, thereby completing the update of the trained scene recognition model in the electronic device.
[0093] This helps to improve the accuracy of holographic scene type recognition and enhances the spatial sound effects of holographic audio.
[0094] In order to extract features of the audio stream, in an optional embodiment, S102 may include:
[0095] The Mel-frequency cepstral coefficient features of the audio stream are extracted to obtain the features of the audio stream.
[0096] It can be understood that here, before feature extraction, data cleaning, file format alignment, data enhancement and spectrum conversion are mainly used to obtain the processed audio stream. It should be noted that the above processing is not limited to this; after the above processing, the processed audio stream is feature extracted by algorithms such as Fbank or MFCC, so as to obtain the characteristics of the audio stream.
[0097] In this way, after obtaining the characteristics of the audio stream, a more accurate holographic scene type of the audio stream can be obtained, thereby improving the accuracy of holographic scene recognition.
[0098] The present application also provides an audio stream recognition method, which is applied in a cloud server. Figure 2 A schematic diagram of an optional audio stream recognition method provided in an embodiment of the present application Figure 2 ,like Figure 2 As shown, the audio stream identification method may include:
[0099] S201: Obtain a sample data set;
[0100] S202: training a scene recognition model using a sample data set to obtain a trained scene recognition model;
[0101] S203: Send the trained scene recognition model to the electronic device.
[0102] In order to enable the electronic device to obtain the trained scene recognition model, here, in S201, the cloud server first obtains a sample data set, wherein the sample data set may include: a collected audio stream and a holographic scene type of the collected audio stream. Here, the collected audio stream may be an audio stream of one application, or may be an audio stream of two or more applications. Here, the embodiment of the present application does not make any specific limitation on this.
[0103] After acquiring the sample data set, in S202, the cloud server uses the sample data set to train a pre-stored scene recognition model, thereby obtaining a trained scene recognition model. In order to enable the electronic device to recognize the holographic scene type, here, in S203, the cloud server sends the trained scene recognition model to the electronic device, so that the electronic device can recognize the holographic scene type of the audio stream.
[0104] In this way, the electronic device can improve the accuracy of holographic scene type recognition while reducing the energy consumption of the electronic device.
[0105] In order to enable the cloud server to optimize the trained scene recognition model in the electronic device, in an optional embodiment, the method may further include:
[0106] Acquire the characteristics of the audio stream and the holographic scene type of the audio stream from the electronic device;
[0107] Using the features of the acquired audio stream and the holographic scene type of the acquired audio stream, the locally trained scene recognition model is trained to obtain new parameters of the locally trained scene recognition model;
[0108] Send new parameters to the electronics.
[0109] It can be understood that the cloud server can obtain the characteristics of the audio stream and the holographic scene type of the identified audio stream from the electronic device, wherein the holographic scene type of the audio stream is the type that is continuously output from the trained scene recognition model and the number of consecutive outputs reaches a preset number of times; that is, the cloud server obtains accurate data from the electronic device, which can be used to retrain the locally trained scene recognition model to obtain new parameters of the trained scene recognition model.
[0110] The cloud server sends the new parameters to the electronic device, so that the electronic device can update the parameters of the locally trained scene recognition model to the new parameters, thereby enabling the electronic device to optimize the locally trained scene recognition model.
[0111] In this way, by continuously optimizing the trained scene recognition model in the electronic device, the accuracy of holographic scene recognition can be further improved, thereby improving the spatial sound effect under holographic audio.
[0112] In addition, in order to improve the accuracy of holographic scene recognition without affecting the holographic scene type recognition, in an optional embodiment, the locally trained scene recognition model is trained using the characteristics of the acquired audio stream and the holographic scene type of the acquired audio stream to obtain new parameters of the locally trained scene recognition model, which may include:
[0113] When the current moment reaches the preset time range, the locally trained scene recognition model is trained using the characteristics of the acquired audio stream and the holographic scene type of the acquired audio stream to obtain new parameters of the locally trained scene recognition model.
[0114] It can be understood that only when the current moment reaches the preset time range, the cloud server uses the characteristics of the acquired audio stream and the holographic scene type of the acquired audio stream to train the locally trained scene recognition model to obtain new parameters of the locally trained scene recognition model.
[0115] That is to say, the cloud server will only train the locally trained scene recognition model in a specific time period to obtain new parameters, which are then used to optimize the trained scene recognition model in the electronic device.
[0116] Among them, the above-mentioned preset time range can be any time period of a day. Usually, a time period at night can be selected to retrain the scene recognition model after local training to prevent too many cloud server resources from being occupied during the day, thereby affecting the work efficiency of the cloud server.
[0117] The following is an example to describe the audio stream recognition method in one or more of the above embodiments.
[0118] Figure 3 A flowchart of an example 1 of an optional audio stream recognition method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the audio stream identification method may include:
[0119] S301: The mobile terminal obtains the audio stream of the application;
[0120] S302: The mobile terminal extracts features from the audio stream to obtain features of the audio stream;
[0121] S303: The mobile terminal inputs the features of the audio stream into the trained artificial intelligence (AI) model to obtain a holographic scene label of the audio stream;
[0122] S304: The mobile terminal outputs a holographic scene tag.
[0123] Specifically, the application in the mobile terminal will write audio stream data to the mobile terminal through the AudioTrack interface. The mobile terminal will extract features from this part of the data. First, data cleaning will be performed, such as: file format alignment, data enhancement, and spectrum conversion. Then, features will be extracted through algorithms such as Fbank or MFCC. After the extracted features are calculated by the trained AI model (equivalent to the above-mentioned trained scene recognition model), the output holographic scene label can be obtained.
[0124] Figure 4 A flowchart of Example 2 of an optional audio stream recognition method provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the audio stream identification method may include:
[0125] S401: The mobile terminal outputs a holographic scene label using the trained AI model;
[0126] S402: The mobile terminal determines whether the holographic scene tag has changed. If yes, execute S403;
[0127] S403: The mobile terminal stores the feature and the holographic scene tag;
[0128] S404: The mobile terminal waits for uploading recent data to the cloud server at night;
[0129] S405: The cloud server receives the data and uses the data to retrain the trained AI model to obtain new parameters of the trained AI model;
[0130] S406: The cloud server periodically publishes new parameters to the mobile terminal.
[0131] Among them, the learning process of the AI model is divided into two parts: model training and nighttime self-learning. Model training is not carried out on the mobile terminal, but on the cloud server. The cloud server trains the AI model through the acquired sample data set, and then pre-installs the trained AI model on the mobile terminal.
[0132] For nighttime self-learning, the mobile terminal uses the trained AI model to output the holographic scene label, and then determines whether the audio stream identified on this end has continuously output the same label for a preset number of times. If so, it waits for uploading recent data to the cloud server at night. The recent data is the characteristics of the audio stream and the holographic scene type of the audio stream. If not, it uploads the data. In this way, after receiving the data, the cloud server uses the data to retrain the trained AI model, obtains the new parameters of the trained AI model, and sends them to the mobile terminal to update the parameters of the locally trained AI model.
[0133] This example proposes a neural network-based model that can infer the playback type of the audio stream in the mobile terminal according to its content, such as music, video, navigation, alarm, notification, incoming call, etc.
[0134] Based on this example, the mobile terminal can asynchronously determine the category of the sound source being played, solving the problem of the application incorrectly marking the audio stream type. At the same time, it can support all music-playing apps to experience the holographic audio effect, providing great convenience and hands-on experience space for third-party development enthusiasts and holographic fans.
[0135] This example is applied to the holographic audio module of the mobile terminal. It uses audio classification technology to perform holographic scene recognition on part of the data content of the input audio source, obtain the holographic scene label of the audio source, and then the holographic algorithm performs audio and video control based on this label and its corresponding priority, so that each audio stream can be distributed according to its scene. While obtaining the holographic scene label, the characteristics of the audio stream and its label are asynchronously saved, and then the AI model is retrained when the mobile phone is charged at night, so as to achieve self-learning of the AI model for user usage habits and improve the accuracy of holographic scene type recognition.
[0136] For example, when a user watches a live broadcast on a music app, after the application writes down the audio stream data, the audio classification technology will regularly sample the content of the audio stream to extract features, and then obtain scene labels through scene recognition. In the subsequent holographic algorithm, the sound and image positions are distributed according to the labels. When the audio stream ends, the useful parts of the features of the audio playback will be stored and cleaned up regularly, and the AI model will be self-learned when the user's phone is idle while charging at night.
[0137] An embodiment of the present application provides an audio stream recognition method, which is applied to an electronic device, including: acquiring an audio stream of the electronic device, performing feature extraction on the audio stream to obtain features of the audio stream, inputting the features of the audio stream into a trained scene recognition model for holographic scene recognition, and obtaining a holographic scene type of the audio stream; that is, in an embodiment of the present application, the holographic scene type of the audio stream can be obtained by inputting the audio stream of the electronic device into a trained scene recognition model for holographic scene recognition, and thus, the holographic scene type of the audio stream is identified by using the trained scene recognition model, so that the determined holographic scene type is more suitable for holographic audio, thereby improving the accuracy of determining the holographic scene type under holographic audio, and further improving the sound effect of the holographic audio.
[0138] Based on the same inventive concept as the above-mentioned embodiment, the embodiment of the present application provides an audio stream recognition device, which is arranged in an electronic device. Figure 5 A schematic diagram of the structure of an optional audio stream recognition device provided in an embodiment of the present application Figure 1 ,like Figure 5 As shown, the audio stream identification device may include:
[0139] A first acquisition module 51 is used to acquire an audio stream of an electronic device;
[0140] An extraction module 52, used to extract features from the audio stream to obtain features of the audio stream;
[0141] The recognition module 53 is used to input the features of the audio stream into the trained scene recognition model to perform holographic scene recognition and obtain the holographic scene type of the audio stream.
[0142] In an optional embodiment, when the number of audio streams is at least two, the extraction module 52 is specifically configured to:
[0143] Extract features from each audio stream in the audio stream to obtain features of each audio stream;
[0144] Accordingly, the identification module 53 is specifically used for:
[0145] The features of each audio stream are input into the trained scene recognition model to perform holographic scene recognition respectively, and the holographic scene type of each audio stream is obtained.
[0146] In an optional embodiment, the device is further used for:
[0147] Based on the holographic scene type of the audio stream, play the spatial audio corresponding to the audio stream.
[0148] In an optional embodiment, when the number of audio streams is one, the apparatus may play the spatial audio corresponding to the audio stream based on the holographic scene type of the audio stream, and may include:
[0149] Based on the holographic scene type of the audio stream, determine the spatial position corresponding to the holographic scene type;
[0150] Play the spatial audio of the audio stream based on the spatial position corresponding to the holographic scene type.
[0151] In an optional embodiment, when the number of audio streams is at least two, the apparatus may play the spatial audio corresponding to the audio stream based on the holographic scene type of the audio stream, and may include:
[0152] Determine the spatial position corresponding to each audio stream based on the priority of the holographic scene type of each audio stream in the audio stream;
[0153] Based on the spatial position corresponding to each audio stream, the spatial audio of each audio stream is played.
[0154] In an optional embodiment, the device is further used for:
[0155] When the trained scene recognition model continuously outputs the same holographic scene type of the audio stream, and the number of continuous outputs reaches a preset number, the characteristics of the audio stream and the holographic scene type of the audio stream are sent to the cloud server;
[0156] Among them, the characteristics of the audio stream and the holographic scene type of the audio stream are used by the cloud server to train the locally trained scene recognition model, obtain new parameters of the trained scene recognition model, and update the locally trained scene recognition model.
[0157] In an optional embodiment, the device is further used for:
[0158] Get the trained scene recognition model from the cloud server.
[0159] In an optional embodiment, the device is further used for:
[0160] Get new parameters of the trained scene recognition model from the cloud server;
[0161] The parameters of the trained scene recognition model are updated to new parameters to obtain the trained scene recognition model again.
[0162] In an optional embodiment, the extraction module 52 is specifically configured to:
[0163] The Mel-frequency cepstral coefficient features of the audio stream are extracted to obtain the features of the audio stream.
[0164] In practical applications, the first acquisition module 51, extraction module 52 and recognition module 53 can be implemented by a processor located on the audio stream recognition device, specifically a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA).
[0165] The present application provides an audio stream recognition device, which is arranged in an electronic device. Figure 6 A schematic diagram of the structure of an optional audio stream recognition device provided in an embodiment of the present application Figure 2 ,like Figure 6 As shown, the audio stream identification device may include:
[0166] The second acquisition module 61 is used to acquire a sample data set; wherein the sample data set includes: a collected audio stream and a holographic scene type of the collected audio stream;
[0167] A training module 62, used to train the scene recognition model using the sample data set to obtain a trained scene recognition model;
[0168] The sending module 63 is used to send the trained scene recognition model to the electronic device.
[0169] In an optional embodiment, the device is further used for:
[0170] Acquire the characteristics of the audio stream and the holographic scene type of the audio stream from the electronic device; wherein the holographic scene type of the audio stream is the type that is continuously output from the trained scene recognition model and the number of consecutive outputs reaches a preset number;
[0171] Using the features of the acquired audio stream and the holographic scene type of the acquired audio stream, the locally trained scene recognition model is trained to obtain new parameters of the locally trained scene recognition model;
[0172] Send new parameters to the electronics.
[0173] In an optional embodiment, the device trains the locally trained scene recognition model using the features of the acquired audio stream and the holographic scene type of the acquired audio stream, and obtains new parameters of the trained scene recognition model, including:
[0174] When the current moment reaches the preset time range, the locally trained scene recognition model is trained using the characteristics of the acquired audio stream and the holographic scene type of the acquired audio stream to obtain new parameters of the locally trained scene recognition model.
[0175] In practical applications, the above-mentioned second acquisition module 61, training module 62 and sending module 63 can be implemented by a processor located on the audio stream recognition device, specifically a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA).
[0176] Figure 7 A schematic diagram of the structure of an optional electronic device provided in an embodiment of the present application, such as Figure 7 As shown, an embodiment of the present application provides an electronic device 700, including:
[0177] A processor 71 and a storage medium 72 storing executable instructions of the processor 71, wherein the storage medium 72 relies on the processor 71 to perform operations through a communication bus 73, and when the instructions are executed by the processor 71, the audio stream recognition method executed in one or more of the above embodiments is executed.
[0178] It should be noted that, in actual application, the components in the electronic device 700 are coupled together via the communication bus 73. It is understood that the communication bus 73 is used to realize the connection and communication between these components. In addition to the data bus, the communication bus 73 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 7 Various buses are labeled as communication buses 73.
[0179] Figure 8 A schematic diagram of the structure of an optional cloud server provided in an embodiment of the present application, such as Figure 8 As shown, the embodiment of the present application provides a cloud server 800, including:
[0180] A processor 81 and a storage medium 82 storing executable instructions of the processor 81, wherein the storage medium 82 relies on the processor 81 to perform operations through a communication bus 83, and when the instructions are executed by the processor 81, the audio stream recognition method executed in one or more of the above embodiments is executed.
[0181] It should be noted that in actual application, the various components in the cloud server 800 are coupled together through the communication bus 83. It is understandable that the communication bus 83 is used to realize the connection and communication between these components. In addition to the data bus, the communication bus 83 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 8 Various buses are labeled as communication buses 83.
[0182] An embodiment of the present application provides a computer storage medium storing executable instructions. When the executable instructions are executed by one or more processors, the processors execute the audio stream recognition method as executed by the control device in one or more of the above embodiments.
[0183] Among them, the computer-readable storage medium can be a ferromagnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and other memories.
[0184] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0185] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0186] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0188] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. A method for identifying an audio stream, characterized in that: The method is applied to an electronic device, comprising: Acquiring an audio stream of the electronic device; Extracting features from the audio stream to obtain features of the audio stream; The features of the audio stream are input into the trained scene recognition model to perform holographic scene recognition, and the holographic scene type of the audio stream is obtained.
2. The method according to claim 1, characterized in that When the number of the audio streams is at least two, extracting features from the audio streams to obtain features of the audio streams includes: Extracting features from each audio stream in the audio streams to obtain features of each audio stream; Accordingly, the step of inputting the features of the audio stream into the trained scene recognition model to perform holographic scene recognition to obtain the holographic scene type of the audio stream includes: The features of each audio stream are input into the trained scene recognition model to perform holographic scene recognition respectively, so as to obtain the holographic scene type of each audio stream.
3. The method according to claim 1, characterized in that The method further comprises: Based on the holographic scene type of the audio stream, the spatial audio corresponding to the audio stream is played.
4. The method according to claim 3, characterized in that When the number of the audio streams is one, the playing of the spatial audio corresponding to the audio stream based on the holographic scene type of the audio stream includes: Based on the holographic scene type of the audio stream, determining a spatial position corresponding to the holographic scene type; Based on the spatial position corresponding to the holographic scene type, the spatial audio of the audio stream is played.
5. The method according to claim 3, characterized in that: When the number of the audio streams is at least two, the playing of the spatial audio corresponding to the audio stream based on the holographic scene type of the audio stream includes: Determine a spatial position corresponding to each audio stream based on a priority of a holographic scene type of each audio stream in the audio streams; Based on the spatial position corresponding to each audio stream, the spatial audio of each audio stream is played.
6. The method according to claim 1, characterized in that The method further comprises: When the trained scene recognition model continuously outputs the same holographic scene type of the audio stream, and the number of continuous outputs reaches a preset number, the characteristics of the audio stream and the holographic scene type of the audio stream are sent to the cloud server; The characteristics of the audio stream and the holographic scene type of the audio stream are used by the cloud server to train the locally trained scene recognition model to obtain new parameters of the trained scene recognition model to update the locally trained scene recognition model.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Get the trained scene recognition model from the cloud server.
8. The method according to claim 7, characterized in that The method further comprises: Get new parameters of the trained scene recognition model from the cloud server; The parameters of the trained scene recognition model are updated to the new parameters to obtain the trained scene recognition model again.
9. The method according to any one of claims 1 to 6, characterized in that: The extracting features from the audio stream to obtain features of the audio stream includes: Mel-frequency cepstral coefficient features are extracted from the audio stream to obtain features of the audio stream.
10. A method for identifying an audio stream, characterized in that: The method is applied to a cloud server and includes: Acquire a sample data set; wherein the sample data set includes: a collected audio stream and a holographic scene type of the collected audio stream; Using the sample data set to train a scene recognition model to obtain a trained scene recognition model; The trained scene recognition model is sent to the electronic device.
11. The method according to claim 10, characterized in that The method further comprises: Acquire the characteristics of the audio stream and the holographic scene type of the audio stream from the electronic device; wherein the holographic scene type of the audio stream is the type that is continuously output from the trained scene recognition model and the number of consecutive outputs reaches a preset number of times; Using the features of the acquired audio stream and the holographic scene type of the acquired audio stream, the locally trained scene recognition model is trained to obtain new parameters of the locally trained scene recognition model; The new parameters are sent to the electronic device.
12. The method according to claim 11, characterized in that The method of training the locally trained scene recognition model by using the features of the acquired audio stream and the holographic scene type of the acquired audio stream to obtain new parameters of the trained scene recognition model includes: When the current moment reaches the preset time range, the locally trained scene recognition model is trained using the characteristics of the acquired audio stream and the holographic scene type of the acquired audio stream to obtain new parameters of the locally trained scene recognition model.
13. An audio stream recognition device, characterized in that: The device is arranged in an electronic device and comprises: A first acquisition module, used to acquire an audio stream of the electronic device; An extraction module, used to extract features from the audio stream to obtain features of the audio stream; The recognition module is used to input the features of the audio stream into the trained scene recognition model to perform holographic scene recognition and obtain the holographic scene type of the audio stream.
14. An audio stream recognition device, characterized in that: The device is arranged on a cloud server and includes: A second acquisition module is used to acquire a sample data set; wherein the sample data set includes: a collected audio stream and a holographic scene type of the collected audio stream; A training module, used to train the scene recognition model using the sample data set to obtain a trained scene recognition model; The sending module is used to send the trained scene recognition model to the electronic device.
15. An electronic device, characterized in that: include: A processor and a storage medium storing instructions executable by the processor, wherein the storage medium relies on the processor to perform operations through a communication bus, and when the instructions are executed by the processor, the audio stream recognition method described in any one of claims 1 to 9 is executed.
16. A cloud server, characterized in that: include: A processor and a storage medium storing instructions executable by the processor, wherein the storage medium relies on the processor to perform operations through a communication bus, and when the instructions are executed by the processor, the audio stream recognition method described in any one of claims 10 to 12 is executed.
17. A computer storage medium, characterized in that: Executable instructions are stored. When the executable instructions are executed by one or more processors, the processors execute the audio stream recognition method as described in any one of claims 1 to 9, or execute the audio stream recognition method as described in any one of claims 10 to 12.
Citation Information
Patent Citations
Sound effect setting method for terminal and terminal
CN106095387A
Systems and methods for spatial audio adjustment
CN108141696A
Scene recognition method and apparatus, storage medium and electronic device
CN108764304A
Audio effect processing method, device and electronic device
CN109218528A
Device and method for audio classification and audio processing
CN109616142A