Scene recognition method, device, terminal, storage medium and program product
Through the two-level scene recognition method, ambient audio is used to identify scenes of terminal devices. First, the muted or non-silent scene changes are judged, and then high-complexity recognition is performed, which solves the problem of high power consumption of terminal devices when scene changes, and achieves efficient and accurate scene recognition.
Patent Information
- Application Number
- CN202111624031.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-28
AI Technical Summary
In the prior art, the terminal device uses a neural network model to identify scenes when frequent scene changes lead to an increase in power consumption, and the image acquisition method recognizes problems such as inaccurate or high power consumption in different scenarios.
The two-level scene recognition method is adopted. First, the muted or non-silent scene changes are judged through the first-level recognition with low computing complexity. If the non-silent scene changes, the second-level recognition with high computing complexity is then used to determine the specific scene type, and the environmental audio is used for identification.
It reduces the power consumption of terminal devices in the scene recognition process, while improving the accuracy and efficiency of recognition, adapting to changes in different scenarios.
Smart Images

Figure CN114299988B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of audio processing, and particularly to an audio scene recognition method, apparatus, terminal, storage medium, and program product. Background Art
[0002] Nowadays, the application of scene recognition is becoming more and more extensive. For example, when the scene where the user is located changes, the terminal needs to perform scene recognition on the scene where the user is located to adjust the information pushed to the user or modify the device parameters to adapt to the scene change.
[0003] In related technologies, the terminal usually directly performs scene recognition based on images or audio through a neural network model. Since a large amount of computing power of the terminal is required for scene recognition through the neural network model, when the scene where the user is located changes frequently, the power consumption of the terminal increases. Summary of the Invention
[0004] Embodiments of the present application provide a scene recognition method, apparatus, terminal, storage medium, and program product, and the technical solutions are as follows:
[0005] On the one hand, embodiments of the present application provide a scene recognition method, and the method includes:
[0006] Obtain real-time ambient audio;
[0007] Perform first-level scene recognition based on the real-time ambient audio to obtain a first scene recognition result, where the first scene recognition result is used to indicate the change situation of a silent scene or a non-silent scene;
[0008] In the case where the first scene recognition result indicates a change in the non-silent scene, perform second-level scene recognition based on the real-time ambient audio to obtain a second scene recognition result, where the second scene recognition result is used to indicate the scene type of the non-silent scene, and the computational complexity of the second-level scene recognition is higher than that of the first-level scene recognition.
[0009] On the other hand, embodiments of the present application provide a scene recognition apparatus, and the apparatus includes:
[0010] An obtaining module, configured to obtain real-time ambient audio;
[0011] A first-level scene recognition module, configured to perform first-level scene recognition based on the real-time ambient audio to obtain a first scene recognition result, where the first scene recognition result is used to indicate the change situation of a silent scene or a non-silent scene;
[0012] A second-level scene recognition module, configured to perform second-level scene recognition on the basis of the real-time environmental audio when the first scene recognition result indicates that a non-silent scene has changed, so as to obtain a second scene recognition result, where the second scene recognition result is used to indicate the scene type of the non-silent scene, and the computational complexity of the second-level scene recognition is higher than that of the first-level scene recognition.
[0013] On the other hand, an embodiment of the present application provides a terminal, where the terminal includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the scene recognition method as described in the above aspect.
[0014] On the other hand, an embodiment of the present application provides a computer-readable storage medium, where at least one instruction, at least one program, a code set or an instruction set is stored in the readable storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the scene recognition method as described in the above aspect.
[0015] On the other hand, an embodiment of the present application provides a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the scene recognition method provided in the above aspect.
[0016] The technical solution provided by the present application may include the following beneficial effects:
[0017] The terminal first performs first-level scene recognition based on the real-time environmental audio, and determines that when a non-silent scene changes, then performs second-level scene recognition based on the real-time environmental audio to determine the specific scene type. Since the computational complexity of the second-level scene recognition is higher than that of the first-level scene recognition, in the embodiment of the present application, the terminal first performs pre-recognition through the first-level scene recognition with lower computational complexity, avoiding the terminal directly performing second-level scene recognition based on the real-time environmental audio, thereby reducing the power consumption of the terminal during scene recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0019] Figure 1 The schematic diagram of the principle of scene recognition provided by an exemplary embodiment of the present application is shown;
[0020] Figure 2 shows a flowchart of a scene recognition method provided by an exemplary embodiment of the present application;
[0021] Figure 3 shows a flowchart of a scene recognition method provided by another exemplary embodiment of the present application;
[0022] Figure 4 shows a flowchart of a scene recognition method provided by another exemplary embodiment of the present application;
[0023] Figure 5 shows a flowchart of a scene recognition method provided by another exemplary embodiment of the present application;
[0024] Figure 6 shows a structural block diagram of a scene recognition device provided by an exemplary embodiment of the present application;
[0025] Figure 7 shows a structural block diagram of a terminal provided by an exemplary embodiment of the present application. Detailed implementation manners
[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0027] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0028] In related technologies, scene recognition applications are becoming increasingly widespread. The terminal adjusts the content pushed to the user or its relevant configuration parameters based on the results of scene recognition, thereby improving the user experience. For example, the terminal indirectly performs scene recognition by collecting sensor data such as the user's movement speed. However, since sensor data such as the user's movement speed may be the same in different scenarios, it is easy to cause inaccurate scene recognition results. Exemplarily, the movement data of the user in a subway station and an airport may be the same. In this case, the terminal using the foregoing method for scene recognition cannot determine the specific scene type. Or the terminal collects the surrounding environment images and performs scene recognition based on deep learning algorithms. Although this method improves the accuracy of scene recognition, it will consume a large amount of computing power of the terminal and increase the power consumption of the terminal.
[0029] Therefore, in order to solve the problems existing in scene recognition in related technologies, in the embodiments of the present application, the terminal performs two-level scene recognition with different computational complexities, reducing the power consumption of the terminal while ensuring the accuracy of terminal scene recognition. In addition, in the embodiments of the present application, scene recognition is performed through environmental audio, which is more sensitive in different scenarios compared to images. As Figure 1 shown, it shows a schematic diagram of the principle of scene recognition provided by an exemplary embodiment of the present application.
[0030] After the terminal device is started, the terminal collects real-time environmental audio and then performs the first-level scene recognition. The first-level scene recognition has a low computational complexity and mainly determines whether the current scene has changed. The terminal first performs a silence judgment on the real-time environmental audio. If the judgment result is a long silence segment and it is determined that the current scene is a silent scene, there is no need to perform the second-level scene recognition with a high computational complexity, thereby reducing the power consumption of the terminal. When the terminal determines that it is a non-silent segment, it extracts the audio features of the real-time environmental audio and judges whether the scene has changed. If the scene has not changed, the terminal updates the relevant parameters of the current scene. If the scene has changed, the terminal performs the second-level scene recognition with a high computational complexity to determine the specific type of the scene and updates the relevant parameters of the scene. The terminal performs pre-recognition of the current scene through the first-level scene recognition with a low computational complexity, and then performs the second-level scene recognition with a high computational complexity when the recognition result is that the non-silent scene has changed, avoiding the terminal directly performing the second-level scene recognition, thereby reducing the power consumption of the terminal. In addition, it should be noted that in the embodiments of the present application, the terminal may be a smart phone, a tablet computer, an e-book reader, a smart wearable device, a laptop computer, a desktop computer, etc., and the embodiments of the present application do not make any limitations in this regard.
[0031] Next, the scene recognition method adopted in the embodiments of the present application will be introduced. Please refer to Figure 2 which shows a flowchart of the scene recognition method provided by an exemplary embodiment of the present application. The method includes:
[0032] Step 210, obtain the real-time ambient audio.
[0033] In a possible implementation, in order to accurately collect the real-time ambient audio, the terminal is provided with microphones in different directions, and the ambient audio is collected in real time through the microphones, and the collected ambient audio is stored in a buffer.
[0034] In the embodiments of the present application, the real-time ambient audio collected by the terminal microphone in different scenarios is different.
[0035] Step 220, perform a first-level scene recognition based on the real-time ambient audio to obtain a first scene recognition result, and the first scene recognition result is used to indicate the change situation of the mute scene or the non-mute scene.
[0036] In the embodiments of the present application, the terminal performs a first-level scene recognition based on the real-time ambient audio. The purpose of the first-level scene recognition is for the terminal to pre-recognize the current scene first, that is, to judge whether the current scene is a mute scene or a non-mute scene. If it is a non-mute scene, it is also necessary to determine whether the non-mute scene has changed, which is convenient for the terminal to perform the subsequent second-level scene recognition. It can be seen that the computational complexity of the first-level scene recognition is relatively low, so the power consumption of the terminal is low.
[0037] Optionally, since the computational complexity of the first-level scene recognition is relatively low and the demand for computing performance is relatively low, the first-level scene recognition can be executed by a DSP (Digital Signal Processor) or an MCU (Micro Controller Unit) in the terminal.
[0038] In a possible implementation, when the terminal determines that the current scene is a mute scene or the non-mute scene has not changed, there is no need to perform the second-level scene recognition, thereby reducing the power consumption of the terminal for scene recognition.
[0039] Optionally, the mute scene can be a library, a hospital, etc., and the embodiments of the present application do not limit this.
[0040] Optionally, the non-mute scene can be a subway station, an airport, a restaurant, an amusement park, etc., and the embodiments of the present application do not limit this.
[0041] Step 230, in the case where the first scene recognition result indicates that the non-mute scene has changed, perform a second-level scene recognition based on the real-time ambient audio to obtain a second scene recognition result, and the second scene recognition result is used to indicate the scene type of the non-mute scene, and the computational complexity of the second-level scene recognition is higher than that of the first-level scene recognition.
[0042] In a possible implementation, when the result of the first-level scene recognition by the terminal indicates a change in the non-silent scene, the terminal reads the real-time environmental audio from the buffer for the second-level scene recognition, and then determines the specific scene type. Therefore, compared with the first-level scene recognition, the computational complexity of the second-level scene recognition is high.
[0043] Optionally, since the computational complexity of the second-level scene recognition is relatively high and the demand for computing performance is high, the second-level scene recognition can be performed by the NPU (Neural Processing Unit) in the terminal.
[0044] In a possible implementation, the terminal modifies the information pushed to the user or the relevant configuration parameters based on the second scene recognition result.
[0045] Optionally, the terminal modifies the information pushed to the user according to the scene type. For example, when the second scene recognition result is a subway station scene, the terminal pushes information such as restaurants and tourist attractions near the subway station to the user. When the second scene recognition result is a restaurant scene, the information pushed by the terminal is changed to shopping malls near the restaurant.
[0046] Optionally, the configuration parameter can be the noise reduction parameter of the terminal filter, etc., and the embodiments of the present application do not limit this.
[0047] Exemplarily, the terminal collects the real-time environmental audio through the microphone and stores it in the buffer. The DSP or MCU in the terminal reads the real-time environmental audio from the buffer and performs the first-level scene recognition based on the real-time environmental audio to obtain the first scene recognition result. When the first scene recognition result indicates a change in the non-silent scene, the NPU in the terminal reads the real-time environmental audio from the buffer to perform the second-level scene recognition and determines that the specific scene is a subway station.
[0048] In summary, in the embodiments of the present application, the terminal first performs the first-level scene recognition based on the real-time environmental audio. When it is determined that there is a change in the non-silent scene, the terminal then performs the second-level scene recognition based on the real-time environmental audio to determine the specific scene type. Since the computational complexity of the second-level scene recognition is higher than that of the first-level scene recognition, in the embodiments of the present application, the terminal first performs pre-recognition through the first-level scene recognition with lower computational complexity, avoiding directly performing the second-level scene recognition based on the real-time environmental audio by the terminal, thereby reducing the power consumption of the terminal during scene recognition.
[0049] In a possible implementation, when the terminal performs the first-level scene recognition based on the real-time environmental audio, it first preprocesses the real-time environmental audio to obtain audio frames, and performs the first-level scene recognition and the second-level scene recognition based on the audio features of the audio frames. Please refer to Figure 3, which shows a flowchart of a scene recognition method provided by another exemplary embodiment of the present application. The method includes:
[0050] Step 310, obtain real-time environmental audio.
[0051] For the implementation manner of this step, please refer to step 210, and the embodiments of the present application will not elaborate on this again.
[0052] Step 320, perform frame splitting processing on the real-time environmental audio to obtain audio frames.
[0053] In a possible implementation manner, the terminal performs frame splitting and windowing processing on the real-time environmental audio to obtain audio frames.
[0054] Audio signals are macroscopically non-stationary and microscopically stationary, with short-time stationarity. Therefore, the terminal can divide the audio signal into segments for processing, that is, perform frame splitting processing. Each segment is called a frame. Usually, the frame length of each frame can be 10 to 30 ms. Additionally, it should be noted that in the embodiments of the present application, the frame length can also be greater than 30 ms.
[0055] When the terminal performs frame splitting processing on the real-time environmental audio, a part of each frame will be repeatedly intercepted. That is, a part of the tail of the previous frame and the head of the current frame are overlapped and then windowing processing is performed. In this way, the global audio signal will not be weakened at both ends of a frame signal due to windowing processing and overly denoised audio data will not be obtained. Therefore, when performing frame splitting processing on the audio signal, there is an overlap between frames to make the audio signal after windowing processing more continuous.
[0056] Optionally, the window function for windowing processing can be a rectangular window, a Hamming window, a Hanning window, etc., and the embodiments of the present application do not limit this.
[0057] Step 330, when the energy of consecutive n audio frames is lower than the energy threshold, determine that the first scene recognition result is a silent scene, where n is a positive integer.
[0058] In the embodiments of the present application, the terminal extracts audio features of the audio frames for the first-level scene recognition.
[0059] Optionally, the audio features can be energy features.
[0060] In a possible implementation, an energy threshold is preset in the terminal. This energy threshold is obtained by extracting and calculating the audio features of the ambient audio in a silent scenario (when in a silent state in different scenarios, the energy of the ambient audio is basically the same). The terminal first determines the magnitude relationship between the audio frame energy and the energy threshold. In a possible implementation, when the audio frame energy is less than the energy threshold, the current scenario may be a silent scenario. To improve the accuracy of the first scenario recognition result, the terminal extracts the energies of consecutive n audio frames for statistical decision-making. That is, when the energies of consecutive n audio frames are lower than the energy threshold, it is determined that the current scenario is a silent scenario.
[0061] Optionally, the method of statistical decision-making can be minimum misjudgment probability criterion decision, minimum loss criterion decision, minimum maximum loss criterion decision, N-P (Neyman-Pearson) decision, etc. The embodiments of this application do not limit this.
[0062] Exemplarily, the terminal extracts the energies of consecutive 5 audio frames. When the energies of consecutive 5 audio frames are all lower than the energy threshold, the terminal determines that the first scenario recognition result is a silent scenario.
[0063] Step 340, when the energy of the audio frame is higher than the energy threshold, extract features from the audio frame to obtain real-time audio frame features.
[0064] In another possible implementation, when the energy of the audio frame is higher than the energy threshold, it indicates that the current scenario may be a non-silent scenario. Furthermore, the terminal determines whether the non-silent scenario has changed. Further, the terminal extracts features from the audio frame to obtain real-time audio frame features.
[0065] Optionally, the real-time audio frame features can be GMM (Gaussian Mixture Model) features.
[0066] Step 350, determine the first scenario recognition result based on the real-time audio frame features and the historical audio frame features. The historical audio frame features are obtained by extracting when performing scenario recognition on historical ambient audio.
[0067] In a possible implementation, the terminal determines whether the non-silent scenario has changed according to the feature similarity between the real-time audio frame features and the historical audio frame features. A similarity threshold is preset in the terminal. When the feature similarity is greater than the similarity threshold, it indicates that the similarity between the real-time audio frame features and the historical audio frame features is relatively high, and further indicates that the non-silent scenario has not changed. When the feature similarity is less than the similarity threshold, it indicates that the similarity between the real-time audio frame features and the historical audio frame features is relatively low, and further indicates that the non-silent scenario may have changed.
[0068] Among them, historical environmental audio is stored in the buffer of the terminal, and historical audio frame features are extracted when the terminal performs scene recognition based on the historical environmental audio.
[0069] Optionally, the historical audio frame features may be GMM (Gaussian Mixture Model) features.
[0070] Step 360: When the first scene recognition result indicates a change in the non - silent scene, perform a second - level scene recognition based on the real - time environmental audio to obtain a second scene recognition result. The second scene recognition result is used to indicate the scene type of the non - silent scene, and the computational complexity of the second - level scene recognition is higher than that of the first - level scene recognition.
[0071] In a possible implementation manner, a pre - trained scene recognition model is preset in the NPU of the terminal. The scene recognition model is used to recognize specific scene types. The NPU reads the real - time environmental audio in the buffer, extracts the audio features of the real - time environmental audio, inputs the audio features into the scene recognition model, and outputs the specific scene type.
[0072] Optionally, the audio features may be Mel Frequency Cepstral Coefficient (MFCC), Linear Predictive Cepstral Coefficient (LPCC), etc. The embodiments of the present application do not limit this.
[0073] Optionally, the scene recognition model may be a Convolutional Neural Networks (CNN) model, a Recurrent Neural Networks (RNN) model, or a combination of the two. The embodiments of the present application do not limit this.
[0074] Step 370: Set noise reduction parameters based on the first scene recognition result or the second scene recognition result, where different scenes correspond to different noise reduction parameters.
[0075] In the related art, in relatively quiet scenes (such as libraries) and relatively noisy scenes (such as subway stations), the terminal uses filters with the same noise reduction parameters, resulting in a reduced noise reduction effect of the terminal in different scenes. Therefore, in the embodiments of the present application, the terminal adjusts the noise reduction parameters of the filter in the terminal based on the first scene recognition result or the second scene recognition result to improve the noise reduction effect of the terminal.
[0076] Exemplarily, taking the terminal as a mobile phone for example, different noise reduction parameters corresponding to different scenarios are preset in the mobile phone. For example, the noise reduction parameter corresponding to the silent scenario (library) is A, and the noise reduction parameter corresponding to the restaurant scenario is B. When the user is in the library, the mobile phone collects the real-time ambient audio for the first-level scenario recognition. The first scenario recognition result is the silent scenario, and the mobile phone adjusts the noise reduction parameter of the filter to A. When the scene where the user is located changes from the library to the restaurant, the mobile phone collects the real-time ambient audio for the first-level scenario recognition and the second-level scenario recognition, determines that the specific scenario type is the restaurant, and the mobile phone adjusts the parameter of the filter to B according to the scenario recognition result, thereby improving the noise reduction effect of the terminal.
[0077] In the embodiment of the present application, the terminal collects the real-time ambient audio, performs the first-level scenario recognition and the second-level scenario recognition based on the audio features of the real-time ambient audio, and reduces the power consumption of the terminal. In addition, the terminal adjusts the noise reduction parameter according to the first scenario recognition result and the second scenario recognition result to adapt to different scenarios, and improves the noise reduction effect of the terminal.
[0078] In a possible implementation manner, when the terminal performs frame processing on the real-time ambient audio through the scene duration, the frame length is adaptively updated, so as to adjust the frequencies of the terminal for performing the first-level scenario recognition and the second-level scenario recognition, and further reduce the power consumption of the terminal. Please refer to Figure 4 , which shows a flowchart of a scenario recognition method provided by another exemplary embodiment of the present application. The method includes:
[0079] Step 401, obtain the real-time ambient audio.
[0080] For the implementation manner of this step, please refer to step 210, and the embodiments of the present application will not elaborate on this again.
[0081] Step 402, determine the frame length based on the scene duration. The scene duration is the duration of the current scene, and the frame length is positively correlated with the scene duration.
[0082] In a possible implementation manner, the terminal determines the frame length for performing frame processing on the real-time ambient audio through the scene duration. Among them, the longer the scene duration, the longer the frame length; the shorter the scene duration, the shorter the frame length. The longer the scene duration, the longer the frame length can reduce the frequencies of the terminal for performing the first-level scenario recognition and the second-level scenario recognition, thereby reducing the power consumption of the terminal. The shorter the scene duration, it indicates that the scene changes frequently. Shortening the frame length can increase the frequencies of the terminal for performing the first-level scenario recognition and the second-level scenario recognition, thereby improving the accuracy of the scenario recognition.
[0083] Optionally, a timer is set in the terminal for determining the scene duration. The terminal determines the scene duration according to the timer and then determines the frame length.
[0084] Exemplarily, when the timer shows 1 hour, it means the duration of the scene is 1 hour, and the terminal determines that the frame length is 10 ms. When the timer shows 5 hours, it means the duration of the scene is 5 hours, and the terminal determines that the frame length is 30 ms.
[0085] Step 403: Perform frame segmentation on the real-time ambient audio based on the frame length to obtain audio frames, where the length of each audio frame is the frame length.
[0086] In the embodiment of the present application, the terminal performs frame segmentation on the real-time ambient audio based on the determined frame length to obtain audio frames.
[0087] Exemplarily, frame segmentation is performed on the real-time ambient audio with 30 ms as the frame length to obtain audio frames, and the length of each audio frame is 30 ms.
[0088] Step 404: When the energy of consecutive n audio frames is lower than the energy threshold, determine that the first scene recognition result is a silent scene, where n is a positive integer.
[0089] For the implementation manner of this step, please refer to Step 330, and the embodiments of the present application will not elaborate on this again.
[0090] Step 405: When the energy of the audio frame is higher than the energy threshold, extract features from the audio frame to obtain real-time audio frame features.
[0091] For the implementation manner of this step, please refer to Step 340, and the embodiments of the present application will not elaborate on this again.
[0092] Step 406: Determine the maximum likelihood ratio between the real-time GMM features and the historical GMM features.
[0093] In the embodiment of the present application, the terminal extracts real-time audio frame features and historical audio frame features, namely real-time GMM features and historical GMM features, based on the real-time ambient audio and the historical ambient audio, and calculates the maximum likelihood ratio between the real-time GMM features and the historical GMM features. The maximum likelihood ratio is used to characterize the feature similarity between the real-time audio frame features and the historical audio frame features.
[0094] Step 407: When the maximum likelihood ratio is greater than the similarity threshold, determine that the first scene recognition result is that the non-silent scene has not changed.
[0095] In a possible implementation manner, a similarity threshold is preset in the terminal. When the above maximum likelihood ratio is greater than the similarity threshold, it indicates that the real-time audio frame features and the historical audio frame features are highly likely to be similar, and the terminal determines that the first scene recognition result is that the non-silent scene has not changed.
[0096] Step 408: Update the historical GMM features based on the maximum likelihood gradient.
[0097] In the embodiment of the present application, although the first scene recognition result is that the non - silent scene has not changed, the terminal collects real - time environmental audio. To ensure the accuracy of the next scene recognition, the terminal updates the historical GMM features.
[0098] In a possible implementation, since the first scene recognition result is that the non - silent scene has not changed and the similarity between the real - time GMM features and the historical GMM features is relatively high, the historical GMM features are updated by the maximum likelihood gradient without re - extracting the real - time GMM features.
[0099] Step 409, update the scene duration.
[0100] When the first scene recognition result is that the non - silent scene has not changed, the terminal updates the scene duration, that is, extends the duration of the timer.
[0101] Since the scene duration is extended, the frame length when the terminal performs frame processing on the real - time environmental audio is extended, thereby reducing the frequency of the terminal's first - level scene recognition and second - level scene recognition, and thus reducing the power consumption of the terminal.
[0102] Step 410, when the maximum likelihood ratio corresponding to m consecutive audio frames is less than the similarity threshold, determine that the first scene recognition result is that the non - silent scene has changed, where m is a positive integer.
[0103] In a possible implementation, when the above - mentioned maximum likelihood ratio is less than the similarity threshold, it indicates that the non - silent scene may have changed. To improve the accuracy of the first scene recognition result, the terminal obtains m consecutive audio frames and performs statistical decision based on the maximum likelihood ratio of the m consecutive audio frames. That is, when the maximum likelihood ratio corresponding to m consecutive audio frames is less than the similarity threshold, it is determined that the current scene is a non - silent scene that has changed.
[0104] Optionally, the method of statistical decision can be minimum misjudgment probability criterion decision, minimum loss criterion decision, minimum - maximum loss criterion decision, N - P (Neyman - Pearson) decision, etc. The embodiment of the present application does not limit this.
[0105] Exemplarily, when the maximum likelihood ratio corresponding to 5 consecutive audio frames is less than the similarity threshold, the terminal determines that the first scene recognition result is that the non - silent scene has changed.
[0106] Step 411, replace the historical GMM features with the real - time GMM features.
[0107] In the embodiment of the present application, since the first scene recognition result is that the non - silent scene has changed, the historical GMM features are replaced with the real - time GMM features.
[0108] Step 412, reset the scene duration.
[0109] In the embodiment of the present application, since the first scene recognition result indicates a change in the non - silent scene, in one possible implementation, the terminal clears the timer and resets the scene duration.
[0110] When the first scene recognition result indicates a change in the non - silent scene, resetting the scene duration shortens the scene duration. At this time, when the terminal performs frame - by - frame processing on the real - time environmental audio, the frame length is shortened, increasing the frequency of the terminal's first - level scene recognition and second - level scene recognition, and improving the accuracy of the terminal's scene recognition.
[0111] Step 413, in the case where the first scene recognition result indicates a change in the non - silent scene, perform second - level scene recognition based on the real - time environmental audio to obtain a second scene recognition result. The second scene recognition result is used to indicate the scene type of the non - silent scene, and the computational complexity of the second - level scene recognition is higher than that of the first - level scene recognition.
[0112] For the implementation manner of this step, please refer to step 230, and the embodiment of the present application will not elaborate on this again.
[0113] In the embodiment of the present application, the terminal adaptively updates the frame length when performing frame - by - frame processing on the real - time environmental audio according to the scene duration, that is, the longer the scene duration, the longer the frame length, and the shorter the scene duration, the shorter the frame length. Furthermore, the frequency of the terminal's first - level scene recognition and second - level scene recognition is adjusted, reducing the power consumption of the terminal while ensuring the accuracy of scene recognition.
[0114] Please refer to Figure 5 , which shows a flowchart of a scene recognition method provided by an exemplary embodiment of the present application. The method includes:
[0115] Step 501, the terminal obtains real - time environmental audio.
[0116] Step 502, the terminal judges the adaptive window length, that is, determines the frame length according to the scene duration.
[0117] Step 503, the terminal performs frame - by - frame processing on the real - time environmental audio to obtain audio frames.
[0118] Step 504, the terminal judges whether the energy of the audio frame is lower than the energy threshold. If so, execute step 505; if not, execute step 506.
[0119] Step 505, mute frame statistical decision, that is, the terminal determines whether the energy of consecutive n audio frames is lower than the energy threshold. If so, it determines that the first scene recognition result is a mute scene. If not, it determines that the first scene recognition result is that the non - mute scene has not changed.
[0120] Step 506, the terminal extracts the real - time GMM features of the real - time environmental audio.
[0121] Step 507, the terminal calculates the feature similarity (maximum likelihood ratio) based on the real - time GMM features and the historical GMM features in Step 508.
[0122] Step 508, the terminal extracts the historical GMM features and determines the scene duration.
[0123] Step 509, the terminal determines whether the feature similarity is greater than the similarity threshold. If so, the first scene recognition result is that the non - mute scene has not changed. If not, it executes Step 510.
[0124] Step 510, the terminal makes a statistical decision based on the audio frame similarity, that is, determines whether the feature similarity corresponding to consecutive m audio frames is less than the similarity threshold. If so, it means that the first scene recognition result is that the non - mute scene has changed. If not, the first scene recognition result is that the non - mute scene has not changed.
[0125] Step 511, when the first scene recognition result is that the non - mute scene has not changed, the terminal updates the historical GMM features by maximum likelihood gradient derivation.
[0126] Step 512, when the first scene recognition result is that the non - mute scene has changed, the terminal performs a second - level scene recognition.
[0127] Step 513, when the result of the first scene recognition is that the non - mute scene has changed, the terminal determines a new GMM model.
[0128] Step 514, the terminal updates the historical GMM features.
[0129] The following is an embodiment of the apparatus of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the embodiment of the apparatus of the present application, please refer to the method embodiment of the present application.
[0130] Please refer to Figure 6 , which shows a structural block diagram of a scene recognition device provided by an exemplary embodiment of the present application. The device includes:
[0131] An acquisition module 601, configured to acquire real - time environmental audio;
[0132] The first-level scene recognition module 602 is configured to perform first-level scene recognition based on the real-time environmental audio to obtain a first scene recognition result, where the first scene recognition result is used to indicate the change situation of a silent scene or a non-silent scene;
[0133] The second-level scene recognition module 603 is configured to perform second-level scene recognition based on the real-time environmental audio to obtain a second scene recognition result when the first scene recognition result indicates that the non-silent scene has changed, where the second scene recognition result is used to indicate the scene type of the non-silent scene, and the computational complexity of the second-level scene recognition is higher than that of the first-level scene recognition.
[0134] Optionally, the first-level scene recognition module 602 includes:
[0135] A frame segmentation processing unit configured to perform frame segmentation processing on the real-time environmental audio to obtain audio frames;
[0136] A first determination unit configured to determine that the first scene recognition result is a silent scene when the energy of consecutive n audio frames is lower than an energy threshold, where n is a positive integer;
[0137] A feature extraction unit configured to perform feature extraction on the audio frame to obtain real-time audio frame features when the energy of the audio frame is higher than the energy threshold;
[0138] A second determination unit configured to determine the first scene recognition result based on the real-time audio frame features and historical audio frame features, where the historical audio frame features are obtained by performing scene recognition on historical environmental audio.
[0139] Optionally, the second determination unit is configured to:
[0140] Determine the feature similarity between the real-time audio frame features and the historical audio frame features;
[0141] Determine that the first scene recognition result is that the non-silent scene has not changed when the feature similarity is greater than a similarity threshold;
[0142] Determine that the first scene recognition result is that the non-silent scene has changed when the feature similarity corresponding to consecutive m audio frames is less than the similarity threshold, where m is a positive integer.
[0143] Optionally, the real-time audio frame features are real-time GMM features, the historical audio frame features are historical GMM features, and the feature similarity is the maximum likelihood ratio between the real-time GMM features and the historical GMM features.
[0144] Optionally, the device further includes:
[0145] A first update module, configured to update the historical GMM features based on the maximum likelihood gradient;
[0146] The apparatus further includes:
[0147] A replacement module, configured to replace the historical GMM features with the real-time GMM features.
[0148] Optionally, the frame processing unit is configured to:
[0149] Determine a frame length based on a scene duration, where the scene duration is the duration of the current scene, and the frame length is positively correlated with the scene duration;
[0150] Perform frame processing on the real-time environmental audio based on the frame length to obtain the audio frames, where the length of the audio frames is the frame length.
[0151] Optionally, the apparatus further includes:
[0152] A second update module, configured to update the scene duration;
[0153] The apparatus further includes:
[0154] A reset module, configured to reset the scene duration.
[0155] Optionally, the first-level scene recognition is performed by an MCU or a DSP, and the second scene recognition is performed by an NPU.
[0156] Optionally, the apparatus further includes:
[0157] A setting module, configured to set noise reduction parameters based on the first scene recognition result or the second scene recognition result, where the noise reduction parameters corresponding to different scenes are different.
[0158] In summary, in the embodiments of the present application, the terminal first performs first-level scene recognition based on the real-time environmental audio. When it is determined that the non-silent scene changes, the terminal then performs second-level scene recognition based on the real-time environmental audio to determine the specific scene type. Since the computational complexity of the second-level scene recognition is higher than that of the first-level scene recognition, in the embodiments of the present application, the terminal first performs pre-recognition through the first-level scene recognition with lower computational complexity, avoiding directly performing second-level scene recognition based on the real-time environmental audio by the terminal, thereby reducing the power consumption of the terminal during scene recognition.
[0159] Please refer to Figure 7, which shows a structural block diagram of a terminal 700 provided by an exemplary embodiment of the present application. The terminal 700 in the present application may include one or more of the following components: a processor 710, a memory 720, and a microphone 730. Among them, the processor 710 is electrically connected to the memory 720 and the microphone 730 respectively.
[0160] The processor 710 may include one or more processing cores. The processor 710 uses various interfaces and lines to connect various parts within the entire terminal 700, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 720, and by calling data stored in the memory 720, it executes various functions of the terminal 700 and processes data. In the embodiments of the present application, the processor 710 includes an MCU or a DSP and an NPU, etc. Among them, the MCU or DSP is used for the first-level scene recognition, and the NPU is used for the second-level scene recognition. In addition, it should be noted that the processor 710 may also include other hardware. For example, the processor 710 may integrate one or several combinations of a CPU, a Graphics Processing Unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, the user interface, and application programs, etc.; the GPU is responsible for the rendering and drawing of the content to be displayed on the touch display screen; the modem is used for processing wireless communication. It can be understood that the above modem may not be integrated into the processor 710 and may be implemented separately through a communication chip.
[0161] The memory 720 may include a Random Access Memory (RAM), and may also include a Read-Only Memory (ROM). Optionally, the memory 720 includes a non-transitory computer-readable storage medium. The memory 720 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 720 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above various method embodiments, etc. The data storage area may also store data created during the use of the terminal 700 (such as a phone book, audio and video data, chat record data) or ambient audio collected by the microphone 730.
[0162] The microphone 730 is used to collect real-time ambient audio and store the collected real-time ambient audio in the memory 720 for the processor 710 to read and process. Optionally, the terminal 700 is provided with multiple microphones 730. Among them, the number of microphones 730 can be 3, 5, 6, etc., and the embodiments of the present application do not limit this. Optionally, the terminal 700 is provided with microphones 730 at different directional positions for accurately collecting real-time ambient audio.
[0163] In the embodiments of the present application, at least one instruction is stored in the memory 720, and the at least one instruction is used to be executed by the processor 710 to execute the scene recognition method as shown in the above embodiments.
[0164] In addition, those skilled in the art can understand that the structure of the terminal 700 shown in the above drawings does not constitute a limitation on the terminal 700. The terminal may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements. For example, the terminal 700 further includes components such as a radio frequency circuit, a shooting component, a sensor (excluding the temperature sensor), an audio circuit, a Wireless Fidelity (WiFi) component, a power supply, and a Bluetooth component, which will not be elaborated here.
[0165] The embodiments of the present application also provide a computer-readable storage medium, which stores at least one program code, and the program code is loaded and executed by the processor to implement the scene recognition method described in each of the above embodiments.
[0166] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the scene recognition method provided in various optional implementation manners of the above aspect.
[0167] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other implementations of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0168] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A scene recognition method, characterized in that, The method includes: Obtaining real-time environmental audio; Based on the MCU or DSP, determining the frame length according to the scene duration, where the scene duration is the duration of the current scene, and the frame length has a positive correlation with the scene duration; Based on the MCU or DSP, performing frame splitting on the real-time environmental audio according to the frame length to obtain audio frames, and the length of the audio frames is the frame length; When the energy of consecutive n audio frames is lower than the energy threshold, determining that the first scene recognition result is a silent scene, where n is a positive integer; When the energy of the audio frame is higher than the energy threshold, based on the MCU or DSP, extracting features from the audio frame to obtain real-time audio frame features; Based on the MCU or DSP, determining the feature similarity between the real-time audio frame features and the historical audio frame features, where the real-time audio frame features are real-time GMM features, the historical audio frame features are historical GMM features, the feature similarity is the maximum likelihood ratio between the real-time GMM features and the historical GMM features, and the historical audio frame features are extracted when performing scene recognition on historical environmental audio; When the feature similarity is greater than the similarity threshold, determining that the first scene recognition result is that the non-silent scene has not changed, and updating the scene duration; When the feature similarity corresponding to consecutive m audio frames is less than the similarity threshold, determining that the first scene recognition result is that the non-silent scene has changed, and resetting the scene duration, where m is a positive integer; When the first scene recognition result indicates that the non-silent scene has changed, based on the NPU, performing second scene recognition on the real-time environmental audio according to the scene recognition model to obtain a second scene recognition result, where the second scene recognition result is used to indicate the scene type of the non-silent scene, and the computational complexity of the second scene recognition is higher than that of the first scene recognition.
2. The method according to claim 1, wherein After determining that the first scene recognition result is that the non-silent scene has not changed when the feature similarity is greater than the similarity threshold, the method further includes: Updating the historical GMM features based on the maximum likelihood gradient; After determining that the first scene recognition result is that the non-silent scene has changed when the feature similarity corresponding to consecutive m audio frames is less than the similarity threshold, the method further includes: Replacing the historical GMM features with the real-time GMM features.
3. The method according to any one of claims 1 to 2, characterized in that, The method further includes: Setting noise reduction parameters based on the first scene recognition result or the second scene recognition result, where different scenes correspond to different noise reduction parameters.
4. A scene recognition device, characterized in that, The device includes: An acquisition module, configured to acquire real-time environmental audio; A first scene recognition module, configured to determine the frame length based on the MCU or DSP according to the scene duration, where the scene duration is the duration of the current scene, and the frame length has a positive correlation with the scene duration; Through the MCU or DSP, frame the real-time environmental audio based on the frame length to obtain audio frames, and the length of the audio frames is the frame length; In the case where the energy of consecutive n audio frames is lower than the energy threshold, determine that the first scene recognition result is a silent scene, where n is a positive integer; In the case where the energy of the audio frame is higher than the energy threshold, through the MCU or DSP, extract features from the audio frame to obtain real-time audio frame features; Through the MCU or DSP, determine the feature similarity between the real-time audio frame features and the historical audio frame features. The real-time audio frame features are real-time GMM features, the historical audio frame features are historical GMM features, the feature similarity is the maximum likelihood ratio between the real-time GMM features and the historical GMM features, and the historical audio frame features are extracted when performing scene recognition on historical environmental audio; In the case where the feature similarity is greater than the similarity threshold, determine that the first scene recognition result is that the non-silent scene has not changed, and update the scene duration; In the case where the feature similarity corresponding to consecutive m audio frames is less than the similarity threshold, determine that the first scene recognition result is that the non-silent scene has changed, and reset the scene duration, where m is a positive integer; A second scene recognition module, configured to, in the case where the first scene recognition result indicates that the non-silent scene has changed, through the NPU, perform second scene recognition on the real-time environmental audio based on a scene recognition model to obtain a second scene recognition result, where the second scene recognition result is used to indicate the scene type of the non-silent scene, and the computational complexity of the second scene recognition is higher than that of the first scene recognition.
5. A terminal, characterized in that, The terminal includes a processor and a memory; at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the scene recognition method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set, or an instruction set is stored in the readable storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the scene recognition method according to any one of claims 1 to 3.
7. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions are executed by a processor, the scene recognition method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Scene recognition method and device
CN110336943A
Volume adjustment method and device of mobile terminal, mobile terminal and storage medium
CN110995933A
Audio processing method and device
CN112543972A
Sound scene recognition method and device, equipment and storage medium
CN112750448A