Audio data processing method and related apparatus
By identifying scenes in audio data processing and combining them with scene-specific event recognition models, the problem of high false recognition rate in audio data event recognition is solved, and the accuracy of recognition is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2022-11-11
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies have a high false recognition rate when using audio data for event recognition, making it difficult to accurately identify events in the environment of hearing-impaired users.
By identifying the scene in which the audio data is located, event recognition is performed using a scene-specific event recognition model. Then, by combining scene probability and model probability, reselection is carried out, and candidate events that do not belong to the scene are eliminated.
It improves the accuracy of audio data event recognition, reduces the false recognition rate, and ensures that the recognition results are more consistent with the actual situation of the user's environment.
Smart Images

Figure CN118042042B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to audio data processing methods and related apparatus. Background Technology
[0002] With the continuous development of computer science and technology, electronic devices such as mobile phones are playing an increasingly important role in people's daily lives. For example, for people with hearing impairments, who cannot hear or clearly hear sounds such as ambient sounds and speech, daily life becomes inconvenient.
[0003] Electronic devices can collect audio data from the environment, input the audio data into an event recognition model for event recognition, and then feed back the event recognition results to hearing-impaired users in a non-voice manner. This can greatly facilitate hearing-impaired users in knowing about events happening in the environment and improve their quality of life.
[0004] However, when using the above method to perform event recognition on audio data, the false recognition rate is high. Summary of the Invention
[0005] This application provides an audio data processing method and related apparatus, which can improve the accuracy of event identification from audio data.
[0006] In a first aspect, embodiments of this application provide an audio data processing method, including:
[0007] Acquire audio data to be processed and identify a first scene, wherein the first scene is the scene in which the sound source that generated the audio data to be processed is located;
[0008] The audio data to be processed is input into the event recognition model for event recognition, resulting in multiple candidate events.
[0009] Among the multiple candidate events mentioned above, the candidate events belonging to the first scenario mentioned above are taken as the event recognition results of the audio data to be processed.
[0010] In this embodiment, the audio data to be processed can be understood as audio data for which an event needs to be identified. In other solutions, after acquiring the audio data to be processed, the electronic device directly inputs the audio data to be processed into an event recognition model for event recognition to obtain the event corresponding to the audio data to be processed. However, in this solution, the electronic device determines the event corresponding to the audio data to be processed based on a first scenario, that is, the electronic device first identifies the first scenario before determining the event corresponding to the audio data to be processed.
[0011] In this embodiment, the first scenario described above can be understood as the key scenario in the following embodiments, namely, the scenario in which the sound source that generates the audio data to be processed, as identified by the electronic device, is located. Here, the sound source can be understood as an object that generates sound, such as a television playing a movie, a laptop playing music, or a barcode scanner that successfully scans a code. In this embodiment, the sound source that generates the audio data to be processed should be understood as a sound source that generates sound, wherein the sound, after processing, becomes the audio data to be processed.
[0012] For ease of understanding, the first scenario described above can be, for example, a home scenario (or living scenario), an office scenario, a subway scenario, etc. As another example, the audio data to be processed can be audio data obtained after processing at least one of the following sounds: the sound of a washing machine operating, the notification sound of turning the air conditioner on or off, and the notification sound output by a microwave oven after heating food.
[0013] In this embodiment, the method for acquiring the audio data to be processed can be referred to the description of step 501 in the following embodiments. The processing procedure for the electronic device to process the audio signal after it has acquired the signal can be referred to the following... Figure 6 The relevant descriptions will not be repeated here. In this embodiment, the audio data to be processed can be understood as one or more frames of audio data, or as a segment of audio data.
[0014] In this embodiment, the order of acquiring the audio data to be processed and recognizing the first scene is not limited. For example, the electronic device may acquire the audio data to be processed first and then recognize the first scene; or the electronic device may recognize the first scene first and then acquire the audio data to be processed. In the second method, the electronic device may periodically recognize scenes and use the latest scene recognition result before the time of acquiring the audio data to be processed as the first scene.
[0015] Understandably, in reality, an event might occur suddenly, such as the air conditioner being turned on abruptly or the doorbell ringing unexpectedly. However, the scenarios in which these events occur are generally not abrupt, because the scenarios themselves have a certain geographical range, and the user's movement does not cause the scenarios to change frequently in a short period of time. Therefore, whether the audio data to be processed is acquired first and then the scenario is identified, or the scenario is identified first and then the audio data to be processed is acquired, the identified scenario can be considered to be the scenario where the sound source that generated the audio data to be processed is located, i.e., the first scenario mentioned above.
[0016] In this embodiment, the event recognition model can be a neural network model, such as a convolutional neural network model, a deep neural network model, or a recurrent neural network model, etc., and this application does not limit this. It is understood that after inputting the audio data to be processed into the event recognition model, the event recognition model can obtain multiple candidate events and the recognition probability of each candidate event. For example, after inputting the audio data to be processed into the event recognition model, the event recognition model's recognition result indicates a 10% probability that the event corresponding to the audio data to be processed is event A, a 15% probability that the event corresponding to the audio data to be processed is event B, and a 3% probability that the event corresponding to the audio data to be processed is event C.
[0017] Although the event recognition model described above can generate multiple candidate events, in this embodiment, the candidate events belonging to the first scenario are taken as the event recognition results of the audio data to be processed. This application improves the accuracy of event recognition from audio data by limiting the scope of the scenario and eliminating events that will not occur or have a low probability of occurring in the first scenario.
[0018] Optionally, the above-mentioned candidate events can be understood as reference events in the following embodiments, for example... Figure 8 Reference events in the illustrated embodiments.
[0019] It is understandable that, under certain special circumstances, there may be multiple events that may occur in a scene, and the sounds produced by different types of events may also be similar. For example, in a home environment, the prompt sound of turning on the air conditioner may be similar to the prompt sound of turning on the television. Therefore, there may be multiple candidate events belonging to the first scene mentioned above. Optionally, the candidate event with the highest recognition probability can be taken as the event recognition result of the audio data to be processed.
[0020] In conjunction with the first aspect, in one possible implementation, the candidate events belonging to the first scenario among the multiple candidate events are taken as the event recognition results of the audio data to be processed, including:
[0021] Obtain the first probability of each of the above candidate events occurring in the first scenario. The first probability is obtained by counting the number of times each of the above candidate events occurs in the first scenario.
[0022] Among the multiple candidate events, the candidate event with the largest result of the calculation of the first probability and the second probability is taken as the event recognition result of the audio data to be processed. The second probability is the recognition probability of each of the candidate events obtained by the event recognition model in the audio data to be processed.
[0023] In this embodiment, the first probability and the second probability have different sources. The first probability is obtained by statistically analyzing the number of times each candidate event occurs in the first scenario, while the second probability is the recognition probability of each candidate event obtained by the event recognition model from the audio data to be processed. In other words, the first probability is obtained by statistically analyzing a large number of scenarios and the events occurring in those scenarios, while the second probability is the probability that the event recognition model considers the audio data to be processed to be a certain event.
[0024] For example, the first probability mentioned above can be obtained by statistically analyzing the occurrence of events in different regions. Taking a shopping mall scenario as an example, one can statistically analyze whether there are escalators, microwave ovens, and barcode scanners in shopping malls in region A, and obtain the probability of escalator sounds, microwave oven sounds, and barcode scanners occurring in the shopping mall scenario. For example, statistically analyzing the occurrence of a certain event in each scenario can yield the following probabilities: Figure 10 The probability distribution shown is shown.
[0025] Alternatively, the aforementioned first probability can be understood as the following text. Figure 8 The second reference probability in the illustrated embodiment can be understood as follows: Figure 8 The first reference probability in the illustrated embodiment.
[0026] In this embodiment, the operation between the first probability and the second probability can be determined according to the actual situation. For example, it can be a direct multiplication, or a multiplication followed by normalization. Taking multiplication as an example, the product of the first probability and the second probability can be understood as follows: Figure 8 The third reference probability in the illustrated embodiment.
[0027] In this embodiment, by reselecting the multiple candidate events based on the first probability of each candidate event occurring in the first scenario, the candidate events belonging to the first scenario can become the event recognition results of the audio data to be processed, thereby improving the accuracy of event recognition.
[0028] To save space, further details of this embodiment can be found in the following text. Figure 8 , Figure 9 , Figure 10 as well as Figure 11 The relevant descriptions will not be repeated here.
[0029] In conjunction with the first aspect, in one possible implementation, the aforementioned event recognition model is a model obtained by training an event recognition model to be trained using audio sample data collected in the aforementioned first scenario; the aforementioned selection of candidate events belonging to the aforementioned first scenario from among the multiple candidate events as the event recognition result of the aforementioned audio data to be processed includes:
[0030] Among the multiple candidate events, the candidate event with the highest recognition probability obtained by the event recognition model from the audio data to be processed is taken as the event recognition result of the audio data to be processed.
[0031] In this embodiment, the event recognition model for performing event recognition on the aforementioned audio data to be processed is a model trained by collecting audio sample data from the aforementioned first scenario. Optionally, in this embodiment, the aforementioned event recognition model may be referred to as the event recognition model corresponding to the aforementioned first scenario, or it may be understood as the model described below. Figure 5 The event recognition model corresponding to the key scenarios in the illustrated embodiment.
[0032] It is understandable that after training the event recognition model by collecting audio sample data from the first scenario, the event recognition model obtained after training will have changes in model parameters compared to the event recognition model before training. This can improve the accuracy of the event recognition model in recognizing events from the audio data in the first scenario, that is, improve the accuracy of recognizing events from the audio data to be processed.
[0033] In conjunction with the first aspect, in one possible implementation, the time interval between the moment when the aforementioned audio data to be processed is acquired and the moment when the aforementioned first scene is identified is less than or equal to a first threshold.
[0034] It is understood that the aforementioned audio data to be processed can be collected by the executing entity itself or obtained from other devices. In this embodiment, the aforementioned first threshold can be determined according to the actual situation, and this application does not limit it. For example, when the aforementioned audio data to be processed is 5 seconds long, the aforementioned first threshold can be any non-zero value less than 10 seconds.
[0035] In this embodiment, the time interval between acquiring the audio data to be processed and identifying the first scene is less than or equal to a first threshold, which can prevent the scene where the sound source of the audio data to be processed is located from not matching the first scene. It is understood that if the time interval between acquiring the audio data to be processed and identifying the first scene is too long, the first scene identified by the electronic device may no longer be the scene where the sound is located.
[0036] In conjunction with the first aspect, in one possible implementation, the identification of the first scenario includes:
[0037] If at least one image is captured by a camera, the above-mentioned at least one image is input into the trained scene recognition model to obtain the above-mentioned first scene; the above-mentioned trained scene recognition model is obtained by training sample images in multiple scenes, including the above-mentioned first scene.
[0038] In one possible implementation, the electronic device can perform scene recognition using at least one image captured by a camera. For example, the trained scene recognition model can be a neural network model, such as a convolutional neural network model, a deep neural network model, or a recurrent neural network model, etc., and this application does not limit this to any particular type.
[0039] Understandably, in the process of training the scene recognition model to obtain the trained scene recognition model, sample images from various scenes can be collected for training. These various scenes can be everyday scenes in life, such as home scenes, office scenes, bus scenes, subway scenes, high-speed rail scenes, airport scenes, shopping mall scenes, coffee shop scenes, and library scenes, etc.
[0040] Understandably, scene recognition based on scene recognition models has a high accuracy rate when at least one image is captured by a camera.
[0041] In conjunction with the first aspect, in one possible implementation, at least one of the aforementioned images is input into a trained scene recognition model to obtain the aforementioned first scene, including:
[0042] Input at least one of the above images into the trained scene recognition model to obtain multiple candidate scenes;
[0043] The scenario that matches at least one of the following data among the multiple candidate scenarios is designated as the first scenario.
[0044] It is understandable that, after inputting at least one of the aforementioned images into the trained scene recognition model for scene recognition, similar to the event recognition model, multiple candidate scenes and their corresponding recognition probabilities can be obtained. For example, after inputting at least one of the aforementioned images into the trained scene recognition model, the model's recognition result indicates a 30% probability that the scene corresponding to the at least one image is scene A, an 18% probability that the scene corresponding to the at least one image is scene B, and an 8% probability that the scene corresponding to the at least one image is scene C.
[0045] In this embodiment, given the aforementioned multiple candidate scenarios, the electronic device can combine other data to determine the scenario. For example, if the electronic device obtains a home address as its location information and a home Wi-Fi network connection, then even if the probability of identifying the home scenario among the multiple candidate scenarios is low, the first scenario can still be considered a home scenario.
[0046] For example, after scene recognition using the trained scene recognition model, the recognition probability of a subway scene is not significantly different from that of a high-speed rail scene. Therefore, before the train departs, the location information can be used to determine whether it is a high-speed rail scene or a subway scene. After the train departs, the speed of the electronic device can be obtained using an accelerometer. Since the speed of a high-speed rail is greater than that of a subway, the speed of the electronic device can be used to determine whether the first scene is a high-speed rail scene or a subway scene.
[0047] In conjunction with the first aspect, in one possible implementation, the above method also includes:
[0048] In the absence of capturing at least one of the above images via a camera, the first scenario is determined based on at least one of the following data: the location information of the electronic device, the network connection object of the electronic device, and the moving speed of the electronic device.
[0049] It is understandable that the camera may be obstructed, preventing the electronic device from capturing images through the camera. Therefore, in the absence of capturing at least one image through the camera, the first scenario is determined based on at least one of the following data: the location information of the electronic device, the network connection object of the electronic device, and the movement speed of the electronic device.
[0050] For example, when a user arrives at a subway station, their workplace, their home, or a shopping mall, the aforementioned first scenario can be determined using the acquired location information. As another example, if an electronic device connects to home Wi-Fi, the aforementioned first scenario can be considered a home scenario; if an electronic device connects to workplace Wi-Fi, the aforementioned first scenario can be considered a workplace scenario.
[0051] Secondly, embodiments of this application provide an audio data processing apparatus, including:
[0052] The acquisition unit is used to acquire the audio data to be processed.
[0053] The recognition unit is used to recognize the first scene, wherein the audio data to be processed is the audio data generated in the first scene.
[0054] The aforementioned recognition unit is also used to input the aforementioned audio data to be processed into an event recognition model for event recognition, thereby obtaining multiple candidate events;
[0055] The determining unit is used to identify the candidate events belonging to the first scenario among the multiple candidate events as the event recognition results of the audio data to be processed.
[0056] Optionally, in the embodiments of this application, the steps performed by the acquisition unit can be performed by a microphone or a communication module, wherein the communication module can be a mobile communication module or a wireless communication module; the steps performed by the identification unit and the determination unit can be performed by a processor.
[0057] Thirdly, embodiments of this application provide an electronic device, which includes a processor and a memory; the memory is coupled to the processor, and the memory is used to store computer program code, which includes computer instructions; the processor calls the computer instructions to execute the method in the first aspect or any possible implementation thereof.
[0058] Fourthly, embodiments of this application provide a chip including logic circuitry and an interface, wherein the logic circuitry and the interface are coupled; the interface is used to input and / or output code instructions, and the logic circuitry is used to execute the code instructions to cause the method in the first aspect or any possible implementation thereof to be executed.
[0059] Fifthly, embodiments of this application disclose a computer program product, which includes program instructions that, when executed by a processor, cause the method in the first aspect or any possible implementation thereof to be executed.
[0060] Sixthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed on a processor, causes the method in the first aspect or any possible implementation thereof to be performed. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are merely some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0062] Figure 1 This is a schematic diagram illustrating the occurrence of the same type of sound in multiple scenarios, as provided in an embodiment of this application.
[0063] Figure 2 This is a schematic diagram of a home scene provided in an embodiment of this application;
[0064] Figure 3 This is a schematic diagram illustrating event recognition based on audio signals, provided in an embodiment of this application.
[0065] Figure 4 This is a schematic diagram illustrating scene-based event recognition of audio signals according to an embodiment of this application;
[0066] Figure 5 This is a flowchart illustrating an audio data processing method provided in an embodiment of this application;
[0067] Figure 6 This is a schematic diagram of an audio feature extraction method provided in an embodiment of this application;
[0068] Figure 7 This is a schematic diagram illustrating an event recognition method that combines images and audio, as provided in an embodiment of this application.
[0069] Figure 8 This is a flowchart illustrating another audio data processing method provided in an embodiment of this application;
[0070] Figure 9 This is a schematic diagram showing the positions of first audio data and second audio data provided in an embodiment of this application;
[0071] Figure 10 This is a schematic diagram of a probability distribution provided in an embodiment of this application;
[0072] Figure 11 This is a schematic diagram illustrating event recognition based on probability distribution, provided in an embodiment of this application.
[0073] Figure 12 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application;
[0074] Figure 13 This is a software structure block diagram of an electronic device 100 provided in an embodiment of this application. Detailed Implementation
[0075] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0076] The terms "first" and "second," etc., used in the specification, claims, and drawings of this application are used to distinguish different objects, rather than to describe a specific order.
[0077] Hearing-impaired users can be understood as those who have difficulty hearing in both ears and cannot hear or clearly hear sounds such as ambient sounds and speech. It's understandable that hearing-impaired users often cannot effectively perceive events in their environment through sound. For example, after a washing machine finishes washing clothes, it usually emits a beeping sound to remind the user that the washing task is complete and to hang the clothes to dry. However, hearing-impaired users cannot respond effectively to these beeping sounds in a timely manner and cannot promptly know that the washing is finished.
[0078] Besides hearing-impaired users being unable to effectively perceive events in their environment, in real life, even hearing-free users may be unable to effectively perceive events in their environment through sound for a period of time due to non-physiological factors. For example, when users wear headphones (especially noise-canceling headphones) to listen to music or watch videos, they almost completely block out external sounds, resulting in an inability to effectively perceive events in their environment through sound during the time they wear the headphones.
[0079] In this embodiment, users who are unable to effectively perceive events in their environment through sound due to physiological or non-physiological factors are collectively referred to as users with hearing impairments. To help users with hearing impairments become aware of events in their environment, electronic devices can collect audio signals from the environment, input these signals into an event recognition model for event recognition, and then feed back the event recognition results to the user with hearing impairments in a non-speech manner. This method greatly facilitates users with hearing impairments in becoming aware of events in their environment, especially those with hearing impairments. Optionally, identifying corresponding events based on audio data can also be called sound event detection.
[0080] In the embodiments of this application, the event recognition model described above may be a convolutional neural network (CNN), a deep neural network (DNN), or a recurrent neural network (RNN), etc., and this application does not limit it to any particular type.
[0081] For example, the aforementioned non-voice methods can be to output text prompts through pop-ups, floating windows, etc. Optionally, vibration can also be used to enhance the prompting effect.
[0082] Understandably, while event recognition models can identify corresponding events based on input audio signals, many sounds in real life are similar, resulting in low accuracy for electronic devices in acquiring audio signals for event recognition.
[0083] For ease of understanding, please refer to the example provided. Figure 1 , Figure 1 This is a schematic diagram illustrating the occurrence of the same type of sound in multiple scenarios, as provided in an embodiment of this application.
[0084] For example, suppose an electronic device receives an audio signal, which is a "ding" sound. In daily life, a "ding" sound can occur in many situations, such as... Figure 1 In (a), the "ding" sound may come from swiping an access card; for example... Figure 1 In (b), the "ding" sound may come from the sound of a barcode scanner scanning a delivery slip; for example... Figure 1 In (c), the "ding" sound may be a notification sound indicating that the microwave oven has completed its task.
[0085] For ease of description, the sound of swiping the access card will be referred to as the access control sound; the sound of scanning a barcode will be referred to as the barcode scanner sound; and the sound of the microwave oven after completing its task will be referred to as the microwave oven sound.
[0086] Because the same type of sound can occur in various scenarios in daily life, electronic devices have a high rate of misidentification when identifying events based on audio signals. For example, an electronic device may misidentify the "ding" sound that should be from a microwave oven as the "ding" sound of swiping an access card, or as the "ding" sound of a barcode scanner.
[0087] To address the aforementioned problems, this application provides an audio data processing method and related apparatus. The method provided can be executed by an electronic device, which can be any electronic device capable of executing the technical solutions disclosed in the method embodiments of this application. Optionally, the electronic device can be any device capable of processing audio data, such as a mobile phone, tablet computer, wearable smart device, etc. It should be understood that the method embodiments of this application can also be implemented by a processor executing computer program code. This application can improve the accuracy of electronic devices in recognizing events based on audio signals.
[0088] For ease of understanding, please refer to the example provided. Figure 2 , Figure 3 as well as Figure 4 ,in, Figure 2 This is a schematic diagram of a home scene provided in an embodiment of this application. Figure 3 This is a schematic diagram illustrating event recognition based on audio signals, provided in an embodiment of this application. Figure 4 This is a schematic diagram illustrating scene-based event recognition of audio signals, provided in an embodiment of this application.
[0089] like Figure 2 The diagram shown can be understood as a home scene illustration, specifically a kitchen. For example, a microwave oven is placed in the kitchen. After the microwave oven finishes heating the food, it emits a "ding" sound to remind the user that the food heating task is complete. Figure 2 As shown in 201.
[0090] It is understandable that, such as Figure 2 After receiving an audio signal including a "ding" sound (hereinafter referred to as the audio signal "ding"), the electronic device 202 will perform event recognition based on the audio signal "ding," that is, identify the event corresponding to the audio signal "ding." For example, Figure 3 The proposed solution can be interpreted as other solutions. Figure 4 The solution shown can be understood as the solution provided in the embodiments of this application.
[0091] For example Figure 3 As shown, in other solutions, after receiving the aforementioned audio signal "ding," the electronic device inputs the audio signal "ding" into an event recognition model for event recognition, obtaining an event recognition result set. For example, the electronic device identifies the audio signal "ding" as having a 35% probability of being a barcode scanner sound, a 20% probability of being an access card sound, and a 15% probability of being a microwave oven sound. Since the barcode scanner sound has the highest probability, in other solutions, the electronic device ultimately considers the audio signal "ding" to be a barcode scanner sound.
[0092] In the solution provided in this application, for example Figure 4 As shown, after receiving the audio signal "ding," the electronic device inputs the audio signal "ding" into the event recognition model corresponding to the home scene to perform event recognition in the home scene, and obtains an event recognition result set. It can be understood that before performing event recognition in the home scene, the electronic device first performs home scene recognition, that is, determines that the electronic device is located in a home scene, and then uses the event recognition model corresponding to the home scene to perform event recognition on the audio signal.
[0093] For example, in this solution, the electronic device identifies the audio signal "ding" as having a 40% probability of being a microwave oven sound, a 10% probability of being an access card sound, and a 5% probability of being a barcode scanner sound. Since the microwave oven sound is the most likely, in this solution, the electronic device ultimately considers the audio signal "ding" to be a microwave oven sound.
[0094] Ultimately, in this plan, if Figure 2 As shown, the electronic device 202 can output the text "The microwave oven's current task has been completed" in the form of a pop-up window after recognizing the sound of the microwave oven, instead of misidentifying the above audio signal as the sound of the barcode scanner completing the scan as other solutions do.
[0095] The above provides an overall overview of this solution. The following section describes the specific process of the method provided in the embodiments of this application. For example, please refer to... Figure 5 , Figure 5 This is a flowchart illustrating an audio data processing method provided in an embodiment of this application. Figure 5 As shown, the above method includes:
[0096] 501: Acquire audio data to be processed and identify key scenes.
[0097] In this embodiment, the audio data to be processed can be understood as audio data from which events need to be identified. In one possible implementation, the electronic device may include a sound acquisition module. The sound acquisition module acquires sound to obtain the aforementioned audio data to be processed. Exemplarily, the sound acquisition module may be one or more microphones; when there are multiple microphones, the multiple microphones may be referred to as a microphone array or array microphone.
[0098] It is understandable that the sound captured by the microphone is an analog audio signal, which can be processed by sampling, quantization, and encoding to obtain the audio data to be processed.
[0099] In another possible implementation, the electronic device can communicate with other devices and use audio data obtained from other devices through the aforementioned communication connection as audio data to be processed.
[0100] In this step, the aforementioned key scenario can be understood as the scenario corresponding to the audio data to be processed identified by the electronic device. In this embodiment, the electronic device can identify the scenario in various ways. For example, the electronic device can identify the scenario through data collected by at least one sensor, wherein the data can be audio data collected by a sound sensor, image data collected by an image sensor, position data collected by a position sensor, and motion data collected by an accelerometer, etc.
[0101] In one possible implementation, the electronic device can identify the scene in real time based on sensor data. If the audio data to be processed is acquired at time A, the latest scene determined before time A is used as the key scene. It is understandable that since scenes correspond to a certain geographical area in reality, scenes are generally not prone to sudden changes. That is, once the electronic device determines a scene, that scene remains valid for a period of time. For example, if a user returns home from get off work, the electronic device determines the current scene as a home scene. Even if the user may go out again, the validity of the home scene will continue for a period of time. Therefore, although the time of determining the key scene is not completely synchronized with the time of acquiring the audio data to be processed, it can be considered that the audio data to be processed acquired at time A is audio data obtained from sounds collected from the key scene.
[0102] In another possible implementation, the electronic device can first acquire the audio data to be processed, and then retrieve sensor data to identify the scene. It is understandable that, similar to the previous description, scenes are generally not prone to sudden changes; therefore, the scene determined after the electronic device acquires the audio data to be processed can be understood as the aforementioned key scene.
[0103] In other words, although the time of identifying the key scene is not completely synchronized with the time of acquiring the audio data to be processed, it can be considered that the audio data to be processed is audio data obtained from the sound collected in the key scene.
[0104] 502: Input the audio data to be processed into the event recognition model corresponding to the key scene to perform event recognition, and obtain the event recognition result of the audio data to be processed; the event recognition model corresponding to the key scene is trained by collecting audio sample data in the key scene.
[0105] In this embodiment of the application, the event recognition model itself can be a neural network model, such as CNN, DNN and RNN. By collecting audio sample data from the key scene and training any of the above event recognition models, the event recognition model corresponding to the key scene can be obtained.
[0106] Understandably, after training the event recognition model using audio sample data collected in key scenarios, the event recognition model obtained after training will have changes in model parameters compared to the event recognition model before training. This can improve the accuracy of the event recognition model corresponding to the key scenario in recognizing events from the audio data to be processed in the key scenario.
[0107] For ease of understanding, by way of example, reuse Figure 3 and Figure 4Suppose the audio data to be processed corresponds to the "ding" sound output by a microwave oven when it completes its task. The electronic device uses another method to perform event recognition on the audio data to be processed. In the event recognition results, the probability of the microwave oven sound is 10%, that is, the electronic device has a 10% probability that the audio data to be processed corresponds to the microwave oven sound.
[0108] In this solution, the electronic device first identifies the key scenario as a home environment. Then, it performs event recognition based on the event recognition model corresponding to the home environment. In the resulting event recognition results, due to the limitation of the scenario's scope, the probability of a microwave oven sound increases from 10% to 40%. That is, in this solution, the electronic device has a 40% probability of recognizing the audio data to be processed as a microwave oven sound. (Comparison) Figure 3 and Figure 4 It can be seen that this solution can improve the accuracy of electronic devices in recognizing events from audio data.
[0109] In this embodiment of the application, for example, before the electronic device acquires the audio data to be processed and inputs it into the event recognition model corresponding to the key scene for event recognition, it may first perform audio feature extraction. Audio feature extraction can be understood as extracting identifiable components from the audio signal to facilitate subsequent event recognition by the event recognition model.
[0110] For ease of understanding, please refer to the example provided. Figure 6 , Figure 6 This is a schematic diagram of an audio feature extraction method provided in an embodiment of this application.
[0111] For example, an electronic device acquires an audio signal to be processed via a microphone. This audio signal is an analog audio signal. Then, the analog audio signal is converted into an electrical signal, and the electrical signal is sampled, quantized, and encoded to obtain a digital audio signal.
[0112] like Figure 6 As shown, the audio digital signal is first divided into frames to obtain audio digital signals in units of frames. For example, each frame is 85ms long, and assuming a sampling rate of 16kHz, then one frame of audio digital signal has 1360 (16000*0.085) sample points.
[0113] Then, each frame of the audio digital signal is windowed, that is, the window function is multiplied by each frame of the audio digital signal to form a windowed audio digital signal. The window function can be a Hamming window, a Hanning window, or a Blackman window, etc. Windowing can make the time-domain signal better meet the periodicity requirements of fast fourier transform (FFT) processing and reduce leakage.
[0114] Finally, Mel frequency cepstrum coefficients (MFCCs) are extracted from the windowed audio digital signal, and the audio feature vector is output. For example, an FFT can be performed on the windowed audio digital signal to obtain the corresponding spectrum, which is then passed through a Mel filter bank to obtain the Mel spectrum. Finally, cepstrum analysis is performed on the Mel spectrum to obtain the Mel frequency cepstrum coefficients (MFCCs), where the MFCCs can be understood as audio features.
[0115] Optionally, the electronic device can perform event recognition in one or more frames as processing units. For example, it can use 10 frames as processing units, that is, recognize one event for every 10 frames of audio data.
[0116] above Figure 6 The processing procedure shown can be either the process of an electronic device inputting the audio data to be processed into the event recognition model corresponding to the key scenario for event recognition, or the process of collecting audio sample data to train the event recognition model to obtain the trained event recognition model.
[0117] For example, common scenarios in daily life can be planned in advance, such as home scenarios, office scenarios, bus scenarios, subway scenarios, shopping mall scenarios, parking lot scenarios, etc. Then, audio sample data is collected from each scenario, and the audio sample data is input into the event recognition model for training, thus obtaining the event recognition model corresponding to each scenario. Multiple audio sample data can be collected for each scenario.
[0118] Understandably, before inputting audio sample data into the event recognition model for training, methods such as... Figure 6 The processing method shown preprocesses the audio sample data. That is, before inputting the audio sample data into the event recognition model for training, it sequentially performs frame segmentation, windowing, and MFCC feature extraction, and then inputs the obtained audio feature vector into the event recognition model for training.
[0119] It's understandable that different users have different life scenarios. For example, office workers mainly commute between work and home, and in their leisure time may visit shopping malls and other entertainment venues. Their daily transportation methods may include subways, buses, and private cars. Hearing-impaired users, on the other hand, may be less likely to frequent shopping malls and other entertainment venues. Therefore, with the user's authorization, electronic devices can further collect audio sample data based on scenarios encountered in the user's daily life to train the initial event recognition model and optimize the model's parameters to further improve the accuracy of event recognition. It should be understood that each time audio sample data is collected for model training, the audio sample data can be processed in various ways, such as... Figure 6 The preprocessing shown.
[0120] The above describes the content related to the event recognition model in the embodiments of this application. Next, the scene recognition in the embodiments of this application will be introduced.
[0121] In one possible implementation, the electronic device can acquire at least one image, perform image recognition on that at least one image, and thereby identify the scene. Optionally, the above method can also be called image scene recognition.
[0122] For example, deep learning technology can be used for image scene recognition. For example, at least one of the above-mentioned images can be input into a deep learning model to achieve scene recognition. The deep learning model can be Places-CNN, DeCAF network model, or multi-resolution CNN, etc., and this application does not limit it to any particular model. For ease of understanding, in the embodiments of this application, the model used for image scene recognition is referred to as an image scene recognition model.
[0123] Therefore, in some embodiments, such as Figure 7 As shown, Figure 7 This is a schematic diagram illustrating an event recognition method that combines images and audio, as provided in an embodiment of this application.
[0124] For example, the electronic device may include a camera that captures at least one image. Alternatively, the electronic device may establish a communication connection with other devices and use the image acquired through the communication connection as the at least one image. For instance, if a user has a home security camera installed in their home, their electronic device can establish a communication connection with the camera while the user is at home, and receive images captured by the camera through this connection. Therefore, optionally, the at least one image may be at least one image from a video stream.
[0125] In this embodiment, the electronic device can input at least one of the aforementioned images into an image scene recognition model for scene recognition, thereby obtaining the scene corresponding to the at least one image. For ease of description and understanding, the scene corresponding to the at least one image is referred to as a key scene. Therefore, the above process can be called key scene recognition.
[0126] After obtaining the key scenarios, such as Figure 7 As shown, the acquired audio data to be processed is input into the event recognition model corresponding to the key scene for event recognition, thereby obtaining the event corresponding to the audio data to be processed. The audio data to be processed can be obtained through a microphone in an electronic device. The event recognition model corresponding to the key scene can be understood as an event recognition model trained based on audio sample data from the key scene; for details, please refer to the preceding text. Figure 6 The relevant information will not be repeated here.
[0127] It is understood that the embodiments of this application are not limited to such Figure 7 The order in which at least one image and the audio data to be processed are acquired can be varied. Alternatively, the at least one image can be acquired first for key scene recognition to obtain the key scene, and then the audio data to be processed can be acquired for event recognition.
[0128] It is understandable that during the use of electronic devices, it may be impossible to acquire at least one of the aforementioned images for key scene recognition, such as when the user is outdoors or the camera of the electronic device is obstructed. In this embodiment, in addition to scene recognition based on images, scene recognition can also be performed using other sensor data.
[0129] For example, location information of electronic devices can be obtained through location sensors, and the scene can be identified through the location information. For instance, when a user arrives at a subway station, or arrives at their office, or arrives at their home, or arrives at a shopping mall, the corresponding scene can be determined through the location information obtained by the location sensor.
[0130] It is understandable that electronic devices can periodically acquire location information for scene recognition, or they can acquire location information for scene recognition after acquiring the audio data to be processed.
[0131] For example, scene recognition can be performed through the network connection of a wireless network sensor, where the network connection can be a Wi-Fi connection. For instance, an electronic device connected to home Wi-Fi can be considered to be in a home scene, and an electronic device connected to office Wi-Fi can be considered to be in a work scene.
[0132] For example, scene recognition can be performed using the speed obtained from an accelerometer. For instance, if a user is on a high-speed train, the speed of the high-speed train is significantly different from the operating speed of buses, private cars, and subways in daily life. If the speed obtained from the accelerometer is within the operating speed range of the high-speed train, the current scene can be considered as a high-speed train.
[0133] Furthermore, scene recognition can be performed using audio data collected by a sound sensor. For example, audio sample data from different scenes can be collected, and then scene labels can be assigned to the audio sample data before inputting it into a neural network model for training. The aforementioned neural network model can be a DNN, CNN, or RNN, etc., and this application does not limit this; the aforementioned scene labels can be, for example, home scenes, bus scenes, subway scenes, and shopping mall scenes. It is understood that the trained neural network model can be used to classify the input audio data into scenes, that is, to perform scene recognition based on the input audio data.
[0134] In the above embodiments, the scene is first identified, and then event recognition is performed on the audio data to be processed based on the event recognition model corresponding to the scene. Alternatively, event recognition can also be performed through the probability distribution between events and scenes.
[0135] It is understandable that there are many different scenarios and events in daily life, and multiple events can occur in a single scenario, while the probability of the same event occurring varies in different scenarios. For example, the sound of an escalator is more likely to occur in subway, shopping mall, and airport scenarios, and less likely to occur in a home scenario (unless some residential buildings have escalators). Similarly, the sound of a microwave oven is more likely to occur in a home or shopping mall scenario, and less likely to occur in a subway scenario (unless there are convenience stores in subway stations).
[0136] Therefore, a probability distribution model can be obtained by statistically analyzing the probability of different events occurring in different scenarios, and event identification can be performed based on this probability distribution model. For example, please refer to [link to example]. Figure 8 , Figure 8 This is a flowchart illustrating another audio data processing method provided in an embodiment of this application, such as... Figure 8 The method shown can be performed by an electronic device, and the method includes:
[0137] 801: Define scenarios and events, where a scenario includes at least one of the events described above.
[0138] In this step, various scenarios and events that occur in daily life can be collected. For example, the scenarios mentioned above can be home scenarios, office scenarios, bus scenarios, subway scenarios, high-speed rail scenarios, airport scenarios, shopping mall scenarios, library scenarios, etc., and the events can be the sounds of microwave ovens, washing machines, induction cookers, escalators, barcode scanners, platform screen doors, access cards, and horns, etc.
[0139] It is understood that an event must occur within a certain scenario. Therefore, in the embodiments of this application, a scenario includes at least one of the aforementioned events. For example, a home scenario may include the sound of a microwave oven and / or a washing machine, while a shopping mall scenario may include the sound of an escalator and / or a platform screen door, etc.
[0140] 802: Collect first audio sample data for classifying the above-mentioned scenarios, and second audio sample data for classifying the above-mentioned events.
[0141] 803: Input the first audio sample data and the second audio sample data into the model to be trained to obtain the trained model; wherein, the first audio sample data includes scene labels and the second audio sample data includes event labels.
[0142] In this embodiment of the application, the first audio sample data can be understood as sample data used for classifying scenes after model training, that is, for scene recognition; the second audio data can be understood as sample data used for classifying events after model training, that is, for event recognition.
[0143] It is understandable that after the model is trained using the first and second audio sample data mentioned above, the trained model can be used to identify both the scene corresponding to the audio data and the event corresponding to the audio data. For example, the audio data for identifying the scene and the audio data for identifying the event can be different audio data.
[0144] In this step, the model to be trained can be a neural network model, such as an RNN, DNN, or CNN, etc., and this application does not limit this. Optionally, before inputting the first audio sample data and the second audio sample data into the model to be trained for training, it can be based on the foregoing... Figure 7 The processing flow shown here processes the audio sample data, which will not be described in detail here.
[0145] In this step, scene labels can be understood as labels used to indicate that the audio sample data is used for scene classification, and event labels can be understood as labels used to indicate that the audio sample data is used for event classification. This application does not limit the specific form of the labels.
[0146] 804: Get the audio data to be processed.
[0147] In this step, the audio data to be processed can be understood as a segment of audio data collected in a real-world scene. The event needs to be identified from this audio data. This segment of audio data may include multiple frames of audio data. Further details of this step can be found in step 501 above, and will not be repeated here.
[0148] 805: Based on the trained model, a key scene is identified from the first audio data in the audio data to be processed, and multiple reference events and a first reference probability corresponding to each reference event are identified from the second audio data in the audio data to be processed. The first audio data and the second audio data are audio data at different time positions in the audio data to be processed. The time interval between the first audio data and the second audio data is less than a threshold A. The first reference probability of the reference event is the probability that the trained model identifies the first audio data as the reference event.
[0149] It should be understood that the aforementioned first audio data and second audio data should be interpreted as audio data at different time positions within the aforementioned audio data to be processed. For example, when processing is done in frames, the aforementioned first audio data could be the 5th frame of the aforementioned audio data to be processed, and the aforementioned second audio data could be understood as the 6th frame of the aforementioned audio data to be processed. The duration of each frame of audio data could be 5 seconds, 10 seconds, etc., and this application does not limit this.
[0150] In this step, the order of the first audio data and the second audio data in the audio data to be processed is not limited. That is, the first audio data can precede or follow the second audio data. For example, when processing is done in frames, the first audio data can be the 8th frame of the audio data to be processed, and the second audio data can be understood as the 7th frame of the audio data to be processed.
[0151] In this step, the time interval between the first audio data and the second audio data is less than a threshold A. The threshold A can be determined based on the actual situation, and may be a duration less than three frames. For example, if each frame is 5 seconds long, then the threshold A could be 10 seconds, meaning the time interval between the first audio data and the second audio data is less than 10 seconds.
[0152] For ease of understanding, please refer to the example provided. Figure 9 , Figure 9This is a schematic diagram showing the positions of first audio data and second audio data provided in an embodiment of this application.
[0153] like Figure 9 As shown in (a), the first audio data can be two adjacent frames of audio data, wherein the first audio data is located before the second audio data. The first audio data is used for scene recognition, and the second audio data is used for event recognition.
[0154] like Figure 9 As shown in (b) and (c), the first audio data can be spaced one frame length before the second audio data. For example... Figure 9 In (b), the first audio data can precede the second audio data; for example... Figure 9 In (c), the first audio data can be located after the second audio data. Similarly, the first audio data is used for scene recognition, and the second audio data is used for event recognition.
[0155] It is understandable that, since the time interval between the first and second audio data is very small, and the scene does not change abruptly in real time, the event identified by the second audio data can be considered to be an event within the scene identified by the first audio data. It should also be understood that the embodiments of this application do not limit the temporal order of the first and second audio data. Even if an event is identified through the second audio data, and then the scene is identified through the first audio data, since the time interval between the first and second audio data is very small, the event identified by the second audio data can still be considered to be an event within the scene identified by the first audio data.
[0156] Understandably, when the first audio data is input into the trained model for scene recognition, the resulting recognition outcome can be multiple candidate scenes and a recognition probability for each candidate scene. For example, in the recognition outcome, the probability that the first audio data corresponds to scene A is 15%, the probability that the first audio data corresponds to scene B is 8%, and the probability that the first audio data corresponds to scene C is 16%. The aforementioned 15%, 8%, and 16% can be understood as recognition probabilities. In this step, the candidate scene with the highest recognition probability can be used as the aforementioned key scene.
[0157] Similarly, when the second audio data is input into the trained model for event recognition, the result can be multiple reference events and a first reference probability for each reference event. For example, in the recognition result, the probability that the second audio data corresponds to event A is 30%, the probability that the second audio data corresponds to event B is 10%, and the probability that the second audio data corresponds to event C is 18%. The above 30%, 10%, and 18% can be understood as the first reference probabilities.
[0158] 806: Multiply the first reference probability and the second reference probability of each of the above multiple reference events to obtain the third reference probability corresponding to each of the above multiple reference events; the second reference probability is the probability of the above reference event occurring in the above key scenario obtained by statistics.
[0159] It is understandable that the first reference probability mentioned above is the probability obtained by the trained model through event recognition of the second audio data, while the second reference probability is obtained through statistics. For example, the second reference probability can be obtained by statistically analyzing the occurrence of events in different regions. For instance, statistics can be collected on whether there are escalators, microwave ovens, and barcode scanners in shopping malls in region A, to obtain the probability of escalator sounds, microwave oven sounds, and barcode scanners occurring in shopping malls.
[0160] For each reference event, the result of multiplying the first reference probability corresponding to the reference event by the second reference probability corresponding to the reference event is taken as the final probability corresponding to the reference event, that is, the third reference probability mentioned above.
[0161] 807: The first reference event among the above multiple reference events is taken as the event recognition result of the second frame audio data, and the first reference event is the reference event with the highest third reference probability among the above multiple reference events.
[0162] For ease of understanding Figure 8 The method illustrated uses home, subway, and bus scenarios as examples; microwave oven sounds, platform screen door sounds, and card reader sounds as events; and a DNN model as an example. The specific steps include:
[0163] (1) Define the scene and events
[0164] The scenarios are home, subway, and bus. For example, it can be represented as:
[0165] Scene = {Home scene, subway scene, bus scene}.
[0166] The events include microwave oven sounds, shielded door sounds, and card reader sounds. For example, this can be represented as:
[0167] Event = {microwave oven sound, shielded door sound, card reader sound}.
[0168] (2) Collect audio sample data
[0169] Audio sample data was collected for classifying home scenarios, subway scenarios, and bus scenarios; as well as audio sample data for microwave oven sounds in home scenarios, platform screen door sounds in subway scenarios, and card reader sounds in bus scenarios.
[0170] (3) Model Training
[0171] The audio sample data in step (2) are labeled. For example, the audio sample data used to classify home scenarios, the audio sample data used to classify subway scenarios, and the audio sample data used to classify bus scenarios are labeled with scenario labels respectively; the audio sample data of microwave oven sound in home scenarios, the audio sample data of platform screen door sound in subway scenarios, and the audio sample data of card reader sound in bus scenarios are labeled with event labels respectively. The data are then input into the DNN for training to obtain the trained DNN.
[0172] (4) Establish a probability distribution model
[0173] For ease of understanding, please refer to the example provided. Figure 10 , Figure 10 This is a schematic diagram of a probability distribution provided in an embodiment of this application.
[0174] like Figure 10 The diagram shows events on the horizontal axis and scenarios on the vertical axis. The probability of an event occurring within a scenario can be represented by the transition probability, i.e., P(event|scenario). For example, Figure 10 The probability 0.7 in the second row and second column indicates that the microwave oven sound occurs with a probability of 0.7 in a home setting, which can be represented as P(microwave oven sound | home setting) = 0.7. For example, Figure 10 The probability 1 in the third row and third column indicates that the probability of the platform screen door sound appearing in the subway scene is 1, which can be expressed as P(platform screen door sound|subway scene) = 0.7.
[0175] (5) Event recognition
[0176] In this embodiment, the probability of an event is P(event) = P(event|scenario) * P(scenario) * P m Where P (scenario) is 1; P m The probability of the model identifying the above events, such as the first reference probability mentioned above.
[0177] Therefore, the final probability of the event is P(event) = P(event|scenario) * P m .
[0178] For example, please refer to Figure 11 , Figure 11This is a schematic diagram illustrating event recognition based on probability distribution, provided in an embodiment of this application.
[0179] like Figure 11 As shown, after acquiring the audio data to be processed, audio data segment A in the audio data to be processed is input into the trained DNN to identify the scene as a home scene, and audio data segment B in the audio data to be processed is input into the trained DNN to identify the events as the sound of the shielded door, the sound of the microwave oven and the sound of the card swipe machine, wherein the probabilities of each of the above events are 20%, 10% and 15%, respectively.
[0180] Understandably, in other schemes, the sound of the shielding door has the highest probability in the DNN recognition after training, so audio data segment B is considered to correspond to the sound of the shielding door.
[0181] In this solution, based on Figure 10 The probability distribution in the data shows that the probability of a microwave oven sound in a home setting is 0.7, the probability of a shielded door sound is 0, and the probability of a card reader sound is 0.1. Therefore, the probability P(shielded door sound) corresponding to audio data segment B is 0, the probability P(microwave oven sound) corresponding to audio data segment B is 0.07, and the probability P(card reader sound) corresponding to audio data segment B is 0.015. Figure 11 As shown, since P (microwave oven sound) is the largest, this scheme ultimately considers audio data segment B to correspond to microwave oven sound, rather than the sound of a shielded door as perceived by the trained DNN model.
[0182] It should be noted that, in the embodiments of this application, the number before the step should be understood as the identifier of the step, which facilitates the description of the solution by reference and increases readability so that the reader can understand the solution, rather than being understood as a limitation on the order of execution of the steps.
[0183] The methods provided in the embodiments of this application have been described above. Next, the electronic devices involved in the embodiments of this application will be described.
[0184] Please see Figure 12 , Figure 12 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application.
[0185] Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an accelerometer sensor 180C, a fingerprint sensor 180D, a temperature sensor 180E, a touch sensor 180F, an ambient light sensor 180G, etc.
[0186] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more than Figure 12 This may involve more or fewer components, or combining certain components, or splitting certain components, or different component arrangements. Figure 12 The components shown can be implemented in hardware, software, or a combination of both.
[0187] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0188] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0189] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the aforementioned memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0190] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0191] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180F, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180F through the I2C interface, enabling the processor 110 and the touch sensor 180F to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.
[0192] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0193] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0194] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.
[0195] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.
[0196] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0197] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0198] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0199] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0200] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0201] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0202] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0203] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0204] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0205] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0206] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0207] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0208] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0209] In some embodiments, the display screen 194 may display the event recognition results of the audio data to be processed, such as text prompts output in the form of pop-up windows.
[0210] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0211] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image or video playback, etc.). The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0212] In this embodiment of the application, the internal memory 121 may include the first storage unit and the second storage unit, and the first storage unit may be referred to as a cache.
[0213] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0214] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0215] The loudspeaker 170A, also known as a horn, is used to convert audio electrical signals into sound signals.
[0216] The receiver 170B, also known as the earpiece, is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0217] Microphone 170C, also known as a microphone or transducer, is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0218] In some embodiments, the sound sensor may be the microphone 170C. For example, an electronic device may acquire the audio data to be processed through the microphone 170C.
[0219] The headphone jack 170D is used to connect wired headphones. The headphone jack 170D can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface.
[0220] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A may be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When a force is applied to pressure sensor 180A, the capacitance between the electrodes changes.
[0221] The electronic device 100 can determine the intensity of pressure based on changes in capacitance. For example, when a touch operation is applied to the display screen 194, the electronic device 100 detects the intensity of the touch operation using the pressure sensor 180A. Also for example, the electronic device 100 can calculate the touch position based on the detection signal from the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS message is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS message is executed.
[0222] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.
[0223] The accelerometer 180C can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic device, and can be applied to applications such as screen orientation switching and pedometers.
[0224] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.
[0225] The fingerprint sensor 180D is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.
[0226] Temperature sensor 180E is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180E to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180E exceeds a threshold, electronic device 100 performs thermal protection by reducing the performance of a processor located near temperature sensor 180E to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 100 heats battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.
[0227] Touch sensor 180F, also known as a "touch panel," can be located on display screen 194. The touch sensor 180F and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180F detects touch operations applied to or near it. Touch sensor 180F can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180F may also be located on the surface of electronic device 100, in a different position than display screen 194.
[0228] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0229] In some embodiments, the mobile communication module 150 or the wireless communication module 160 can receive audio data sent by other electronic devices, and the processor 110 can call computer instructions stored in the internal memory 121 to use the audio data sent by the other electronic devices as audio data to be processed.
[0230] In other embodiments, the processor 110 may call computer instructions stored in the internal memory 121 to implement the audio data processing method provided in the embodiments of this application.
[0231] For example, the processor 110 can call computer instructions stored in the internal memory 121 to obtain at least one image captured by the camera 193, obtain acceleration data detected by the accelerometer, obtain network connection status (such as Wi-Fi connection status) in the mobile communication module 150, obtain location information determined by the mobile communication module 150, etc., to perform scene recognition and obtain scene recognition results.
[0232] For example, the processor 110 can call computer instructions stored in the internal memory 121 to acquire audio data to be processed through the microphone 170C, or the mobile communication module 150, or the wireless communication module 160, and then obtain the event recognition result based on the scene recognition result.
[0233] It is understood that the software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.
[0234] Please see Figure 13 , Figure 13 This is a software structure block diagram of an electronic device 100 provided in an embodiment of this application.
[0235] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system can be divided into four layers, from top to bottom: the application layer, the application framework layer, the system runtime library layer, and the kernel layer. The descriptions of each layer are as follows:
[0236] First, the application layer may include a series of application packages. For example, the application packages in the application layer may include applications such as camera, gallery, calendar, call, map, navigation, browser, Bluetooth, music, video, and SMS.
[0237] Secondly, the application framework layer can provide application programming interfaces (APIs) and programming frameworks for applications within the application layer. The application framework layer can include some predefined functions.
[0238] For example, the application framework layer may include an activity manager, a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc. Among them:
[0239] The Activity Manager can be used to manage the lifecycle of individual applications and the usual navigation back functionality.
[0240] A window manager can be used to manage window programs. For example, a window manager can obtain the screen size of an electronic device 100, lock the screen, capture the screen, and determine whether a status bar is present, etc.
[0241] Content providers can be used to store and retrieve data, and make that data accessible to applications, enabling different applications to access or share data. For example, the data mentioned above may include videos, images, audio, phone calls made and received, browsing history and bookmarks, and phone books, etc.
[0242] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0243] The phone manager is used to provide communication functions for electronic device 100, such as managing call status (including answering and hanging up calls).
[0244] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, etc.
[0245] The notification manager allows applications to display notifications in the status bar. These notifications can be used to convey informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify of download completion, message alerts, etc. The notification manager can also appear as an icon or scrolling text bar in the top status bar, such as notifications from background applications, or as a dialog box on the screen. Examples include displaying text messages in the status bar, emitting alert sounds, vibrating electronic devices, and flashing indicator lights.
[0246] Furthermore, the system runtime layer can include system libraries and the Android runtime. Specifically:
[0247] The Android runtime consists of the core libraries and the virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core libraries comprise two parts: one part contains the functionalities that Java calls, and the other part contains the core Android libraries. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0248] System libraries can be understood as the support for the application framework, serving as a crucial link between the application framework layer and the kernel layer. The system layer can include multiple functional modules, such as a surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), and 2D graphics engines (e.g., SGL). Among these:
[0249] The surface manager can be used to manage the display subsystem, such as when multiple applications are running on electronic device 100, it manages the interaction between display and access operations. The surface manager can also be used to provide the blending of two-dimensional and three-dimensional layers for multiple applications.
[0250] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0251] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0252] A 2D graphics engine can be understood as a graphics engine for 2D drawing.
[0253] Finally, the kernel layer can be understood as an abstraction layer between hardware and software. The kernel layer can include system services such as security, memory management, process management, power management, network protocol management, and driver management. For example, the kernel layer may include display drivers, camera drivers, audio drivers, and sensor drivers.
[0254] In some embodiments, the application layer may further include an audio data processing module, which is used to implement the audio data processing method provided in the embodiments of this application.
[0255] It is understood that in other embodiments, the audio data processing module described above may also be located at other levels of the layered architecture described above, such as the system layer, etc., which is not limited here.
[0256] This application also provides a computer-readable storage medium storing computer code that, when executed on a computer, causes the computer to perform the methods described in the above embodiments.
[0257] This application also provides a computer program product, which includes computer code or a computer program that, when run on a computer, causes the methods described in the above embodiments to be executed.
[0258] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0259] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0260] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0261] It should also be understood that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the above claims.
Claims
1. An audio data processing method, characterized in that, The method includes: Acquire the audio data to be processed; If at least one image is captured by a camera, the at least one image is input into the trained scene recognition model to obtain multiple candidate scenes; The scenario that matches at least one of the following data among the multiple candidate scenarios is designated as the first scenario; the trained scene recognition model is obtained by training sample images from multiple scenarios, including the first scenario, which is the scenario where the sound source generating the audio data to be processed is located, and the first scenario is any one of the following scenarios: home scenario, office scenario, bus scenario, subway scenario, high-speed rail scenario, airport scenario, shopping mall scenario, coffee shop scenario, and library scenario; The audio data to be processed is input into an event recognition model for event recognition, resulting in multiple candidate events; Obtain a first probability of each candidate event occurring in the first scenario, wherein the first probability is obtained by counting the number of times each candidate event occurs in the first scenario; The candidate event with the largest result of the calculation of the first probability and the second probability among the multiple candidate events is taken as the event recognition result of the audio data to be processed. The second probability is the recognition probability of each candidate event obtained by the event recognition model in the audio data to be processed.
2. The method according to claim 1, characterized in that, The event recognition model is a model obtained by training the event recognition model to be trained by collecting audio sample data in the first scene.
3. The method according to claim 1 or 2, characterized in that, The time interval between the moment when the audio data to be processed is acquired and the moment when the first scene is identified is less than or equal to a first threshold.
4. The method according to claim 1, characterized in that, The method further includes: In the absence of capturing at least one image via a camera, the first scenario is determined based on at least one of the following: the location information of the electronic device, the network connection object of the electronic device, and the moving speed of the electronic device.
5. An electronic device, characterized in that, The device includes a processor and a memory, the memory being used to store a computer program, the computer program including program instructions, and the processor being configured to invoke the program instructions such that the method as described in any one of claims 1-4 is executed.
6. A chip, characterized in that, It includes logic circuitry and an interface, the logic circuitry and the interface being coupled; the interface is used to input and / or output code instructions, and the logic circuitry is used to execute the code instructions to cause the method of any one of claims 1-4 to be performed.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the method as described in any one of claims 1-4 to be performed.