Event detection method and device, electronic equipment and storage medium

By acquiring uplink audio data during voice calls and using a multilayer perceptron neural network model to detect target sounds, the problem of untimely event detection caused by microphone channel occupancy is solved, enabling timely handling of emergencies during calls.

CN122073597APending Publication Date: 2026-05-22BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-22

Smart Images

  • Figure CN122073597A_ABST
    Figure CN122073597A_ABST
Patent Text Reader

Abstract

The invention relates to an event detection method and device, electronic equipment and a storage medium. The event detection method comprises the following steps: when the electronic equipment executes a voice call based on uplink audio data acquired by a microphone, acquiring the uplink audio data; determining whether target sound data exists in the uplink audio data, wherein the target sound data is sound data generated when a corresponding target event occurs; and in response to target sound data existing in the uplink audio data, determining that an event for correspondingly generating the target sound data is detected, and executing a function corresponding to the event. According to the invention, the target sound data is detected in the voice communication process, the instantaneity of event detection is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of electronic equipment technology, and in particular to an event detection method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the continuous development of computer technology, in the event of an emergency during a call, the user often loses the ability to independently call for help due to delayed event detection. For example, if a user is unconscious in a serious car accident and has lost the ability to call for help independently, their life may be in danger if the current event cannot be monitored in time.

[0003] In related technologies, audio data is acquired through recording or voice wake-up to detect relevant events, such as collision sound recognition. However, due to the occupation of the microphone channel during a call, the microphone cannot acquire recording data or data for voice wake-up, thus making it impossible to detect the current event in a timely manner and take appropriate measures in case of an emergency. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this disclosure provides an event detection method, apparatus, electronic device, and storage medium.

[0005] According to a first aspect of the present disclosure, an event detection method is provided, comprising: acquiring the uplink audio data during a voice call performed by an electronic device based on uplink audio data collected by a microphone; determining whether target sound data exists in the uplink audio data, wherein the target sound data is sound data generated when a target event occurs; and, in response to the presence of target sound data in the uplink audio data, determining that an event corresponding to the generation of the target sound data has been detected, and performing a function corresponding to the event.

[0006] In one embodiment, before acquiring the uplink audio data, the method further includes: determining that a target application is running; wherein the target application is used to trigger the process of acquiring the uplink audio data and determining target sound data when the electronic device is detected to be performing a voice call, and the target application is used to perform a function corresponding to the event.

[0007] In one embodiment, determining that an event corresponding to the generation of the target sound data has occurred includes: in response to the target application acquiring the target sound data, determining that an event corresponding to the generation of the target sound data has occurred.

[0008] In one embodiment, the method further includes: triggering the target application to acquire the target sound data in response to the sensor of the electronic device detecting that the acceleration change data of the electronic device is greater than the acceleration change threshold; or, reporting the target sound data to the target application in response to the presence of target sound data in the uplink audio data.

[0009] In one embodiment, the method further includes: in response to the presence of target sound data in the uplink audio data, saving the target sound data for a set duration; wherein the set duration is determined based on the duration required for the target application to acquire the target sound data.

[0010] In one embodiment, determining whether target sound data exists in the uplink audio data includes: copying the uplink audio data in real time; denoising the copied uplink audio data and inputting it into a multilayer perceptron neural network model to classify the uplink audio data, and determining whether target sound data exists in the uplink audio data based on the classification results; wherein the multilayer perceptron neural network model is used to classify the uplink audio data, and the classification results include target sound data and / or non-target sound data.

[0011] In one embodiment, the noise reduction processing of the copied uplink audio data includes: removing the target sound data played by the electronic device from the copied uplink audio data.

[0012] In one embodiment, the target sound data is the collision sound in a car accident scene, and the function is a distress call function.

[0013] According to a second aspect of the present disclosure, an event detection apparatus is provided, comprising: an acquisition unit, configured to acquire uplink audio data during a voice call performed by an electronic device based on uplink audio data collected by a microphone; a detection unit, configured to determine whether target sound data exists in the uplink audio data, wherein the target sound data is sound data generated when a target event occurs; and a processing unit, configured to, in response to the presence of target sound data in the uplink audio data, determine that an event corresponding to the generation of the target sound data has been detected, and execute a function corresponding to the event.

[0014] In one embodiment, the acquisition unit is further configured to: determine that a target application is running before acquiring the uplink audio data; wherein the target application is configured to trigger the process of acquiring the uplink audio data and determining target sound data when the electronic device is detected to be performing a voice call, and the target application is configured to perform a function corresponding to the event.

[0015] In one embodiment, the processing unit determines that an event corresponding to the generation of the target sound data has occurred in the following manner: in response to the target application acquiring the target sound data, it determines that an event corresponding to the generation of the target sound data has occurred.

[0016] In one embodiment, the processing unit is further configured to: trigger the target application to acquire the target sound data in response to the sensor of the electronic device detecting that the acceleration change data of the electronic device is greater than the acceleration change threshold; or, report the target sound data to the target application in response to the presence of target sound data in the uplink audio data.

[0017] In one embodiment, the processing unit is further configured to: in response to the presence of target sound data in the uplink audio data, save the target sound data for a set duration; wherein the set duration is determined based on the duration required for the target application to acquire the target sound data.

[0018] In one embodiment, the detection unit determines whether target sound data exists in the uplink audio data in the following manner: the uplink audio data is copied in real time; the copied uplink audio data is denoised and then input into a multilayer perceptron neural network model to classify the uplink audio data, and the existence of target sound data is determined based on the classification result; wherein, the multilayer perceptron neural network model is used to classify the uplink audio data, and the classification result includes target sound data and / or non-target sound data.

[0019] In one embodiment, the detection unit performs noise reduction processing on the copied uplink audio data by removing the target sound data played by the electronic device from the copied uplink audio data.

[0020] In one embodiment, the target sound data is the collision sound in a car accident scene, and the function is a distress call function.

[0021] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to perform an event detection method according to the first aspect or any embodiment of the first aspect.

[0022] According to a fourth aspect of the present disclosure, a storage medium is provided that stores instructions which, when executed by a processor, enable the execution of an event detection method according to the first aspect or any of the embodiments of the first aspect.

[0023] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: by acquiring uplink audio data during a voice call to detect the corresponding sound data for the event, the target event can be detected directly during the call. Furthermore, when the target sound data is detected, the corresponding event function is executed, ensuring timely processing of the event under the current circumstances.

[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0026] Figure 1 This is a flowchart illustrating an event detection method according to an exemplary embodiment.

[0027] Figure 2 This is a flowchart illustrating a method for determining whether a target sound exists, according to an exemplary embodiment.

[0028] Figure 3 This is a flowchart illustrating a method for determining detected target sound data according to an exemplary embodiment.

[0029] Figure 4 This is a flowchart illustrating the detection of target sound data according to an exemplary embodiment.

[0030] Figure 5 This is a flowchart illustrating an event detection method according to an exemplary embodiment.

[0031] Figure 6 This is a schematic diagram illustrating an event detection according to an exemplary embodiment.

[0032] Figure 7 This is a block diagram of an event detection device according to an exemplary embodiment.

[0033] Figure 8 This is a block diagram of an event detection device according to an exemplary embodiment.

[0034] Figure 9 This is a block diagram illustrating an event detection device according to an exemplary embodiment. Detailed Implementation

[0035] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure.

[0036] The event detection method provided in this disclosure is applied to scenarios where a traffic accident is detected and an emergency call is made. For example, this disclosure is applied to a scenario where a user detects a traffic accident during a call but is unable to call for help independently.

[0037] In related technologies, for example, in the event of a serious car accident, if a user making a voice call inside the vehicle is unconscious or injured and thus unable to independently call for help, the technology typically uses a microphone installed on the electronic device to acquire recordings or voice wake-up data for collision sound recognition to detect the accident. However, if a car accident occurs while the user is on a call, the microphone's audio channel is occupied by the call path, thus disabling the function of acquiring recordings or voice wake-up data. This prevents the acquisition of microphone data through recordings or voice wake-up, rendering the car accident detection solution ineffective during the call.

[0038] In view of this, this disclosure addresses situations where users are unable to independently call for help due to sudden events during a call, such as a car accident. It utilizes uplink audio data from the call's uplink channel for event detection, thereby expanding the methods for event detection in the event of a sudden event and enabling timely event detection.

[0039] In this embodiment of the disclosure, the following is adopted: Figure 1 Event detection is performed in the manner shown.

[0040] Figure 1 This is a flowchart illustrating an event detection method according to an exemplary embodiment. Figure 1 As shown, it includes the following steps.

[0041] In step S11, during a voice call performed by the electronic device based on the uplink audio data collected by the microphone, uplink audio data is acquired.

[0042] In this embodiment of the disclosure, if the microphone of the electronic device is in voice call mode, the uplink audio data during the voice call is acquired in real time.

[0043] In step S12, it is determined whether target sound data exists in the uplink audio data.

[0044] In this embodiment of the disclosure, the acquired uplink audio data is analyzed, and the presence of target sound data in the uplink audio data is determined based on the analysis results.

[0045] The target sound data refers to the sound data generated when the target event occurs. For example, if a car accident occurs during a voice call, the target sound is the sound of the car accident that occurs when the car accident event takes place.

[0046] In step S13, in response to the presence of target sound data in the uplink audio data, it is determined that an event corresponding to the generation of target sound data has occurred, and the corresponding event function is executed.

[0047] According to an exemplary embodiment of this disclosure, by analyzing uplink audio data in real time, target events can be detected directly during a call without interrupting the voice call, thus saving time in detecting target events. Furthermore, when target sound data is detected, the corresponding event function is executed, ensuring that the user receives timely feedback regarding the relevant event.

[0048] In this embodiment of the disclosure, the target sound data involved can be the collision sound in a car accident scene, or the sound of other types of collisions occurring in a sudden state during a user's voice call, and the function can be a distress call function.

[0049] In this embodiment of the disclosure, a target application for detecting target events can also be set. The detection of an event that generates target sound data is determined by opening the target application, and the corresponding functional event detection is performed. For example, the target application can be opened based on user demand, or it can be automatically triggered when an event occurs. This disclosure does not limit the method of opening the target application.

[0050] The target application can be understood as a car accident detection application, which is used to trigger the process of acquiring uplink audio data and determining target sound data when an electronic device is detected to be making a voice call. The target application is used to perform the corresponding event functions.

[0051] For example, taking the detection of car accident events as an example, this explains how to enable a car accident detection application.

[0052] In one example, the car accident detection application opens the Audio Digital Signal Processor (ADSP). It uses the car accident detection algorithm within the ADSP's call module to perform car accident detection, such as by setting parameters ("CarAccidentDetect=ON") in the SetParameter function. The SetParameter function ultimately calls the Audio HAL (Audio Server) through the Audio Server module. The Audio HAL then passes the parameters from SetParameter to the ADSP via the interface corresponding to the car accident detection thread, thereby opening the car accident detection algorithm module within the ADSP's call path.

[0053] According to an exemplary embodiment of this disclosure, by pre-activating the target application, it is ensured that the subsequent target application can successfully detect the occurrence of the event, thereby providing a prerequisite for the subsequent event detection method.

[0054] In this embodiment of the disclosure, taking a car accident collision that occurs during a user's voice call as an example, the method of event detection using the target application described above is further explained.

[0055] Figure 2 This is a flowchart illustrating a method for determining the presence of a target sound according to an exemplary embodiment. Figure 2 As shown, it includes the following steps.

[0056] In step S21, the upstream audio data is copied in real time.

[0057] In this embodiment of the disclosure, uplink audio data during a user's voice call is copied in real time.

[0058] In step S22, the copied uplink audio data is denoised and then input into a multilayer perceptron neural network model to classify the uplink audio data and determine whether the target sound data exists in the uplink audio data based on the classification results.

[0059] Among them, the multilayer perceptual neural network model is used to classify the uplink audio data, and the classification results include target sound data and / or non-target sound data.

[0060] In one example, high-frequency interference in the uplink audio data is first filtered to obtain low-pass filtered audio data. Then, a preset loudness is used to filter the -15dB loudness data, further filtering the low-pass filtered audio data to determine the target loudness audio data. Next, the signal characteristics of the target loudness audio data are converted into a frequency domain audio signal. Based on the frequency domain audio signal, a logarithmic amplitude spectrum is obtained. Target frequency information is selected from the logarithmic amplitude spectrum to determine the spectral characteristics of the uplink audio data. Based on these spectral characteristics, the frequency band information of the number of target frames preceding and following the peak time frame is selected. This information is then classified using a multilayer perceptron neural network to determine whether the target sound data exists in the uplink audio data.

[0061] According to an exemplary embodiment of this disclosure, by acquiring copied uplink audio data in real time, the immediacy of detecting a target event during a voice call is improved. Furthermore, analyzing the copied uplink audio data reduces the impact on the original call data, thereby preserving the user's original call experience.

[0062] In this embodiment of the disclosure, before copying the uplink audio data, the target sound data played by the electronic device can be removed from the copied uplink audio data, thereby reducing the impact of the sound data played by the electronic device itself on the copied uplink audio data.

[0063] In one example, the uplink audio data is converted into a digital signal by a signal conversion module (codec) and then fed to the Acoustic Echo Cancellation (AEC) module in the ADSP. The AEC can eliminate the sound played by the electronic device itself, preventing the electronic device from playing the sound of a car accident collision, which could cause the algorithm to make a misjudgment.

[0064] In this embodiment of the disclosure, when the electronic device obtains audio data after removing the influence of the target sound data played by the electronic device, the classification result of the analysis of the uplink audio data can control the target application to obtain the target sound data and thus detect whether the target sound exists.

[0065] Figure 3 This is a flowchart illustrating a method for determining detected target sound data according to an exemplary embodiment. For example... Figure 3 As shown, it includes the following steps.

[0066] In step S31, the target application acquires the target sound data.

[0067] In step S32, in response to the target application acquiring the target sound data, it is determined that an event corresponding to the generation of the target sound data has been detected.

[0068] In this embodiment of the disclosure, the target application obtains the target sound data and determines that the event corresponding to the generation of the target sound data has occurred, thereby avoiding the situation where the target application obtains sound data in real time and saving power consumption in determining the occurrence of the event corresponding to the generation of the target sound data.

[0069] In this embodiment of the disclosure, the target application can also acquire target sound data by detecting sensor data.

[0070] The sensor can be understood as an application component used to trigger the target application to acquire the target sound data.

[0071] Figure 4 This is a flowchart illustrating the detection of target sound data according to an exemplary embodiment. For example... Figure 4 As shown, it includes the following steps.

[0072] In step S41, target sound data is acquired.

[0073] In step S42, in response to the sensor of the electronic device detecting that the acceleration change data of the electronic device is greater than the acceleration change threshold, the target application is triggered to acquire the target sound data.

[0074] In step S43, in response to the presence of target sound data in the uplink audio data, the target sound data is reported to the target application.

[0075] According to an exemplary embodiment of this disclosure, the target application is triggered to acquire target sound data via the sensors of an electronic device, avoiding unnecessary power consumption caused by the target application indiscriminately acquiring sound data. Furthermore, the target application can also acquire target sound data when the sensors of the electronic device detect an acceleration change data greater than an acceleration change threshold, without requiring additional hardware and saving hardware costs.

[0076] In this embodiment of the disclosure, in response to the presence of target sound data in the uplink audio data, the target sound data is saved for a set duration.

[0077] The set duration is determined based on the time required for the target application to acquire the target sound data.

[0078] In one example, if a car crash sound is detected, the sound is saved for 5 seconds before being updated with the latest detection result. Saving it for 5 seconds gives the crash detection application sufficient time to obtain the audio detection results.

[0079] In another example, if the crash detection application does not obtain the audio detection result within 5 seconds, it means that the large-range accelerometer sensor has not been triggered to report, and the audio detection result is invalid.

[0080] In an exemplary embodiment of this disclosure, when the car accident detection application is running, if the user is detected to be in a call, the audio data during the user's call is copied in real time. The copied audio data is then classified to determine whether the classification results contain car accident collision sounds. If so, the car accident collision sounds are retained. Simultaneously, the system control processor (SCP) is used to detect whether a physical collision exists. If so, the car accident detection application is triggered to retrieve the previously retained car accident collision sounds and then initiate a car accident emergency call.

[0081] In this embodiment of the disclosure, the event detection method involved in the present disclosure is further illustrated by taking the example of performing car accident detection and event detection through a car accident detection application.

[0082] Figure 5 This is a flowchart illustrating an event detection method according to an exemplary embodiment. Figure 5 As shown, it includes the following steps.

[0083] In step S51, a traffic accident detection is performed.

[0084] In this embodiment of the disclosure, for example, the user can manually activate the car accident detection application to perform car accident detection according to their needs.

[0085] In step S52, it is determined whether the call is in progress.

[0086] In this embodiment of the disclosure, the system determines whether the user is currently in a call state based on the vehicle accident detection application. If so, steps S53 and S55 are executed; otherwise, the vehicle accident detection ends.

[0087] In step S53, the car accident detection algorithm in the ADSP call module is activated.

[0088] In this embodiment of the disclosure, for example, the car accident detection algorithm in the ADSP call module is enabled by setting SetParameter("CarAccidentDetect=ON"), the implementation of which has been explained in this disclosure and will not be repeated here.

[0089] In step S54, the car accident collision sound detection algorithm is run in the call module of the ADSP.

[0090] In this embodiment, the microphone analog signal data is converted into a digital signal by a signal conversion module (codec) and then fed to the data preprocessing (acoustic echo cancellation, AEC) module in the ADSP. The AEC module can eliminate the sound played by the electronic device itself, preventing the electronic device from playing the sound of a car accident collision, which could cause the algorithm to misjudge. The data processed by the AEC module is then fed to the car accident collision sound detection algorithm. The car accident collision sound detection module makes a copy of the audio data. The original audio data is given directly to other modules of the call algorithm without any processing, so as to avoid affecting the call effect. The car accident collision sound detection algorithm performs a 200Hz low-pass filter on the copied audio data, then filters it by -15db loudness, performs a 2048-point FFT and converts it into a logarithmic amplitude spectrum. After taking the frequency components within 200Hz (12 points), it obtains the peak corresponding to the time frame before and after the time frame, for a total of 5 frames (12*5=60) of data to obtain the classifier input, and then performs classification through MLP to output the results.

[0091] In one example, if the output indicates the presence of a car crash sound, the output is stored in the algorithm and updated with the latest detection result after 5 seconds. Retaining the test result for 5 seconds allows the car crash detection application sufficient time to obtain the audio detection results.

[0092] In another example, if the crash detection application does not obtain the audio detection result within 5 seconds, it means that the large-range accelerometer sensor has not been triggered to report, and the audio detection result is invalid.

[0093] In step S55, the large-range acceleration sensor on the SCP side detects a car accident collision.

[0094] In this embodiment of the present disclosure, when a car accident occurs, the large-range acceleration sensor will detect a large-range acceleration change and consider it a suspected car accident. At this time, step S56 is executed to send the sensor data to the car accident detection application. If no accident is detected, the process ends.

[0095] In step S56, the sensor detects a collision and uploads the result to the crash detection application.

[0096] In step S57, the crash detection application obtains the audio crash detection results.

[0097] In this embodiment of the disclosure, for example, the car accident detection result is obtained by setting getParameter("CarAccidentDetectResult"). getParameter is ultimately called by AudioServer to Audio HAL. Audio HAL will call the ADSP through the aurisys_get_parameter interface to obtain the car accident collision sound detection result.

[0098] In step S58, the collision sound detection algorithm in the ADSP returns the collision detection result.

[0099] In step S59, the vehicle accident detection application automatically calls for rescue based on the obtained vehicle accident collision sound detection results.

[0100] In step S510, the vehicle accident detection ends.

[0101] In this embodiment of the disclosure, the following methods are adopted: Figure 6 The method shown further illustrates the above-mentioned method for event detection using a car accident detection application. Figure 6 This is a schematic diagram illustrating an event detection according to an exemplary embodiment.

[0102] In this embodiment of the disclosure, if a call is detected in the car accident detection application, the car accident detection application calls the car accident detection algorithm (setParameter("OnCallCarAccidentDetect=ON")) to open the car accident detection algorithm module in the call path of the ADSP. setParameter ultimately calls the Audio HAL via the AudioServer. The Audio HAL then calls the parameters passed by setParameter to the ADSP through the aurisys_set_parameter interface, thereby opening the car accident detection algorithm module in the call path of the ADSP.

[0103] The microphone analog signal data is converted into a digital signal by the codec and then sent to the ADSP's AEC module. The AEC module can eliminate the sound played by the electronic device itself, preventing the electronic device from playing collision sounds that could cause misjudgments in the collision sound detection algorithm. The AEC module then sends the processed data to the collision sound detection algorithm.

[0104] The car accident collision sound detection module copies the received audio data. The original audio data is sent directly to other modules of the call algorithm without any processing to avoid affecting the call quality. The car accident collision sound detection algorithm performs a 200Hz low-pass filter on the copied audio data, then filters it by -15dB loudness, performs a 2048-point FFT and converts it to a logarithmic amplitude spectrum. After taking the frequency components within 200Hz (12 points), it obtains the peak corresponding to the time frame before and after the peak, a total of 5 frames (12*5=60) of data to obtain the classifier input. Then, it is classified and output by a multilayer perceptron (MLP) neural network model. If the output result indicates the presence of a car accident collision sound, the output result is saved in the algorithm and updated to the latest detection result after 5 seconds. The test result is retained for 5 seconds to allow the upper layer application sufficient time to obtain the audio detection result. If the upper layer application does not obtain the audio detection result within 5 seconds, it means that the large-range accelerometer sensor has not been triggered to report, and the audio detection result is invalid.

[0105] When a car accident occurs, a large-range accelerometer will detect significant changes in acceleration, indicating a suspected accident. The sensor will then report this data to the accident detection application. Upon receiving the suspected accident report, the application will call `getParameter("CarAccidentDetectResult")` to retrieve the collision sound detection results. `getParameter` calls the AudioServer, which in turn calls the Audio HAL. The Audio HAL then uses the `aurisys_get_parameter` interface to call the ADSP to obtain the collision sound detection results. At this point, the accident detection application has completed acquiring the collision sound detection results and will automatically call for roadside assistance based on the results.

[0106] According to an exemplary embodiment of this disclosure, recording and voice wake-up are both turned off during a voice call. At this time, the microphone is exclusively used by the call application and microphone data cannot be obtained through recording or voice wake-up; only ambient sound can be obtained through the call channel, making it impossible to perform collision sound detection during a call. This disclosure performs collision sound detection by copying data from the voice call, thus realizing the process of event detection during a voice call. Furthermore, a multilayer perceptron neural network model is used to classify the voice call data and output the results, improving the recognition accuracy.

[0107] Based on the same concept, embodiments of this disclosure also provide an event detection device.

[0108] It is understood that the event detection device provided in this disclosure includes hardware structures and / or software modules corresponding to each function in order to achieve the above-mentioned functions. In conjunction with the units and algorithm steps of the various examples disclosed in this disclosure, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of this disclosure.

[0109] Figure 7 This is a block diagram illustrating an event detection device according to an exemplary embodiment. (Refer to...) Figure 7 The device 100 includes an acquisition unit 101, a detection unit 102, and a processing unit 103.

[0110] The acquisition unit 101 is used to acquire uplink audio data during a voice call performed by an electronic device based on uplink audio data collected by a microphone.

[0111] The detection unit 102 is used to determine whether target sound data exists in the uplink audio data, where the target sound data is the sound data generated when the target event occurs.

[0112] The processing unit 103 is configured to, in response to the presence of target sound data in the uplink audio data, determine that an event corresponding to the generation of target sound data has occurred, and execute the corresponding event function.

[0113] In one embodiment, the acquisition unit 101 is further configured to: determine that the target application is open before acquiring the uplink audio data; wherein the target application is configured to trigger the process of acquiring uplink audio data and determining target sound data when the electronic device is detected to be performing a voice call, and the target application is configured to perform the function of the corresponding event.

[0114] In one embodiment, the processing unit 103 determines that an event corresponding to the generation of target sound data has occurred in the following manner: in response to the target application obtaining target sound data, it determines that an event corresponding to the generation of target sound data has occurred.

[0115] In one embodiment, the processing unit 103 is further configured to: trigger the target application to acquire target sound data in response to the sensor of the electronic device detecting that the acceleration change data of the electronic device is greater than the acceleration change threshold; or, report the target sound data to the target application in response to the presence of target sound data in the uplink audio data.

[0116] In one embodiment, the processing unit 103 is further configured to: in response to the presence of target sound data in the uplink audio data, save the target sound data for a set duration; wherein the set duration is determined based on the duration required for the target application to acquire the target sound data.

[0117] In one embodiment, the detection unit 102 determines whether target sound data exists in the uplink audio data in the following manner: real-time copying of uplink audio data; after denoising the copied uplink audio data, inputting it into a multilayer perceptron neural network model to classify the uplink audio data, and determining whether target sound data exists in the uplink audio data based on the classification results; wherein, the multilayer perceptron neural network model is used to classify the uplink audio data, and the classification results include target sound data and / or non-target sound data.

[0118] In one embodiment, the detection unit 102 performs noise reduction processing on the copied uplink audio data in the following manner: removing the target sound data played by the electronic device from the copied uplink audio data.

[0119] In one implementation, the target sound data is the collision sound in a car accident scene, and its function is to call for help.

[0120] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0121] Figure 8 This is a block diagram illustrating an event detection device 200 according to an exemplary embodiment. For example, device 200 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0122] Reference Figure 8 The device 200 may include one or more of the following components: processing component 202, memory 204, power component 206, multimedia component 208, audio component 210, input / output (I / O) interface 212, sensor component 214, and communication component 216.

[0123] Processing component 202 typically controls the overall operation of device 200, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 202 may include one or more modules to facilitate interaction between processing component 202 and other components. For example, processing component 202 may include a multimedia module to facilitate interaction between multimedia component 208 and processing component 202.

[0124] Memory 204 is configured to store various types of data to support the operation of device 200. Examples of such data include instructions for any application or method operating on device 200, contact data, phonebook data, messages, pictures, videos, etc. Memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0125] The power supply component 206 provides power to the various components of the device 200. The power supply component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 200.

[0126] Multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 208 includes a front-facing camera and / or a rear-facing camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0127] Audio component 210 is configured to output and / or input audio signals. For example, audio component 210 includes a microphone (MIC) configured to receive external audio signals when device 200 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 204 or transmitted via communication component 216. In some embodiments, audio component 210 also includes a speaker for outputting audio signals.

[0128] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0129] Sensor assembly 214 includes one or more sensors for providing status assessments of various aspects of device 200. For example, sensor assembly 214 may detect the on / off state of device 200, the relative positioning of components such as the display and keypad of device 200, changes in the position of device 200 or a component of device 200, the presence or absence of user contact with device 200, the orientation or acceleration / deceleration of device 200, and temperature changes of device 200. Sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 214 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 214 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0130] Communication component 216 is configured to facilitate wired or wireless communication between device 200 and other devices. Device 200 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 216 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 216 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0131] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0132] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, which can be executed by a processor 220 of the device 200 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0133] Figure 9 This is a block diagram illustrating an event detection device 300 according to an exemplary embodiment. For example, device 300 may be provided as a server. (Refer to...) Figure 9 The device 300 includes a processing component 322, which further includes one or more processors, and memory resources represented by memory 332 for storing instructions, such as application programs, that can be executed by the processing component 322. The application programs stored in memory 332 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 322 is configured to execute instructions to perform the methods described above.

[0134] Device 300 may also include a power supply component 326 configured to perform power management of device 300, a wired or wireless network interface 350 configured to connect device 300 to a network, and an input / output (I / O) interface 352. Device 300 may operate on an operating system stored in memory 332, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0135] It is understood that in this disclosure, "multiple" refers to two or more, and other quantifiers are similar. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0136] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.

[0137] It is further understood that the terms “center,” “longitudinal,” “lateral,” “front,” “rear,” “up,” “down,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” and “outer,” etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this embodiment and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation.

[0138] It can be further understood that, unless otherwise specified, "connection" includes both direct connections where no other components exist between the two parties and indirect connections where other components exist between them.

[0139] It is further understood that although operations are described in a specific order in the accompanying drawings in the embodiments of this disclosure, this should not be construed as requiring these operations to be performed in the specific order or serial order shown, or requiring all of the shown operations to be performed to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.

[0140] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0141] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An event detection method, characterized in that, The method includes: During a voice call performed by an electronic device based on uplink audio data collected by a microphone, the uplink audio data is acquired. Determine whether target sound data exists in the uplink audio data, wherein the target sound data is the sound data generated when the target event occurs; In response to the presence of target sound data in the uplink audio data, it is determined that an event corresponding to the generation of the target sound data has occurred, and the function corresponding to the event is executed.

2. The method according to claim 1, characterized in that, Before acquiring the upstream audio data, the method further includes: Ensure the target application is running; The target application is used to trigger the process of acquiring the uplink audio data and determining the target sound data when the electronic device is detected to be making a voice call, and the target application is used to perform the function corresponding to the event.

3. The method according to claim 1, characterized in that, The determination that an event corresponding to the generation of the target sound data has occurred includes: In response to the target application acquiring the target sound data, it is determined that an event corresponding to the generation of the target sound data has been detected.

4. The method according to claim 3, characterized in that, The method further includes: In response to the electronic device's sensors detecting an acceleration change data greater than an acceleration change threshold, the target application is triggered to acquire the target sound data; or In response to the presence of target sound data in the uplink audio data, the target sound data is reported to the target application.

5. The method according to claim 3 or 4, characterized in that, The method further includes: In response to the presence of target sound data in the uplink audio data, the target sound data is saved for a set duration; The set duration is determined based on the time required for the target application to acquire the target sound data.

6. The method according to claim 1, characterized in that, Determining whether target sound data exists in the uplink audio data includes: The upstream audio data is copied in real time; After denoising the copied uplink audio data, it is input into a multilayer perceptron neural network model to classify the uplink audio data and determine whether the target sound data exists in the uplink audio data based on the classification results. The multilayer perceptual neural network model is used to classify the uplink audio data, and the classification results include target sound data and / or non-target sound data.

7. The method according to claim 6, characterized in that, The noise reduction process for the copied uplink audio data includes: Remove the target sound data played by the electronic device from the copied uplink audio data.

8. The method according to any one of claims 1 to 7, characterized in that, The target sound data is the collision sound in a car accident scenario, and the function is the emergency call function.

9. An event detection device, characterized in that, The device includes: The acquisition unit is used to acquire the uplink audio data during a voice call performed by an electronic device based on uplink audio data collected by a microphone; The detection unit is used to determine whether target sound data exists in the uplink audio data, wherein the target sound data is the sound data generated when a target event occurs; The processing unit is configured to, in response to the presence of target sound data in the uplink audio data, determine that an event corresponding to the generation of the target sound data has occurred, and execute a function corresponding to the event.

10. The apparatus according to claim 9, characterized in that, The acquisition unit is also used before acquiring the uplink audio data: Ensure the target application is running; The target application is used to trigger the process of acquiring the uplink audio data and determining the target sound data when the electronic device is detected to be making a voice call, and the target application is used to perform the function corresponding to the event.

11. The apparatus according to claim 9, characterized in that, The processing unit determines that an event corresponding to the generation of the target sound data has occurred using the following method: In response to the target application acquiring the target sound data, it is determined that an event corresponding to the generation of the target sound data has been detected.

12. The apparatus according to claim 11, characterized in that, The processing unit is also used for: In response to the electronic device's sensors detecting an acceleration change data greater than an acceleration change threshold, the target application is triggered to acquire the target sound data; or In response to the presence of target sound data in the uplink audio data, the target sound data is reported to the target application.

13. The apparatus according to claim 11 or 12, characterized in that, The processing unit is also used for: In response to the presence of target sound data in the uplink audio data, the target sound data is saved for a set duration; The set duration is determined based on the time required for the target application to acquire the target sound data.

14. The apparatus according to claim 9, characterized in that, The detection unit determines whether target sound data exists in the uplink audio data using the following method: The upstream audio data is copied in real time; After denoising the copied uplink audio data, it is input into a multilayer perceptron neural network model to classify the uplink audio data and determine whether the target sound data exists in the uplink audio data based on the classification results. The multilayer perceptual neural network model is used to classify the uplink audio data, and the classification results include target sound data and / or non-target sound data.

15. The apparatus according to claim 14, characterized in that, The detection unit performs noise reduction processing on the copied uplink audio data in the following manner: Remove the target sound data played by the electronic device from the copied uplink audio data.

16. The apparatus according to any one of claims 9 to 15, characterized in that, The target sound data is the collision sound in a car accident scenario, and the function is the emergency call function.

17. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the event detection method according to any one of claims 1-8.

18. A storage medium, characterized in that, The storage medium stores instructions that, when executed by a processor, enable the execution of the event detection method according to any one of claims 1-8.