Headback detection processing method and device, computer readable medium and electronic equipment
By setting the ear return channel in the terminal device and performing voice detection processing, volume-related events are identified to identify abnormalities in the hardware ear return function and performing repair operations, the problem that the hardware ear return function in the prior art cannot perceive its actual operating effect is solved, effectively detecting and repairing the ear return function, and improving the user experience.
Patent Information
- Application Number
- CN202311514495.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the hardware ear return function has a low delay but its actual operating effect cannot be sensed, resulting in the inability to repair when the function is abnormal, affecting the user experience.
By setting an ear return channel between the audio acquisition element and the audio output element of the terminal device, audio data is acquired for voice detection processing, volume-related events are identified to identify abnormalities in the hardware ear return function, and repair operations are performed.
The actual operation effect detection of the ear return function is realized, and it is timely repaired when abnormalities are abnormal, significantly improving the user experience of the ear return function.
Smart Images

Figure CN119996916A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer and communication technology, and in particular to an ear return detection processing method, device, computer readable medium and electronic device. Background Art
[0002] In-ear monitoring is a type of headphone device, generally used for sound detection in stage performances and recording studios. It can feed back sound signals to performers in real time, helping them to better hear their own performances or singing, so as to better control the details and quality of the performance. With the development of technology, some terminal devices also have an in-ear monitoring function, which is mainly used to collect the sound of the terminal device and then play it on the terminal device. For example, it can be used for anchors to listen to accompaniment, correct pronunciation, and other scenarios.
[0003] Related technologies have proposed a hardware ear return implementation solution, which is to establish a fast channel for audio data in the terminal device, so that the audio data collected by the microphone can be output through the speaker without being transferred through the application layer. Although this solution has a low latency, since the audio data does not pass through the application layer, the actual operation effect of the ear return function cannot be perceived, and thus the ear return function cannot be repaired when an abnormality occurs, which greatly affects the user experience of the ear return function. Summary of the invention
[0004] The embodiments of the present application provide an ear return detection and processing method, device, computer-readable medium and electronic device, which realize the detection of the actual operating effect of the ear return function, and can promptly repair the ear return function when an abnormality occurs, greatly improving the user experience of the ear return function.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to one aspect of an embodiment of the present application, an ear return detection and processing method for a terminal device is provided, wherein an ear return channel is arranged between an audio acquisition element and an audio output element of the terminal device, and the ear return channel is used to directly transmit the audio data collected by the audio acquisition element to the audio output element, so as to realize a hardware ear return function of the terminal device, and the ear return detection and processing method comprises: when the hardware ear return function of the terminal device is turned on, acquiring the audio data collected by the audio acquisition element, and performing voice detection processing on the audio data; if it is detected that the audio data contains human voice, performing abnormal identification processing on the hardware ear return function according to the volume-related events of the terminal device; if it is identified that the hardware ear return function has an abnormality, performing an ear return repair operation for the terminal device.
[0007] According to one aspect of an embodiment of the present application, an ear return detection and processing device for a terminal device is provided, wherein an ear return channel is arranged between an audio acquisition element and an audio output element of the terminal device, and the ear return channel is used to directly transmit the audio data collected by the audio acquisition element to the audio output element to realize a hardware ear return function of the terminal device, and the ear return detection and processing device comprises: an acquisition unit, configured to acquire the audio data collected by the audio acquisition element when the hardware ear return function of the terminal device is turned on, and perform voice detection processing on the audio data; an abnormality recognition unit, configured to perform abnormality recognition processing on the hardware ear return function according to a volume-related event of the terminal device if it is detected that the audio data contains a human voice; and a processing unit, configured to execute an ear return repair operation for the terminal device if it is recognized that the hardware ear return function has an abnormality.
[0008] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit is configured to: obtain the ear return volume set for the terminal device; if the ear return volume exceeds a first set threshold, it is determined that an abnormality is identified in the hardware ear return function.
[0009] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit is configured to: obtain the system volume set for the terminal device; if the system volume exceeds a second set threshold, it is determined that an abnormality is identified in the hardware ear return function.
[0010] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit is configured to: detect the volume adjustment operation; if the number of times the volume is increased within the set time period exceeds a third set threshold, it is determined that an abnormality is identified in the hardware ear return function.
[0011] In some embodiments of the present application, based on the aforementioned scheme, the detecting of the volume adjustment operation includes at least one of the following: detecting an adjustment operation for the system volume of the terminal device, and detecting an adjustment operation for the earphone return volume.
[0012] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit is configured to: detect the closing and opening operations of the ear return function of the terminal device; if the number of times the ear return function is detected to be closed and opened reaches a set number within a set time period, it is determined that an abnormality is identified in the hardware ear return function.
[0013] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit is configured as: according to multiple volume-related events of the terminal device, respectively perform abnormality identification processing on the hardware ear return function to obtain identification results corresponding to each volume-related event; if there are a set number of volume-related events corresponding to the identification results that the hardware ear return function has an abnormality, it is determined that the hardware ear return function has an abnormality.
[0014] In some embodiments of the present application, based on the aforementioned scheme, the acquisition unit is configured to: convert the audio frames contained in the audio data into frequency domain signals; calculate the energy ratio of the signal part in the frequency domain signal whose frequency is lower than the frequency threshold according to the frequency domain signals obtained by converting the audio frames; if the energy ratio is greater than the set ratio, it is determined that the audio frame contains human voice.
[0015] In some embodiments of the present application, based on the aforementioned solution, the processing unit is configured to: close the ear return channel between the audio collection element and the audio output element; and after closing the ear return channel, reopen the ear return channel.
[0016] In some embodiments of the present application, based on the aforementioned scheme, the processing unit is configured to: switch the hardware ear return function to a software ear return function, and the software ear return function obtains the audio data collected by the audio collection element through the application layer, and transmits the audio data to the audio output element through the application layer.
[0017] In some embodiments of the present application, based on the aforementioned solution, the acquisition unit is configured to: acquire the audio data collected by the audio acquisition element through an application layer, and perform voice detection processing on the audio data in the application layer.
[0018] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the ear return detection processing method for a terminal device as described in the above embodiment is implemented.
[0019] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more computer programs, wherein when the one or more computer programs are executed by the one or more processors, the electronic device implements the ear return detection processing method for a terminal device as described in the above embodiments.
[0020] According to one aspect of an embodiment of the present application, a computer program product is provided, the computer program product comprising a computer program, the computer program being stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the ear return detection processing method for a terminal device provided in the above-mentioned various optional embodiments.
[0021] In the technical solutions provided in some embodiments of the present application, when the hardware ear return function of the terminal device is turned on, the audio data collected by the audio collection element is obtained, and when it is detected that the audio data contains human voice, the hardware ear return function is identified and processed according to the volume-related events of the terminal device. When it is identified that the hardware ear return function has an abnormality, the ear return repair operation for the terminal device is performed, so that when the hardware ear return function is started by the terminal device, the abnormality of the hardware ear return function can be identified and processed through the volume-related events of the terminal device, thereby realizing the detection of the actual operating effect of the ear return function, and the ear return function can be repaired in time when an abnormality occurs, which greatly improves the user experience of the ear return function.
[0022] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A schematic diagram showing a hardware ear return function implemented in a terminal device is shown;
[0024] Figure 2 A schematic diagram showing an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied;
[0025] Figure 3 A flowchart of an ear return detection processing method for a terminal device according to an embodiment of the present application is shown;
[0026] Figure 4 A structural diagram of a hardware ear return detection module according to an embodiment of the present application is shown;
[0027] Figure 5 A structural diagram of a volume detection module according to an embodiment of the present application is shown;
[0028] Figure 6 A flowchart of a method for detecting abnormality of hardware ear-return function according to an embodiment of the present application is shown;
[0029] Figure 7 A block diagram of an ear return detection processing device for a terminal device according to an embodiment of the present application is shown;
[0030] Figure 8A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0031] The exemplary embodiments are now described in a more comprehensive manner with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be understood as being limited to these examples; on the contrary, the purpose of providing these embodiments is to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0032] In addition, the features, structures or characteristics described in the present application may be combined in one or more embodiments in any suitable manner. In the following description, there are many specific details so that the embodiments of the present application can be fully understood. However, those skilled in the art will appreciate that when implementing the technical scheme of the present application, all the detailed features in the embodiments may not be needed, one or more specific details may be omitted, or other methods, elements, devices, steps, etc. may be adopted.
[0033] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0034] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0035] It should be noted that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0036] It is understandable that the present application can display a prompt interface or pop-up window before and during the process of collecting relevant data (such as audio data collected by the audio collection element), and the prompt interface or pop-up window is used to prompt the user that the relevant data is currently being collected, so that the present application only starts to execute the relevant steps of obtaining the relevant data after obtaining the user's confirmation operation on the prompt interface or pop-up window, otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the relevant data are terminated, that is, the relevant data is not obtained. In other words, all data collected by the present application are collected with the user's consent and authorization, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0037] In audio and video application scenarios, earphone monitoring is often used. Earphone monitoring is a type of headphone device, generally used for sound detection in stage performances and recording studios. It can feed back sound signals to performers in real time, helping them to better hear their own performances or singing, so as to better control the details and quality of the performance. With the development of technology, some terminal devices also have earphone monitoring functions, which are mainly used to collect the sound of the terminal device and then play it on the terminal device. For example, it can be used for anchors to listen to accompaniment and correct pronunciation.
[0038] In the related art, there are two implementation schemes for the ear return function on the terminal device. One is the ear return function implemented by software (referred to as software ear return function), which specifically obtains the audio data collected by the microphone by the application layer, and then plays it out through the control of the application layer. The advantage of the software ear return function is good compatibility, but because it needs to be transferred through the application layer, the delay is relatively high. Another implementation scheme of the ear return function on the terminal device is the ear return function implemented by hardware (referred to as hardware ear return function), as shown in the following figure. Figure 1 As shown, a fast channel between the microphone and the speaker can be established in the hardware abstraction layer (HAL) within the system, so that the audio data collected by the microphone does not need to be transferred through the application layer, but is directly sent to the speaker for playback. The advantage of the hardware ear return solution is low latency, but the compatibility of the hardware ear return solution is poor, and it is prone to problems such as low sound or no sound. At the same time, since the application layer cannot obtain the data played by the ear return, it is impossible to perceive the actual operating effect of the hardware ear return (this is because the audio data collected by the microphone does not pass through the application layer, so it is impossible to perceive whether the ear return has problems such as low sound or no sound), so it cannot be repaired, which greatly affects the user experience of the ear return function.
[0039] Based on this, the technical solution of the embodiment of the present application proposes a new ear return detection and processing solution for a terminal device, which is specifically applied to a terminal device with a hardware ear return function. An ear return channel is arranged between the audio acquisition element (such as a microphone) and the audio output element (such as a speaker) of the terminal device. The ear return channel is used to directly transmit the audio data collected by the audio acquisition element to the audio output element to realize the hardware ear return function of the terminal device.
[0040] Specifically, if Figure 2 As shown, a hardware ear return detection module is provided in the application layer, and the hardware ear return detection module is used to detect whether the hardware ear return function of the terminal device is abnormal. When the hardware ear return function of the terminal device is turned on, the acquisition module of the application layer can collect the audio data recorded by the microphone through the acquisition thread of the framework layer (Framework), and then the hardware ear return detection module performs voice detection processing on the audio data. If it is detected that the audio data contains human voice, the hardware ear return function can be abnormally identified and processed according to the volume-related events of the terminal device (such as the size of the system volume, the size of the set ear return volume, the volume adjustment operation, etc.). If the hardware ear return function is identified to be abnormal, the ear return repair operation for the terminal device can be performed (such as restarting the hardware ear return function, or switching to the software ear return function, etc.).
[0041] It should be noted that the process of abnormal identification and processing of the hardware ear return function can be processed by the hardware ear return detection module on the terminal device or by a server connected to the terminal device. If it is processed by the server, the terminal device needs to transmit the audio data collected by the microphone and the volume-related events of the terminal device to the server for abnormal identification and processing.
[0042] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal device can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart home, vehicle terminal, aircraft, etc., but is not limited to this. The terminal device and the server can be directly or indirectly connected by wired or wireless communication, and this application does not limit this.
[0043] The implementation details of the technical solution of the embodiment of the present application are described in detail below:
[0044] Figure 3 A flowchart of an ear return detection processing method for a terminal device according to an embodiment of the present application is shown. An ear return channel is provided between the audio acquisition element and the audio output element of the terminal device. The ear return channel is used to directly transmit the audio data collected by the audio acquisition element to the audio output element to realize the hardware ear return function of the terminal device. The ear return detection processing method for a terminal device can be executed by an electronic device, which can be a terminal device or a server, etc. Figure 3 As shown, the ear return detection processing method for the terminal device includes at least steps S310 to S330, which are described in detail as follows:
[0045] In step S310, when the hardware ear return function of the terminal device is turned on, the audio data collected by the audio collection element is obtained, and voice detection processing is performed on the collected audio data.
[0046] In some optional embodiments, a touch button for turning on the hardware ear return function may be provided on the terminal device, and the user may turn on the hardware ear return function of the terminal device by triggering the touch button, or may turn off the hardware ear return function of the terminal device by triggering the touch button. The hardware ear return function of the terminal device is that the audio acquisition element (such as a microphone) of the terminal device transmits the collected audio data directly to the audio output element (such as a speaker) through the transmission channel between the audio output element (such as a speaker) without passing through the application layer, so that the transmission delay of the audio data can be reduced.
[0047] In some optional embodiments, the audio data collected by the audio collection element can be obtained through the application layer, and the audio data can be processed by voice detection in the application layer. Figure 2 As shown, the audio data acquired by the microphone (ie, the audio acquisition element) can be collected through the collection thread of the framework layer (Framework), and then the audio data can be subjected to voice detection processing in the application layer.
[0048] In one embodiment of the present application, the process of performing voice detection processing on audio data can be to detect whether the audio data contains human voice. If it contains human voice, it means that the audio acquisition element has collected the audio data that needs to be transmitted to the ear return device. Optionally, when detecting whether the audio data contains human voice, the detection can be performed by analyzing the spectral characteristics, energy characteristics, etc. of the audio data. For example, the spectral characteristics of the audio can be obtained by calculating the short-time Fourier transform of the audio data, and then classified and detected according to the spectral characteristics of the human voice. Specifically, the audio frame contained in the audio data can be converted into a frequency domain signal; then, according to the frequency domain signal obtained by the audio frame conversion, the energy ratio of the signal part with a frequency lower than the frequency threshold in the frequency domain signal is calculated. If the energy ratio is greater than the set ratio, it is determined that the audio frame contains human voice. This is because the human voice is mainly distributed below 2kHz, while the noise is above 2kHz, so it is possible to identify whether the audio data contains human voice by extracting the signal energy below 2kHz. Among them, the frequency threshold can be a value less than or equal to 2kHz.
[0049] In some optional embodiments, human voices can also be detected based on machine learning. Specifically, a machine learning model can be trained using a labeled audio data set to allow the machine learning model to learn the characteristics of human voices, and then the model can be used to detect new audio data. For example, a support vector machine, a naive Bayes classifier, a deep learning model, etc. can be used.
[0050] Continue to refer to Figure 3 As shown, in step S320, if it is detected that the audio data contains human voice, an abnormal identification process is performed on the hardware ear return function according to the volume-related events of the terminal device.
[0051] In some optional embodiments, the volume-related event may be the ear return volume set for the terminal device. Then, when the hardware ear return function is subjected to abnormal identification processing according to the volume-related event of the terminal device, the ear return volume set for the terminal device may be obtained. If the ear return volume set for the terminal device exceeds the first set threshold, it may be determined that an abnormality has been identified in the hardware ear return function. This is because under normal circumstances, the ear return volume set by the user for the terminal device is within a certain range. If it exceeds the range, it means that the ear return volume set by the user is large, which means that the normal volume can no longer meet the needs of the user. In fact, it means that the sound of the ear return is very small, or even the ear return sound cannot be heard. At this time, it can be considered that the hardware ear return function has an abnormality.
[0052] In some optional embodiments, the volume-related event may be the system volume set for the terminal device. Then, when the hardware ear return function is subjected to abnormal identification processing according to the volume-related event of the terminal device, the system volume set for the terminal device may be obtained. If the system volume set for the terminal device exceeds the second set threshold, it may be determined that the hardware ear return function is abnormal. This is because under normal circumstances, the system volume set by the user for the terminal device is within a certain range. If it exceeds the range, it means that the system volume set by the user is relatively large, which means that the sound of the ear return may be very small or even inaudible. Therefore, the user will increase the system volume. In this case, it can be considered that the hardware ear return function is abnormal.
[0053] In some optional embodiments, the volume-related event may be a volume adjustment operation. Then, when the hardware ear return function is subjected to abnormal identification processing according to the volume-related events of the terminal device, the volume adjustment operation may be detected. If the number of times the volume is increased within the set time period exceeds the third set threshold, it is determined that an abnormality has been identified in the hardware ear return function. This is because if the user frequently triggers the volume increase operation, it means that the ear return sound may be very small, or even the ear return sound cannot be heard, so the user will keep turning up the volume. In this case, it can be considered that the hardware ear return function has an abnormality. It should be noted that the volume adjustment operation may be an adjustment operation for the system volume of the terminal device, or an adjustment operation for the ear return volume, or an adjustment operation for the system volume of the terminal device and an adjustment operation for the ear return volume.
[0054] In some optional embodiments, the turning off and on of the ear return function of the terminal device can also be detected. If the number of times the ear return function is turned off and on reaches a set number within a set time, it is determined that an abnormality has been identified in the hardware ear return function. This is because after turning on the hardware ear return function, if the user frequently turns off and on the ear return function in a short period of time, it means that the actual effect of the ear return function is not ideal, so the user wants to try to repair it by restarting. Therefore, in this case, it can be determined that there is an abnormality in the hardware ear return function.
[0055] In some optional embodiments, when performing abnormal identification processing on the hardware ear return function, it is also possible to perform abnormal identification processing on the hardware ear return function respectively according to multiple volume-related events of the terminal device to obtain identification results corresponding to each volume-related event. If the identification results corresponding to a set number of volume-related events indicate that the hardware ear return function is abnormal, then it is determined that the hardware ear return function is abnormal. For example, the abnormality detection of the hardware ear return function can be performed through all or part of the ear return volume, system volume, volume adjustment operation, the number of times the ear return function is turned off and on, etc. in the above embodiments. If a set number (for example, greater than or equal to 2) of detection results indicate that the hardware ear return function is abnormal, then it is determined that the hardware ear return function is abnormal.
[0056] Continue to refer to Figure 3 As shown, in step S330, if it is identified that the hardware ear return function is abnormal, an ear return repair operation for the terminal device is performed.
[0057] In some optional embodiments, performing the ear return modification operation on the terminal may be closing the ear return channel between the audio acquisition element and the audio output element, and reopening the ear return channel after closing the ear return channel. That is, in this embodiment, the ear return repair operation is implemented by restarting.
[0058] In some optional embodiments, performing an ear return modification operation on a terminal may be to switch a hardware ear return function to a software ear return function, wherein the software ear return function obtains audio data collected by an audio acquisition element through an application layer, and transmits the audio data to an audio output element through an application layer. That is, in this embodiment, the ear return repair operation is implemented by switching to a software ear return function. It should be noted that the restart method and the switching to a software ear return function method may be selected; or the repair may be performed by restarting first, and if the ear return abnormality cannot be repaired, the repair may be performed by switching to a software ear return function; or the repair may be performed by switching to a software ear return function first, and if the ear return abnormality cannot be repaired, the repair may be performed by restarting.
[0059] It can be seen that the technical solution of the above-mentioned embodiments of the present application can realize the abnormal identification and processing of the hardware ear return function through the volume-related events of the terminal device, thereby realizing the detection of the actual operating effect of the ear return function, and can promptly repair the ear return function when an abnormality occurs, thereby greatly improving the user experience of the ear return function.
[0060] The following combination Figure 2 ,as well as Figures 4 to 6 The implementation details of the technical solution of the embodiment of the present application are elaborated in detail:
[0061] Reference Figure 2 As shown, in the embodiment of the present application, a hardware ear return detection module is set in the application, and the hardware ear return detection module is used to detect whether the hardware ear return function is normal. Specifically, when the hardware ear return function is turned on, the acquisition module can obtain the audio data collected by the microphone through the acquisition thread, and then send the audio data to the hardware ear return detection module, and then the hardware ear return detection module outputs the detection result of whether the hardware ear return function is normal.
[0062] In one embodiment of the present application, Figure 4 As shown, the hardware ear return detection module includes a Vad (Voiceactivity detection) detection module, a volume detection module and a button detection module. When the hardware ear return function is turned on, the acquisition module obtains the audio data collected by the microphone, and then detects the collected data through the Vad detection module to confirm whether it contains human voice. When a data frame containing human voice appears, the volume event and the button event are obtained, and the volume detection module and the button detection module are used to judge whether the user's ear return data is normal based on the volume event and the button event, thereby deducing whether the hardware ear return function is normal. When the hardware ear return detection module finds that the hardware ear return function is abnormal, it executes the operation of restarting the ear return or switching to the software ear return function to implement the repair processing of the ear return function.
[0063] The input of the Vad detection module is an audio data frame, which is generally a 20ms or 10ms data frame. The Vad detection module then analyzes the data frame, for example, by converting the data frame into a frequency domain signal, and then calculating the ratio of the energy of the frequency domain signal below 2kHz to the total energy of the frame data. When it is greater than a certain threshold, the data frame is considered to contain human voice, and the information is recorded in the data frame structure. This is because human voice is mainly distributed below 2kHz, while noise is above 2kHz. By extracting the audio energy below 2kHz and then comparing it with the threshold, it is possible to identify whether it contains human voice.
[0064] In one embodiment of the present application, Figure 5As shown, the volume detection module may include an ear return volume module and a system volume module. The ear return volume module is used to provide the user with an entry for setting the ear return volume. When the user needs to adjust the ear return volume, the ear return volume can be adjusted by setting the ear return volume entry. Optionally, in an embodiment of the present application, the adjustment range of the hardware ear return volume can be expanded from 0-150 (the value is only an example) to greater than 150, such as the ear return volume that the user can set from the adjustment interval of 0-150 to 0-160. When the volume set by the user is within the range of 0-150, the ear return volume can be adjusted synchronously through the interface provided by the hardware ear return, so that the ear return volume will also change synchronously in theory. At this time, 150 is also the upper limit of the adjustment of the hardware ear return volume. When the volume set by the user exceeds 150, the application layer can record the event, no longer operate the hardware ear return volume interface, and the hardware ear return does not support setting a volume exceeding 150. In this case, it can be considered that the hardware ear return function is abnormal.
[0065] In one embodiment of the present application, the system volume module is mainly used to adjust the system volume of the terminal device. The user can adjust the volume by pressing a button (either a physical case or a virtual case), and the application layer can also obtain the current system volume. The system volume is essentially the gain of the playback data, and the ear return volume will also be correspondingly gained through the configuration of the system volume. For example, if the system volume is increased, the ear return volume will also be increased synchronously; if the system volume is lowered, the ear return volume will also be reduced synchronously.
[0066] In one embodiment of the present application, the key detection module mainly detects key events generated by the trigger operation of the volume key. Optionally, the key events include volume up key trigger events and volume down key trigger events. The user's intention to adjust the volume can also be obtained by detecting key events.
[0067] Based on the introduction of the above embodiments, Figure 6 The flowchart of the method for detecting abnormality of hardware ear return function according to one embodiment of the present application is shown, and includes the following steps:
[0068] S601, reading the collected audio data, that is, the reading collection module obtains the audio data collected by the microphone through the collection thread.
[0069] S602, perform Vad detection on the audio data to determine whether the audio data frame contains human voice, if it contains human voice, execute S603; if it does not contain human voice, no judgment is made, that is, S603 is not executed.
[0070] S603, detect the volume of the hardware ear return function, that is, detect the volume set for the hardware ear return function. If it exceeds 150 (the value is only an example, which represents the maximum value of the set volume range), set the hardware ear return abnormal flag; if it does not exceed 150, it is considered that the hardware ear return volume is normal. Then execute S604.
[0071] This is because when the hardware earphone volume is set to more than 150, it means that the volume limit of the hardware earphone can no longer meet the user's needs. In fact, it means that the user's earphone sound is very small, or even completely unable to hear the earphone sound. At this time, the hardware earphone function is considered abnormal.
[0072] S604, detecting the system playback volume, that is, detecting the system volume of the terminal device. If it exceeds the threshold, the system volume abnormality flag is set; if it does not exceed the threshold, the system volume is considered normal. Then execute S605.
[0073] Since headphones are worn in the in-ear monitoring scenario, a moderate volume is often enough to meet the listening needs of the playback volume. When the volume is adjusted beyond the threshold or even reaches the maximum, the in-ear monitoring sound heard by the user is often weak, or even no sound is heard at all. At this time, it can be considered that the hardware in-ear monitoring function is abnormal.
[0074] S605, detecting key events, that is, detecting volume up key events and volume down key events of the terminal device. If the number of times the volume up key is triggered exceeds the threshold, the volume key event abnormal flag is set; if it does not exceed the threshold, the volume key event is considered normal. Then execute S606.
[0075] In the earphone monitoring scenario, if the user frequently triggers the volume up button, even when the system volume has reached the upper limit, the volume up button is still triggered. Often, the earphone monitoring sound heard by the user at this time is relatively weak, or even no earphone monitoring sound is heard at all. At this time, it is considered that the hardware earphone monitoring function is abnormal.
[0076] S606, counting the number of abnormal events, if the number of abnormal events exceeds 1, then the hardware ear return function is considered abnormal; if the number of abnormal events does not exceed 1, then the hardware ear return function is considered normal. When it is determined that the hardware ear return function is abnormal, the application layer can be notified to restart the hardware ear return function, or switch to the software ear return function.
[0077] It should be noted that in Figure 6In the process shown, detection and processing are performed in the order of detecting the hardware earphone return volume, the system playback volume and the key events. In other embodiments of the present application, the detection order of these three detection methods can be arbitrarily exchanged, or can also be performed simultaneously. That is, in the embodiments of the present application, the detection order of the hardware earphone return volume, the system playback volume and the key events is not strictly limited, and can be set arbitrarily according to needs during actual execution.
[0078] The technical solution of the above-mentioned embodiments of the present application enables the application layer to indirectly sense whether the hardware ear return function is abnormal by detecting volume events and key events, and when the hardware ear return function is abnormal, it can quickly initiate repair measures, thereby improving the user experience of the ear return.
[0079] The following describes an embodiment of the device of the present application, which can be used to execute the ear return detection processing method for a terminal device in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the ear return detection processing method for a terminal device in the above embodiment of the present application.
[0080] Figure 7 A block diagram of an ear return detection and processing device for a terminal device according to an embodiment of the present application is shown, wherein an ear return channel is arranged between the audio acquisition element and the audio output element of the terminal device, and the ear return channel is used to directly transmit the audio data collected by the audio acquisition element to the audio output element to realize the hardware ear return function of the terminal device.
[0081] Reference Figure 7 As shown, an ear return detection processing device 700 for a terminal device according to an embodiment of the present application includes: an acquisition unit 702, an abnormality identification unit 704 and a processing unit 706.
[0082] Among them, the acquisition unit 702 is configured to acquire the audio data collected by the audio acquisition element when the hardware ear return function of the terminal device is turned on, and perform voice detection processing on the audio data; the abnormality recognition unit 704 is configured to perform abnormality recognition processing on the hardware ear return function according to the volume-related events of the terminal device if it is detected that the audio data contains human voice; the processing unit 706 is configured to execute the ear return repair operation for the terminal device if it is recognized that the hardware ear return function has an abnormality.
[0083] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit 704 is configured to: obtain the ear return volume set for the terminal device; if the ear return volume exceeds a first set threshold, it is determined that an abnormality is identified in the hardware ear return function.
[0084] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit 704 is configured to: obtain the system volume set for the terminal device; if the system volume exceeds a second set threshold, it is determined that an abnormality is identified in the hardware ear return function.
[0085] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit 704 is configured to: detect the volume adjustment operation; if the number of times the volume is increased within the set time period exceeds a third set threshold, it is determined that an abnormality is identified in the hardware ear return function.
[0086] In some embodiments of the present application, based on the aforementioned scheme, the detecting of the volume adjustment operation includes at least one of the following: detecting an adjustment operation for the system volume of the terminal device, and detecting an adjustment operation for the earphone return volume.
[0087] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit 704 is configured to: detect the closing and opening operations of the ear return function of the terminal device; if the number of times the ear return function is detected to be closed and opened reaches a set number within a set time period, it is determined that an abnormality is identified in the hardware ear return function.
[0088] In some embodiments of the present application, based on the aforementioned scheme, the abnormality identification unit 704 is configured as: according to multiple volume-related events of the terminal device, respectively perform abnormality identification processing on the hardware ear return function to obtain identification results corresponding to each volume-related event; if there are a set number of volume-related events corresponding to the identification results that the hardware ear return function has an abnormality, it is determined that the hardware ear return function has an abnormality.
[0089] In some embodiments of the present application, based on the aforementioned scheme, the acquisition unit 702 is configured to: convert the audio frames contained in the audio data into frequency domain signals; calculate the energy ratio of the signal part with a frequency lower than a frequency threshold in the frequency domain signal according to the frequency domain signals obtained by converting the audio frames; if the energy ratio is greater than the set ratio, it is determined that the audio frame contains human voice.
[0090] In some embodiments of the present application, based on the aforementioned solution, the processing unit 706 is configured to: close the ear return channel between the audio collection element and the audio output element; and after closing the ear return channel, reopen the ear return channel.
[0091] In some embodiments of the present application, based on the aforementioned scheme, the processing unit 706 is configured to: switch the hardware ear return function to a software ear return function, and the software ear return function obtains the audio data collected by the audio collection element through the application layer, and transmits the audio data to the audio output element through the application layer.
[0092] In some embodiments of the present application, based on the aforementioned solution, the acquisition unit 702 is configured to: acquire the audio data collected by the audio acquisition element through the application layer, and perform voice detection processing on the audio data in the application layer.
[0093] Figure 8 A schematic diagram of the structure of a computer system of an electronic device suitable for implementing an embodiment of the present application is shown, and the electronic device may be a terminal device or a server in the aforementioned embodiment.
[0094] It should be noted that Figure 8 The computer system 800 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0095] like Figure 8 As shown, the computer system 800 may include a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage part 808 to the random access memory (RAM) 803, such as executing the method described in the above embodiment. In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, ROM 802 and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0096] The following components can be connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read therefrom is installed into the storage section 808 as needed.
[0097] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program is used to perform the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 809, and / or installed from a removable medium 811. When the computer program is executed by a central processing unit (CPU) 801, various functions defined in the system of the present application are executed.
[0098] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a computer program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0099] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and a computer program.
[0100] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0101] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more computer programs, and when the above one or more computer programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.
[0102] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0103] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable an electronic device to execute the method according to the implementation methods of the present application.
[0104] For example, the electronic device can be a terminal device, and the terminal device can execute Figure 3 The ear return detection processing method for the terminal device is shown.
[0105] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.
[0106] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. An ear return detection processing method for a terminal device, characterized in that: An ear return channel is provided between the audio collection element and the audio output element of the terminal device, and the ear return channel is used to directly transmit the audio data collected by the audio collection element to the audio output element to realize the hardware ear return function of the terminal device. The ear return detection processing method includes: When the hardware ear return function of the terminal device is turned on, the audio data collected by the audio collection element is obtained, and the audio data is subjected to voice detection processing; If it is detected that the audio data contains human voice, an abnormality recognition process is performed on the hardware ear return function according to the volume-related events of the terminal device; If it is identified that the hardware earphone return function is abnormal, an earphone return repair operation is performed on the terminal device.
2. The ear return detection processing method according to claim 1, characterized in that: The abnormality recognition and processing of the hardware ear return function is performed according to the volume-related events of the terminal device, including: Obtaining the earphone return volume set for the terminal device; If the earphone return volume exceeds a first set threshold, it is determined that an abnormality has occurred in the hardware earphone return function.
3. The ear return detection processing method according to claim 1, characterized in that: The abnormality recognition and processing of the hardware ear return function is performed according to the volume-related events of the terminal device, including: Obtaining the system volume set for the terminal device; If the system volume exceeds a second set threshold, it is determined that an abnormality exists in the hardware ear return function.
4. The ear return detection processing method according to claim 1, characterized in that: The abnormality recognition and processing of the hardware ear return function is performed according to the volume-related events of the terminal device, including: Detect volume adjustment operation; If the number of times the volume is turned up within the set time period exceeds a third set threshold, it is determined that an abnormality has been identified in the hardware ear return function.
5. The ear return detection processing method according to claim 4, characterized in that: Detecting a volume adjustment operation includes at least one of the following: Detecting an adjustment operation on the system volume of the terminal device and detecting an adjustment operation on the earphone return volume.
6. The ear return detection processing method according to claim 1, characterized in that: The abnormality recognition and processing of the hardware ear return function is performed according to the volume-related events of the terminal device, including: Detecting the turning off and on operation of the ear return function of the terminal device; If it is detected that the ear return function is turned off and on for a set number of times within a set time, it is determined that an abnormality has been identified in the hardware ear return function.
7. The ear return detection processing method according to claim 1, characterized in that: The abnormality recognition and processing of the hardware ear return function is performed according to the volume-related events of the terminal device, including: According to the multiple volume-related events of the terminal device, respectively perform abnormal identification processing on the hardware ear return function to obtain identification results corresponding to each volume-related event; If the recognition results corresponding to a set number of volume-related events are that the hardware ear return function has an abnormality, it is determined that the hardware ear return function has an abnormality.
8. The ear return detection processing method according to claim 1, characterized in that: Performing voice detection processing on the audio data includes: Converting the audio frames contained in the audio data into frequency domain signals; Calculating, according to the frequency domain signal obtained by converting the audio frame, an energy ratio of a signal portion whose frequency is lower than a frequency threshold in the frequency domain signal; If the energy ratio is greater than a set ratio, it is determined that the audio frame contains human voice.
9. The ear return detection processing method according to claim 1, characterized in that: Performing an ear return repair operation on the terminal device includes: Closing the ear return channel between the audio collection element and the audio output element; After closing the ear return channel, reopen the ear return channel.
10. The ear-return detection processing method according to claim 1, characterized in that: Performing an ear return repair operation on the terminal device includes: The hardware ear return function is switched to a software ear return function, wherein the software ear return function obtains the audio data collected by the audio collection element through the application layer, and transmits the audio data to the audio output element through the application layer.
11. The ear-return detection processing method according to any one of claims 1 to 10, characterized in that: Acquiring audio data collected by the audio collection element and performing voice detection processing on the audio data includes: The audio data collected by the audio collection element is acquired through the application layer, and voice detection processing is performed on the audio data in the application layer.
12. An ear return detection processing device for a terminal device, characterized in that: An ear return channel is provided between the audio collection element and the audio output element of the terminal device, and the ear return channel is used to directly transmit the audio data collected by the audio collection element to the audio output element to realize the hardware ear return function of the terminal device. The ear return detection processing device includes: an acquisition unit, configured to acquire the audio data collected by the audio acquisition element when the hardware ear return function of the terminal device is turned on, and perform voice detection processing on the audio data; an abnormality identification unit, configured to perform abnormality identification processing on the hardware ear return function according to a volume-related event of the terminal device if it is detected that the audio data contains human voice; The processing unit is configured to perform an ear-return repair operation for the terminal device if an abnormality is identified in the hardware ear-return function.
13. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the ear-return detection processing method for a terminal device as described in any one of claims 1 to 11 is implemented.
14. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more computer programs, which, when executed by the one or more processors, enables the electronic device to implement the ear return detection processing method for a terminal device as described in any one of claims 1 to 11.
15. A computer program product, characterized in that The computer program product includes a computer program, which is stored in a computer-readable storage medium. The processor of the electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device performs the ear return detection processing method for a terminal device as described in any one of claims 1 to 11.