Earphone control method, electronic device, and readable medium

CN116546381BActive Publication Date: 2026-09-22HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310633342.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2026-09-22
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

[0004]为解决由于TWS耳机开启了降噪模式,导致用户错过重要提醒的问题,本申请实施例提供了一种耳机的控制方法、电子设备及可读介质

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116546381B_ABST
    Figure CN116546381B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of wireless devices, and discloses a control method of earphones, an electronic device and a readable medium. The control method of earphones can be used in the process in which the earphones are in a noise reduction mode. When it is determined, based on a preset audio template, that environmental sound of an environment in which a user is currently located includes an important reminder that the user may need to hear, the earphones are controlled to switch from the noise reduction mode to a pass-through mode, so that the user can hear environmental sound from the outside world, and missing the important reminder by the user is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless device technology, and in particular to a control method for headphones, an electronic device, and a readable medium. Background Technology

[0002] With the development of true wireless stereo (TWS) earbuds, prolonged use of TWS earbuds has become commonplace. Among them, such as... Figure 1 As shown, the TWS earbuds 10, earbuds A and B, are independent of each other and can achieve wireless separation of the left and right channels without the need for a cable connection. During the use of TWS earbuds, users often prefer to turn on noise cancellation mode (i.e., a mode that reduces ambient noise) in many scenarios to improve the wearing experience and obtain higher quality music in noisy environments.

[0003] However, in some scenarios, the noise cancellation mode on TWS earbuds may cause users to miss important notifications. For example, while riding the subway, if noise cancellation is enabled, users might miss the station announcements and miss their stop. Summary of the Invention

[0004] To address the issue of users missing important notifications due to noise cancellation mode being enabled in TWS earphones, embodiments of this application provide an earphone control method, an electronic device, and a readable medium.

[0005] In a first aspect, embodiments of this application provide a control method for headphones, applied to an electronic device. The control method for headphones includes: the headphones being in noise cancellation mode; detecting that the matching degree between ambient sound data in the environment where the headphones are located and an audio template meets the matching conditions; and controlling the headphones to switch from noise cancellation mode to pass-through mode.

[0006] It is understood that the ambient sound data of the environment in which the headphones are located corresponds to the audio data of the user's current environment in this embodiment of the application.

[0007] In this embodiment, while the headphones are in noise-canceling mode, audio data of the user's current environment is collected in real time and matched with a preset audio template to determine whether the audio data of the user's current environment includes an important reminder that the user needs to hear. For example, if the matching degree between the audio data and the preset audio template is greater than a first threshold, it can be determined that the audio data includes an important reminder that the user needs to hear. At this time, controlling the headphones to enter pass-through mode can effectively prevent the user from missing important reminders.

[0008] In one possible implementation, detecting that the matching degree between the ambient sound data in the environment where the headphones are located and the audio template meets the matching condition includes: detecting that the matching degree between the ambient sound data in the environment where the headphones are located and the audio template is greater than a first threshold.

[0009] It is understandable that the first threshold can be set according to actual needs, such as 0.75 or 0.8.

[0010] In one possible implementation, detecting that the matching degree between the ambient sound data and the audio template in the environment where the headphones are located meets the matching condition includes: detecting that the matching value determined based on the ambient sound data and the audio template, and that the first event probability determined by the ambient sound data and the scene data of the environment where the headphones are located is greater than a second threshold, wherein the first event probability is used to reflect the possibility of the ambient sound data matching the audio template.

[0011] It can be understood that the first event probability corresponds to the event occurrence probability p in this application embodiment, and the second threshold corresponds to the decision threshold θ in this application embodiment. When the first event probability is greater than the second threshold, it indicates that the ambient sound data in the environment where the headphones are located is the external ambient sound that the user needs to hear in the environment where the headphones are located. For example, when the first event probability is greater than the second threshold, it is determined that the ambient sound data in the environment where the headphones are located is the airport check-in announcement that the user needs to hear.

[0012] In one possible implementation, scene data includes location data and time data.

[0013] It is understood that the location data corresponds to the location information in the embodiments of this application, and the time data corresponds to the time information in the embodiments of this application. The location information can be the longitude, latitude, and altitude information of the user's current location obtained by the positioning system of the electronic device. The time information can be the current date, the user's dwell time, and the user's movement time.

[0014] In one possible implementation, the ambient sound data includes music features and / or speech features, the audio template includes music features and / or speech features, and detecting that the matching degree between the ambient sound data in the environment where the headphones are located and the audio template is greater than a first threshold includes: detecting that the matching degree between any music feature of the ambient sound data in the environment where the headphones are located and the corresponding music feature in the audio template is greater than the first threshold; and / or detecting that the matching degree between any speech feature of the ambient sound data in the environment where the headphones are located and the corresponding speech feature in the audio template is greater than the first threshold.

[0015] In one possible implementation, the method of determining the first event probability based on the matching value determined by the ambient sound data and the audio template, and the scene data of the environment in which the headphones are located, includes: determining the matching value based on the matching degree of the ambient sound data and the audio template; inputting the ambient sound data and scene data into a preset scene model and a first event model to determine the second event probability, wherein the second event probability is used to reflect the degree of matching between the ambient sound data and the preset event corresponding to the first event model; and determining the first event probability based on the matching value and the second event probability.

[0016] It can be understood that the second event probability corresponds to the event occurrence probability t in the embodiments of this application, and the first event model corresponds to the typical event model in the embodiments of this application. Based on the second event probability, the probability that the event corresponding to the ambient sound data is an event in which the user needs to hear external ambient sound in the current environment can be determined. For example, by inputting the ambient sound data and scene data into a preset scene model and the first event model, the probability that the event corresponding to the current scene is the airport check-in announcement is determined to be 0.8.

[0017] In this embodiment, ambient sound data and scene data are input into a preset scene model and a first event model to determine the probability of a second event, i.e., the probability that the event corresponding to the ambient sound data is an event in which the user needs to hear external ambient sounds in the current environment. The probability of the second event is combined with the matching value to calculate the probability of a first event, so as to further determine the likelihood that the event corresponding to the ambient sound data is an event in which the user needs to hear external ambient sounds in the current environment, thereby improving the accuracy of the judgment result and avoiding switching to pass-through mode when it is not necessary.

[0018] In one possible implementation, ambient sound data and scene data are input into a preset scene model and a first event model to determine the probability of the second event, including: inputting ambient sound data and scene data into a preset scene model to determine the current scene corresponding to the environment in which the headphones are located; and inputting ambient sound data and scene data into the first event model corresponding to the current scene to determine the probability of the second event.

[0019] It is understandable that the pre-set scenario model is a classification model. The pre-set scenario model includes at least one typical scenario in which a user needs to hear external ambient sounds, such as the check-in scenario at an airport, the ticket inspection scenario at a train station, the triage and calling scenario at a hospital, and the lost and found scenario in a shopping mall.

[0020] In one possible implementation, determining the probability of the first event based on the matching value and the probability of the second event includes: obtaining the weight corresponding to the matching value; and determining the probability of the first event based on the matching value, the weight corresponding to the matching value, and the probability of the second event.

[0021] It can be understood that the matching value includes a first quantity and a second quantity. The first quantity corresponds to the voice matching value s in this application embodiment, and the second quantity corresponds to the music matching value m in this application embodiment. That is, the matching value includes the voice matching value s and the music matching value m. The weights corresponding to the matching values ​​correspond to the weight w1 corresponding to the music feature (i.e., the music matching value m) and the weight w2 corresponding to the voice feature (i.e., the voice matching value s). The probability of the first event = the probability of the event occurring p = t + w1·m + w2·s.

[0022] In one possible implementation, determining a matching value based on the matching degree between ambient sound data and an audio template includes: determining a first quantity based on the difference between the number of speech features in the ambient sound data whose matching degree with the corresponding speech features in the audio template is greater than a first threshold and the number of speech features in the ambient sound data whose matching degree with the corresponding speech features in the audio template is less than or equal to the first threshold; determining a second quantity based on the difference between the number of music features in the ambient sound data whose matching degree with the corresponding music features in the audio template is greater than the first threshold and the number of music features in the ambient sound data whose matching degree with the corresponding music features in the audio template is less than or equal to the first threshold; and determining a matching value based on the first and second quantities.

[0023] It can be understood that the first quantity corresponds to the voice matching value s in this application embodiment, and the second quantity corresponds to the music matching value m in this application embodiment. The initial values ​​of the first quantity and the second quantity are 0.

[0024] For example, speech features include tone envelope and zero-crossing rate. If the speech matching degree of tone envelope meets the first threshold, but the speech matching degree of zero-crossing rate does not meet the first threshold, then the speech matching value s = 0 + 1 - 1 = 0.

[0025] For example, the musical features of audio data include spectral contrast, spectral flatness, and tone centroid. If the musical matching degree of spectral contrast and spectral flatness meets the first threshold, but the musical matching degree of tone centroid does not meet the first threshold, then the musical matching value m = 0 + 1 + 1 - 1 = 1.

[0026] In one possible implementation, when the time it takes for the control headphones to enter the pass-through mode reaches a third threshold, the control headphones close the pass-through mode, wherein the third threshold is determined based on the template length time of the audio template corresponding to the ambient sound data.

[0027] For example, the template length of the audio template corresponding to the ambient sound data is 30 seconds. When the time for the headphones to enter the pass-through mode reaches 30 seconds, the headphones are controlled to switch back from the pass-through mode to the noise cancellation mode.

[0028] In some embodiments, when the time required for the headphones to enter pass-through mode meets the condition of (template length time of the audio template corresponding to the ambient sound data + buffer time), the headphones are switched back from pass-through mode to noise cancellation mode. The user can set the buffer time according to their specific needs.

[0029] For example, if the audio template corresponding to the ambient sound data has a template length of 30 seconds and a buffer time of 10 seconds, then when the headphones have been in pass-through mode for 40 seconds, the headphones will switch back to noise cancellation mode.

[0030] In this embodiment, automatic switching from noise reduction mode to pass-through mode and then back to noise reduction mode is achieved, which improves the user experience.

[0031] In one possible implementation, the method for obtaining the audio template includes: obtaining target audio, wherein the target audio includes audio data that meets preset conditions in various scenarios; extracting audio features of the target audio, wherein the audio features include music features and / or speech features; and creating an audio template based on the audio features of the target audio.

[0032] It is understandable that the target audio can be audio from typical scenarios where users need to hear ambient sounds, such as check-in audio at an airport, ticket checking audio at a train station, triage and queuing audio at a hospital, and lost and found audio from a shopping mall.

[0033] In one possible implementation, extracting audio features from the target audio includes: dividing the target audio into multiple audio segments; converting the time-domain signals of the multiple audio segments into frequency-domain signals to obtain the spectral information of the multiple audio segments; and determining the audio features of each frame in the multiple audio segments based on the spectral information of the multiple audio segments.

[0034] In one possible implementation, the musical features include at least one of spectral contrast, spectral flatness, tonality centroid, and chromaticity variation; the speech features include at least one of Mel frequency cepstrum, tonality envelope, zero-crossing rate, and power spectrum variation.

[0035] In one possible implementation, the method for obtaining the pre-set scene model includes: acquiring ambient sound data and scene data of the target scene; and training the pre-set scene model based on the ambient sound data and scene data of the target scene, wherein the pre-set scene model is used to determine the current scene.

[0036] It is understandable that a pre-defined scenario model can determine the current scenario corresponding to the input data based on the input data. For example, it can determine the top 3 scenarios with the highest probability of occurrence corresponding to the input data.

[0037] In one possible implementation, the method for obtaining the first event model includes: acquiring ambient sound data of the target scene and a preset scene model; training the first event model based on the ambient sound data of the target scene and the preset scene model, wherein the first event model is used to determine the degree of matching between the ambient sound data in the current scene and the preset event.

[0038] It is understood that the first event model corresponds to the typical event model in the embodiments of this application. The first event model can determine the probability of the event corresponding to the ambient sound data in the current scene, such as determining that the probability of the event corresponding to the current scene is the airport check-in announcement is 0.8.

[0039] Secondly, embodiments of this application provide an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for implementing any of the headphone control methods provided by the first aspect and various possible implementations of the first aspect.

[0040] Thirdly, embodiments of this application provide a readable medium storing instructions that, when executed on an electronic device, cause the electronic device to implement any of the headphone control methods provided in the first aspect and various possible implementations of the first aspect. Attached Figure Description

[0041] Figure 1 According to an embodiment of this application, a structural schematic diagram of a TWS earphone is shown;

[0042] Figure 2 According to an embodiment of this application, a schematic diagram of interface changes for template recording is shown;

[0043] Figure 3 According to an embodiment of this application, a flowchart of a method for controlling headphones is shown;

[0044] Figure 4 According to an embodiment of this application, a flowchart of a method for controlling headphones is shown;

[0045] Figure 5 According to an embodiment of this application, a flowchart of an audio template acquisition method is shown;

[0046] Figure 6 According to an embodiment of this application, a schematic diagram of an audio feature change is shown;

[0047] Figure 7 According to an embodiment of this application, a flowchart of a model acquisition method is shown;

[0048] Figure 8According to an embodiment of this application, a schematic diagram of a model acquisition method is shown;

[0049] Figure 9 According to an embodiment of this application, a flowchart of a method for controlling headphones is shown;

[0050] Figure 10 According to an embodiment of this application, a schematic diagram of the structure of an electronic device 10 is shown. Detailed Implementation

[0051] The illustrative embodiments of this application include, but are not limited to, a method for controlling headphones, an electronic device, and a readable medium.

[0052] It is understood that the technical solution of this application is applicable to electronic devices that can control the operating mode of headphones, such as, but not limited to, mobile phones, smartwatches, televisions, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of this invention do not impose any restrictions on the specific type of electronic device.

[0053] For ease of description, the headphones described below are TWS headphones by default, and the user is the wearer of the TWS headphones. It is understood that the headphones mentioned in this application can also be any other headphones that can achieve mode switching.

[0054] The two headphone modes mentioned in this application will be explained below.

[0055] Noise cancellation mode: The headphones emit noise with an amplitude similar to but opposite phase to the ambient noise, thereby reducing ambient noise.

[0056] Hear-through (HT) mode: also known as directional sound enhancement mode. In this mode, the headphones can directionally pick up and amplify external signals, such as ambient sounds.

[0057] In some embodiments, to address the issue of users missing important notifications due to noise cancellation mode being activated on the headphones, the headphones can be controlled to switch between noise cancellation and pass-through modes based on user behavior. For example, if it is determined that the user is speaking, it can be assumed that the user is likely chatting with someone. In this case, if the headphones are currently in noise cancellation mode, the headphones can be controlled to switch from noise cancellation mode to pass-through mode to amplify external sounds, allowing the user to communicate without removing the headphones. If it is determined that the user is not speaking, it can be assumed that the user is likely listening to music or similar content. If the headphones are currently in pass-through mode, the headphones can be controlled to switch from pass-through mode to noise cancellation mode to reduce ambient noise and improve the user's auditory experience.

[0058] However, the above methods can still cause users to miss important reminders in other scenarios. For example, if a user is waiting for their flight alone in a noisy airport terminal, the electronic device, knowing the user is not speaking, might assume the user is listening to music through headphones and therefore switch the headphones to noise-canceling mode to reduce ambient noise. This could cause the user to miss check-in information from the terminal's announcement system.

[0059] To address the aforementioned issues, this application provides a method for controlling headphones. Users can first preset multiple audio templates in an electronic device and store these audio templates in headphones and / or the electronic device. The audio templates can include feature information of audio data corresponding to important reminders that the user needs to hear, such as feature information of airport check-in announcements, train station ticket inspection announcements, hospital triage and queuing announcements, and shopping mall lost and found announcements.

[0060] In this way, while the headphones are in noise-canceling mode, they can collect audio data of the user's current environment in real time and match this data with a preset audio template to determine whether the audio data in the user's current environment includes an important reminder that the user needs to hear. For example, if the match between the audio data and the preset audio template is greater than a first threshold, it can be determined that the audio data includes an important reminder that the user needs to hear. At this point, controlling the headphones to enter pass-through mode can effectively prevent the user from missing important reminders.

[0061] In some embodiments, to further improve the accuracy of mode switching judgment, while the headphones are in noise-canceling mode, in addition to real-time acquisition of audio data corresponding to the user's current environment, scene data can also be acquired to determine the current scene. Scene data can be data that characterizes the user's current environment; for example, scene data can include the user's location information and the time information of the user at each location. For example, if the user is at an airport and has a 30-minute layover, the current scene can be determined as the user waiting for their flight. After determining the current scene, the matching degree between the acquired audio data and the audio template can be combined to determine whether the audio data of the user's current environment (i.e., ambient sound) includes important reminders that the user needs to hear in their current scene. For example, the probability p of the event that the audio data of the current environment includes important reminders that the user needs to hear can be determined by the matching degree and scene data. When the event probability p meets a second threshold, the headphones are controlled to enter pass-through mode. This improves the accuracy of mode switching judgment and effectively prevents the user from missing important reminders.

[0062] In some embodiments, the way to switch the headphones from pass-through mode back to noise cancellation mode is as follows: when the time for the headphones to enter pass-through mode meets the template length time of the audio template corresponding to the audio data, the headphones are controlled to switch from pass-through mode back to noise cancellation mode.

[0063] In some embodiments, the method for presetting audio templates in electronic devices and / or headphones can be as follows: extracting the music features and / or speech features of the audio data corresponding to important reminders in typical events, and using the music features and / or speech features as audio templates. Typical events can include airport check-in announcements, train station ticket announcements, hospital triage announcements, and shopping mall lost and found announcements, etc. For example, the user records the audio data corresponding to important reminders they need to hear in their current situation (such as airport announcements), extracts the frame features of each frame of the audio data, and uses the extracted music features and / or speech features as audio templates.

[0064] Musical features include spectral contrast, spectral flatness, spectral centroid, and chroma variation. Speech features include mel frequency cepstrum coefficient (MFCC), pitch envelope, zero-crossing rate, and power spectrum variation.

[0065] In some embodiments, the matching degree between the audio data of the user's environment and the audio template can be obtained by: extracting frame features of each frame of the audio data to obtain music features and / or speech features; inputting the music features into the music template to obtain music matching degree; and / or inputting the speech features into the speech template to obtain speech matching degree.

[0066] In some embodiments, determining the event probability p of the audio data in the current environment including an important reminder that the user needs to hear at their current location, based on matching degree and scene data, can be achieved by: inputting the audio data and scene data into a typical event model to determine the event probability t of the audio data in the current environment including an important reminder that the user needs to hear at their current location; determining the matching value of the music features and the voice features based on the matching degree of the music features and the voice features; and finally, determining the event probability p of the audio data in the current environment including an important reminder that the user needs to hear at their current location based on the matching value, the weight corresponding to the matching value, and the event probability t.

[0067] The method described above for determining the matching value of music features and speech features based on their matching degree can include: determining the final matching value based on whether the matching degree of each music feature and each speech feature satisfies a first threshold. For example, if the matching degree of each music feature and each speech feature satisfies the first threshold, the corresponding matching value (initially 0) is incremented by 1; otherwise, the matching value is decremented by 1.

[0068] For example, musical features include spectral contrast, spectral flatness, and tonal centroid. If the musical match in terms of spectral contrast meets the first threshold, then the musical match value m = 0 + 1 = 1. If the musical match in terms of spectral flatness meets the first threshold, then the musical match value m = 1 + 1 = 2. If the musical match in terms of tonal centroid does not meet the first threshold, then the musical match value m = 2 - 1 = 1.

[0069] In some embodiments, the method for creating the above-mentioned typical event model includes: acquiring voice information, music information, location information, and time information of typical scenarios for important reminders that the user needs to hear; training based on the voice information, music information, location information, and time information to obtain a preset scenario model for determining the current scenario; and retraining the preset scenario model, voice information, and music information to obtain a typical event model that can determine the degree of matching between the current scenario and typical events in typical scenarios.

[0070] In some embodiments, the audio template can be obtained by real-time recording generated by the user in the corresponding scenario, as described below. Figure 1 The headphones shown Figure 2 The mobile phone interface shown introduces how users can record audio templates using their mobile phone 20.

[0071] In some embodiments, when a user is waiting at an airport, such as Figure 2 As shown in interface a, users can click the "Add" button 202 on the recording interface 201 of phone 20. (See also...) Figure 2 As shown in interface b, in response to a user-triggered new operation, mobile phone 20 displays template 5 on the recording interface 201. When mobile phone 20 detects that the user clicks the record button 203, it responds to this operation by sending a sound capture command to the headset. The headset receives the sound capture command and, based on this command, captures sound through the headset's microphone. For example, the captured sound might be an airport announcement such as "A notification tone (e.g., ding ding ding) + Attention passengers traveling to XXXX, check-in for your flight XXXX is now beginning. Please proceed to the check-in counter. Thank you." Simultaneously, the capture duration is recorded, and the capture progress is calculated based on the capture duration and the preset total capture duration. The corresponding progress is displayed on the capture progress bar. Figure 2 As shown in interface c, if the acquisition progress is 79%, the phone 20 can display 79% on the acquisition progress bar of template 5 in the recording interface 201. Figure 2 As shown in the d interface, if the acquisition progress is 100%, it can be displayed on the acquisition progress bar.

[0072] After the user finishes recording, the mobile phone 20 can store the above audio template and control the headphones to enter noise cancellation mode. When the headphones' microphone recognizes information with the audio characteristics of template 5, such as "Attention passengers traveling to XXXX, check-in for your flight XXXX is now open. Please proceed to the check-in counter. Thank you," the headphones will automatically enter pass-through mode to prevent the user from missing check-in information, and will automatically enter noise cancellation mode after the broadcast ends.

[0073] In some embodiments, after the user finishes recording, the headphones can store the aforementioned audio template and enter noise reduction mode. When the headphones' microphone recognizes information with the audio characteristics of template 5, such as "Attention passengers traveling to XXXX, check-in for your flight XXXX is now beginning. Please proceed to the check-in counter. Thank you," the headphones enter pass-through mode to prevent the user from missing check-in information, and automatically enter noise reduction mode after the broadcast ends.

[0074] The following is a detailed description of the headphone control method provided in the embodiments of this application. The headphone control method of the embodiments of this application is applied to electronic devices. Figure 3 This illustration shows a schematic diagram of a headphone control method according to an embodiment of this application. The headphone control method includes:

[0075] 301: Get data to be processed.

[0076] In this embodiment, the data to be processed may include audio data collected via the microphone of the headphones and scene data collected in real time by the electronic device. The audio data may include music data and voice data. Music data may include songs, piano pieces, violin pieces, etc. Voice data may include broadcast information, chat messages, crosstalk, skits, etc. Scene data may include location information and time information. Location information may be the longitude, latitude, and altitude of the user's current location obtained by the positioning system of the electronic device. Time information may include the current date, the user's dwell time, and the user's movement time.

[0077] For example, the audio data in the pending data includes airport announcements captured by the microphone of the headset connected to the mobile phone: "A notification tone (such as ding ding ding) + Attention passengers traveling to XXXX, check-in for your flight XXXX is now open. Please proceed to the check-in counter. Thank you." The scene data in the pending data includes: the mobile phone's global positioning system (GPS) indicating that user A remained stationary for 3 minutes at a location of 116 degrees longitude, 39 degrees latitude, and 26 meters altitude at 11:20 AM on April 5, 20xx.

[0078] 302: Determine the musical and speech features of the audio data in the data to be processed.

[0079] In this embodiment of the application, audio data (i.e., ambient sound signal x(t)) can be collected in real time through the microphone of the earphone or the microphone of the electronic device (such as a mobile phone), and the audio data can be divided into time-domain frames by the electronic device so that the signal length of each frame meets 10ms or 20ms. Then, the feature extraction of the current frame is performed on each frame to obtain the music features and speech features of each frame of audio data.

[0080] In some embodiments, the specific method for determining the musical and speech features of audio data can be as follows: First, the audio data is divided into frames with a length of 20-30 milliseconds and an overlap rate of 50%, thereby converting the audio data into a two-dimensional matrix for subsequent processing. Second, each frame of the audio data is processed using a window function (such as a Hamming window) to reduce spectral energy leakage. Spectral energy leakage refers to the frequency obtained using FFT, which is not the actual frequency of the original signal but a altered frequency (similar to energy leakage from one frequency to other frequencies). Then, a Fast Fourier Transform is performed on each frame to convert the time-domain signal of each frame into a frequency-domain signal, obtaining spectral information. Finally, the musical and speech features of the spectral information are extracted.

[0081] The categories of music and speech features that can be extracted are described below. As shown in Table 1, music features include spectral contrast, spectral flatness, tonality centroid, and chromaticity variation; speech features include MFCC, tonality envelope, zero-crossing rate, and power spectrum variation.

[0082] Table 1

[0083] Spectral contrast MFCC Spectral flatness Tonal envelope Regulating Mind Zero crossing rate Color variation Power spectral variation

[0084] The following section introduces the methods for determining spectral contrast, spectral flatness, tone centroid, chromaticity variation, MFCC, tone envelope, zero-crossing rate, and power spectrum variation.

[0085] In some embodiments, spectral contrast can be determined by calculating the logarithmic ratio of the amplitude of each frequency band in the spectral information to the average amplitude.

[0086] In some embodiments, spectral flatness can be obtained by calculating the ratio of the geometric mean to the arithmetic mean of the power spectrum of each frame in the spectral information.

[0087] In some embodiments, the centroid of tuning can be obtained by calculating the centroid position of the power spectrum in each frame of the spectrum information.

[0088] In some embodiments, the chromaticity variation can be obtained by calculating the square root of the sum of squares of the amplitudes of each pitch in the spectral information.

[0089] In some embodiments, the power spectral density of each frame in the spectral information can be filtered through a set of Mel filters, and then the logarithm of the filtered power spectral density (e.g., base 10) can be taken, and finally a discrete cosine transform (DCT) can be performed to obtain the MFCC of each frame.

[0090] In some embodiments, the spacing tone envelope can be obtained by calculating the fundamental frequency envelope of each frame in the spectrum information.

[0091] In some embodiments, the zero-crossing rate can be obtained by calculating the zero-crossing rate of each frame in the spectral information, i.e., the number of zero-crossing points of the signal from positive to negative or from negative to positive.

[0092] In some embodiments, the power spectrum variation characteristics can be obtained by calculating the difference in the power spectrum of each frame in the spectral information.

[0093] In this embodiment of the application, the extracted music features and speech features can be converted into a representation as shown in the vector feature matrix (1) to obtain music features as shown in the vector feature matrix (2) and speech features as shown in the vector feature matrix (3), respectively.

[0094]

[0095]

[0096]

[0097] In the vector feature matrices (1), (2), and (3), i represents the i-th template, M represents the M-th feature, N is the number of frequency points, and P is the total number of frames.

[0098] In some embodiments, the power spectral density can also be determined by calculating the power spectral density of each frame in the spectral information, i.e., the square of the modulus of the Fourier transform result.

[0099] In some embodiments, the FFT method can be used to calculate the short-time spectrum of each frame in the spectrum information, and then the average of the short-time spectrum over the entire time range can be calculated to obtain the long-time average spectrum (Ltas).

[0100] In some embodiments, a spectrogram can be determined based on the power spectral density, and a pitch contour can be determined based on the tone centroid. Power spectral density, Ltas, and pitch contour are considered musical characteristics.

[0101] 303: Determine the matching degree based on music features, speech features, and preset audio templates.

[0102] In this embodiment, the matching degree includes music matching degree and speech matching degree. The method for determining the matching degree includes: inputting music features and speech features into a preset audio template in an electronic device, so as to determine the music matching degree of the music features and the speech matching degree of the speech features in the audio data based on the music features and speech features in the audio template. The training method for the audio template is described in... Figure 5 The embodiments shown are described in detail and will not be repeated here.

[0103] It is understandable that the vector matrix of each music feature and the vector matrix of each speech feature in the audio data have corresponding vector matrices of music features and speech features in the audio template.

[0104] In some embodiments, the method of determining the music matching degree of music features in audio data based on music features in audio templates includes: after obtaining the vector matrix of music features in audio data in the format shown in the vector feature matrix (4), calculating the similarity between the two vector matrices (i.e., the vector matrix of music features in audio templates and the vector matrix of music features in audio data) using the Pearson correlation coefficient method, and using the similarity between the two vector matrices as the music matching degree.

[0105] If the music features of the current frame are known to be the feature vectors shown in the vector feature matrix (4), then... The similarity is calculated based on the vector matrix of the current frame and the vector matrix of the music features in the first frame of the audio template. The calculation process is as follows:

[0106]

[0107]

[0108] In some embodiments, the method of determining the speech matching degree of speech features in audio data based on speech features in audio template includes: after obtaining the vector matrix of speech features in audio data in the format shown in the above vector feature matrix (4), calculating the similarity between two feature vectors (i.e., the vector matrix of speech features in audio template and the vector matrix of speech features in audio data) by using the Pearson correlation coefficient method, and taking the similarity between the two vector matrices as the speech matching degree.

[0109] If the speech features of the current frame are known to be the feature vectors shown in the vector feature matrix (4), then... The similarity is calculated based on the vector matrix of the current frame and the vector matrix of the speech features of the first frame in the speech template, as shown in the example below:

[0110]

[0111] In some embodiments, the similarity between the music features of the current frame and the music features of the audio template, and the similarity between the speech features of the current frame and the speech features of the audio template can also be calculated by calculating the distance between vectors.

[0112] For example, if the music features of the current frame are dense features, the similarity between the music features and the music template is calculated by using the Euclidean or Mahalanobis distance between the music features of the current frame and the music features of the music template. If the music features of the current frame are sparse features, the similarity between the music features and the music template is calculated by using the cosine similarity between the music features of the current frame and the music features of the music template.

[0113] If the speech features of the current frame are dense, the similarity between the speech features of the current frame and the speech features of the audio template is calculated by using the Euclidean or Mahalanobis distance. If the speech features of the current frame are sparse, the similarity between the speech features of the current frame and the speech features of the audio template is calculated by using the cosine similarity.

[0114] 304: Determine whether the matching degree meets the first threshold.

[0115] In this embodiment, it can be determined whether the music matching degree and / or the voice matching degree meet a first threshold. When either the music matching degree or the voice matching degree meets the first threshold, proceed to step 305: determine the matching value of the data to be processed based on the matching degree. When neither the music matching degree nor the voice matching degree meets the first threshold, proceed to step 302: determine the matching degree based on the music features and voice features.

[0116] For example, when the music matching degree and / or speech matching degree are greater than the first threshold λ1, proceed to step 305: determine the matching value of the data to be processed based on the matching degree. When both the music matching degree and speech matching degree are less than or equal to the first threshold λ1, proceed to step 302: determine the matching degree based on music features and speech features. λ1 can be set according to actual needs, such as 0.75 or 0.8, and is not specifically limited here.

[0117] 305: Determine the matching value of the data to be processed based on the matching degree.

[0118] In this embodiment, the matching degree includes music matching degree and voice matching degree; therefore, the corresponding matching value includes music matching value and voice matching value. The method of determining the matching value of the data to be processed based on the matching degree may include: determining the music matching value based on the music matching degree, and determining the voice matching value based on the voice matching degree.

[0119] In some embodiments, the method for determining a music matching value based on the music matching degree may include: when the music matching degree of each music feature of the audio data meets a first threshold, incrementing the music matching value m (i.e., the music counter m, with an initial value of 0) by 1; and decrementing the music matching value m by 1 when the music matching degree of each music feature of the audio data does not meet the first threshold.

[0120] For example, the musical features of audio data include spectral contrast, spectral flatness, and tonal centroid. If the musical matching degree based on spectral contrast meets the first threshold, then the musical matching value m = 0 + 1 = 1. If the musical matching degree based on spectral flatness meets the first threshold, then the musical matching value m = 1 + 1 = 2. If the musical matching degree based on tonal centroid does not meet the first threshold, then the musical matching value m = 2 - 1 = 1.

[0121] For example, the musical features of audio data include spectral contrast, spectral flatness, and tone centroid. If the musical matching degree of spectral contrast and spectral flatness meets the first threshold, but the musical matching degree of tone centroid does not meet the first threshold, then the musical matching value m = 0 + 1 + 1 - 1 = 1.

[0122] In some embodiments, the method for determining a voice matching value based on the voice matching degree may include: when the voice matching degree of each voice feature of the audio data meets a first threshold, incrementing the voice matching value s (i.e., the voice counter s, with an initial value of 0) by 1; and decrementing the voice matching value s by 1 when the voice matching degree of each voice feature of the audio data does not meet the first threshold.

[0123] For example, speech features include MFCC, tonality envelope, and zero-crossing rate. If the speech matching degree of MFCC meets the first threshold, then the speech matching value s = 0 + 1 = 1. If the speech matching degree of tonality envelope meets the first threshold, then the speech matching value s = 1 + 1 = 2. If the speech matching degree of zero-crossing rate does not meet the first threshold, then the speech matching value s = 2 - 1 = 1.

[0124] For example, speech features include tone envelope and zero-crossing rate. If the speech matching degree of tone envelope meets the first threshold, but the speech matching degree of zero-crossing rate does not meet the first threshold, then the speech matching value s = 0 + 1 - 1 = 0.

[0125] 306: Determine the probability p of the occurrence of typical events in the current scene based on the matching value and scene data.

[0126] In this embodiment of the application, the method for determining the event occurrence probability p of a typical event in the current scene based on the matching value and scene data may include: acquiring a preset scene model and a typical event model; determining the event occurrence probability t based on the preset scene model, the typical event model, audio data, and scene data; then acquiring the weights corresponding to the music matching value and the voice matching value; and determining the event occurrence probability p of the typical event in the current scene based on each weight, the matching value, and the event occurrence probability t. The training methods for the preset scene model and the typical event model are described in detail below. Figure 7 The embodiments shown are described in detail and will not be repeated here.

[0127] In some embodiments, the method for determining the probability t of an event may include: inputting music features, speech features, and scene data into a preset scene model to determine the current scene. Then, inputting the music features, speech features, and scene data into a typical event model included in the current scene to determine the probability t of a typical event through the typical event model. The method for determining the probability t of an event is as shown in formula (1):

[0128]

[0129] For example, the music features, speech features, and scene data of the data to be processed include: Passenger A stayed at a location with longitude of 116 degrees, latitude of 39 degrees, and altitude of 26 meters for 3 minutes without moving at 11:20 AM on April 5, 20xx, and has two audio features, audio_1 and audio_2 (audio features are represented as digital vectors). The input features corresponding to the music features, speech features, and scene data can be represented as:

[0130] VF_1 = [116, 39, 26, 1, 3, 20xx04051120, [audio_1], [audio_2]]. The input features corresponding to the music features, speech features, and scene data are input into the typical event model to obtain the event occurrence probability t.

[0131] In some embodiments, the method for determining the event occurrence probability p of a typical event in the current scene includes: first obtaining the decision function as shown in formula (2), the weight w1 corresponding to the music feature, and the weight w2 corresponding to the speech feature. Then, based on the decision function, the weight w1 corresponding to the music feature, the weight w2 corresponding to the speech feature, the music matching value m (i.e., the music counter m), the speech matching value s (i.e., the speech counter s), and the event occurrence probability t (i.e., the event flag t), the event occurrence probability p of the typical event in the current scene is determined.

[0132]

[0133] 307: Determine if the probability p of the event meets the decision threshold θ. If the result is yes, proceed to 308: Control the headphones to enter pass-through mode. If the result is no, proceed to 302: Determine the music and speech features of the audio data in the data to be processed.

[0134] In this embodiment, the decision threshold θ corresponds to the second threshold mentioned above, and the decision threshold θ is close to the event occurrence probability t. As shown in formula (2), when the event occurrence probability p satisfies the decision threshold θ, that is, when the event occurrence probability p is greater than the decision threshold θ, the decision result is 1, and the headphones are controlled to enter the pass-through mode. When the event occurrence probability p does not satisfy the decision threshold θ, that is, when the event occurrence probability p is less than or equal to the decision threshold θ, the decision result is 0, and the music features and speech features of the audio data in the data to be processed are continuously determined.

[0135] 308: Controls the headphones to enter pass-through mode.

[0136] In this embodiment, when the time it takes for the control headphones to enter the pass-through mode meets the template length time of the audio template corresponding to the audio data, the control headphones switch back from the pass-through mode to the noise reduction mode.

[0137] For example, if the template length of the audio template corresponding to the data to be processed is 30 seconds, when the time for the headphones to enter the pass-through mode reaches 30 seconds, the headphones are controlled to switch from the pass-through mode back to the noise cancellation mode.

[0138] In some embodiments, when the time required for the headphones to enter pass-through mode meets the condition of (the template length time of the audio template corresponding to the audio data + the buffer time), the headphones are switched back from pass-through mode to noise cancellation mode. The user can set the buffer time according to the actual situation.

[0139] For example, if the template length of the feature template corresponding to the data to be processed is 30 seconds and the buffer time is 10 seconds, then when the time the headphones spend in pass-through mode reaches 40 seconds, the headphones will be controlled to switch back from pass-through mode to noise cancellation mode.

[0140] This application embodiment achieves automatic activation of the headphone's pass-through mode by acquiring audio data and scene data, preventing users from missing important notifications due to noise cancellation being enabled. It also automatically deactivates the headphone's pass-through mode, ensuring a comfortable headphone wearing experience for the user.

[0141] Figure 4 This illustration shows a schematic diagram of another headphone control method according to an embodiment of this application. Figure 3 The difference is that, Figure 4 The control method determines the headphone mode switching based on the matching degree of audio data. Figure 4 The control methods for the headphones shown include:

[0142] 401: Get data to be processed.

[0143] The method for obtaining the data to be processed described above can be found in [link to relevant documentation]. Figure 3 Step 301 shown will not be repeated here.

[0144] 402: Determine the musical and speech features of the audio data in the data to be processed.

[0145] The method for determining the musical and speech features of audio data in the data to be processed is described above. Figure 3 Step 302 shown will not be repeated here.

[0146] 403: Determine the matching degree based on musical and speech features.

[0147] The method for determining the matching degree based on musical and speech features mentioned above can be found in [link to relevant documentation]. Figure 3 Step 303 shown will not be repeated here.

[0148] 404: Determine if the matching degree meets the first threshold.

[0149] In this embodiment, it can be determined whether the music matching degree and / or voice matching degree meet a first threshold. When either the music matching degree or the voice matching degree meets the first threshold, the process proceeds to step 405: controlling the headphones to enter the pass-through mode. When neither the music matching degree nor the voice matching degree meets the first threshold, the process proceeds to step 402: determining the music features and voice features of the audio data in the data to be processed.

[0150] 405: Controls the headphones to enter pass-through mode.

[0151] The method for controlling the headphones to enter pass-through mode is described above. Figure 3 Step 308 shown will not be repeated here.

[0152] This application embodiment, based on the matching degree between music features, voice features, and audio templates, implements an automatic headphone pass-through mode, preventing users from missing important reminders due to noise cancellation being enabled. It also implements an automatic headphone pass-through mode, ensuring a comfortable headphone wearing experience for the user.

[0153] Figure 5 This illustration shows a schematic diagram of the method for extracting music features and speech features from an audio template in the headphone control method provided in an embodiment of this application. The method includes:

[0154] 501: Get the target audio.

[0155] In this embodiment, the electronic device can pre-store a certain number of target audio files, and the user can also record target audio files according to actual needs. The target audio files can be audio from typical scenarios where the user needs to hear ambient sounds, such as check-in audio at an airport, ticket checking audio at a train station, triage and queuing audio at a hospital, and lost and found audio from a shopping mall.

[0156] The following section introduces methods for users to record target audio according to their actual needs.

[0157] When users are waiting at the airport, such as Figure 2 As shown in interface a, users can click the "Add" button 202 on the recording interface 201 of phone 20. (See also...) Figure 2As shown in interface b, in response to a user-triggered new operation, mobile phone 20 displays template 5 on the recording interface 201. When mobile phone 20 detects that the user clicks the record button 203, it responds to this operation by sending a sound capture command to the headset. The headset receives the sound capture command and, based on this command, captures sound through the headset's microphone. For example, the captured sound might be an airport announcement such as "A notification tone (e.g., ding ding ding) + Attention passengers traveling to XXXX, check-in for your flight XXXX is now beginning. Please proceed to the check-in counter. Thank you." Simultaneously, the capture duration is recorded, and the capture progress is calculated based on the capture duration and the preset total capture duration. The corresponding progress is displayed on the capture progress bar. Figure 2 As shown in interface c, if the acquisition progress is 79%, the phone 20 can display 79% on the acquisition progress bar of template 5 in the recording interface 201. Figure 2 As shown in the d interface, if the acquisition progress is 100%, it can be displayed on the acquisition progress bar. After the user finishes recording, the electronic device can control the headphones to enter noise cancellation mode.

[0158] In some embodiments, after the user finishes recording, the user can also add keywords to the template in the settings interface of the mobile phone 20, such as adding the keyword "flight number: AB1234" to the template 5, so that the user can directly obtain the check-in broadcast of the flight they are taking.

[0159] 502: Extract musical and speech features from the target audio.

[0160] In this embodiment, the target audio can first be divided into frames with a length of 20-30 milliseconds and an overlap rate of 50%, thereby converting the target audio into a two-dimensional matrix for easier subsequent processing. Then, each frame of the target audio is processed using a window function (such as a Hamming window) to reduce spectral energy leakage. Next, a Fast Fourier Transform is performed on each frame to convert the time-domain signal of each frame into a frequency-domain signal, obtaining spectral information. Finally, the musical and speech features of the spectral information are extracted.

[0161] like Figure 3 As shown in Table 1 of the illustrated embodiment, the extracted music features may include spectral contrast, spectral flatness, tonality centroid, chromaticity variation, spectrogram, tonality, and Ltas; speech features may include MFCC, tonality envelope, zero-crossing rate, and power spectrum variation. Figure 6 As shown, Figure 6 The diagram shows the frequency variation of the spectrogram over time, the frequency variation of the tonality over time, and the sound pressure level (SPL) of Ltas over frequency. Figure 6This includes a schematic diagram showing the frequency variation over time in the spectrogram with a maximum harmonic-to-noise ratio (HNR) of 9.73 dB. It also shows the frequency variation over time in the pitch spectrum with a vocal fold (VF) of 70.7%, jitter of 1.4%, shimmer of 30.4%, a mean of 129 Hz, and a standard deviation (SD) of 60.1 Hz. Finally, it shows the sound pressure level (SPL) variation in the LTAs spectrum with a maximum sound signal frequency of 973 Hz, a minimum of 75 Hz, a breakdown elemental (BED) duration of 12 dB, and a center of gravity (CoG) of 1185 Hz.

[0162] The methods for extracting spectral contrast, spectral flatness, tone centroid, chromaticity variation, spectrogram, tone and Ltas, MFCC, tone envelope, zero-crossing rate, and power spectrum variation are described in step 302 and will not be repeated here.

[0163] 503: Determine the audio template based on musical and speech features.

[0164] In this embodiment of the application, matrix-form music features and speech features can be used as audio templates.

[0165] In some embodiments, music features can be converted into a representation as shown in the vector feature matrix (1) above, resulting in music features as shown in the vector feature matrix (2) above. Speech features can be converted into a representation as shown in the vector feature matrix (1) above, resulting in speech features as shown in the vector feature matrix (3) above.

[0166] In this way, the music features and speech features in the data to be processed can be input into the audio template. Based on the music features and speech features in the audio template, the music matching degree between the music features in the data to be processed and the corresponding music features in the audio template and the speech matching degree between the corresponding speech features in the audio template can be determined. Based on the music matching degree and the speech matching degree, it can be determined whether to control the headphones to enter the pass-through mode.

[0167] The training method of a typical event model in the headphone control method provided in this application embodiment will be described in detail below, such as... Figure 7 As shown, the training methods for typical event models include:

[0168] 701: Obtain audio and non-audio information of the target scene.

[0169] It is understood that the target scenario in this application embodiment can be a typical scenario where the user needs to hear ambient sounds, such as airport check-in, train station ticket inspection, hospital triage and queuing, and shopping mall lost and found notices. Audio information can be the aforementioned audio data, including music data and voice data. Non-audio information can be the aforementioned scenario data, which can be data that characterizes the user's current situation. For example, scenario data can include the user's location information and the time information of the user at each location.

[0170] In this embodiment of the application, the electronic device can pre-store a certain amount of audio and non-audio information of the target scene, and can also collect audio and non-audio information of a specific scene according to the user's actual needs.

[0171] The following section introduces methods for collecting audio and non-audio information from specific scenarios based on users' actual needs.

[0172] For example, a user records the boarding announcement at gate A of airport A. The electronic device obtains the geographical location information of airport A through positioning software and retrieves statutory holiday information from the cloud. The boarding announcement at gate A is treated as audio information, the geographical location of airport A is treated as location information, and the statutory holidays are treated as time information.

[0173] 702: Train on audio and non-audio information to obtain a pre-defined scene model.

[0174] In the embodiments of this application, such as Figure 8 As shown, audio and non-audio information can be trained to obtain a pre-defined scene model. This pre-defined scene model is a classification model and includes at least one typical scenario where a user needs to hear ambient sounds, such as airport check-in, train station ticket checking, hospital triage and queuing, and shopping mall lost and found notices. Along with obtaining the pre-defined scene model, the probability of occurrence for each scenario can also be obtained.

[0175] The following section describes the above method using the airport scenario described above, taking the training of a pre-defined scene model based on non-audio information such as location and audio information as an example. First, the longitude of Airport A is obtained: 123 degrees, latitude: 45 degrees, altitude: 67 meters. Restaurant M at Airport A is located 1 degree away from its longitude coordinates and at an altitude of 68 meters. The audio data of Restaurant M is obtained as audio_1, ..., audio_N. Based on the feature vector VF_s = [124, 45, 68, ...] based on the location information of Restaurant M and the audio data of Restaurant M, a pre-defined scene model of Restaurant M at Airport A is obtained.

[0176] 703: Train audio information and a pre-set scene model to obtain a typical event model.

[0177] It is understandable that a scenario (such as an airport) may include at least one pre-set scenario model. The typical events of different pre-set scenario models are not necessarily the same. For example, the typical event of the pre-set scenario model of restaurant M in airport A may be the food pick-up announcement, and the typical event of the pre-set scenario model of gate A in airport A may be the check-in announcement.

[0178] In the embodiments of this application, such as Figure 8 As shown, by training audio information and a pre-defined scene model, the typical event model can output the probability t of occurrence of typical events (i.e., events included in the target scene). The typical event model is a regression model that includes at least one possible event in the target scene. For example, for the scene model corresponding to an airport, the corresponding typical event model includes check-in events, missing person announcements, and delay announcements.

[0179] For example, audio data and a pre-defined scene model can be represented as: VF_e = [scene_1, scene_2, scene_3, time_1, time_2, audio_1, ..., audio_N]; where scene_1, scene_2, and scene_3 represent the output labels of the first three categories of scenes output by the pre-defined scene model, respectively. The first three categories of scenes are based on the top 3 scenes with the highest probability of occurrence in the pre-defined scene model. Time_1 indicates whether it is a holiday or rest day, time_2 is the real-time time, and audio_1 to audio_N represent the feature vectors of N typical events.

[0180] In the practical application of the typical event model, the probability t of the event corresponding to the data to be processed can be determined by inputting the data to be processed into the typical event model. For example, the data to be processed may include: Passenger A stayed at a location with longitude: 116 degrees, latitude: 39 degrees, altitude: 26 meters for 3 minutes without moving at 11:20 am on April 5, 20xx, and has two audio features, audio_1 and audio_2 (audio features are represented as digital vectors). The input features of the data to be processed can be represented as: VF_1 = [116, 39, 26, 1, 3, 20xx04051120, [audio_1], [audio_2]]. t = event)(scene)(VF_1(1:x)), VF_1(x:end)); where event represents the trained typical event model, scene represents the pre-set scene model; VF_1(1:x) refers to the first to the xth features of the input features, used to input into the pre-set scene model; VF_1(x:end) refers to the xth to the last feature, used to input into the typical event model.

[0181] Therefore, by using the typical event model, the probability t of the occurrence of the target scene (or the typical events included in the target scene, or the typical events of the current scene) can be obtained, so as to determine whether to switch the operating mode of the headphones based on the probability t of the occurrence of the event.

[0182] The following is a detailed description of another headphone control method provided in the embodiments of this application, such as... Figure 9 As shown, the headphone control method includes the following process:

[0183] First, the method for obtaining audio templates will be introduced based on steps 901-904.

[0184] 901: Template recording.

[0185] In this embodiment, users can record templates according to their actual needs. The templates can include audio from typical scenarios where users need to hear ambient sounds, such as check-in audio at an airport, ticket checking audio at a train station, triage audio at a hospital, and lost and found audio from a shopping mall.

[0186] The following describes how users can record templates. For example, when a user is waiting at an airport... Figure 2 As shown in interface a, users can click the "Add" button 202 on the recording interface 201 of phone 20. (See also...) Figure 2As shown in interface b, in response to a user-triggered new operation, mobile phone 20 displays template 5 on the recording interface 201. When mobile phone 20 detects that the user clicks the record button 203, it responds to this operation by sending a sound capture command to the headset. The headset receives the sound capture command and, based on this command, captures sound through the headset's microphone. For example, the captured sound might be an airport announcement such as "A notification tone (e.g., ding ding ding) + Attention passengers traveling to XXXX, check-in for your flight XXXX is now beginning. Please proceed to the check-in counter. Thank you." Simultaneously, the capture duration is recorded, and the capture progress is calculated based on the capture duration and the preset total capture duration. The corresponding progress is displayed on the capture progress bar. Figure 2 As shown in interface c, if the acquisition progress is 79%, the phone 20 can display 79% on the acquisition progress bar of template 5 in the recording interface 201. Figure 2 As shown in the d interface, if the acquisition progress is 100%, it can be displayed on the acquisition progress bar. After the user finishes recording, the phone controls the headphones to enter noise cancellation mode.

[0187] In some embodiments, during the user's template recording process, the headphones can be in noise cancellation mode or pass-through mode.

[0188] 902: Template feature extraction.

[0189] In this embodiment of the application, feature extraction can be performed on the template to obtain music features and speech features.

[0190] In this embodiment, the recorded template can first be divided into frames with a length of 20-30 milliseconds and an overlap rate of 50%, thereby converting the recorded template into a two-dimensional matrix for easier subsequent processing. Each frame of the target audio is processed using a window function (such as a Hamming window) to reduce spectral energy leakage. Then, a Fast Fourier Transform is performed on each frame to convert the time-domain signal of each frame into a frequency-domain signal, obtaining spectral information. Finally, musical and speech features of the spectral information are extracted. The extracted musical features include spectral contrast, spectral flatness, tonality centroid, and chromaticity variation; the speech features include MFCC, tonality envelope, zero-crossing rate, and power spectrum variation.

[0191] The methods for extracting spectral contrast, spectral flatness, tone centroid, chromaticity variation, MFCC, tone envelope, zero-crossing rate, and power spectrum variation are described in step 502 and will not be repeated here.

[0192] 903: Obtain musical characteristics.

[0193] 904: Obtain speech features.

[0194] The following describes the headphone mode switching process based on steps 905-911.

[0195] 905: Frame feature extraction.

[0196] In this application embodiment, the frame feature is the frame feature of the frame signal in the data to be processed obtained by the electronic device, and the frame feature includes music features and voice features.

[0197] The method for extracting frame features as described above is described in step 402, and will not be repeated here.

[0198] 906: Frame feature matching.

[0199] In this embodiment, frame feature matching includes matching the music features in the frame features with the music features extracted in step 902 to obtain a music matching degree, and matching the speech features in the frame features with the speech features extracted in step 902 to obtain a speech matching degree. The methods for determining the music matching degree and the speech matching degree are the same as those in step 303 above, and will not be repeated here.

[0200] 907: Update feature matching flags.

[0201] In this embodiment of the application, the feature matching flags include a music counter m for music features and a voice counter s for speech features.

[0202] The update methods for the music counter m and the voice counter s are described below.

[0203] In some embodiments, the method for determining the music counter m includes: if each music matching degree of the music feature meets a first threshold, then the music matching value m (i.e., the music counter m, with an initial value of 0) is incremented by 1; if each music matching degree of the music feature does not meet the first threshold, then the music matching value m is decremented by 1.

[0204] For example, musical features include spectral contrast, spectral flatness, and tonal centroid. If the musical match degree for spectral contrast meets the first threshold, then the musical match value m = 0 + 1 = 1. If the musical match degree for spectral flatness meets the first threshold, then the musical match value m = 1 + 1 = 2. If the musical match degree for tonal centroid does not meet the first threshold, then the musical match value m = 2 - 1 = 1.

[0205] In some embodiments, the method for determining the voice counter s includes: if each voice matching degree of the voice feature meets a first threshold, then the voice matching value s (i.e., the voice counter s, with an initial value of 0) is incremented by 1; if each voice matching degree of the voice feature does not meet the first threshold, then the voice matching value s is decremented by 1.

[0206] For example, speech features include spectral contrast, spectral flatness, and tone centroid. If the speech matching degree for judging spectral contrast meets the first threshold, then the speech matching value s = 0 + 1 = 1. If the speech matching degree for judging spectral flatness meets the first threshold, then the speech matching value s = 1 + 12. If the speech matching degree for judging tone centroid does not meet the first threshold, then the speech matching value s = 2 - 1 = 1.

[0207] 908: Scene detection.

[0208] In this embodiment of the application, the scene detection method includes: acquiring scene data, and inputting music features, speech features, and scene data into a preset scene model, and the preset scene model outputting the top three scenes with the highest probability of occurrence corresponding to the music features, speech features, and scene data.

[0209] For example, the music features, speech features, and scene data include: Passenger A stayed at a location with longitude of 116 degrees, latitude of 39 degrees, and altitude of 26 meters for 3 minutes without moving at 11:20 AM on April 5, 20xx, and has two audio features, audio_1 and audio_2 (audio features are represented as digital vectors). The input features corresponding to the music features, speech features, and scene data can be represented as: VF_1[116, 39, 26, 1, 3, 20xx04051120, [audio_1], [audio_2]]. Then, the input features corresponding to the music features, speech features, and scene data are input into a pre-set scene model to obtain the top three scenes with the highest probability of occurrence corresponding to the music features, speech features, and scene data.

[0210] 909: Event Detection.

[0211] In this embodiment of the application, the event detection method includes: inputting music features, voice features, and scene data into a typical event model, and the typical event model outputs the event occurrence probability t. The method for determining the event occurrence probability t is as shown in the above formula (1).

[0212] For example, the music features, speech features, and scene data include: Passenger A stayed at a location with longitude of 116 degrees, latitude of 39 degrees, and altitude of 26 meters for 3 minutes without moving at 11:20 AM on April 5, 20xx, and has two audio features, audio_1 and audio_2 (audio features are represented as digital vectors). The input features corresponding to the music features, speech features, and scene data can be represented as: VF_1 = [116, 39, 26, 1, 3, 20xx04051120, [audio_1], [audio_2]]. Then, the input features corresponding to the music features, speech features, and scene data are input into a typical event model to obtain the event occurrence probability t.

[0213] 910: Determine if D(m, s, e) > threshold. If the result is yes, proceed to 911: Enable pass-through mode. If the result is no, proceed to 905: Frame feature extraction.

[0214] It can be understood that judging D(m, s, e) > threshold means judging whether D(m, s, e) is greater than the threshold, where the threshold corresponds to the second threshold mentioned above, i.e., the decision threshold θ.

[0215] In this embodiment of the application, the decision function, the weight w1 corresponding to the music feature, and the weight w2 corresponding to the speech feature can be obtained first, as shown in the above formula (2). Then, based on the decision function, the weight w1 corresponding to the music feature (i.e., the music matching value m), the weight w2 corresponding to the speech feature (i.e., the speech matching value s), the music matching value (i.e., the music counter m), the speech matching value (i.e., the speech counter s), and the event flag e (i.e., the event occurrence probability t), it is determined that D(m, s, e) > the threshold; wherein, the calculation method of D(m, s, e) is the same as that in step 307 above, and will not be repeated here.

[0216] It is understood that in the embodiments of this application, the pass-through mode is turned on when a typical event occurs in the typical scenario model.

[0217] 911: Enable pass-through mode.

[0218] In this embodiment of the application, the method for opening the transparent transmission is the same as step 308 above, and will not be repeated here.

[0219] This application embodiment achieves automatic activation of the headphone's pass-through mode by acquiring frame features, preventing users from missing important notifications due to noise cancellation being enabled. This application embodiment also achieves automatic deactivation of the headphone's pass-through mode, ensuring a superior headphone wearing experience for the user.

[0220] The hardware structure of the electronic device 10 mentioned in this application will be described below. For example... Figure 10 As shown, the electronic device 10 may include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, buttons 101, and a display screen 102, etc.

[0221] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 10. In other embodiments of this application, the electronic device 10 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0222] Processor 110 may include one or more processing units, such as processing modules or circuits of a central processing unit (CPU), graphics processing unit (GPU), digital signal processing (DSP), microprocessor (MCU), artificial intelligence (AI) processor, or field-programmable gate array (FPGA). Different processing units may be independent devices or integrated within one or more processors. Processor 110 may include storage units for storing instructions and data. In some embodiments, the storage unit in processor 110 is a cache memory 180.

[0223] It is understood that the headphone control method in this embodiment can be executed by the processor 110 of the corresponding electronic device. The power module 140 may include a power supply, a power management component, etc. The power supply may be a battery. The power management component manages the charging of the power supply and the power supply to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module receives charging input from a charger; the power management module connects to the power supply and the processor 110. The power management module receives input from the power supply and / or the charging management module to supply power to the processor 110, display screen 102, camera 170, and wireless communication module 120, etc.

[0224] The mobile communication module 130 may include, but is not limited to, antennas, power amplifiers, filters, and low-noise amplifiers (LNAs). The mobile communication module 130 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the electronic device 10. The mobile communication module 130 can receive electromagnetic waves via the antenna, filter and amplify the received electromagnetic waves, and then transmit them to a modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the mobile communication module 130 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 and at least some modules of the processor 110 may be housed in the same device.

[0225] The wireless communication module 120 may include an antenna, which enables the transmission and reception of electromagnetic waves. The wireless communication module 120 can provide solutions for wireless communication applications on the electronic device 10, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The electronic device 10 can communicate with networks and other devices through wireless communication technologies.

[0226] It is understood that, in this embodiment of the application, when the electronic device is a receiving device, it can receive video and audio data from other electronic devices in the recording group, as well as the recording content tags and time stamp information corresponding to each data, through the wireless communication module. And when the electronic device is a transmitting device, it can send video and audio data, as well as the recording content tags and time stamp information corresponding to each data, to other electronic devices in the recording group through the wireless communication module.

[0227] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the electronic device 10 may also be located in the same module.

[0228] The display screen 102 is used to display human-computer interaction interfaces, images, videos, etc. The display screen 102 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a quantum dot light-emitting diode (QLED), etc.

[0229] The sensor module 190 may include proximity sensors, pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0230] Audio module 150 is used to convert digital audio information into analog audio signal output, or to convert analog audio input into digital audio signal. Audio module 150 can also be used for encoding and decoding audio signals. In some embodiments, audio module 150 may be located in processor 110, or some functional modules of audio module 150 may be located in processor 110. In some embodiments, audio module 150 may include a speaker, earpiece, microphone, and headphone jack. Camera 170 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to image signal processing (ISP) to convert it into a digital image signal. Electronic device 10 can implement shooting functions through ISP, camera 170, video codec, graphics processing unit (GPU), display screen 102, and application processor, etc.

[0231] Interface module 160 includes an external storage interface, a USB interface, and a subscriber identification module (SIM) card interface. The external storage interface can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the electronic device 10. The external storage card communicates with the processor 110 through the external storage interface to perform data storage. The universal serial bus interface is used for communication between the electronic device 10 and other electronic devices. The subscriber identification module card interface is used to communicate with the SIM card installed in the electronic device 10, for example, to read or write phone numbers stored in the SIM card.

[0232] In some embodiments, the electronic device 10 further includes buttons 101, a motor, and indicators. The buttons 101 may include volume buttons, a power button, etc. The motor is used to generate a vibration effect in the electronic device 10, for example, vibrating when the user's electronic device 10 is called to prompt the user to answer the call. The indicators may include laser indicators, radio frequency indicators, LED indicators, etc.

[0233] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0234] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.

[0235] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0236] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, compact disc-read-only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagated signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0237] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0238] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0239] It should be noted that in the examples and description of this patent, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Although this application has been illustrated and described with reference to certain preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application.

Claims

1. A method for controlling headphones, characterized in that, The control method for the headphones, applied to electronic devices, includes: In response to the user's input of a new template in the recording interface, the newly added first template to be recorded is displayed in the recording interface; In response to the user's recording operation on the first template to be recorded, a command to collect sound is sent to the headphones, so that the headphones collect audio data for a preset total duration. Once the data acquisition is complete, the audio data is stored based on the preset total duration, and the first audio template corresponding to the first template to be recorded is stored. The headphones are then controlled to enter noise reduction mode. A matching value is determined based on the matching degree between the ambient sound data of the environment in which the headphones are located and the audio templates stored in the electronic device. The audio templates stored in the electronic device include the first audio template. The ambient sound data and the scene data of the environment in which the headphones are located are input into a preset scene model and a first event model to determine the probability of a second event. The probability of the second event is used to reflect the degree of matching between the ambient sound data and the preset event corresponding to the first event model. The scene data includes location data and time data. The time data includes the current date, the user's dwell time, and the user's movement time. Based on the matching value and the second event probability, a first event probability is determined, wherein the first event probability is used to reflect the probability that the ambient sound data matches the audio template; If the probability of the first event is greater than the second threshold, the headphones are controlled to switch from the noise cancellation mode to the pass-through mode. The step of determining the matching value based on the matching degree between the ambient sound data of the environment in which the headphones are located and the audio templates stored in the electronic device includes: The musical and speech features in the ambient sound data are determined. The musical features include spectral contrast, spectral flatness, tonality centroid, and chromaticity variation. The speech features include Mel frequency cepstral, tonality envelope, zero-crossing rate, and power spectrum variation. Determine the music matching degree between each music feature in the ambient sound data and the corresponding music feature of the audio template, and determine the speech matching degree between each speech feature in the ambient sound data and the corresponding speech feature of the audio template; Determine whether the music matching degree corresponding to each of the music features and / or the speech matching degree corresponding to each of the speech features meets a first threshold. If any matching degree among the music matching degree corresponding to each of the music features and / or the speech matching degree corresponding to each of the speech features meets the first threshold, determine the music matching value based on the music matching degree corresponding to each of the music features and determine the speech matching value based on the speech matching degree corresponding to each of the speech features. The step of determining the music matching value based on the music matching degree corresponding to each of the music features includes: when the music matching degree corresponding to each of the music features meets the first threshold, incrementing the music matching value by 1 to obtain an updated music matching value; when the music matching degree corresponding to each of the music features does not meet the first threshold, decrementing the music matching value by 1 to obtain an updated music matching value, wherein the initial value of the music matching value is a preset value. The step of determining the voice matching value based on the voice matching degree corresponding to each of the voice features includes: when the voice matching degree corresponding to each of the voice features meets the first threshold, incrementing the voice matching value by 1 to obtain the updated voice matching value; when the voice matching degree of each of the voice features does not meet the first threshold, decrementing the voice matching value by 1 to obtain the updated voice matching value, wherein the initial value of the voice matching value is a preset value. The step of inputting the ambient sound data and the scene data of the environment in which the headphones are located into a preset scene model and a first event model to determine the probability of the second event includes: The music features, speech features, and scene data of the ambient sound data are input into the preset scene model to determine the current scene corresponding to the environment in which the headphones are located; The music features of the ambient sound data, the speech features of the ambient sound data, and the scene data are input into the first event model corresponding to the current scene to determine the probability of the second event; Determining the probability of the first event based on the matching value and the second event probability includes: Obtain the weights corresponding to the voice matching values ​​and the music matching values; The first event probability is determined based on the voice matching value, the weight corresponding to the voice matching value, the music matching value, the weight corresponding to the music matching value, and the second event probability.

2. The headphone control method according to claim 1, characterized in that, The method includes: When the time for the control earphone to enter the pass-through mode reaches a third threshold, the control earphone to close the pass-through mode, wherein the third threshold is determined based on the template length time of the audio template corresponding to the ambient sound data.

3. The headphone control method according to claim 1 or 2, characterized in that, The method for obtaining the first audio template includes: Extract audio features from the audio data of the preset total duration, wherein the audio features include music features and / or speech features; The first audio template is created based on the audio features of the audio data with the preset total duration.

4. The headphone control method according to claim 3, characterized in that, The extraction of audio features from the audio data of the preset total duration includes: The audio data of the preset total duration is divided to obtain multiple audio segments; The time-domain signals of the multiple audio segments are converted into frequency-domain signals to obtain the spectral information of the multiple audio segments; Based on the spectral information of the multiple audio segments, the audio features of each frame in the multiple audio segments are determined.

5. The headphone control method according to claim 1, characterized in that, The method for obtaining the pre-set scene model includes: Acquire ambient sound data and scene data for the target scene; Based on the ambient sound data and scene data of the target scene, the preset scene model is trained, wherein the preset scene model is used to determine the current scene.

6. The headphone control method according to claim 1, characterized in that, The method for obtaining the first event model includes: Acquire the ambient sound data of the target scene and the preset scene model; Based on the ambient sound data of the target scene and the preset scene model, the first event model is trained, wherein the first event model is used to determine the degree of matching between the ambient sound data in the current scene and the preset event.

7. An electronic device, characterized in that, It includes: a memory for storing instructions executed by one or more processors of the electronic device, and the processor being one of the one or more processors of the electronic device for executing the control method of the headphones according to any one of claims 1 to 6.

8. A readable medium, characterized in that, The readable medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the headphone control method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Earphone control method and device and earphone

    CN112770214A

  • Earphone control method and earphone

    CN116033312A

  • Switching control method and system of wireless earphone and wireless earphone

    CN116112839A

  • Switching control method and system of wireless earphone and wireless earphone

    CN116193315A