Earphone mode switching method, earphone and computer readable storage medium

By recognizing when a user speaks to others on the headphones and automatically switching to pass-through mode, the problem of not being able to hear external sounds while wearing headphones is solved, enabling convenient conversation and a better auditory experience.

CN117956332BActive Publication Date: 2026-05-01HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2022-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

When users turn on active noise cancellation while wearing headphones, they cannot hear external sounds clearly, making it cumbersome to talk to others and affecting the user experience.

Method used

By acquiring sound signals from the microphone on the headphones, the system can identify whether the user wearing the headphones is talking to someone else and automatically switch to pass-through mode to improve the ability to acquire external sounds.

Benefits of technology

Users can hear external sounds without manual operation, improving the convenience and auditory experience of communicating with others.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117956332B_ABST
    Figure CN117956332B_ABST
Patent Text Reader

Abstract

The application provides a switching method of earphone mode, an earphone and a computer readable storage medium. The switching method of earphone mode comprises the following steps: acquiring a sound signal collected by a microphone on the earphone; in the case that the current working mode of the earphone is a first mode and a first user wearing the earphone is talking with a second user according to the sound signal, switching the earphone from the first mode to a second mode, and the ability of the earphone to acquire external sound in the second mode is greater than the ability of the earphone to acquire external sound in the first mode. Therefore, when the user listens to audio through the earphone, if there is a need to talk with others, the external sound can be heard without manual operation, and the user can better talk with others.
Need to check novelty before this filing date? Find Prior Art

Description

Method for switching headphone modes, headphones, and computer-readable storage media Technical Field

[0001] This application relates to the field of smart wearable devices, and more particularly to a method for switching headphone modes, headphones, and a computer-readable storage medium. Background Technology

[0002] As headphone technology continues to develop, headphone functions are becoming increasingly diverse. For example, some headphones feature active noise cancellation (ANC), which allows users to eliminate external noise and better hear the sound within the headphones. However, when users activate ANC, other noise reduction features, or are focused solely on listening to the audio playing, they may find it difficult to hear external sounds clearly. When they need to communicate with others, they must remove the headphones, pause audio playback, or adjust the headphone's operating mode, which is cumbersome and negatively impacts the user experience. Summary of the Invention

[0003] This application provides a method for switching headphone modes, a headphone, and a computer-readable storage medium, which solves the problem of cumbersome operation when a user needs to communicate with others while wearing headphones to listen to audio.

[0004] To achieve the above objectives, this application adopts the following technical solution:

[0005] In a first aspect, a method for switching headphone modes is provided, applied to headphones, comprising: acquiring sound signals collected by a microphone on the headphones; when the current working mode of the headphones is a first mode, and it is determined from the sound signals that a first user wearing the headphones is talking to a second user, switching the headphones from the first mode to a second mode, wherein the headphones have a greater ability to acquire external sounds in the second mode than in the first mode.

[0006] In the above embodiments, when the headphones are currently in the first operating mode, their ability to acquire external sounds is poor, and the user may not be able to hear external sounds clearly. If, based on the sound signal, it is determined that the first user wearing the headphones is speaking to the second user, the headphones are switched to the second mode, allowing them to acquire external sounds better. In this way, the user can hear external sounds without manual operation, thus enabling better conversation with others.

[0007] In one embodiment, determining that a first user wearing the earphone is speaking to a second user based on the sound signal includes: if the sound signal contains the voice signal of the first user wearing the earphone, then determining that the first user wearing the earphone is speaking to the second user, thereby improving the accuracy of identifying whether the first user is speaking.

[0008] In one embodiment, determining that the first user wearing the headphones is talking to the second user based on the sound signal includes: if the sound signal contains the voice signal of the first user wearing the headphones, and the voice signal of the first user contains a first wake word, then it is determined that the first user wearing the headphones is talking to the second user, thereby avoiding the probability of misidentifying a user singing as a user talking to someone else.

[0009] In one embodiment, the method further includes: acquiring vibration signals collected by sensors on the earphones to further determine whether the first user is speaking.

[0010] In one embodiment, when the current operating mode of the earphone is a first mode, and it is determined from the sound signal that a first user wearing the earphone is talking to a second user, switching the earphone from the first mode to a second mode includes: when the current operating mode of the earphone is a first mode, and it is determined from the sound signal and the vibration signal that a first user wearing the earphone is talking to a second user, switching the earphone from the first mode to the second mode improves the accuracy of identifying when a first user is talking to a second user.

[0011] In one embodiment, after acquiring the sound signal collected by the microphone on the earphone, the method further includes: when the current working mode of the earphone is a first mode, the sound signal contains the voice signal of the second user, and the voice signal of the second user contains a second wake word, switching the earphone from the first mode to the second mode, thereby improving the accuracy of recognizing that the first user is talking to the second user.

[0012] In one embodiment, the microphone on the earphone includes a feedforward microphone, a feedback microphone, and a call microphone. The feedforward microphone is located on the side of the earphone furthest from the ear, and the feedback microphone is located on the side of the earphone closest to the ear. The earphone is used to determine, based on the sound signal collected by the feedback microphone, that the first user is speaking to the second user. The earphone is also used to determine, based on the sound signals collected by the feedforward microphone and the call microphone, that the sound signal contains the voice signal of the second user. Collecting the corresponding sound signal according to the position of the microphone in the earphone can improve the accuracy of subsequent sound signal recognition.

[0013] In one embodiment, the method further includes: determining the propagation direction of the second user's voice signal based on the sound signals collected by the feedforward microphone and the call microphone, and collecting the sound signal based on the propagation direction, thereby improving the accuracy of speech recognition.

[0014] In one embodiment, after switching the headphones from the first mode to the second mode, the method further includes: if no conversation between the first user and the second user is detected within a first duration, and the second user's voice signal is not present in the sound signal, switching the headphones from the second mode back to the first mode, so that the user can continue to listen to the audio played in the headphones without manual operation when not needing to talk to others.

[0015] In one embodiment, the method further includes: receiving setting information for the first duration sent by an electronic device. Setting the first duration on the electronic device allows the user to conveniently set and view the first duration.

[0016] In one embodiment, before acquiring the sound signal collected by the microphone on the earphone, the method further includes: acquiring a wake word recognition model sent by an electronic device, the wake word recognition model being trained from text including the second wake word and / or speech including the second wake word;

[0017] Correspondingly, after acquiring the sound signal collected by the microphone on the earphone, the method further includes: when it is determined that the sound signal contains the voice signal of the second user, inputting the voice signal of the second user into the wake word recognition model to obtain a recognition result output by the wake word recognition model indicating whether the second wake word is contained. Recognizing the voice signal of the second user through the wake word recognition model can improve recognition accuracy.

[0018] In one embodiment, before obtaining the wake-word recognition model sent by the electronic device, the method further includes: sending test speech collected by the microphone on the earpiece to the electronic device, wherein the electronic device is used to train a classification model based on the test speech to obtain the wake-word recognition model. Collecting speech through the earpiece and training the wake-word recognition model by the electronic device can improve the accuracy of the obtained wake-word recognition model.

[0019] In one embodiment, before acquiring the sound signal collected by the microphone on the headphones, the method further includes: receiving an instruction from an electronic device to activate the automatic headphone mode switching function, thereby activating the automatic headphone mode switching function according to user needs.

[0020] In one embodiment, acquiring the sound signal collected by the microphone on the earphone includes: when the earphone is worn on both ears, the probability that the user cannot hear external sounds is relatively high. Acquiring the sound signal collected by the microphone on the earphone is used to determine whether the earphone needs to be switched to a second mode, thereby improving the intelligence level of the earphone.

[0021] In one embodiment, the headphones include a main earpiece and a secondary earpiece, and the sound signal is collected by a microphone on the main earpiece.

[0022] In one embodiment, when the main earphone is not in operation or the microphone on the main earphone is in an abnormal working state, the microphone on the earphone is instructed to collect the sound signal, thereby enabling the collection of a more accurate sound signal.

[0023] In one embodiment, the method further includes: when the current operating mode of the headphones is a first mode, and it is determined from the sound signal that a first user wearing the headphones is talking to a second user, reducing the volume of the audio being played in the headphones, or instructing the headphones to stop playing audio, so that the user can hear external sound signals more easily and communicate better with others.

[0024] Secondly, a headphone mode switching device is provided, applied to headphones, the device comprising:

[0025] The communication module is used to acquire sound signals collected by the microphone on the earphone;

[0026] The processing module is configured to switch the headphones from the first mode to a second mode when the current working mode of the headphones is the first mode and the sound signal determines that the first user wearing the headphones is talking to the second user. The headphones have a greater ability to acquire external sounds in the second mode than in the first mode.

[0027] In one embodiment, the processing module is specifically used for:

[0028] If the sound signal contains the voice signal of the first user wearing the headphones, then it is determined that the first user wearing the headphones is speaking to the second user.

[0029] In one embodiment, the processing module is specifically used for:

[0030] If the sound signal contains the voice signal of the first user wearing the headphones, and the voice signal of the first user contains a first wake-up word, then it is determined that the first user wearing the headphones is talking to the second user.

[0031] In one embodiment, the communication module is further configured to:

[0032] The vibration signal collected by the sensor on the earphone is acquired.

[0033] In one embodiment, the processing module is further configured to:

[0034] If the current operating mode of the headphones is the first mode, and it is determined from the sound signal and the vibration signal that the first user wearing the headphones is talking to the second user, the headphones are switched from the first mode to the second mode.

[0035] In one embodiment, the processing module is further configured to:

[0036] When the current working mode of the headphones is the first mode, the sound signal contains the voice signal of the second user, and the voice signal of the second user contains a second wake-up word, the headphones are switched from the first mode to the second mode.

[0037] In one embodiment, the microphone on the earphone includes a feedforward microphone, a feedback microphone, and a call microphone. The feedforward microphone is located on the side of the earphone away from the ear, and the feedback microphone is located on the side of the earphone closer to the ear. The earphone is used to determine, based on the sound signal collected by the feedback microphone, that the first user is speaking to the second user, and the earphone is used to determine, based on the sound signals collected by the feedforward microphone and the call microphone, that the sound signal contains the voice signal of the second user.

[0038] In one embodiment, the processing module is further configured to:

[0039] The propagation direction of the second user's voice signal is determined based on the sound signals collected by the feedforward microphone and the call microphone, and the sound signal is collected based on the propagation direction.

[0040] In one embodiment, the processing module is further configured to:

[0041] If no conversation between the first user and the second user is detected within the first time period, and the second user's voice signal is not present in the sound signal, the earphones will be switched from the second mode to the first mode.

[0042] In one embodiment, the communication module is further configured to:

[0043] Receive the setting information for the first duration sent by the electronic device.

[0044] In one embodiment, the communication module is further configured to:

[0045] Obtain a wake word recognition model sent by an electronic device, wherein the wake word recognition model is trained from text including the second wake word and / or speech including the second wake word;

[0046] Correspondingly, after acquiring the sound signal collected by the microphone on the earphone, the method further includes:

[0047] When it is determined that the sound signal contains the voice signal of the second user, the voice signal of the second user is input into the wake word recognition model to obtain the recognition result of whether the wake word recognition model contains the second wake word.

[0048] In one embodiment, the communication module is further configured to:

[0049] The electronic device sends test voice collected by the microphone on the earphone to the electronic device, and the electronic device uses the test voice to train a classification model to obtain the wake word recognition model.

[0050] In one embodiment, the communication module is further configured to:

[0051] Receives an instruction from an electronic device to activate the automatic headphone mode switching function.

[0052] In one embodiment, the communication module is specifically used for:

[0053] When the headphones are worn on both ears, the sound signals collected by the microphone on the headphones are acquired.

[0054] In one embodiment, the headphones include a main earpiece and a secondary earpiece, and the sound signal is collected by a microphone on the main earpiece.

[0055] In one embodiment, when the main earphone is in a non-operating state or the microphone on the main earphone is in an abnormal operating state, the microphone on the secondary earphone is instructed to collect the sound signal.

[0056] In one embodiment, the processing module is further configured to:

[0057] If the current operating mode of the headphones is the first mode, and it is determined from the sound signal that the first user wearing the headphones is talking to the second user, the volume of the audio being played in the headphones is reduced, or the headphones are instructed to stop playing audio.

[0058] Thirdly, a headset is provided, including a processor for executing a computer program stored in a memory to implement the headset mode switching method as described in the first aspect above.

[0059] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the headphone mode switching method as described in the first aspect above.

[0060] Fifthly, a chip is provided, the chip including a processor and a memory coupled thereto, the processor executing a computer program or instructions stored in the memory to implement the headphone mode switching method as described in the first aspect above.

[0061] Sixthly, a computer program product is provided, which, when running on a terminal device, causes the terminal device to execute the headphone mode switching method described in the first aspect above.

[0062] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0063] Figure 1 is a schematic diagram of the hardware structure of an earphone provided in an embodiment of this application;

[0064] Figure 2 is a schematic diagram of the internal structure of an earphone provided in an embodiment of this application;

[0065] Figure 3 is an application scenario diagram of an earphone provided in an embodiment of this application;

[0066] Figure 4 is a flowchart illustrating a headphone mode switching method provided in an embodiment of this application;

[0067] Figure 5 is a schematic diagram of a principle for identifying whether a first user wearing headphones is speaking, according to an embodiment of this application.

[0068] Figure 6 is a flowchart of a method for recognizing a second wake word according to an embodiment of this application;

[0069] Figure 7 is a flowchart of a wake word recognition model for headphones provided in an embodiment of this application;

[0070] Figure 8 is a schematic diagram of the settings page for turning on headphones in a scenario provided by an embodiment of this application;

[0071] Figure 9 is a schematic diagram of the headphone settings page in another scenario provided by an embodiment of this application;

[0072] Figure 10 is a schematic diagram of the settings page of an AI pass-through mode provided in an embodiment of this application;

[0073] Figure 11 is a scenario diagram of setting a wake word according to an embodiment of this application;

[0074] Figure 12 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0075] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0076] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0077] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0078] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0079] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0080] The headphone mode switching method provided in this application is applied to headphones. The headphones have at least active noise cancellation and hearthrough (HT) functionality. The active noise cancellation function reduces unwanted external noise while wearing the headphones, allowing the user to better hear the audio played within the headphones. The hearthrough function transmits sound from the external environment, achieving the same effect as if the user were not wearing headphones.

[0081] Headphones can be over-ear headphones, ear-hook headphones, neckband headphones, or in-ear headphones. In-ear headphones also include in-ear headphones (or canine headphones) or semi-in-ear headphones. Headphones consist of two sound-emitting units that hang on the ears. The headphone that fits the left ear is called the left headphone, and the headphone that fits the right ear is called the right headphone. The left and right headphone have similar structures.

[0082] Figure 1 shows a schematic diagram of an optional hardware structure for the headphones 100.

[0083] As shown in Figure 1, the earphone 100 (left earphone or right earphone) includes: processor 110, memory 120, interface device 130, communication device 140, microphone 150, speaker 160 and sensor 170.

[0084] The processor 110 may include one or more processing units, such as a central processing unit (CPU), a microprocessor (MCU), etc.

[0085] The memory 120 can be used to store computer executable program code, which includes instructions. Examples include ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. The processor 110 executes various functional applications and data processing of the headset 100 by running instructions stored in the internal memory 121 and / or instructions stored in memory located within the processor.

[0086] The interface device 130 includes, for example, various bus interfaces, such as serial bus interfaces, parallel bus interfaces, etc.

[0087] The communication device 140 is used for wired or wireless communication with external devices.

[0088] Microphone 150 is used to convert received audio signals into electrical signals. Microphone 150 can be an analog microphone or a digital microphone. Microphone 150 may include a feed-forward microphone (FF), a feed-back microphone (FB), and a talk microphone.

[0089] The speaker 160 is used to convert electrical signals into sound signals and output them.

[0090] Sensor 170 can be an accelerometer or a bone conduction sensor.

[0091] Taking in-ear headphones as an example, the left or right earphone includes a rubber tip that can be inserted into the ear canal, an ear cup that fits close to the ear, and an earphone stem suspended from the ear cup. The rubber tip directs sound into the ear canal, and the ear cup contains components such as a battery, speaker, and sensors. The earphone stem can be equipped with a microphone, physical buttons, etc., and can be shaped like a cylinder, cuboid, or ellipse.

[0092] As shown in Figure 2, in one embodiment, the feedforward microphone 151 is located on the outside of the earphone, the feedback microphone 152 is located on the inside of the earphone, and the call microphone 153 is located at the bottom of the earphone stem. When the user wears the earphone, the feedforward microphone 151 is located on the side away from the ear, the feedback microphone 152 is located on the side closer to the ear, and the call microphone 153 is located on the side away from the ear and close to the user's mouth.

[0093] A speaker 160 is located between a feedforward microphone 151 and a feedback microphone 152. A sensor 170 is disposed in the ear canal for detecting vibration signals of the auricle.

[0094] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the earphone 100. In other embodiments of this application, the earphone 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0095] As shown in Figure 3, when user A is wearing headphones 100 (left and / or right) to listen to music, or when headphones 100 are in active noise cancellation mode, if user A needs to communicate with the outside world, such as with user B, user A typically needs to remove headphones 100 (at least one of the left or right headphones), pause music playback, lower the music volume, or switch the headphones 100's operating mode from active noise cancellation to pass-through mode to ensure that user A can hear user B's speech and thus have a good conversation with user B. This operation is cumbersome and does not provide a good user experience.

[0096] To address the aforementioned issues, this application provides a method for switching headphone modes. By acquiring sound signals collected by the microphone on the headphones, and when the headphones are currently in a first mode with poor external sound acquisition capabilities, and the sound signal indicates that the first user wearing the headphones is speaking with a second user, the headphones are switched from the first mode to a second mode with stronger external sound acquisition capabilities, allowing external sounds to pass through. This eliminates the need for manual operation, enabling users to clearly hear external sounds, facilitating effective communication, and enhancing the user's auditory experience of sounds of interest.

[0097] The following is an exemplary description of the headphone mode switching method provided in the embodiments of this application.

[0098] As shown in Figure 4, an embodiment of this application provides a method for switching headphone modes, including:

[0099] S401: Acquire the sound signal collected by the microphone on the earphone.

[0100] The microphone used to collect sound signals can be any one or more of the following: a feedforward microphone, a feedback microphone, or a call microphone on the headset.

[0101] In one embodiment, when it is detected that the headphones are worn in both ears, it is determined that the user may be unable to hear external sounds. The sound signal collected by the microphone on the headphones is then acquired, and subsequent steps determine whether the headphones' operating mode needs to be switched to pass-through mode. In the scenario where the headphones are worn in both ears, the sound signal can be collected by the microphone on one earbud or by the microphones on both earbuds simultaneously. If the sound signal is collected by the microphones on both earbuds simultaneously, the sound signal collected by one earbud can be used to subsequently determine whether the first user wearing the headphones is speaking to the second user, or the combined sound signal from both earbuds can be used to subsequently determine whether the first user wearing the headphones is speaking to the second user.

[0102] The headset includes a master headset and a slave headset. The master headset can be the default headset (e.g., the left headset or the right headset) or one of the headsets set by the user. In scenarios where both headsets are worn and the sound signal is collected by the microphone on one of the headsets, the microphone on the master headset can be used by default.

[0103] In scenarios where the microphone on the main earphone is used to collect sound signals, if the microphone on the main earphone is malfunctioning (e.g., unable to collect sound or the volume of the collected sound signal is too low), the system will switch from the secondary earphone to the main earphone, and the microphone on the secondary earphone will collect the sound signal, or the microphone on the secondary earphone will directly collect the sound signal. For example, in scenarios where the user wears headphones in both ears, if the user covers the microphone on the main earphone with their hand or the microphone on the main earphone is obstructed by clothing or hat, the volume of the sound signal collected by the microphone on the main earphone will be too low. In this case, the microphone on the secondary earphone will collect the sound signal, thereby improving the reliability of the received sound signal.

[0104] In one embodiment, when a user wears either earphone, the earphone acquires the sound signal collected by its microphone, and then determines whether to switch the earphone's operating mode to a second mode based on subsequent steps. When the user wears only one earphone, the sound signal can be collected by the microphone on the worn earphone (master or slave earphone), while the unworn earphone (i.e., the earphone in a non-operating state) does not collect a sound signal.

[0105] S402: When the current working mode of the earphone is the first mode, and it is determined from the sound signal that the first user wearing the earphone is talking to the second user, the earphone is switched from the first mode to the second mode, wherein the earphone's ability to acquire external sound in the second mode is greater than its ability to acquire external sound in the first mode.

[0106] The first mode can be a preset noise cancellation mode, an off mode, or another mode that reduces external noise. The second mode is a pass-through mode, in which the headphones can normally receive or amplify external sound signals, or filter out ambient sounds from the external sound signals.

[0107] In one embodiment, the earphone, upon determining that the current operating mode is a first mode, acquires sound signals collected by the microphone, and then determines whether a first user wearing the earphone is speaking to a second user based on the sound signals. If it is determined that the first user is speaking to the second user, the earphone switches from the first mode to the second mode. If it is determined that the first user is not speaking to the second user, no switching occurs. Here, the second user refers to anyone other than the first user.

[0108] In another embodiment, the headset's microphone can continuously collect sound signals. If it is determined from the sound signals that the first user is speaking to the second user, it then determines whether the current operating mode is the first mode. If the current operating mode is the first mode, the headset is switched from the first mode to the second mode. If the current operating mode is not the first mode, no switching occurs.

[0109] In one embodiment, when the current operating mode of the headphones is the first mode, regardless of what audio is being played in the headphones, it is determined whether the first user is speaking to the second user based on the collected sound signal.

[0110] In another embodiment, when the headset is currently in the first operating mode and is not in a call state, it is determined whether the first user is speaking to the second user based on the collected sound signal, thereby avoiding the impact of headset mode switching on call quality. The headset can determine whether it is in a non-call state based on a preset identifier stored in the headset during a call or based on the identifier of the application playing audio obtained from the electronic device, wherein the electronic device is communicatively connected to the headset.

[0111] In one embodiment, after acquiring an audio signal, if the earphone determines that the audio signal contains the voice signal of a first user, then it is determined that the first user wearing the earphone is speaking to a second user. Exemplarily, the earphone can determine whether the audio signal contains a voice signal after acquiring it. The earphone can determine whether the audio signal contains a voice signal based on the frequency distribution range of the audio signal. For example, if the audio signal contains a signal with a frequency range of 300Hz to 3400Hz, it is determined that the audio signal contains a voice signal. After determining that the audio signal contains a voice signal, the features of the voice signal are extracted and compared with the features of a pre-stored voice signal. The pre-stored voice signal can be pre-stored in the earphone or recorded by the first user. If the features match, it is determined that the audio signal contains the voice signal of the first user, and thus it is determined that the first user is speaking to the second user. Alternatively, after acquiring the audio signal, if it is determined that the audio signal contains a voice signal, the earphone can further determine the propagation direction of the voice signal. If the propagation direction of the voice signal is inward of the earphone, it is determined that the audio signal contains the voice signal of the first user, and thus it is determined that the first user is speaking to the second user.

[0112] In one embodiment, after acquiring an audio signal, if the earphone determines that the audio signal contains the voice signal of a first user, and the first user's voice signal contains a first wake-up word, then it determines that the first user is speaking to the second user. The first wake-up word can be a commonly used phrase set by the first user for conversations with others, such as "hello," "work," or "weather." This method reduces the probability of misidentifying the first user singing as speaking to the second user, thereby improving the accuracy of audio signal recognition.

[0113] Specifically, when the earphone determines that the sound signal contains the first user's voice signal, it inputs the voice signal into the wake word recognition model and obtains the recognition result of whether the wake word recognition model contains the first wake word. The wake word recognition model can be a model pre-stored on the earphone, or it can be a model obtained by the earphone through training on the text or speech input by the first user that includes the first wake word, or it can be obtained by an electronic device through training on the text or speech input by the first user that includes the first wake word, and then sending the wake word recognition model to the earphone.

[0114] In one embodiment, after acquiring an audio signal, if the earphone determines that the audio signal contains the voice signal of a first user, and the electronic device communicatively connected to the earphone is not currently in singing mode, then it is determined that the first user is speaking to the second user. The earphone can determine whether the electronic device is in singing mode based on the name of the currently running application on the electronic device and the data transmission status between the electronic device and the earphone. For example, if the currently running application on the electronic device is a preset application, and the data transmission status between the electronic device and the earphone is that the electronic device receives audio signals sent by the earphone, then it is determined that the electronic device is in singing mode; otherwise, it is determined that the electronic device is not in singing mode. The preset application can be... Applications such as [examples of such applications]. By using the methods described above, the probability of misidentifying a user singing as a user talking to a second user can be reduced, thereby improving the accuracy of sound signal recognition.

[0115] In one embodiment, it can be determined whether a first user is speaking to a second user based on the sound signal collected by any one of the feedforward microphone, feedback microphone, and call microphone. For example, one microphone can be configured to collect the sound signal. If the collected sound signal is determined to include a speech signal, and the speech signal propagates inwards from the earpiece, then it is determined whether the first user is speaking to the second user. As another example, three microphones or two of them can be configured to collect the sound signal. When any microphone collects a sound signal, the collected sound signal is identified. If the identification result determines that the sound signal includes a speech signal, and the speech signal propagates inwards from the earpiece, then it is determined whether the first user is speaking to the second user. The earpiece can determine the propagation direction of the speech signal based on the intensity of the speech signals received from multiple directions or based on the time it takes to receive the speech signals from each direction.

[0116] In one embodiment, the earphones also acquire vibration signals collected by sensors on the earphones to determine the speaking state of the first user. Specifically, if the earphones are currently in a first operating mode and it is determined, based on the sound signal and vibration signal, that the first user wearing the earphones is speaking to the second user, the earphones switch from the first mode to the second mode. For example, after acquiring a sound signal, if the earphones determine that the sound signal includes a speech signal and the vibration signal of the sensor is consistent with a pre-set reference vibration signal, then it determines that the first user is speaking to the second user and switches the earphones from the first mode to the second mode. Here, speaking can cause ear movement, and the reference vibration signal can be a pre-collected vibration signal from when the first user is speaking. The earphones can determine that the first user is speaking to the second user when a sound signal is acquired by a pre-set microphone for acquiring sound signals, the sound signal includes a speech signal, and the vibration signal acquired by the sensor is consistent with the reference vibration signal. Alternatively, the first user can determine that the first user is speaking to the second user when a sound signal is acquired by any one of the three microphones, the sound signal includes a speech signal, and the vibration signal acquired by the sensor is consistent with the reference vibration signal.

[0117] In one embodiment, as shown in Figure 5, the feedforward microphone, feedback microphone, call microphone, and sensor are all in working order. When all three microphones collect sound signals, the sound signal collected by one of the microphones and the vibration signal collected by the sensor are input into the user self-talk detection model. The user self-talk detection model identifies the sound signal and the vibration signal. If the identification result determines that the sound signal includes a speech signal, and the vibration signal collected by the sensor is consistent with the reference vibration signal, then it is determined that the first user is talking to the second user. Alternatively, when all three microphones collect sound signals, the sound signals collected by the three microphones and the vibration signal collected by the sensor are input into the user self-talk detection model. The user self-talk detection model fuses the sound signals collected by the three microphones and identifies the fused sound signal. If the identification result determines that the sound signal includes a speech signal, and the vibration signal collected by the sensor is consistent with the reference vibration signal, then it is determined that the first user is talking to the second user. Otherwise, it is determined that the first user is not talking to the second user. Optionally, the headset can input the sound signals from the three microphones and the vibration signals from the sensors into the user self-talk detection model. The model then filters the sound signals from the feedforward microphone, feedback microphone, and talk microphone based on the audio signal being played from the feedback microphone, resulting in a filtered audio signal. The model then fuses the filtered sound signals and identifies the fused signal. Based on the identification result, it determines whether the first user wearing the headset is speaking to the second user, thereby improving the accuracy of sound signal recognition.

[0118] In one embodiment, if it is determined that the first user wearing the headphones is talking to the second user, the headphones are switched from the first mode to the second mode regardless of whether audio is playing in the headphones.

[0119] In another embodiment, if it is determined that the first user wearing the headphones is talking to the second user, and audio is playing in the headphones, the headphones are switched from the first mode to the second mode. If the volume of the audio in the headphones is lower than a set value or the audio is paused, the headphones' operating mode is not switched. If audio is detected playing in the headphones and the volume is higher than a set value, it indicates that the audio playing in the headphones is affecting the user's ability to hear external sounds. If it is detected that the first user wearing the headphones is still speaking, the headphones are switched from the first mode to the second mode.

[0120] In one embodiment, while switching from the first mode to the second mode, if audio (e.g., music or video) is being played in the headphones, the headphones can also reduce the volume of the playing audio or stop playing the audio.

[0121] After switching the headphones from mode one to mode two, if no conversation between the first and second users is detected within a first duration, the headphones will switch back to mode one, restoring the original operating mode so that users can continue listening to the audio. The first duration can be a default value or set by the user. Users can input the first duration using buttons on the headphones or following prompts. Alternatively, users can input the first duration on an electronic device (such as a mobile phone) connected to the headphones, which will then send the duration to the headphones.

[0122] In the above embodiments, when the headphones are currently in a first mode with poor external sound acquisition capability, and it is determined from the sound signal collected by the microphone on the headphones that the first user wearing the headphones is talking to the second user, the headphones are switched from the first mode to a second mode with strong external sound acquisition capability, allowing the headphones to transmit external sounds. In this way, without manual operation, the user can clearly hear external sounds, thus enabling better communication with others, and also improving the user's auditory experience of sounds of interest.

[0123] In one embodiment, in addition to having the function of switching the headphones from a first mode to a second mode when it is recognized that the first user wearing the headphones is talking to the second user (i.e., user self-speaking recognition function), the headphones also have the function of switching the headphones from a first mode to a second mode when the voice signal of the second user is recognized from the sound signal and the voice signal contains a second wake-up word (i.e., wake-up word recognition function).

[0124] The second wake-up word can be the user's name, title, or some commonly used greetings. If the earphone determines from the sound signal received by the microphone that the sound signal contains the voice signal of the second user, and the voice signal of the second user contains the second wake-up word, it means that the second user is talking to the first user wearing the earphone. The first user has a need to communicate with others, so the earphone is switched from the first mode to the second mode, so that the first user wearing the earphone can hear the outside sounds clearly without removing the earphone, and thus can have a good conversation with others.

[0125] In one embodiment, the determination of whether a second user's voice signal is present in the sound signal collected by the earphones can be made only when both earphones are detected to be worn. When both earphones are worn, the determination can be made based on the sound signal collected by one earphone, or it can be based on the combined sound signal from both earphones.

[0126] In one embodiment, when a user wears either earphone, the earphone acquires the sound signal collected by the microphone on the earphone, and then determines whether the sound signal contains the voice signal of the second user.

[0127] In one embodiment, when a user wears either earphone and a preset audio is playing, the earphone acquires a sound signal from its microphone and determines whether the sound signal contains the voice signal of a second user. The preset audio can be a default setting for the earphones or a user-defined setting. For example, the preset audio could be news audio, meeting audio, etc. The earphone can obtain the application containing the currently playing audio from the electronic device and determine whether the currently playing audio is the preset audio based on the application. If the currently playing audio is the preset audio, it indicates that the first user is highly focused on listening to the audio and may not notice external sounds. Therefore, after receiving the sound signal, it determines whether the sound signal contains the voice signal of the second user. If the sound signal does not contain the voice signal of the second user, and the second user's voice signal contains a second wake-up word, it indicates that the second user is speaking to the first user. The earphone then switches to a second mode to facilitate better conversation between the user and others.

[0128] For any one earphone (left or right), it can be determined whether the sound signal contains the voice signal of the second user based on the sound signal collected by any one of the microphones: the feedforward microphone, the feedback microphone, and the call microphone. For example, it can be set to collect the sound signal using one of the microphones. If it is determined that the collected sound signal contains a voice signal, and the voice signal propagates outward from the earphone, then it is determined whether the sound signal contains the voice signal of the second user. As another example, it can be set to collect the sound signal using three microphones or two of them. When any microphone collects a sound signal, if it is determined that the collected sound signal contains a voice signal, and the voice signal propagates outward from the earphone, then it is determined whether the sound signal contains the voice signal of the second user.

[0129] The headphones can also determine whether the audio signal contains the voice signal of a second user based on the audio signals simultaneously collected by the feedforward microphone, feedback microphone, and call microphone. For example, if all three microphones collect audio signals, and it is determined that the audio signal contains a voice signal based on the audio signal collected by one of the microphones or based on the audio signal obtained by fusing the audio signals collected by the three microphones, and the direction of propagation of the voice signal is outward from the headphones, then it is determined whether the audio signal contains the voice signal of a second user. Optionally, when the headphones determine that all three microphones have collected audio signals, they filter the audio signals collected by the feedforward microphone, feedback microphone, and call microphone based on the audio signal of the currently playing audio collected by the feedback microphone. Then, the headphones fuse the filtered audio signals and determine whether the audio signal contains a voice signal based on the fused audio signal. The headphones can also filter and identify only the audio signal collected by the feedforward microphone or only the audio signal collected by the call microphone to determine whether the audio signal contains the voice signal of a second user.

[0130] In one embodiment, the earphone can determine that a first user wearing the earphone is speaking to a second user based on the sound signal collected by the feedback microphone, and determine that the sound signal collected by the feedforward microphone and the call microphone contains the second user's voice signal. For example, the earphone determines that the first user wearing the earphone is speaking to the second user when it determines that the sound signal collected by the feedback microphone contains a voice signal; or the earphone determines that the first user wearing the earphone is speaking when it determines that the sound signal collected by the feedback microphone contains a voice signal and the direction of propagation of the voice signal is inward of the earphone; or the earphone determines that the first user wearing the earphone is speaking when it determines that the sound signal collected by the feedback microphone contains a voice signal and the vibration signal collected by the sensor is consistent with the reference vibration signal. If the earphone determines that both the feedforward microphone and the call microphone have picked up sound signals, and the sound signal picked up by one of the microphones contains a speech signal, or if the sound signal after the two microphones are fused contains a speech signal, then the earphone determines that the speech signal is the second user's speech signal, i.e., the sound signal contains the second user's speech signal; or if the earphone determines that the sound signal contains a speech signal, and the direction of propagation of the speech signal is outward from the earphone, then the earphone determines that the sound signal contains the second user's speech signal; or if the earphone determines that the sound signal contains a speech signal, and the vibration signal picked up by the vibration sensor is inconsistent with the reference vibration signal, then the earphone determines that the sound signal contains the second user's speech signal.

[0131] In one embodiment, when the earphone determines that the sound signal contains the voice signal of a second user, and that the voice signal of the second user contains a second wake-up word, it determines the location of the second user based on the sound signal collected by the feedforward microphone and simultaneously determines the location of the second user based on the sound signal collected by the call microphone. Then, based on the two locations and the distance between the feedforward microphone and the call microphone, the position of the second user is determined. Based on the position of the second user, the propagation direction of the second user's voice signal is determined, and the sound signal is collected based on the propagation direction. That is, the propagation direction of the sound signal is used as the pickup direction of the feedforward microphone and the call microphone, so that the feedforward microphone and the call microphone only collect sound signals in the pickup direction, thereby improving the accuracy of sound signal analysis.

[0132] In one embodiment, when the earphone determines that the sound signal contains the voice signal of a second user, it inputs the voice signal of the second user into a wake-up word recognition model to obtain a recognition result from the wake-up word recognition model indicating whether the second wake-up word is included. For example, as shown in Figure 6, when the earphone determines that both the feedforward microphone and the call microphone have collected sound signals, it inputs the sound signals collected by the feedforward microphone and the call microphone into a Voice Activity Detection (VAD) module. The VAD module fuses the sound signals collected by the feedforward microphone and the call microphone to determine whether the fused sound signal contains a voice signal. If the fused sound signal contains a voice signal, it determines whether the sound signal contains the voice signal of the second user based on the propagation direction of the voice signal or the vibration signal collected by the vibration sensor. If the sound signal contains the voice signal of the second user, the VAD module filters out ambient sounds from the sound signal to obtain the voice signal of the second user, and inputs the voice signal of the second user into the wake-up word recognition model to obtain a recognition result from the wake-up word recognition model indicating whether the second wake-up word is included. Optionally, the VAD module can also acquire the audio signal of the playing audio from the electronic device. If it is determined that the fused audio signal contains the second user's voice signal, the module filters out ambient sounds and the audio signal of the playing audio to obtain the second user's voice signal. Specifically, the VAD module can extract the voice signal and filter ambient sounds by extracting audio signals within a preset frequency band.

[0133] In one embodiment, the wake word recognition model can be a model pre-stored on the headphones, or it can be a model obtained by training the headphones on a second wake word input by the user.

[0134] In another embodiment, the wake-up word recognition model can be trained by a communication terminal (e.g., an electronic device or a server) connected to the headset and sent to the headset. For example, taking an electronic device as the communication terminal, as shown in Figure 7, when the electronic device receives a user's instruction to input a second wake-up word, it receives both the text and speech input by the user, including the second wake-up word. The text is then converted to speech using a Text-to-Speech (TTS) method. Based on the speech obtained from the text and the speech input by the user including the second wake-up word, the electronic device trains a classification model to obtain the wake-up word recognition model. The electronic device then sends the wake-up word recognition model to the headset.

[0135] It is understandable that electronic devices can train a classification model based solely on the speech obtained from the text to obtain a second wake-up word recognition model, or they can train a classification model based solely on the speech input by the user that includes the second wake-up word to obtain a second wake-up word recognition model.

[0136] The speech containing the second wake-up word can be captured by an electronic device or by headphones and then sent to the electronic device. For example, the second wake-up word can be spoken by another person while the user is wearing headphones. When the other person speaks the second wake-up word, the microphone on the headphones captures the speech signal and sends it to the electronic device as test speech. Alternatively, the second wake-up word can be spoken by a first user wearing headphones. When the first user speaks the second wake-up word, the microphone on the headphones captures the speech signal and sends it to the electronic device as test speech. After receiving the test speech, the electronic device can train a classification model to obtain a second wake-up word recognition model.

[0137] In one embodiment, when the wake word recognition function is enabled in the headphones, if no voice signal of the second user is detected within a first duration, that is, no other person is detected speaking, the headphones are switched from the second mode to the first mode.

[0138] In one embodiment, when both the user self-speaking recognition function and the second wake-up word recognition function are enabled, if no speech is detected between the first user and the second user within a first duration, and the second user's voice signal is not present in the audio signal, the headset is switched from the second mode to the first mode. Alternatively, if no voice signal is detected within a first duration when both the user self-speaking recognition function and the second wake-up word recognition function are enabled, the headset is switched from the second mode to the first mode.

[0139] In one embodiment, if the user self-speaking recognition function and the second wake-up word recognition function are both enabled, and no speech is detected between the first user and the second user within a first duration, the headset is switched from the second mode to the first mode. Alternatively, if the user self-speaking recognition function and the second wake-up word recognition function are both enabled, and no voice signal from the second user is detected within a first duration, the headset is switched from the second mode to the first mode.

[0140] In the above embodiments, when the headphones are currently in a first mode where their ability to acquire external sounds is relatively poor, if the headphones determine from the sound signal received by the microphone that the sound signal contains the voice signal of the second user and that the voice signal contains the second wake-up word, it indicates that the second user is speaking to the first user wearing the headphones. This suggests that the first user wearing the headphones has a need to communicate with others, and the headphones are switched from the first mode to the second mode, allowing the headphones to transmit external sounds. In this way, without manual operation, the user can clearly hear external sounds, enabling them to have good conversations and communication with others, while also enhancing the user's auditory experience of sounds of interest.

[0141] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0142] In one embodiment, the headphones and electronic device communicate via any of the following methods: Bluetooth, Wi-Fi, 5G, etc. The user can configure headphone parameters on the electronic device. For example, the user can input an instruction to activate the automatic headphone mode switching function on the electronic device, causing the headphones to automatically switch from a first mode to a second mode based on the instruction. The user can choose to enable either the function of recognizing the user's own voice or the function of recognizing a second wake-up word from external sounds on the electronic device. The user can set a first duration for reverting the headphones from the second mode to the first mode on the electronic device. The user can input text or voice, including a second wake-up word, on the electronic device.

[0143] The electronic device can be a mobile phone, tablet computer, handheld computer, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, media player, wearable device, or other device that can be held / operated with one hand. This application does not impose any special limitations on the specific form / type of the electronic device. The aforementioned electronic device includes, but is not limited to, devices equipped with… Devices running Harmony OS or other operating systems.

[0144] The following section uses a mobile phone as an example to introduce the interaction scenarios between electronic devices and headphones.

[0145] In one scenario, as shown in Figure 8(a), when the phone detects a command to open the Bluetooth page, it displays the Bluetooth page on the screen. The Bluetooth page includes controls for turning Bluetooth on or off, the phone's name, and controls for opening received files. Users can turn the phone's Bluetooth function on or off using these controls. When Bluetooth is on, the Bluetooth page also displays the names of paired devices and available devices. For example, the phone's name is "Tom's Phone," and paired devices include "Tom's Headphones" and "Tom's Watch." Available devices include "Tom's Computer" and "Tom's Speaker." The Bluetooth page also displays settings controls corresponding to each paired device. When the phone detects a click on the settings control corresponding to "Tom's Headphones," it opens the headphone's settings page, as shown in Figure 8(b).

[0146] As shown in Figure 8(b), the headset's settings page includes the name of the headset connected to the phone via Bluetooth, for example, the headset's name is "Tom's Headset." Users can rename the headset by clicking its name. The headset's settings page also includes controls for turning call audio on or off, turning media audio on or off, turning Bluetooth auto-connection on or off, and settings for synchronizing Bluetooth device volume with the phone. Users can use these controls to enable or disable the corresponding headset functions. The headset's settings page also includes multiple noise control modes. For example, the noise control modes include noise cancellation mode, off mode, and pass-through mode. Noise cancellation mode reduces ambient noise, allowing users to better hear the audio played in the headset; off mode disables noise cancellation; and pass-through mode (i.e., the second mode) allows ambient noise to pass through, allowing users to hear ambient sounds while listening to audio. The headset's settings page also displays information about whether AI pass-through mode is enabled. When AI pass-through mode is enabled, the phone sends an instruction to the earphones to activate the automatic earphone mode switching function. Based on this instruction, and with noise control in noise cancellation or off mode, the earphones determine whether the user needs to communicate with others based on the sound signal collected by the microphone. If the user does need to communicate, the noise control mode is switched to pass-through mode. Specifically, if the earphones determine that the first user wearing the earphones is speaking to a second user, or if the sound signal contains the second user's voice signal and a preset second wake-up word, the earphones determine that the first user needs to communicate and switch the noise control mode to pass-through mode. When AI pass-through mode is disabled, the earphones will not switch the noise control mode to pass-through mode based on the sound signal collected by the microphone.

[0147] In another scenario, as shown in Figure 9(a), when the phone detects a command to open the Smart Life application, it displays the Smart Life application page on the screen. The homepage of the Smart Life application displays the devices currently managed by the phone and their connection status. For example, the devices currently managed by the phone include TV A, TV B, speaker A, router A, and Tom's headphones. TV A is connected to the phone, TV B is not connected, speaker A is not connected, router A is connected, and Tom's headphones are connected. When the phone detects a click on Tom's headphones, it opens the headphones' settings page as shown in Figure 9(b).

[0148] As shown in Figure 9(b), the headphone settings page includes the current battery level of the headphones. For example, the left earbud (L) has a current battery level of 100%, the right earbud (R) has a current battery level of 100%, and the charging case has a current battery level of 55%. The headphone settings page also includes multiple noise control modes. For example, the noise control modes include noise cancellation mode, off mode, and pass-through mode. The headphone settings page also displays information on whether the AI ​​pass-through mode is enabled.

[0149] In both scenarios described above, as shown in Figure 8(b) or Figure 9(b), the mobile phone can open the AI ​​pass-through mode settings page when it detects that the user clicks on the AI ​​pass-through mode.

[0150] As shown in Figure 10(a), the AI ​​pass-through mode settings page includes controls to enable or disable the AI ​​pass-through mode, allowing users to turn it on or off. The AI ​​pass-through mode settings page also includes controls to enable or disable the user-defined voice recognition function, allowing users to turn the headset's user-defined voice recognition function on or off. The AI ​​pass-through mode settings page also displays the wake-word recognition function's on / off status, indicating whether the wake-word recognition function is enabled.

[0151] In one embodiment, if an operation to enable AI pass-through mode is detected, the phone defaults to enabling the user's self-speaking recognition function but disables the wake-word recognition function; that is, the phone instructs the headset to only enable the user's self-speaking recognition function. The user can enable the wake-word recognition function in the wake-word recognition settings interface. In another embodiment, if an operation to enable AI pass-through mode is detected, the phone defaults to enabling both the user's self-speaking recognition function and the wake-word recognition function simultaneously; that is, the phone instructs the headset to enable both the user's self-speaking recognition function and the wake-word recognition function simultaneously.

[0152] If the phone detects that the AI ​​pass-through mode has been turned off, it instructs the earphones to simultaneously disable both the user-initiated speech recognition and wake-up word recognition functions. If the phone detects that both user-initiated speech recognition and wake-up word recognition functions are disabled, it sets the AI ​​pass-through mode to off. With the AI ​​pass-through mode off, the phone instructs the earphones to disable the automatic switching of earphone modes.

[0153] Upon detecting a downward swipe operation, the AI ​​pass-through mode settings page swipes down to display more content. For example, as shown in Figure 10(b), the AI ​​pass-through mode settings page also includes a control to enable or disable the external noise detection function. Users can use this control to enable or disable the external noise detection function, and the electronic device sends the enable or disable instruction information to the headphones. When the external noise detection function is disabled, after the headphones switch to pass-through mode according to the AI ​​pass-through mode, if no conversation between the first user wearing the headphones and the second user is detected within a first time period, the noise cancellation mode of the headphones reverts to the mode before switching to pass-through mode. When the external noise detection function is enabled, after the headphones switch to pass-through mode according to the AI ​​pass-through mode, if no conversation between the first user wearing the headphones and the second user is detected within a first time period, and no voice signal from the second user is detected, the noise cancellation mode of the headphones reverts to the mode before switching to pass-through mode.

[0154] The settings page for the AI ​​pass-through mode also includes a setting option for the first duration. For example, as shown in Figure 10(b), the first duration setting option is a slider. Users can set the first duration by sliding the slider, and the electronic device will send the first duration to the earphone. For instance, the first duration setting option can be any duration between 5s and 20s, and the current first duration is 10s. In other embodiments, the first duration setting option can also be multiple selection controls. Users can set the corresponding duration as the first duration by selecting one of the selection controls. For example, the first durations corresponding to the multiple selection controls are 5s, 10s, 15s, etc.

[0155] When a wake word recognition operation is detected, the phone opens the wake word recognition settings page as shown in Figure 11(a). The wake word recognition settings page includes controls to enable or disable the wake word recognition function. Users can enable or disable the wake word recognition function by clicking these controls, and the phone will send the enable or disable instruction information to the headset. The wake word recognition settings page also includes text input and voice input options. Text input is used to receive text input by the user that includes the second wake word, and voice input is used to receive voice input by the user that includes the second wake word. If the phone only receives text input by the user that includes the second wake word, a wake word recognition model is trained based on the text containing the second wake word, and the wake word recognition model is sent to the headset. If the phone only receives voice input by the user that includes the second wake word, a wake word recognition model is trained based on the voice containing the second wake word, and the wake word recognition model is sent to the headset. If the phone receives both text and voice input that include the second wake word, a wake word recognition model is trained based on both, and the wake word recognition model is sent to the headset.

[0156] When the phone detects a click on text input, it opens the text input page shown in Figure 11(b). The text input page includes a previously entered second wake-up word or text containing the second wake-up word. For example, the previously entered text containing the second wake-up word or the text identifier is Input Text 1, Input Text 2, or Input Text 3. When the phone detects that the user clicks the enable control 111 on the right, it enables the corresponding text containing the second wake-up word. Then, the phone trains a wake-up word recognition model based on the enabled text containing the second wake-up word. When the phone detects that the user long-presses a previously entered text containing the second wake-up word or the text identifier, it can delete the previously entered text containing the second wake-up word. For example, if the phone detects that the user long-presses Input Text 1, it deletes Input Text 1. The text input interface also includes a control for adding text. When the phone detects that the user clicks the add text control, it receives the text entered by the user and uses the entered text as the newly added text containing the second wake-up word. When the newly added text containing the second wake-up word is enabled, the phone trains a wake-up word recognition model based on this text.

[0157] When the phone detects a click on voice input, it opens the voice input page shown in Figure 11(c). The voice input page includes already recorded voice messages containing a second wake-up word. For example, the file names of the already recorded voice messages containing the second wake-up word are "Voice Message 1," "Voice Message 2," and "Voice Message 3," respectively. When the phone detects a user clicking on a file name, it plays the corresponding voice message. Similar to the text input page, when the phone detects a user clicking the enable control on the right, it activates the corresponding voice message containing the second wake-up word. The phone then trains a wake-up word recognition model based on the activated voice message containing the second wake-up word. When the phone detects a user long-pressing a file name, it can delete the voice message corresponding to that file name. The voice input interface also includes a control for adding voice messages. When the phone detects a user clicking the add voice message control, it receives the user's input voice message and uses it as a newly added voice message containing the second wake-up word. With the newly added voice message containing the second wake-up word activated, the phone trains a wake-up word recognition model based on that voice message.

[0158] For example, when an action of clicking the "Add Voice" control is detected, the phone opens the "Add Voice" page on the right. When the phone detects that the user has long-pressed the "Long-press to Record Voice" control, it receives the user's recorded voice and uses it as the newly added voice including the second wake-up word. The phone can receive voice recordings with the second wake-up word from the user via the phone or via headphones. When recording voice via headphones, the user can speak, the headphones can record the voice and send it to the phone, and the phone will use the received voice as the voice including the second wake-up word. Alternatively, the headphones can collect external sound while the user is wearing them and send it to the phone, which will then use the received voice as the voice including the second wake-up word.

[0159] In the above embodiments, by setting headphone-related parameters on the electronic device, users can more intuitively understand the headphone settings and easily modify them, thereby improving the user experience. At the same time, setting headphone-related parameters on the electronic device can reduce the number of buttons on the headphones and save headphone power consumption.

[0160] For example, Figure 12 shows a schematic diagram of the structure of an electronic device 200.

[0161] As shown in Figure 12, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an accelerometer sensor 280E, a distance sensor 280F, a proximity sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0162] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0163] Processor 210 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0164] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0165] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0166] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0167] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0168] The wireless communication module 260 can provide solutions for wireless communication applications on the electronic device 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 210. The wireless communication module 260 can also receive signals to be transmitted from processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0169] Electronic device 200 implements display functions through GPU, display screen 294, and application processor.

[0170] The external storage interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external storage interface 220 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0171] Internal memory 221 can be used to store computer executable program code, which includes instructions. Internal memory 221 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 200 (such as audio data, phonebook, etc.). Furthermore, internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 210 executes various functional applications and data processing of electronic device 200 by running instructions stored in internal memory 221 and / or instructions stored in memory disposed in the processor.

[0172] Electronic device 200 can implement audio functions such as music playback and recording through audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, and application processor.

[0173] The audio module 270 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 may be located in the processor 210, or some functional modules of the audio module 270 may be located in the processor 210.

[0174] The speaker 270A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 200 can listen to music or make hands-free calls through the speaker 270A.

[0175] The receiver 270B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 200 answers a telephone call or voice message, the receiver 270B can be brought close to the ear to listen to the voice.

[0176] Microphone 270C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 270C, inputting the sound signal into microphone 270C. Electronic device 200 may have at least one microphone 270C. In some embodiments, electronic device 200 may have two microphones 270C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 200 may also have three, four, or more microphones 270C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0177] The 270D headphone jack is used to connect wired headphones. The 270D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0178] Touch sensor 280K, also known as a "touch device," can be located on display screen 294. The touch sensor 280K and display screen 294 together form a touchscreen, also known as a "touchscreen." Touch sensor 280K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 294. In other embodiments, touch sensor 280K may also be located on the surface of electronic device 200, in a different position than display screen 294.

[0179] Buttons 290 include a power button, volume buttons, etc. Buttons 290 can be mechanical buttons or touch-sensitive buttons. Electronic device 200 can receive button input and generate key signal inputs related to user settings and function control of electronic device 200.

[0180] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0181] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0182] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0184] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0185] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0186] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0187] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for switching headphone modes, applied to headphones, characterized in that, The method includes: in a scenario where the headphones are worn on both ears, and the main earphone is in an abnormal working state while the secondary earphone is not in an abnormal working state, acquiring a first sound signal collected by the microphone on the secondary earphone, wherein the abnormal working state includes a state where sound cannot be collected or a state where the volume of the collected sound signal is less than or equal to a preset value; or, when neither the main earphone nor the secondary earphone is in an abnormal working state, acquiring a second sound signal collected by the microphone on the main earphone, and acquiring a third sound signal collected by the microphone on the secondary earphone, and fusing the second sound signal and the third sound signal to obtain a first sound signal; acquiring vibration signals collected by sensors on the headphones. The method involves: acquiring a fourth audio signal from the feedback microphone that is currently playing audio; performing playback sound filtering on the audio signals acquired by the feedforward microphone, feedback microphone, and call microphone in the first audio signal based on the fourth audio signal to obtain a fifth audio signal, and fusing the filtered fifth audio signals to obtain a sixth audio signal; wherein the microphones on the headset include a feedforward microphone, a feedback microphone, and a call microphone; when the headset is currently in a first operating mode and the sixth audio signal indicates that the first user wearing the headset is speaking to the second user, the headset is switched from the first mode to a second mode, wherein the headset's ability to acquire external sounds in the second mode is greater than its ability to acquire external sounds in the first mode; wherein determining that the first user wearing the headset is speaking to the second user based on the sixth audio signal includes: determining that the first user is speaking to the second user based on the sixth audio signal containing the first user's voice signal and the vibration signal being consistent with a pre-acquired reference vibration signal, wherein the reference vibration signal is a pre-acquired vibration signal of the first user speaking; after acquiring the audio signal acquired by the microphones on the headset, the method further includes: when the headset is currently in a first operating mode and the sixth audio signal contains the second user's voice signal. When the second user's voice signal contains a second wake-up word, the headset is switched from the first mode to the second mode; wherein, when it is determined that the sound signal contains the second user's voice signal and the second user's voice signal contains a second wake-up word, the location of the second user is determined based on the sound signal collected by the feedforward microphone, and the location of the second user is also determined based on the sound signal collected by the call microphone. The position of the second user is determined based on the two locations and the distance between the feedforward microphone and the call microphone. The propagation direction of the second user's voice signal is determined based on the position of the second user. Based on the propagation direction, the feedforward microphone and the call microphone collect the sound signal.

2. The method according to claim 1, characterized in that, The step of determining that the first user wearing the earphone is talking to the second user based on the sixth sound signal includes: if the sixth sound signal contains the voice signal of the first user wearing the earphone, and the voice signal of the first user contains a first wake-up word, then it is determined that the first user wearing the earphone is talking to the second user.

3. The method according to claim 1 or 2, characterized in that, The feedforward microphone is located on the side of the earphone furthest from the ear, and the feedback microphone is located on the side of the earphone closest to the ear; the earphone is used to determine, based on the sound signal collected by the feedback microphone, that the first user is speaking to the second user, and the earphone is used to determine, based on the sound signals collected by the feedforward microphone and the call microphone, that the sound signal contains the voice signal of the second user.

4. The method according to claim 1 or 2, characterized in that, After switching the headphones from the first mode to the second mode, the method further includes: if no conversation between the first user and the second user is detected within a first duration, and the sixth sound signal does not contain the voice signal of the second user, then switching the headphones from the second mode back to the first mode.

5. The method according to claim 4, characterized in that, The method further includes: receiving setting information for the first duration sent by an electronic device.

6. The method according to claim 1 or 2, characterized in that, Before acquiring the sound signal collected by the microphone on the earphone, the method further includes: acquiring a wake-up word recognition model sent by an electronic device, wherein the wake-up word recognition model is trained by text including the second wake-up word and / or speech including the second wake-up word; correspondingly, after acquiring the sound signal collected by the microphone on the earphone, the method further includes: when it is determined that the sixth sound signal contains the speech signal of the second user, inputting the speech signal of the second user into the wake-up word recognition model to obtain a recognition result output by the wake-up word recognition model indicating whether the second wake-up word is included.

7. The method according to claim 6, characterized in that, Before obtaining the wake-up word recognition model sent by the electronic device, the method further includes: sending test speech collected by the microphone on the earphone to the electronic device, wherein the electronic device is used to train a classification model based on the test speech to obtain the wake-up word recognition model.

8. The method according to claim 1 or 2, characterized in that, Before acquiring the sound signal collected by the microphone on the headphones, the method further includes: receiving an instruction from an electronic device to activate the automatic headphone mode switching function.

9. The method according to claim 1 or 2, characterized in that, The method further includes: when the current operating mode of the headphones is the first mode, and the sixth sound signal determines that the first user wearing the headphones is talking to the second user, reducing the volume of the audio being played in the headphones, or instructing the headphones to stop playing audio.

10. An earphone, characterized in that, Includes a processor for executing a computer program stored in a memory to implement the headphone mode switching method as described in any one of claims 1 to 9.

11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the headphone mode switching method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Headset with hear-through mode, and operating method of the same

    CN106937194A

  • Earphone control method and device and earphone

    CN112770214A

  • Audio playing method, wireless earphone and computer readable storage medium

    CN113038337A

  • Volume control method and device and Bluetooth earphone

    CN114979896A