Human voice signal response method, control device and computer-readable storage medium
By obtaining the target sound source direction and location information of the target object of the human voice signal, identifying whether the human voice signal is a human voice signal emitted by the user, and determining the response strategy, the problem of misrecognition of speech recognition is solved and the user experience is improved.
Patent Information
- Application Number
- CN202110926557.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-12
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-08-12
AI Technical Summary
In speech recognition technology, the acoustic characteristics of the human voice signals actively sent by the user and the human voice signals emitted by electronic devices are similar, resulting in a high possibility of misrecognition and easily leading to operations that violate the user's intentions, affecting the user experience.
The human voice signal is collected through the sound pickup device, and the control device obtains the target sound source direction and the position information of the target object. The target sound source direction and the position information of the target object are used to identify whether the human voice signal is a human voice signal emitted by the user, and determine the response strategy based on the recognition result, including responding to the human voice signal or not, to reduce the possibility of misidentification.
It reduces the possibility of misrecognition in voice recognition, reduces operations that go against user intentions, improves user experience, and has strong practicality and ease of use.
Smart Images

Figure CN115706680B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition, and in particular to a human voice signal response method, a control device, and a computer-readable storage medium. Background Art
[0002] With the development of voice recognition technology, more and more smart home devices can be controlled by users through voice or realize voice interaction with users. However, the human voice signal for voice recognition may be the human voice signal emitted by the user, or it may also be the human voice signal emitted by other electronic devices with speakers.
[0003] Since the acoustic characteristics of the human voice signals actively sent by the user and the human voice signals sent by electronic devices are relatively similar, when the latter scenario occurs, the possibility of misrecognition of voice recognition is higher, and operations that go against the user's intentions are likely to occur, affecting the user's experience. Summary of the Invention
[0004] The embodiments of the present application provide a human voice signal response method, a control device, and a computer-readable storage medium, which solve the technical problems that the existing voice recognition technology has a high possibility of misrecognition and is prone to operations that violate the user's intentions.
[0005] To achieve the above objectives, the present invention adopts the following technical solutions:
[0006] In a first aspect, a method for responding to a human voice signal is provided, comprising:
[0007] When the sound pickup device collects a human voice signal, the control device obtains the target sound source direction of the human voice signal;
[0008] The control device acquires position information of a target object, where the target object includes a user and / or a speaker device, where the speaker device is an electronic device having a speaker;
[0009] The control device determines a response strategy for the human voice signal according to the target sound source direction and the position information of the target object.
[0010] It should be noted that the sound pickup device refers to an electronic device with a microphone, which can collect surrounding sound signals.
[0011] A speaker device refers to an electronic device with a speaker, which can emit sound signals through the speaker.
[0012] A control device is an electronic device that can control various smart home devices. The control device can be a smart home device, or it can also be a cloud server.
[0013] For example, the control device may be a smart home device such as a smart speaker or a smart screen; or, the control device may be a server installed in the cloud.
[0014] In the human voice signal response method provided in the embodiment of the present application, the sound pickup device can collect surrounding sound signals and detect whether there is a human voice signal therein; alternatively, the sound pickup device can collect surrounding sound signals and send the sound signals to the control device, which then detects whether there is a human voice signal in the above sound signals.
[0015] When the sound signal collected by the sound pickup device includes a human voice signal, the control device can obtain the target sound source direction of the human voice signal.
[0016] Furthermore, the control device may also obtain location information of a target object, where the target object includes a user and / or a speaker device.
[0017] By using the target sound source direction and the position information of the target object, the control device can identify to a certain extent whether the human voice signal is a human voice signal emitted by the user, and determine the corresponding response strategy accordingly.
[0018] In summary, in the human voice signal method provided in the embodiment of the present application, the control device will identify whether the human voice signal is a human voice signal emitted by the user based on the direction of the target sound source and the position information of the target object before responding to the human voice signal, and determine the corresponding response strategy. This can reduce the possibility of misidentification, reduce operations that violate the user's intentions, improve the user experience, and has strong practicality and ease of use.
[0019] In a possible implementation manner of the first aspect, the response strategy for the human voice signal includes responding to the human voice signal and not responding to the human voice signal.
[0020] It should be noted that the above-mentioned response strategy for the human voice signal may include responding to the human voice signal or not responding to the human voice signal.
[0021] That is, when the control device recognizes that the human voice signal is likely to be a human voice signal emitted by the user, the control device may respond to the human voice signal.
[0022] When the control device identifies that the human voice signal is likely to be a human voice signal emitted by the speaker device, in order to reduce the possibility of misidentification, the control device may not respond to the human voice signal, thereby reducing operations that violate the user's intention.
[0023] In a possible implementation of the first aspect, the human voice signal includes a command word; and the responding includes:
[0024] The control device executes the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
[0025] It should be noted that, when the above-mentioned human voice signal includes a command word, if the control device is provided with a function corresponding to the command word, the control device may execute the operation corresponding to the command word.
[0026] Alternatively, if the control device does not have the function corresponding to the command word, the control device sends a control instruction to the smart home device corresponding to the command word. After receiving the control instruction, the smart home device performs the operation corresponding to the command word.
[0027] For example, assuming the control device is a smart speaker, and the command word in the human voice signal is "play music", the smart speaker has the function corresponding to the command word and can perform the operation corresponding to "play music".
[0028] Assume the control device is a home hub, and the command word in the human voice signal is "turn on the air conditioner." In this case, the home hub does not have the function corresponding to the command word, so the home hub can send a control instruction to the smart air conditioner.
[0029] After receiving the control command, the smart air conditioner executes the operation corresponding to "turn on the air conditioner".
[0030] In a possible implementation of the first aspect, the control device determines, based on the target sound source direction and the position information of the target object, a response strategy for the human voice signal, including:
[0031] The control device determines, based on the target sound source direction and the position information of the target object, a degree of deviation between the position information of the target object and the target sound source direction;
[0032] The control device determines a response strategy for the human voice signal according to the degree of deviation.
[0033] It should be noted that, after obtaining the target sound source direction and the position information of the target object, the control device may calculate the degree of deviation between the position information of the target object and the target sound source direction.
[0034] By using the degree of deviation, the control device can identify whether the human voice signal is a human voice signal emitted by the user, and then determine a corresponding response strategy to reduce the possibility of misidentification.
[0035] In a possible implementation of the first aspect, the control device determines, according to the degree of deviation, a response strategy for the human voice signal, including:
[0036] When the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the user, the control device responds to the human voice signal;
[0037] When the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the speaker device, the control device does not respond to the human voice signal.
[0038] It should be noted that, after calculating the deviation degree, the control device can directly identify whether the human voice signal is a human voice signal emitted by the user according to the deviation degree.
[0039] When the control device determines that the human voice signal is a human voice signal emitted by the user according to the deviation degree, the control device should respond to the human voice signal.
[0040] When the control device determines that the human voice signal is a human voice signal emitted by the speaker device based on the above-mentioned deviation degree, the control device may not respond to the human voice signal, thereby reducing the possibility of misrecognition and reducing operations that violate the user's intention.
[0041] In a possible implementation of the first aspect, the target object includes a user and a speaker device; the deviation degree includes a first deviation degree between the position information of the speaker device and the direction of the target sound source, and a second deviation degree between the position information of the user and the direction of the target sound source;
[0042] The control device determines a response strategy for the human voice signal according to the degree of deviation, including:
[0043] The control device calculates a negative consistency coefficient according to the first deviation degree;
[0044] The control device calculates a positive consistency coefficient according to the second deviation degree;
[0045] The control device determines a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient.
[0046] It should be noted that, when the target object includes a user and a speaker device, the deviation degree may include a first deviation degree between the position information of the speaker device and the target sound source direction, and a second deviation degree between the position information of the user and the target sound source direction.
[0047] The control device can then calculate a negative consistency coefficient based on the first deviation degree. The negative consistency coefficient and the first deviation degree are negatively correlated. A greater first deviation degree results in a smaller negative consistency coefficient, indicating a greater deviation between the speaker device and the target sound source direction. A smaller first deviation degree results in a larger negative consistency coefficient, indicating a closer proximity between the speaker device and the target sound source direction.
[0048] The control device may also calculate a positive consistency coefficient based on the second deviation degree, wherein the positive consistency coefficient and the second deviation degree are negatively correlated. A greater second deviation degree results in a smaller positive consistency coefficient, indicating a greater deviation between the user and the target sound source direction. A smaller second deviation degree results in a larger positive consistency coefficient, indicating a closer proximity between the user and the target sound source direction.
[0049] After obtaining the positive consistency coefficient and the negative consistency coefficient, the control device can determine whether the above-mentioned human voice signal is a human voice signal emitted by the user based on the positive consistency coefficient and the negative consistency coefficient, and then determine the response strategy of the human voice signal to reduce the possibility of misrecognition.
[0050] In a possible implementation of the first aspect, the human voice signal includes a command word; and the control device determines a response strategy for the human voice signal based on the positive consistency coefficient and the negative consistency coefficient, including:
[0051] The control device obtains the vocabulary confidence and confidence threshold corresponding to the command word;
[0052] The control device adjusts the vocabulary confidence according to the positive consistency coefficient and the negative consistency coefficient, or adjusts the confidence threshold;
[0053] If the adjusted vocabulary confidence is greater than the confidence threshold, or the vocabulary confidence is greater than the adjusted confidence threshold, the control device performs the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
[0054] It should be noted that when the human voice signal includes a command word, the control device can obtain the vocabulary confidence and confidence threshold corresponding to the command word. The vocabulary confidence refers to the degree of credibility of the recognized command word.
[0055] Then, the control device may adjust the vocabulary confidence or the confidence threshold according to the positive consistency coefficient and the negative consistency coefficient.
[0056] When the control device adjusts the lexical confidence of the command word according to the positive consistency coefficient and the negative consistency coefficient, the positive consistency coefficient is positively correlated with the adjusted lexical confidence, and the negative consistency coefficient is negatively correlated with the adjusted lexical confidence.
[0057] That is to say, the larger the positive consistency coefficient is, the more likely it is that the above-mentioned human voice signal is issued by the user, and the control device can appropriately increase the vocabulary confidence; the smaller the positive consistency coefficient is, the less likely it is that the above-mentioned human voice signal is issued by the user, and the control device can appropriately lower the vocabulary confidence.
[0058] The larger the negative consistency coefficient is, the more likely it is that the above-mentioned human voice signal is emitted by the speaker device, and the control device can appropriately lower the vocabulary confidence; the smaller the negative consistency coefficient is, the less likely it is that the above-mentioned human voice signal is emitted by the speaker device, and the control device can appropriately increase the vocabulary confidence.
[0059] When the control device adjusts the confidence threshold of the command word according to the positive consistency coefficient and the negative consistency coefficient, the positive consistency coefficient is negatively correlated with the adjusted confidence threshold, and the negative consistency coefficient is positively correlated with the adjusted confidence threshold.
[0060] That is to say, the larger the positive consistency coefficient is, the smaller the adjusted confidence threshold is; the smaller the positive consistency coefficient is, the larger the adjusted confidence threshold is.
[0061] The larger the negative consistency coefficient, the larger the adjusted confidence threshold; the smaller the negative consistency coefficient, the smaller the adjusted confidence threshold.
[0062] After adjusting the lexical confidence or the confidence threshold of the command word, the control device may compare the lexical confidence of the command word with the confidence threshold.
[0063] When the vocabulary confidence is less than or equal to the confidence threshold, it indicates that the command word recognition fails, and the control device may not respond to the command word.
[0064] When the vocabulary confidence is greater than the confidence threshold, it means that the command word is recognized successfully, and the control device can execute the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word. After receiving the control instruction, the smart home device executes the operation corresponding to the command word.
[0065] In summary, in the human voice signal response method provided in the embodiment of the present application, the control device can adjust the vocabulary confidence or confidence threshold of the command word according to the positive consistency coefficient and the negative consistency coefficient.
[0066] Then, the control device performs command word recognition based on the adjusted vocabulary confidence / confidence threshold, which can improve the accuracy of command word recognition, reduce the possibility of misrecognition, and improve the user experience.
[0067] In a possible implementation of the first aspect, the control device determines, according to the positive consistency coefficient and the negative consistency coefficient, a response strategy for the human voice signal, including:
[0068] When the positive consistency coefficient and the negative consistency coefficient meet a preset correction rule, the control device updates the target sound source direction to the direction where the user is located;
[0069] The control device controls the sound pickup device to perform an audio directional enhancement operation on the updated target sound source direction, and the sound pickup device is an electronic device with a microphone.
[0070] It should be noted that after the control device obtains the positive consistency coefficient and the negative consistency coefficient, if the positive consistency coefficient and the negative consistency coefficient meet the preset correction rules, it indicates that the above-mentioned human voice signal is likely to be emitted by the user. In this case, the control device can update the target sound source direction to the direction of the user.
[0071] Afterwards, the control device can control the sound pickup device to perform an audio directional enhancement operation on the updated target sound source direction, thereby improving the sound pickup effect of the sound pickup device in the direction of the user.
[0072] In a possible implementation manner of the first aspect, the preset correction rule is that the positive consistency coefficient is greater than the negative consistency coefficient;
[0073] Alternatively, the preset correction rule is that the positive consistency coefficient is greater than a first deviation threshold, and the negative consistency coefficient is less than a second deviation threshold;
[0074] Alternatively, the preset correction rule is that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than a third deviation threshold.
[0075] It should be noted that the content of the preset correction rule can be set according to actual needs. For example, the preset correction rule can be that the positive consistency coefficient is greater than the negative consistency coefficient; or, the preset correction rule can be that the positive consistency coefficient is greater than the first deviation threshold, and the negative consistency coefficient is less than the second deviation threshold; or, the preset correction rule can be that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than the third deviation threshold; or, the preset correction rule can also be other content. The embodiments of the present application do not limit the specific content of the preset correction rule.
[0076] In a possible implementation of the first aspect, the control device determines, according to the positive consistency coefficient and the negative consistency coefficient, a response strategy for the human voice signal, including:
[0077] The control device adjusts a signal ratio corresponding to a target audio channel according to the positive consistency coefficient and the negative consistency coefficient, where the target audio channel is an audio channel corresponding to the target sound source direction.
[0078] It should be noted that the sound pickup device may be provided with multiple microphones, each corresponding to an audio channel. After the multiple microphones on the sound pickup device collect sound signals, the sound signals collected by each microphone are respectively converged through the corresponding audio channel to form a beam.
[0079] Therefore, after acquiring the target sound source direction, the control device may determine the audio channel of the sound pickup device in the target sound source direction as the target audio channel.
[0080] Then, the control device may adjust the signal ratio of the target audio channel signal in the beam according to the positive consistency coefficient and the negative consistency coefficient.
[0081] The above signal ratio is positively correlated with the positive consistency coefficient, and the above signal ratio is negatively correlated with the negative consistency coefficient.
[0082] That is to say, the larger the positive consistency coefficient, the larger the signal ratio; the smaller the positive consistency coefficient, the smaller the signal ratio. The larger the negative consistency coefficient, the smaller the signal ratio; the smaller the negative consistency coefficient, the larger the signal ratio.
[0083] When the human voice signal is likely to be a human voice signal emitted by the user, the control device increases the signal ratio of the target audio channel according to the positive consistency coefficient and the negative consistency coefficient, so that the sound pickup device can more clearly collect the audio signal in the direction of the target sound source.
[0084] When the human voice signal is more likely to be a human voice signal emitted by the speaker device, the control device lowers the signal ratio of the target audio channel through the positive consistency coefficient and the negative consistency coefficient, so that the sound pickup device can reduce interference in the direction of the target sound source.
[0085] In a possible implementation manner of the first aspect, the target object includes a user; and after the control device obtains the target sound source direction of the human voice signal, the method further includes:
[0086] When the control device fails to obtain the user's location information, the control device does not respond to the human voice signal.
[0087] It should be noted that when the control device does not obtain the user's location information, it means that the user is not in the current area and the above human voice signal is emitted by the speaker device. Therefore, the control device may not respond to the above human voice signal, reducing operations that violate the user's intentions.
[0088] In a second aspect, a human voice signal response device is provided, which is applied to a control device, comprising:
[0089] A sound source direction module is used to obtain the target sound source direction of the human voice signal when the sound pickup device detects and collects the human voice signal;
[0090] A location information module, configured to obtain location information of a target object, wherein the target object includes a user and / or a speaker device, wherein the speaker device is an electronic device having a speaker;
[0091] The response strategy module is used to determine the response strategy of the human voice signal according to the direction of the target sound source and the position information of the target object.
[0092] In a possible implementation manner of the second aspect, the response strategy for the human voice signal includes responding to the human voice signal and not responding to the human voice signal.
[0093] In a possible implementation of the second aspect, the human voice signal includes a command word; and the responding includes:
[0094] The control device executes the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
[0095] In a possible implementation of the second aspect, the response strategy module includes:
[0096] a deviation submodule, configured to determine a degree of deviation between the position information of the target object and the direction of the target sound source based on the direction of the target sound source and the position information of the target object;
[0097] The strategy submodule is used to determine a response strategy for the human voice signal according to the degree of deviation.
[0098] In a possible implementation of the second aspect, the policy submodule includes:
[0099] a response submodule, configured to respond to the human voice signal when the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the user;
[0100] The silent submodule is configured to not respond to the human voice signal when the control device determines, based on the deviation degree, that the human voice signal is a human voice signal emitted by the speaker device.
[0101] In a possible implementation of the second aspect, the target object includes a user and a speaker device; the deviation degree includes a first deviation degree between the position information of the speaker device and the direction of the target sound source, and a second deviation degree between the position information of the user and the direction of the target sound source;
[0102] The strategy submodule includes:
[0103] A negative coefficient submodule, configured to calculate a negative consistency coefficient based on the first deviation degree;
[0104] a forward coefficient submodule, configured to calculate a forward consistency coefficient according to the second deviation degree;
[0105] The comprehensive decision submodule is used to determine a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient.
[0106] In a possible implementation of the second aspect, the human voice signal includes a command word; and the comprehensive decision submodule includes:
[0107] A confidence submodule, configured to obtain a vocabulary confidence and a confidence threshold corresponding to the command word;
[0108] an adjusting submodule, configured to adjust the vocabulary confidence according to the positive consistency coefficient and the negative consistency coefficient, or to adjust the confidence threshold;
[0109] An operation submodule is configured to execute the operation corresponding to the command word, or send a control instruction to the smart home device corresponding to the command word, if the adjusted word confidence is greater than the confidence threshold, or the word confidence is greater than the adjusted confidence threshold.
[0110] In a possible implementation of the second aspect, the comprehensive decision-making submodule includes:
[0111] a direction updating submodule, configured to update the target sound source direction to the direction of the user when the positive consistency coefficient and the negative consistency coefficient meet a preset correction rule;
[0112] The directional sound pickup submodule is used to control the sound pickup device to perform audio directional enhancement operations on the updated target sound source direction, and the sound pickup device is an electronic device with a microphone.
[0113] In a possible implementation manner of the second aspect, the preset correction rule is that the positive consistency coefficient is greater than the negative consistency coefficient;
[0114] Alternatively, the preset correction rule is that the positive consistency coefficient is greater than a first deviation threshold, and the negative consistency coefficient is less than a second deviation threshold;
[0115] Alternatively, the preset correction rule is that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than a third deviation threshold.
[0116] In a possible implementation of the second aspect, the comprehensive decision-making submodule includes:
[0117] A ratio adjustment submodule is used to adjust the signal ratio corresponding to the target audio channel according to the positive consistency coefficient and the negative consistency coefficient, where the target audio channel is the audio channel corresponding to the target sound source direction.
[0118] In a possible implementation manner of the second aspect, the target object includes a user; and the apparatus further includes:
[0119] The rejection response module is used to not respond to the human voice signal when the control device fails to obtain the user's location information.
[0120] In a third aspect, an embodiment of the present application provides a control device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement a method as described in the first aspect and any one of the possible implementations of the first aspect when executing the computer program.
[0121] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is configured to store a computer program, and when the computer program is executed by a processor, it implements the method as described in the first aspect and any one of the possible implementation methods of the first aspect.
[0122] In a fifth aspect, an embodiment of the present application provides a computer program product, which is configured to, when running on a control device, enable the control device to execute the method described in the first aspect and any one of the possible implementations of the first aspect.
[0123] In a sixth aspect, an embodiment of the present application provides a chip system, comprising a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory to implement a method as described in the first aspect and any one of the possible implementation methods of the first aspect.
[0124] The beneficial effects of the embodiments of the present application are:
[0125] In the vocal signal response method of an embodiment of the present application, when a sound pickup device collects a vocal signal, a control device obtains the target sound source direction of the vocal signal and the location information of a target object, which may include a user and / or a speaker device. The control device can then identify whether the vocal signal originated from a user based on the target sound source direction and the location information of the target object, and thereby determine a response strategy for the vocal signal.
[0126] The human voice signal response method of the embodiment of the present application determines the response strategy of the human voice signal based on the target sound source direction and the position information of the target object before responding to the human voice signal. This can reduce the possibility of misidentification, reduce operations that violate the user's intentions, improve the user experience, and has strong practicality and ease of use. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] Figure 1 A system architecture diagram of a smart home system provided in an embodiment of the present application;
[0128] Figure 2 A schematic diagram of an electronic device provided in an embodiment of the present application;
[0129] Figure 3 A schematic diagram of a scenario provided in an embodiment of the present application;
[0130] Figure 4 Another scenario diagram provided for an embodiment of the present application;
[0131] Figure 5 Another scenario diagram provided for an embodiment of the present application;
[0132] Figure 6 Another scenario diagram provided for an embodiment of the present application;
[0133] Figure 7 Another scenario diagram provided for an embodiment of the present application;
[0134] Figure 8 Another scenario diagram provided for an embodiment of the present application;
[0135] Figure 9 Another scenario diagram provided for an embodiment of the present application;
[0136] Figure 10 Another scenario diagram provided for an embodiment of the present application;
[0137] Figure 11 Another scenario diagram provided for an embodiment of the present application;
[0138] Figure 12Another scenario diagram provided for an embodiment of the present application;
[0139] Figure 13 Another scenario diagram provided for an embodiment of the present application;
[0140] Figure 14 Another scenario diagram provided for an embodiment of the present application;
[0141] Figure 15 Another scenario diagram provided for an embodiment of the present application;
[0142] Figure 16 Another scenario diagram provided for an embodiment of the present application;
[0143] Figure 17 A schematic diagram of a microphone array provided in an embodiment of the present application;
[0144] Figure 18 A flowchart of a method for responding to a human voice signal provided in an embodiment of the present application;
[0145] Figure 19 A schematic structural diagram of a human voice signal response device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0146] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0147] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0148] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0149] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0150] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0151] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0152] The term “plurality” used in the present specification may be understood as “more than two” or “at least two”.
[0153] With the development of voice recognition technology, more and more smart home devices can be controlled by users through voice or realize voice interaction with users. For example, users can wake up smart home devices with wake-up commands such as "Xiaoyi, Xiaoyi" and control smart home devices with control commands such as "Play music" and "Stop playing" to perform corresponding operations.
[0154] However, when performing speech recognition, the voice signal for speech recognition may be a voice signal emitted by a user, or may be a voice signal emitted by other electronic devices with speakers.
[0155] For example, when a smart speaker detects a human voice signal "turn on music", the human voice signal may be a command given by the user to the smart speaker, or it may be a human voice signal emitted by a TV when playing a program.
[0156] Since the acoustic characteristics of the human voice signals actively sent by the user and the human voice signals sent by electronic devices are relatively similar, when the latter scenario occurs, the possibility of misrecognition of voice recognition is higher, and operations that go against the user's intentions are likely to occur, affecting the user's experience.
[0157] In view of this, an embodiment of the present application provides a human voice signal response method, which can determine a response strategy based on the location information of the user and / or the speaker device when a human voice signal is detected, thereby reducing the probability of misrecognition of speech recognition to a certain extent, or the probability of miscontrol and mistriggering, improving the user experience, and having strong ease of use and practicality.
[0158] First, see Figure 1 . Figure 1 A smart home system to which the embodiments of the present application are applicable is exemplarily shown.
[0159] like Figure 1 As shown, the smart home system may include: one or more smart home devices 101 ( Figure 1 3 are shown), and one or more control devices 102 ( Figure 1 1 is shown).
[0160] The above-mentioned smart home devices 101 may include one or more electronic devices such as computers, mobile phones, tablets, smart TVs, smart large screens, smart speakers, smart air conditioners, sweeping robots, dishwashers, smart lamps, smart door locks, smart curtains, routers, home hubs, etc.
[0161] Among them, the home hub can refer to a home central control electronic device dedicated to controlling smart home devices throughout the house, or the home hub can also refer to a software that can be installed in the user's existing smart home devices (such as routers, smart large screens, etc.), so that the existing smart home devices can realize the control and intelligent decision-making of smart home devices throughout the house based on their own existing functions.
[0162] Furthermore, the smart home devices 101 may be located in the same local area network (LAN), and the smart home devices 101 may exchange data via the LAN. Alternatively, the smart home devices 101 may exchange data via one or more of the following connection methods: Bluetooth, Wi-Fi, or a universal serial bus (USB).
[0163] Among the above-mentioned smart home devices 101, some smart home devices 101 (such as cameras, millimeter wave radars, laser radars, ultrasonic radars, pyroelectric infrared sensors, etc.) can scan and / or detect the environment. In the following description, these electronic devices with environmental detection functions are defined as detection devices.
[0164] Among the above-mentioned smart home devices 101, some smart home devices 101 (such as smart TVs, smart speakers, etc.) have speakers and can play audio signals externally. In the following description, these electronic devices with speakers are defined as speaker devices.
[0165] Among the above-mentioned smart home devices 101, some smart home devices 101 (such as smart speakers, tablet computers, etc.) are equipped with microphones and can collect external audio signals. In the following description, these electronic devices equipped with microphones are defined as sound pickup devices.
[0166] The control device 102 refers to an electronic device that can control the smart home device 101. The control device 102 can be the smart home device 101, and / or the control device 102 can also be a remote server.
[0167] For example, in some scenarios, users can set the smart TV as the control device 102 and control other smart home devices 101 through the smart TV; in other scenarios, users can set the smart speaker as the control device 102 and control other smart home devices 101 through the smart speaker; in other scenarios, users can set the home hub as the control device 102 and control other smart home devices 101 through the home hub; in other scenarios, users can set the remote server as the control device 102 and control the smart home devices 101 in the local area network through the remote server and the wide area network.
[0168] refer to Figure 2 , Figure 2 A structural diagram of an electronic device 200 provided in an embodiment of the present application is exemplarily shown. The electronic device 200 may be the above-mentioned smart home device 101, and / or, the electronic device 200 may also be the above-mentioned control device 102.
[0169] The electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, an earphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, an air pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.
[0170] The processor 210 may include one or more processing units, for example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0171] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.
[0172] Processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 210 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 210. If processor 210 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 210 latency, and thus improves system efficiency.
[0173] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0174] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 210 may include multiple I2C bus lines. The processor 210 may be coupled to the touch sensor 280K, charger, flash, camera 293, etc. via different I2C bus interfaces. For example, the processor 210 may be coupled to the touch sensor 280K via the I2C interface, enabling communication between the processor 210 and the touch sensor 280K via the I2C bus interface, thereby implementing the touch function of the electronic device 200.
[0175] The I2S interface can be used for audio communication. In some embodiments, the processor 210 can include multiple I2S buses. The processor 210 can be coupled to the audio module 270 via the I2S bus to enable communication between the processor 210 and the audio module 270. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.
[0176] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 270 and the wireless communication module 260 can be coupled via a PCM bus interface. In some embodiments, the audio module 270 can also transmit audio signals to the wireless communication module 260 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0177] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 210 and the wireless communication module 260. For example, the processor 210 communicates with the Bluetooth module in the wireless communication module 260 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the UART interface, enabling the function of playing music through Bluetooth headphones.
[0178] The MIPI interface can be used to connect the processor 210 to peripheral devices such as the display screen 294 and the camera 293. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 210 and the camera 293 communicate via the CSI interface to implement the camera function of the electronic device 200. The processor 210 and the display screen 294 communicate via the DSI interface to implement the display function of the electronic device 200.
[0179] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 210 to the camera 293, the display 294, the wireless communication module 260, the audio module 270, the sensor module 280, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0180] USB interface 230 is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. USB interface 230 can be used to connect a charger to charge electronic device 200, or to transfer data between electronic device 200 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices.
[0181] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0182] The charging management module 240 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 240 can receive charging input from the wired charger via the USB interface 230. In some wireless charging embodiments, the charging management module 240 can receive wireless charging input via the wireless charging coil of the electronic device 200. While charging the battery 242, the charging management module 240 can also provide power to the electronic device via the power management module 241.
[0183] The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240 and provides power to the processor 210, the internal memory 221, the display 294, the camera 293, and the wireless communication module 260. The power management module 241 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 241 can also be set in the processor 210. In other embodiments, the power management module 241 and the charging management module 240 can also be set in the same device.
[0184] The wireless communication function of the electronic device 200 can be implemented through the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modem processor and the baseband processor.
[0185] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0186] The mobile communication module 250 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 200. The mobile communication module 250 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 250 can be set in the processor 210. In some embodiments, at least some of the functional modules of the mobile communication module 250 can be set in the same device as at least some of the modules of the processor 210.
[0187] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 270A, the receiver 270B, etc.) or displays an image or video through the display screen 294. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 210 and be set in the same device as the mobile communication module 250 or other functional modules.
[0188] The wireless communication module 260 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 200. The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 260 can also receive the signal to be sent from the processor 210, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0189] In some embodiments, antenna 1 of electronic device 200 is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, so that electronic device 200 can communicate with a network and other devices via wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0190] Electronic device 200 implements display functionality through a GPU, display screen 294, and an application processor. A GPU is a microprocessor for image processing that connects display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 210 may include one or more GPUs that execute program instructions to generate or modify display information.
[0191] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 200 may include one or N display screens 294, where N is a positive integer greater than one.
[0192] The electronic device 200 can implement a shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, and an application processor.
[0193] The ISP processes data fed back by camera 293. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 293.
[0194] The camera 293 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 200 may include 1 or N cameras 293, where N is a positive integer greater than 1.
[0195] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 200 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0196] Video codecs are used to compress or decompress digital video. Electronic device 200 may support one or more video codecs. This allows electronic device 200 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0197] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 200, such as image recognition, face recognition, speech recognition, and text comprehension.
[0198] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 via the external memory interface 220 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0199] The internal memory 221 can be used to store computer executable program codes, which include instructions. The internal memory 221 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 200 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 221 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device 200 by running instructions stored in the internal memory 221 and / or instructions stored in a memory provided in the processor.
[0200] The electronic device 200 can implement audio functions such as music playback and recording through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headphone jack 270D, and the application processor.
[0201] The audio module 270 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 270 can also be used to encode and decode audio signals. In some embodiments, the audio module 270 can be provided in the processor 210, or some functional modules of the audio module 270 can be provided in the processor 210.
[0202] The speaker 270A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 200 can listen to music or listen to hands-free calls through the speaker 270A.
[0203] The receiver 270B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 200 receives a call or a voice message, the user can place the receiver 270B close to the ear to hear the voice.
[0204] The microphone 270C, also known as a "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 270C to input the sound signal into the microphone 270C. The electronic device 200 can be provided with at least one microphone 270C. In other embodiments, the electronic device 200 can be provided with two microphones 270C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 200 can also be provided with three, four or more microphones 270C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.
[0205] The headphone jack 270D is used to connect a wired headphone and can be a USB interface 230 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0206] Pressure sensor 280A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 280A can be located on display screen 294. There are many types of pressure sensors 280A, such as resistive, inductive, and capacitive. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When force is applied to pressure sensor 280A, the capacitance between the electrodes changes. Electronic device 200 determines the intensity of the pressure based on this change in capacitance. When a touch operation is applied to display screen 294, electronic device 200 detects the touch intensity based on pressure sensor 280A. Electronic device 200 can also calculate the touch location based on the detection signal from pressure sensor 280A. In some embodiments, touch operations applied to the same touch location but with different touch intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, a command to view short messages is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to a short message application icon, a command to create a new short message is executed.
[0207] The gyroscope sensor 280B can be used to determine the motion posture of the electronic device 200. In some embodiments, the angular velocity of the electronic device 200 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 280B. The gyroscope sensor 280B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 280B detects the angle of the electronic device 200 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 200 through reverse movement to achieve anti-shake. The gyroscope sensor 280B can also be used for navigation and somatosensory game scenes.
[0208] The air pressure sensor 280C is used to measure air pressure. In some embodiments, the electronic device 200 calculates the altitude using the air pressure value measured by the air pressure sensor 280C to assist in positioning and navigation.
[0209] The magnetic sensor 280D includes a Hall sensor. The electronic device 200 can use the magnetic sensor 280D to detect the opening and closing of the flip case. In some embodiments, when the electronic device 200 is a flip phone, the electronic device 200 can detect the opening and closing of the flip cover based on the magnetic sensor 280D. Based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.
[0210] Accelerometer 280E can detect the magnitude of acceleration of electronic device 200 in all directions (generally three axes). When electronic device 200 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.
[0211] The distance sensor 280F is used to measure distance. The electronic device 200 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the electronic device 200 can use the distance sensor 280F to measure distance to achieve fast focusing.
[0212] The proximity light sensor 280G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device 200 emits infrared light outward through the light emitting diode. The electronic device 200 uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 200. When insufficient reflected light is detected, the electronic device 200 can determine that there is no object near the electronic device 200. The electronic device 200 can use the proximity light sensor 280G to detect that the user is holding the electronic device 200 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 280G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.
[0213] Ambient light sensor 280L is used to sense ambient light brightness. Electronic device 200 can adaptively adjust the brightness of display screen 294 based on the perceived ambient light brightness. Ambient light sensor 280L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 280L can also work with proximity light sensor 280G to detect whether electronic device 200 is in a pocket to prevent accidental touches.
[0214] The fingerprint sensor 280H is used to collect fingerprints. The electronic device 200 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering calls, etc.
[0215] The temperature sensor 280J is used to detect temperature. In some embodiments, the electronic device 200 uses the temperature detected by the temperature sensor 280J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 280J exceeds a threshold, the electronic device 200 reduces the performance of the processor located near the temperature sensor 280J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 200 heats the battery 242 to prevent the electronic device 200 from shutting down abnormally due to low temperature. In other embodiments, when the temperature is lower than another threshold, the electronic device 200 boosts the output voltage of the battery 242 to prevent abnormal shutdown due to low temperature.
[0216] The touch sensor 280K is also referred to as a "touch-sensitive device." The touch sensor 280K can be disposed on the display screen 294. The touch sensor 280K and the display screen 294 form a touch screen, also referred to as a "touch screen." The touch sensor 280K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 294. In other embodiments, the touch sensor 280K can also be disposed on the surface of the electronic device 200, at a location different from that of the display screen 294.
[0217] The bone conduction sensor 280M can acquire vibration signals. In some embodiments, the bone conduction sensor 280M can acquire vibration signals from the vibrating bones of the human body's vocal cords. The bone conduction sensor 280M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 280M can also be set in headphones to form bone conduction headphones. The audio module 270 can parse the voice signal based on the vibration signal of the vibrating bones of the vocal cords acquired by the bone conduction sensor 280M to implement voice functions. The application processor can parse heart rate information based on the blood pressure signals acquired by the bone conduction sensor 280M to implement heart rate detection functions.
[0218] The buttons 290 include a power button, a volume button, and the like. The buttons 290 may be mechanical buttons or touch buttons. The electronic device 200 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 200.
[0219] Motor 291 can generate vibration prompts. Motor 291 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 294, motor 291 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0220] The indicator 292 may be an indicator light, which may be used to indicate the charging status, power level change, messages, missed calls, notifications, etc.
[0221] The SIM card interface 295 is used to connect a SIM card. The SIM card can be connected to and disconnected from the electronic device 200 by inserting it into or removing it from the SIM card interface 295. The electronic device 200 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 295 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 295 can also be compatible with different types of SIM cards. The SIM card interface 295 can also be compatible with external memory cards. The electronic device 200 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the electronic device 200 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 200 and cannot be separated from the electronic device 200.
[0222] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0223] The following will be based on Figure 1 The smart home system shown, and Figure 2 The structure of the electronic device in the smart home system shown is combined with specific application scenarios to describe in detail the human voice signal response method provided in the embodiment of the present application.
[0224] 1. Initialize the device location.
[0225] During the initialization phase, the control device can obtain the location information of each smart home device in an appropriate manner.
[0226] In some possible implementations, the control device may obtain the location information of each smart home device in response to a user's configuration operation.
[0227] For example, assume that the control device is a mobile phone and the smart home device is a smart speaker.
[0228] After the smart speaker is connected to the local area network, the user can open the "Smart Life" application on the mobile phone and set the room where the smart speaker is located in "Smart Life".
[0229] exist Figure 3In scenario (a), user interface 301 may be the category management interface for the "Smart Life" application. User interface 301 includes room controls 3011, 3012, and 3013. Room control 3011 may correspond to "personal devices," room control 3012 may correspond to "living room," and room control 3013 may correspond to "kitchen."
[0230] The user can click on the room control 3012 in the user interface 301 to enter the user interface 302 .
[0231] exist Figure 3 In scenario (b), user interface 302 may be the settings page for the "living room." User interface 302 includes device controls 3021, 3022, operation controls 3023, and 3024. Device control 3021 represents "smart TV," device control 3022 represents "robot vacuum," operation control 3023 represents "add device," and operation control 3024 represents "delete device."
[0232] The user can click on control 3023 in user interface 302. A selection box 3025 will pop up in user interface 302. Selection box 3025 includes three options: "Smart Speaker," "Desktop Computer," and "Tablet." The user can click on "Smart Speaker" to set the smart speaker's location to "Living Room." The phone, responding to the user's configuration, determines that the smart speaker is located in the living room.
[0233] In some other possible implementations, the control device may detect the environment through a detection device to obtain the location information of each smart home device.
[0234] As mentioned above, smart home devices may include detection equipment, which may include one or more types of electronic devices such as cameras, millimeter wave radars, lidars, ultrasonic radars, and pyroelectric infrared sensors.
[0235] During the initialization phase, the control device may send an environmental detection instruction to the detection device.
[0236] When the detection device receives the environment detection instruction, the detection device can perform the environment detection operation to obtain the surrounding environment data.
[0237] Afterwards, the detection device can process the environmental data locally, identify the device type and location information of the smart home devices contained in the environmental data, and send the device type and location information of each detected smart home device to the control device.
[0238] Alternatively, the detection device may also directly send the environmental data to the control device. When the control device receives the environmental data, the control device processes the environmental data and identifies the device type and location information of the smart home device contained in the environmental data.
[0239] In addition, when the smart home system includes multiple detection devices, the control device can control one of the detection devices to perform an environment detection operation, or the control device can also control multiple detection devices to perform an environment detection operation.
[0240] When the control device controls multiple detection devices to perform an environment detection operation, the detection results corresponding to the multiple detection devices may include multiple different location information of the same smart home device.
[0241] At this time, the control device may select the location information with the highest location confidence from multiple different location information of the same smart home device as the location information of the smart home device.
[0242] The location confidence level is used to indicate the reliability of the location information. When the detection device / control device processes the location information of the smart home device, it can also obtain the location confidence level corresponding to the location information.
[0243] For example, Figure 4 As shown, it is assumed that a home hub 11, a smart speaker 12, a camera 13 and a millimeter wave radar 14 are set in the room. Among them, the home hub 11 is a control device, and the camera 13 and the millimeter wave radar 14 are detection devices.
[0244] When the home hub 11 enters the initialization phase in response to the user's operation, the home hub 11 sends an environment detection instruction to the camera 13 and the millimeter-wave radar 14 .
[0245] When the millimeter-wave radar 14 receives the environment detection instruction, the millimeter-wave radar 14 transmits a millimeter-wave detection signal and receives a millimeter-wave reflection signal.
[0246] like Figure 5 As shown, the millimeter-wave radar 14 processes the millimeter-wave reflected signal and identifies the smart speaker 12 and camera 13. The millimeter-wave radar 14 then sends the detection results to the home hub 11. The detection results include the device type, location information, and location confidence level of the smart speaker 12 and camera 13.
[0247] When the camera 13 receives the environment detection instruction, it captures the environment image and then sends the environment image to the home hub 11.
[0248] After receiving the environmental image, the home hub 11 performs image processing on the environmental image to identify the smart speaker 12, the location information of the smart speaker 12, and the location confidence corresponding to the location information.
[0249] At this point, both the camera 13 and the millimeter-wave radar 14 detect the location information of the smart speaker 12. The millimeter-wave radar 14 detects that the smart speaker 12 is at the first location, and the camera 13 detects that the smart speaker 12 is at the second location. Therefore, the home hub 11 can obtain the location confidence level corresponding to the first location and the location confidence level corresponding to the second location.
[0250] Assuming that the position confidence corresponding to the first position is 0.84, the position confidence corresponding to the second position is 0.59, and the position confidence of the first position is higher than the position confidence of the second position, the home hub 11 uses the first position detected by the millimeter wave radar 14 as the position information of the smart speaker 12.
[0251] In some other possible implementations, the control device may also obtain the location information of each smart home device through other methods. The embodiment of the present application does not limit the method in which the control device obtains the location information.
[0252] After obtaining the location information of each smart home device, the control device can record the location information of each smart home device in the same physical coordinate system.
[0253] The physical coordinate system may be any one of a two-dimensional rectangular coordinate system, an oblique coordinate system, a polar coordinate system, a three-dimensional rectangular coordinate system, a spherical coordinate system, and the like.
[0254] Furthermore, the origin of the physical coordinate system may be the location of the control device, or the location of any smart home device, or any other location. The embodiment of the present application does not limit the manner in which the origin of the physical coordinate system is set.
[0255] For example, Figure 6 As shown, assuming the control device receives the detection results sent by the millimeter-wave radar, the control device can use the location of the millimeter-wave radar as the origin to establish a two-dimensional rectangular coordinate system. Furthermore, the control device can also display the location of each smart home device in the two-dimensional rectangular coordinate system based on the location information of each smart home device detected by the millimeter-wave radar.
[0256] for example, Figure 6 Point O in the figure represents the origin, which is the location of the millimeter-wave radar. Point A1 represents the location of the smart speaker, point B1 represents the location of the smart TV, and point C1 represents the location of the camera.
[0257] During the initialization phase, the control device may establish a sound source coordinate system in addition to the physical coordinate system. The control device may then establish a mapping relationship between the physical coordinate system and the sound source coordinate system.
[0258] Specifically, the control device may establish a sound source coordinate system. The sound source coordinate system may be any of a two-dimensional rectangular coordinate system, an oblique coordinate system, a polar coordinate system, a three-dimensional rectangular coordinate system, a spherical coordinate system, etc. The origin of the sound source coordinate system may be the position of the sound pickup device.
[0259] Furthermore, the control device may send a broadcast instruction to the speaker device and a sound pickup instruction to the sound pickup device.
[0260] When the speaker device receives the broadcast instruction, the speaker device may play a specific test audio.
[0261] The content of the test audio can be configured based on actual circumstances. For example, in some embodiments, the test audio can be a voice announcement; in other embodiments, the test audio can be a segment of white noise; in still other embodiments, the test audio can be a specific piece of music. The present application does not limit the content of the test audio.
[0262] When the sound pickup device receives the sound pickup instruction, it can monitor the sound in the environment. When the sound pickup device monitors a specific test audio, it obtains the sound source direction of the test audio and sends the sound source direction of the test audio to the control device.
[0263] After the control device obtains the sound source direction of the test audio, the control device can determine the mapping direction of the 0-degree direction of the pickup device (which can also be understood as the 0-degree direction of the sound source coordinate system) in the physical coordinate system based on the sound source direction of the test audio, and establish the mapping relationship between the above-mentioned physical coordinate system and the sound source coordinate system based on the above-mentioned mapping direction to complete the initialization stage.
[0264] For example, Figure 7 As shown, it is assumed that the control device is a home hub 21, the speaker device is a smart TV 22, and the sound pickup device is a smart speaker 23.
[0265] When the home hub 21 enters the initialization phase, the home hub 21 sends a broadcast instruction to the smart TV 22 , and the home hub 21 sends a sound pickup instruction to the smart speaker 23 .
[0266] When the smart TV 22 receives the broadcast instruction, the smart TV 22 plays the audio of "Xiaoyi, Xiaoyi" through the speaker.
[0267] When the smart speaker 23 receives the sound pickup instruction, the smart speaker 23 monitors the sound in the environment. When the smart speaker 23 monitors the audio of "Xiaoyi, Xiaoyi", the smart speaker 23 obtains the sound source direction of the audio and sends the sound source direction of the audio to the home hub 21.
[0268] After receiving the sound source direction of the above audio, the home hub 21 establishes a polar coordinate system with the smart speaker 23 as the origin, and establishes a mapping relationship between the physical coordinate system and the sound source coordinate system according to the sound source direction of the above audio.
[0269] See also Figure 8 and Figure 9 , Figure 8 Where O1 represents the origin of the polar coordinate system, O2 represents the origin of the physical coordinate system, and the X-axis and Y-axis are the coordinate axes of the physical coordinate system. Figure 8 The polar coordinate system and physical coordinate system in are extracted separately, and we can get Figure 9 (a) The polar coordinate system of the scene representation, and Figure 9 (b) Physical coordinate system of scene representation. Figure 9 Midpoint A2 represents the position of the smart speaker 23 , and point B2 represents the position of the smart TV 22 .
[0270] like Figure 9 (a) scenario and Figure 9 As shown in scenario (b), assuming that the smart speaker 23 detects the sound source direction as 53 degrees, the control device can rotate the line A2-B2 in the physical coordinate system 53 degrees clockwise along point A2 to find dotted line 1. The direction pointed by the arrow on dotted line 1 is the mapping direction of the 0-degree direction of the smart speaker 23 in the physical coordinate system.
[0271] Afterwards, the control device may determine the mapping relationship between the physical coordinate system and the sound source coordinate system according to the dotted line 1 .
[0272] See also Figure 10 After establishing the mapping relationship between the physical coordinate system and the sound source coordinate system, assuming that the smart speaker 23 detects that the sound source direction of the audio signal to be measured is 79 degrees, the control device can rotate the dotted line 1 in the physical coordinate system counterclockwise along point A2 by 79 degrees according to the above mapping relationship to obtain the dotted line 2.
[0273] At this time, the direction of the arrow of the dotted line 2 is the sound source direction of the audio signal to be measured in the physical coordinate system.
[0274] 2. Analysis of sound source properties.
[0275] After completing the initialization phase, the sound pickup device can collect sound signals from the surrounding environment.
[0276] When the sound pickup device collects a sound signal, the sound pickup device can locally detect whether there is a human voice signal in the sound signal, and send the target sound source direction of the human voice signal to the control device when a human voice signal is detected.
[0277] Alternatively, when the sound pickup device collects a sound signal, the sound pickup device may directly send the sound signal to the control device. After receiving the sound signal, the control device detects whether there is a human voice signal in the sound signal and obtains the target sound source direction of the human voice signal when a human voice signal is detected.
[0278] After obtaining the target sound source direction, the control device can determine a response strategy for the human voice signal according to the position information of each speaker device and the current position information of the user.
[0279] Specifically, after acquiring the target sound source direction, the control device may map the target sound source direction from the sound source coordinate system to the physical coordinate system, or may also map the position information of each speaker device from the physical coordinate system to the sound source coordinate system.
[0280] Afterwards, the control device may calculate a first deviation degree between the position information of each speaker device and the target sound source direction, and calculate a negative consistency coefficient according to the first deviation degree.
[0281] The negative consistency coefficient is negatively correlated with the first deviation degree. That is, the smaller the deviation between the speaker device's position information and the target sound source direction, the greater the negative consistency coefficient; and the greater the deviation between the speaker device's position information and the target sound source direction, the smaller the negative consistency coefficient.
[0282] When there are multiple speaker devices, the control device may obtain multiple negative consistency coefficients and may use the largest negative consistency coefficient.
[0283] Furthermore, the control device may also send a user positioning instruction to the detection device.
[0284] When the detection device receives a user positioning instruction, the detection device may perform a person detection operation to detect whether there is a user in the current area.
[0285] For example, assuming the detection device is a camera, the camera can capture an image of the environment upon receiving a user positioning instruction. The camera can then process the image of the environment to identify whether the user is present in the image of the environment and the location of the user.
[0286] Assuming the detection device is a millimeter-wave radar, upon receiving a user positioning command, the millimeter-wave radar transmits a millimeter-wave detection signal and receives a millimeter-wave reflected signal. The millimeter-wave radar then processes the millimeter-wave reflected signal to identify whether a user is in the current area and the user's location.
[0287] If the detection device does not detect the user in the current area, the detection device may send a positioning error instruction to the control device. The positioning error instruction is used to inform the control device that the user is not found in the current area.
[0288] If the detection device detects a user in the current area, the detection device may send the user's location information to the control device.
[0289] If all the data returned to the control device by the detection devices indicates a positioning error, it indicates that the user is not in the current area. In other words, the human voice signal detected by the sound pickup device is coming from the speaker device, not the user. In this case, the control device may not respond to the human voice signal, or may perform other pre-set operations.
[0290] For example, in some embodiments, when the control device receives a positioning error instruction, the control device may remain silent and wait for the next detection of a human voice signal; in other embodiments, when the control device receives a positioning error message, the control device may control the speaker device to emit a specific prompt sound, which is used to inform or remind the user to perform a second confirmation.
[0291] When the detection device receives the user's position information, the detection device may calculate a second deviation degree between the user's position information and the target sound source direction, and calculate a positive consistency coefficient based on the second deviation degree.
[0292] The positive consistency coefficient is negatively correlated with the second deviation degree. That is, the smaller the deviation between the user's location information and the target sound source direction, the greater the positive consistency coefficient; and the greater the deviation between the user's location information and the target sound source direction, the smaller the positive consistency coefficient.
[0293] The specific calculation method of the negative consistency coefficient and the positive consistency coefficient can be set according to the actual scenario.
[0294] In some possible implementations, the control device may represent the degree of deviation by a deviation angle between the position information of the speaker device / user and the direction of the target sound source.
[0295] At this time, the control device can determine the first line segment according to the position information of the sound pickup device and the position information of the speaker device, calculate the angle between the first line segment and the target sound source direction, and obtain a first deviation angle value between the speaker device and the target sound source direction.
[0296] The smaller the first deviation angle value is, the higher the degree of overlap between the loudspeaker device and the target sound source direction is, the more likely the human voice signal is emitted by the loudspeaker device, and the larger the negative consistency coefficient is.
[0297] The larger the first deviation angle value is, the greater the deviation between the loudspeaker device and the target sound source direction is, the less likely the human voice signal is emitted by the loudspeaker device, and the smaller the negative consistency coefficient is.
[0298] Furthermore, the control device may determine a second line segment based on the position information of the sound pickup device and the user, calculate an angle between the second line segment and the target sound source direction, and obtain a second deviation angle value between the user and the target sound source direction.
[0299] The smaller the second deviation angle value is, the higher the degree of overlap between the user and the target sound source direction is, the more likely the above-mentioned human voice signal is emitted by the user, and the larger the positive consistency coefficient is.
[0300] The larger the second deviation angle value is, the greater the deviation between the user and the target sound source direction is, the less likely the above-mentioned human voice signal is to be emitted by the user, and the smaller the positive consistency coefficient is.
[0301] For example, Figure 11 As shown, it is assumed that point O represents the origin of the physical coordinate system, point A3 represents the position of the sound pickup device, point B3 represents the position of the speaker device, point C3 represents the position of the user, and dotted line 3 represents the direction of the target sound source.
[0302] When calculating the forward consistency coefficient, the control device may draw a line segment A3-C3 and calculate the size of the angle 1 between the line segment A3-C3 and the dotted line 3. The larger the angle 1, the smaller the forward consistency coefficient; the smaller the angle 1, the larger the forward consistency coefficient.
[0303] When calculating the negative consistency coefficient, the control device can draw a line segment A3-B3 and calculate the size of the angle 2 between the line segment A3-B3 and the dotted line 3. The larger the angle 2, the smaller the negative consistency coefficient; the smaller the angle 2, the larger the negative consistency coefficient.
[0304] In some other possible implementations, the control device may use the distance from the position information of the speaker device / user to the direction of the target sound source to represent the above-mentioned degree of deviation.
[0305] At this time, the control device may draw a first perpendicular line in the direction of the target sound source according to the position information of the speaker device, and calculate the length of the first perpendicular line.
[0306] The shorter the length of the first vertical line is, the higher the degree of overlap between the loudspeaker device and the target sound source direction is, the more likely the human voice signal is emitted by the loudspeaker device, and the larger the negative consistency coefficient is.
[0307] The longer the length of the first vertical line is, the greater the deviation between the loudspeaker device and the target sound source direction is, the less likely the human voice signal is emitted by the loudspeaker device, and the smaller the negative consistency coefficient is.
[0308] Furthermore, the control device may also draw a second perpendicular line in the direction of the target sound source according to the user's position information, and calculate the length of the second perpendicular line.
[0309] The shorter the length of the second vertical line is, the higher the degree of overlap between the user and the target sound source direction is, the more likely the above-mentioned human voice signal is emitted by the user, and the larger the positive consistency coefficient is.
[0310] The longer the second vertical line is, the greater the deviation between the user and the target sound source direction is, the less likely the human voice signal is to be emitted by the user, and the smaller the positive consistency coefficient is.
[0311] For example, Figure 12 As shown, it is assumed that point O represents the origin of the coordinate system, point A4 represents the position of the sound pickup device, point B4 represents the position of the speaker device, point C4 represents the position of the user, and dotted line 4 represents the direction of the target sound source.
[0312] When calculating the forward consistency coefficient, the control device can use point C4 as the starting point and draw a perpendicular line 1 to the dotted line 4. The longer the perpendicular line 1 is, the smaller the forward consistency coefficient is; the shorter the perpendicular line 1 is, the larger the forward consistency coefficient is.
[0313] When calculating the negative consistency coefficient, the control device can use point B4 as the starting point and draw a perpendicular line 2 to the dotted line 4. The longer the perpendicular line 2 is, the smaller the negative consistency coefficient is; the shorter the perpendicular line 2 is, the larger the negative consistency coefficient is.
[0314] In other possible implementations, the control device may also use other parameters to represent the above-mentioned deviation degree and calculate the consistency coefficient based on the deviation degree. The embodiment of the present application does not limit the specific calculation method of the above-mentioned positive consistency coefficient and negative consistency coefficient.
[0315] After obtaining the positive consistency coefficient and the negative consistency coefficient, the control device may determine a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient.
[0316] In some possible implementations, the human voice signal includes a command word, which may be a wake-up command word, or a control command word.
[0317] When the control device detects a command word in a human voice signal, it can also obtain a lexical confidence level of the command word. The lexical confidence level is used to indicate the accuracy of the command word detection.
[0318] After the control device obtains the positive consistency coefficient and the negative consistency coefficient, the control device can adjust the above-mentioned vocabulary confidence according to the positive consistency coefficient and the negative consistency coefficient, or the control device can also adjust the confidence threshold corresponding to the command word according to the positive consistency coefficient and the negative consistency coefficient.
[0319] When the control device adjusts the lexical confidence of the command word according to the positive consistency coefficient and the negative consistency coefficient, the positive consistency coefficient is positively correlated with the adjusted lexical confidence, and the negative consistency coefficient is negatively correlated with the adjusted lexical confidence.
[0320] That is to say, the larger the positive consistency coefficient is, the more likely it is that the above-mentioned human voice signal is issued by the user, and the control device can appropriately increase the vocabulary confidence; the smaller the positive consistency coefficient is, the less likely it is that the above-mentioned human voice signal is issued by the user, and the control device can appropriately lower the vocabulary confidence.
[0321] The larger the negative consistency coefficient is, the more likely it is that the above-mentioned human voice signal is emitted by the speaker device, and the control device can appropriately lower the vocabulary confidence; the smaller the negative consistency coefficient is, the less likely it is that the above-mentioned human voice signal is emitted by the speaker device, and the control device can appropriately increase the vocabulary confidence.
[0322] The manner in which the control device adjusts the vocabulary confidence according to the positive consistency coefficient and the negative consistency coefficient can be set according to actual needs.
[0323] For example, in some embodiments, the control device may adjust the vocabulary confidence by using formula (1).
[0324] T=t*f / g (1)
[0325] Among them, T represents the vocabulary confidence after adjustment, t represents the vocabulary confidence before adjustment, f represents the positive consistency coefficient, and g represents the negative consistency coefficient.
[0326] In other embodiments, the control device may also adjust the vocabulary confidence by using formula (2).
[0327] T=t+fg (2)
[0328] In other embodiments, the control device may also adjust the vocabulary confidence by using formula (3).
[0329] T=t*fg (3)
[0330] In other embodiments, the control device may also adjust the vocabulary confidence in other ways. The embodiment of the present application does not limit the specific way in which the control device adjusts the vocabulary confidence.
[0331] When the control device adjusts the confidence threshold of the command word according to the positive consistency coefficient and the negative consistency coefficient, the positive consistency coefficient is negatively correlated with the adjusted confidence threshold, and the negative consistency coefficient is positively correlated with the adjusted confidence threshold.
[0332] That is to say, the larger the positive consistency coefficient is, the smaller the adjusted confidence threshold is; the smaller the positive consistency coefficient is, the larger the adjusted confidence threshold is.
[0333] The larger the negative consistency coefficient, the larger the adjusted confidence threshold; the smaller the negative consistency coefficient, the smaller the adjusted confidence threshold.
[0334] The manner in which the control device adjusts the confidence threshold according to the positive consistency coefficient and the negative consistency coefficient can be set according to actual needs.
[0335] For example, in some embodiments, the control device may adjust the confidence threshold using formula (4).
[0336] D=d*g / f (4)
[0337] Where D represents the confidence threshold after adjustment, and d represents the confidence threshold before adjustment.
[0338] In other embodiments, the control device may also adjust the confidence threshold using formula (5).
[0339] D=d+gf (5)
[0340] In other embodiments, the control device may also adjust the confidence threshold using formula (6).
[0341] D=d*gf (6)
[0342] In other embodiments, the control device may also adjust the confidence threshold in other ways. The embodiments of the present application do not limit the specific way in which the control device adjusts the confidence threshold.
[0343] After adjusting the lexical confidence or the confidence threshold of the command word, the control device may compare the lexical confidence of the command word with the confidence threshold.
[0344] When the vocabulary confidence is less than or equal to the confidence threshold, it indicates that the command word recognition fails, and the control device may not respond to the command word.
[0345] When the vocabulary confidence is greater than the confidence threshold, it means that the command word is recognized successfully, and the control device can execute the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word. After receiving the control instruction, the smart home device executes the operation corresponding to the command word.
[0346] For example, see Figure 13 Assume that a room is equipped with a smart TV 31, a smart speaker 32, a smart air conditioner 33, and a camera 34. The smart speaker 32 is a control device and a sound pickup device, the smart TV 31 is a speaker device, and the camera 34 is a detection device. Furthermore, the adjustment formula for word confidence is formula (1), and the confidence threshold for each command word is 0.9.
[0347] At a certain moment, smart TV 31 is playing a program, and one of the program audio clips is "It's too hot today! Turn on the air conditioner!" At this point, the smart TV 31, an electronic device, emits a human voice signal, but the user may not have made a human voice signal and may not want to turn on the air conditioner. If the control device mistakenly controls the air conditioner to turn on based on the command word "turn on the air conditioner" in the human voice signal contained in the program audio, the result will not meet the user's expectations.
[0348] Applying the method provided in the embodiments of the present application can effectively avoid such misidentification, or miscontrol and mistriggering.
[0349] Specifically, the smart speaker 32 may execute the following steps in response to the program audio:
[0350] At this time, the smart speaker 32 detects the human voice signal and recognizes the command word "turn on the air conditioner" in the human voice signal. Therefore, the smart speaker 32 obtains the target sound source direction of the human voice signal and sends a user positioning instruction to the camera 34.
[0351] After receiving the user's positioning instruction, the camera 34 captures the environment image in the room, recognizes the environment image, and obtains the user's location information.
[0352] The camera 34 then transmits the user's location information to the smart speaker 32.
[0353] After receiving the user's location information, the smart speaker 32 maps the user's location information to a pre-identified physical coordinate system, calculates a positive consistency coefficient based on the user's location information and the target sound source direction, and calculates a negative consistency coefficient based on the location information of the smart TV 31 and the target sound source direction.
[0354] like Figure 14 As shown, Figure 14is a pre-identified physical coordinate system. Point O represents the location of camera 34, point A5 represents the location of smart air conditioner 33, point B5 represents the location of smart speaker 32, point C5 represents the location of smart TV 31, and point D5 represents the user's location. The arrow on dashed line 5 indicates the mapping direction of the 0-degree direction of smart speaker 32 in the physical coordinate system.
[0355] Assuming that the smart speaker 32 detects that the target sound source direction is 33 degrees, the smart speaker 32 can rotate the dotted line 5 33 degrees counterclockwise along point B5 to obtain a dotted line 6. The arrow direction of the dotted line 6 represents the target sound source direction.
[0356] Then, the smart speaker 32 can calculate the angle 3 between the line segment B5-D5 and the dotted line 5, and calculate the positive consistency coefficient based on the angle 3; and the smart speaker 32 can calculate the angle 4 between the line segment B5-C5 and the dotted line 5, and calculate the negative consistency coefficient based on the angle 4.
[0357] Assume that the positive consistency coefficient calculated by the smart speaker 32 is 0.7, the negative consistency coefficient is 0.9, and the vocabulary confidence of the command word "turn on the air conditioner" recognized by the smart speaker 32 is 0.92.
[0358] At this time, the smart speaker 32 can substitute the values of the positive consistency coefficient, the negative consistency coefficient and the vocabulary confidence into formula (1), and obtain the adjusted vocabulary confidence of 0.92*0.7 / 0.9≈0.716.
[0359] Then, the smart speaker 32 compares the adjusted vocabulary confidence of 0.716 with the above-mentioned confidence threshold of 0.9. The vocabulary confidence is less than the confidence threshold, indicating that the command word "turn on the air conditioner" is recognized incorrectly, and the smart speaker 32 does not respond to the command word.
[0360] It can be seen from the above examples that when the control device detects a command word, it can adjust the vocabulary confidence or confidence threshold of the command word according to the positive consistency coefficient and the negative consistency coefficient.
[0361] Then, the control device performs command word recognition based on the adjusted vocabulary confidence / confidence threshold, thereby improving the accuracy of command word recognition, reducing the possibility of misrecognition of command words due to the sound of the electronic device, and improving the user experience.
[0362] In some other possible implementations, the control device may correct the direction of the target sound source according to the positive consistency coefficient and the negative consistency coefficient.
[0363] Specifically, after obtaining the positive consistency coefficient and the negative consistency coefficient, the control device may determine whether the positive consistency coefficient and the negative consistency coefficient meet a preset correction rule.
[0364] If the positive consistency coefficient and the negative consistency coefficient meet the preset correction rule, the control device can determine the direction of the user according to the user's position information and update the target sound source direction to the direction of the user.
[0365] If the positive consistency coefficient and the negative consistency coefficient do not satisfy the preset correction rule, the control device does not correct the target sound source direction.
[0366] The preset correction rule can be set according to the actual scenario. For example, in some embodiments, the preset correction rule can be that the positive consistency coefficient is greater than the negative consistency coefficient.
[0367] In other embodiments, the preset correction rule may be that the positive consistency coefficient is greater than a first deviation threshold, and the negative consistency coefficient is less than a second deviation threshold.
[0368] The first deviation threshold and the second deviation threshold are preset parameters.
[0369] In other embodiments, the preset correction rule may be that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than a third deviation threshold. The third deviation threshold is a preset parameter, and the third deviation threshold is greater than 0.
[0370] In other embodiments, the preset correction rules may also be other rules, and the embodiments of the present application do not limit the specific content of the preset correction rules.
[0371] After correcting the target sound source direction, the control device may send the corrected target sound source direction to the sound pickup device, and control the sound pickup device to perform an audio directional enhancement operation on the corrected target sound source direction.
[0372] For example, see Figure 15 , assume that a home hub 41, a smart TV 42, a smart speaker 43, and a millimeter-wave radar 44 are set up in a room. Among them, the home hub 41 is a control device, the smart speaker 43 is a sound pickup device, the smart TV 42 is a speaker device, and the millimeter-wave radar 44 is a detection device.
[0373] At a certain moment, the user speaks to the smart speaker 43, "Xiaoyi, Xiaoyi, let's chat."
[0374] At this time, the smart speaker 43 detects the human voice signal, identifies the human voice signal, and obtains the target sound source direction.
[0375] Then, the smart speaker 43 sends the target sound source direction to the home hub 41.
[0376] After receiving the direction of the target sound source, the home hub 41 sends a user positioning instruction to the millimeter wave radar 44.
[0377] After receiving the user's positioning instruction, the millimeter wave radar 44 transmits a millimeter wave detection signal to the surrounding area and receives the millimeter wave reflected signal. The millimeter wave radar 44 identifies the millimeter wave reflected signal and obtains the user's location information.
[0378] Afterwards, the millimeter wave radar 44 transmits the user's location information to the home hub 41.
[0379] After receiving the user's location information, the home hub 41 maps the user's location information to a pre-identified physical coordinate system, calculates a positive consistency coefficient based on the user's location information and the target sound source direction, and calculates a negative consistency coefficient based on the location information of the smart TV 42 and the target sound source direction.
[0380] like Figure 16 As shown, Figure 16 is a pre-identified physical coordinate system. Point O represents the location of millimeter-wave radar 44, point A6 represents the location of smart speaker 43, point B6 represents the location of home hub 41, point C6 represents the location of smart TV 42, and point D6 represents the user's location. The arrow direction of dashed line 7 represents the mapping direction of the 0-degree direction of smart speaker 43 in the physical coordinate system.
[0381] Assuming that the smart speaker 43 detects that the target sound source direction is 33 degrees, the home hub 41 can rotate the dotted line 7 33 degrees counterclockwise along point A6 to obtain a dotted line 8. The arrow direction of the dotted line 8 represents the target sound source direction.
[0382] Then, the home hub 41 can calculate the angle 5 between the line segment A6-D6 and the dotted line 8, and calculate the positive consistency coefficient based on the angle 5; and the home hub 41 can calculate the angle 6 between the line segment A6-C6 and the dotted line 8, and calculate the negative consistency coefficient based on the angle 6.
[0383] Assume that the positive consistency coefficient calculated by the home hub 41 is 0.9 and the negative consistency coefficient is 0.3. At this time, the home hub 41 can determine whether the positive consistency coefficient and the negative consistency coefficient meet the preset calibration rules.
[0384] Assuming that the preset correction rule is that the positive consistency coefficient is greater than the negative consistency coefficient, since the positive consistency coefficient calculated by the home hub 41 is 0.9, which is greater than the negative consistency coefficient of 0.3, it means that the positive consistency coefficient and the negative consistency coefficient meet the preset correction rule, and the above-mentioned human voice signal is likely to be a human voice signal emitted by the user.
[0385] At this point, the home hub 41 can update the target sound source direction using angle 5. Assuming angle 5 is 5 degrees, this means that the angle between line segment A6-D6 and dashed line 7 is 38 degrees (i.e., 33 degrees plus 5 degrees). The user is located at 38 degrees in the polar coordinate system, and the target sound source direction is updated from 33 degrees to 38 degrees.
[0386] Then, the home hub 41 sends the updated target sound source direction to the smart speaker 43, and the smart speaker 43 performs an audio directional enhancement operation on the 38-degree direction.
[0387] It can be seen from the above examples that the control device can determine whether the positive consistency coefficient and the negative consistency coefficient meet the preset correction rules, so as to identify whether the human voice signal is emitted by the user. For example, in the above examples, the preset correction rule may be that the positive consistency coefficient is greater than the negative consistency coefficient; in other examples not shown, the preset correction rule may also be that the positive consistency coefficient is greater than the first deviation threshold, and the negative consistency coefficient is less than the second deviation threshold; in other examples not shown, the preset correction rule may also be that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than the third deviation threshold. In other examples not shown, the preset correction rule may also be other correction rules.
[0388] When the positive consistency coefficient and the negative consistency coefficient meet the preset calibration rules, it means that the above human voice signal is likely to be emitted by the user. When the positive consistency coefficient and the negative consistency coefficient do not meet the calibration rules, it means that the above human voice signal is likely to be emitted by the speaker device.
[0389] Therefore, when the positive consistency coefficient and the negative consistency coefficient meet the preset correction rules, the control device can update the target sound source position direction to the direction where the user is located. The control device can then control the sound pickup device to perform an audio directional enhancement operation on the corrected target sound source direction, thereby improving the sound pickup effect of the sound pickup device in the direction where the user is located.
[0390] In some other possible implementations, the control device may directionally adjust the sound pickup effect in the direction of the target sound source according to the positive consistency coefficient and the negative consistency coefficient.
[0391] The sound pickup device may be equipped with multiple microphones, each corresponding to an audio channel. After the multiple microphones on the sound pickup device collect sound signals, the sound signals collected by each microphone are respectively converged through the corresponding audio channel to form a beam.
[0392] Therefore, after acquiring the target sound source direction, the control device may determine the audio channel of the sound pickup device in the target sound source direction as the target audio channel.
[0393] Then, the control device may adjust the signal ratio of the target audio channel signal in the beam according to the positive consistency coefficient and the negative consistency coefficient.
[0394] The above signal ratio is positively correlated with the positive consistency coefficient, and the above signal ratio is negatively correlated with the negative consistency coefficient.
[0395] That is to say, the larger the positive consistency coefficient, the larger the signal ratio; the smaller the positive consistency coefficient, the smaller the signal ratio. The larger the negative consistency coefficient, the smaller the signal ratio; the smaller the negative consistency coefficient, the larger the signal ratio.
[0396] The specific adjustment method of the above signal ratio can be set according to actual needs.
[0397] In some embodiments, the relationship between the signal ratio and the positive consistency coefficient and the negative consistency coefficient can be expressed by formula (7):
[0398] K=k*f / g (7)
[0399] Where K represents the signal ratio after the update, and k represents the signal ratio before the update.
[0400] In other embodiments, the relationship between the signal ratio and the positive consistency coefficient and the negative consistency coefficient can also be expressed by formula (8):
[0401] K=k+fg (8)
[0402] In other embodiments, the relationship between the signal ratio and the positive consistency coefficient and the negative consistency coefficient can also be expressed by formula (9):
[0403] K=k*fg (9)
[0404] In other embodiments, the control device may also adjust the signal ratio in other ways. The embodiments of the present application do not limit the specific way in which the control device adjusts the signal ratio.
[0405] For example, see Figure 15 , assume that a home hub 41, a smart TV 42, a smart speaker 43, and a millimeter-wave radar 44 are set up in a room. Among them, the home hub 41 is a control device, the smart speaker 43 is a sound pickup device, the smart TV 42 is a speaker device, and the millimeter-wave radar 44 is a detection device.
[0406] like Figure 17 As shown, it is assumed that the microphone array of the smart speaker 43 includes 6 microphones, namely Figure 17Microphone 431, microphone 432, microphone 433, microphone 434, microphone 435, and microphone 436 are shown in FIG. Since the positions of the microphones are different, the primary sound pickup areas corresponding to the microphones are different. For example, the primary sound pickup area corresponding to microphone 431 is area 1, the primary sound pickup area corresponding to microphone 432 is area 2, the primary sound pickup area corresponding to microphone 433 is area 3, the primary sound pickup area corresponding to microphone 434 is area 4, the primary sound pickup area corresponding to microphone 435 is area 5, and the primary sound pickup area corresponding to microphone 436 is area 6.
[0407] During the sound pickup process of the smart speaker 43, each microphone corresponds to an audio channel, and the sound signals collected by each microphone are converged through the corresponding audio channel to form a beam. The signal ratio corresponding to each audio channel represents the proportion of the sound signal of that audio channel in the beam. The signal ratio of the audio channel corresponding to each microphone is a default of 16.7%. The signal ratio adjustment formula can use the above formula (7).
[0408] At a certain moment, the user speaks to the smart speaker 43, "Xiaoyi, Xiaoyi, let's chat."
[0409] At this time, the smart speaker 43 detects the human voice signal, identifies the human voice signal, and obtains the target sound source direction.
[0410] Then, the smart speaker 43 sends the target sound source direction to the home hub 41.
[0411] After receiving the direction of the target sound source, the home hub 41 sends a user positioning instruction to the millimeter wave radar 44.
[0412] After receiving the user's positioning instruction, the millimeter wave radar 44 transmits a millimeter wave detection signal to the surrounding area and receives the millimeter wave reflected signal. The millimeter wave radar 44 identifies the millimeter wave reflected signal and obtains the user's location information.
[0413] The millimeter wave radar 44 then transmits the user's location information to the home hub 41.
[0414] After receiving the user's location information, the home hub 41 maps the user's location information to a pre-identified physical coordinate system, calculates a positive consistency coefficient based on the user's location information and the target sound source direction, and calculates a negative consistency coefficient based on the location information of the smart TV 42 and the target sound source direction.
[0415] like Figure 16 As shown, Figure 16is a pre-identified physical coordinate system. Point O represents the location of millimeter-wave radar 44, point A6 represents the location of smart speaker 43, point B6 represents the location of home hub 41, point C6 represents the location of smart TV 42, and point D6 represents the user's location. The arrow direction of dashed line 7 represents the mapping direction of the 0-degree direction of smart speaker 43 in the physical coordinate system.
[0416] Assuming that the smart speaker 43 detects that the target sound source direction is 33 degrees, the home hub 41 can rotate the dotted line 7 33 degrees counterclockwise along point A6 to obtain a dotted line 8. The arrow direction of the dotted line 8 represents the target sound source direction.
[0417] Then, the home hub 41 can calculate the angle 5 between the line segment A6-D6 and the dotted line 8, and calculate the positive consistency coefficient based on the angle 5; and the home hub 41 can calculate the angle 6 between the line segment A6-C6 and the dotted line 8, and calculate the negative consistency coefficient based on the angle 6.
[0418] Assume that the positive consistency coefficient calculated by family hub 41 is 0.9 and the negative consistency coefficient is 0.3.
[0419] At this time, the home hub 41 substitutes the values of the positive consistency coefficient, the negative consistency coefficient and the default signal ratio into formula (7), and obtains an updated signal ratio of 16.7%*0.9 / 0.3=50.1%.
[0420] The home hub 41 then sends the updated signal ratio to the smart speaker 43.
[0421] Assuming that the target audio channel corresponding to the target sound source direction is the audio channel corresponding to microphone 432, after the smart speaker 43 receives the updated signal ratio, it increases the signal ratio of the audio channel corresponding to microphone 432 to 50.1% and lowers the signal ratio of the audio channels corresponding to other microphones to 9.98%.
[0422] It can be seen from the above examples that the control device can adjust the signal ratio of the target audio channel according to the positive consistency coefficient and the negative consistency coefficient.
[0423] When the human voice signal is likely to be a human voice signal emitted by the user, the control device increases the signal ratio of the target audio channel according to the positive consistency coefficient and the negative consistency coefficient, so that the sound pickup device can more clearly collect the audio signal in the direction of the target sound source.
[0424] When the human voice signal is more likely to be a human voice signal emitted by the speaker device, the control device lowers the signal ratio of the target audio channel through the positive consistency coefficient and the negative consistency coefficient, so that the sound pickup device can reduce interference in the direction of the target sound source.
[0425] In other possible implementations, the control device may also implement other response strategies based on the positive consistency coefficient and the negative consistency coefficient in combination with the needs of the actual scenario. The embodiment of the present application does not limit the specific content of the above response strategy.
[0426] In addition, it should be noted that although in the various embodiments described above, the control device determines the response strategy based on both the positive consistency coefficient and the negative consistency coefficient, in some other possible embodiments, the control device may also determine the response strategy based on the positive consistency coefficient or the negative consistency coefficient alone.
[0427] For example, when the control device obtains the vocabulary confidence of the command word, the control device can adjust the vocabulary confidence based on the positive consistency coefficient using formula (10). The expression of formula (10) is:
[0428] T=t+f (10)
[0429] Alternatively, the control device may also adjust the above-mentioned word confidence based solely on the negative consistency coefficient through formula (11), and the expression of formula (11) is:
[0430] T=tg (11)
[0431] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0432] Hereinafter, another method for responding to a human voice signal provided by an embodiment of the present application will be described in detail from the perspective of a control device. Figure 18 The human voice signal response method provided in this embodiment includes:
[0433] S181. When a sound pickup device collects a human voice signal, the control device obtains a target sound source direction of the human voice signal, where the sound pickup device is an electronic device with a microphone.
[0434] In this embodiment, the sound pickup device refers to an electronic device including a microphone, and the speaker device refers to an electronic device including a speaker.
[0435] The sound pickup device can collect sound signals from the surrounding environment, detect whether there is a human voice signal in the sound signal, and send the sound source direction of the human voice signal to the control device when a human voice signal is detected.
[0436] Alternatively, the sound pickup device may also collect surrounding sound signals and send the sound signals to the control device, which then detects whether a human voice signal is present in the sound signals.
[0437] When the control device detects that a human voice signal exists in the sound signal, the control device obtains the target sound source direction of the human voice signal.
[0438] S182: The control device obtains location information of a target object, where the target object includes a user and / or a speaker device, where the speaker device is an electronic device with a speaker.
[0439] After the control device obtains the target sound source direction of the human voice signal, the control object can obtain the position information of the target object.
[0440] The target object may include location information of the user and / or the speaker device.
[0441] Specifically, the control device may detect the location information of the user and the speaker device through electronic devices with scanning and / or detection functions, such as millimeter wave radars and cameras.
[0442] Furthermore, when the target object includes a user, if the control device detects location information of the user in the current area, the control device may perform subsequent steps according to the location information of the user.
[0443] If the control device fails to detect the user's location information in the current area, it means that the user is not in the current area, the above-mentioned human voice signal is not a human voice signal emitted by the user, and the control device may not respond to the human voice signal.
[0444] S183: The control device determines a response strategy for the human voice signal according to the target sound source direction and the position information of the target object.
[0445] After obtaining the target sound source direction and the target object's position information, the control device can determine whether the above-mentioned human voice signal is a human voice signal emitted by the user based on the target sound source direction and the target object information, and determine a corresponding response strategy accordingly.
[0446] Specifically, the control device determines a response strategy for the human voice signal based on the target sound source direction and the position information of the target object, including:
[0447] The control device determines the degree of deviation between the position information of the target object and the direction of the target sound source based on the direction of the target sound source and the position information of the target object;
[0448] The control device determines the response strategy for the human voice signal according to the degree of deviation.
[0449] It should be noted that the above deviation degree indicates the deviation between the target object and the target sound source direction. A larger deviation degree indicates a greater difference between the target object and the target sound source direction; a smaller deviation degree indicates a closer relationship between the target object and the target sound source direction.
[0450] Therefore, the control device can determine the relationship between the above-mentioned human voice signal and the target object according to the degree of deviation, and then determine the corresponding response strategy.
[0451] In some possible implementations, the control device determines a response strategy for the human voice signal based on the degree of deviation, including:
[0452] When the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the user, the control device responds to the human voice signal;
[0453] When the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the speaker device, the control device does not respond to the human voice signal.
[0454] It should be noted that the control device may directly determine the relationship between the target object and the above-mentioned human voice signal according to the degree of deviation.
[0455] For example, when the target object includes a speaker device, if the first deviation degree between the speaker device and the target object is greater, it means that the above-mentioned human voice signal is less likely to be emitted by the speaker device; if the first deviation degree between the speaker device and the target object is smaller, it means that the above-mentioned human voice signal is more likely to be emitted by the speaker device.
[0456] When the target object is a user, the greater the second deviation between the user and the target object, the less likely the human voice signal is to be emitted by the user; and the smaller the second deviation between the user and the target object, the more likely the human voice signal is to be emitted by the user.
[0457] When the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the user, the control device may respond to the human voice signal.
[0458] When the control device determines that the human voice signal is a human voice signal emitted by the speaker device based on the degree of deviation, the control device may not respond to the human voice signal, thereby reducing the possibility of misidentification and reducing the impact of erroneous operation on the user.
[0459] In addition, the above response strategy may include responding to the human voice signal and not responding to the human voice signal.
[0460] Furthermore, the control device's response to the voice signal can be set according to the actual scenario. For example, when the voice signal includes a command word, the control device's response to the voice signal can be:
[0461] The control device executes the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
[0462] For example, assuming the control device is a smart speaker, and the command word in the human voice signal is "play music", the smart speaker has the function corresponding to the command word and can perform the operation corresponding to "play music".
[0463] Assume the control device is a home hub, and the voice command is "Turn on the air conditioner." Since the home hub doesn't have the functionality for this command, it can instead send a control command to the smart air conditioner. Upon receiving the command, the smart air conditioner executes the action "Turn on the air conditioner."
[0464] In some other possible implementations, the target object includes a user and a speaker device; the deviation degree includes a first deviation degree between the position information of the speaker device and the direction of the target sound source, and a second deviation degree between the position information of the user and the direction of the target sound source;
[0465] The control device determines the response strategy for the human voice signal based on the degree of deviation, including:
[0466] The control device calculates a negative consistency coefficient based on the first deviation degree;
[0467] The control device calculates a positive consistency coefficient based on the second deviation degree;
[0468] The control device determines a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient.
[0469] It should be noted that the negative consistency coefficient is negatively correlated with the first deviation degree. A larger first deviation degree results in a smaller negative consistency coefficient, indicating a greater deviation between the speaker device and the target sound source direction. A smaller first deviation degree results in a larger negative consistency coefficient, indicating a closer alignment between the speaker device and the target sound source direction.
[0470] The positive consistency coefficient is negatively correlated with the second deviation degree. A larger second deviation degree results in a smaller positive consistency coefficient, indicating a greater deviation between the user and the target sound source direction. A smaller second deviation degree results in a larger positive consistency coefficient, indicating a closer alignment between the user and the target sound source direction.
[0471] After obtaining the positive consistency coefficient and the negative consistency coefficient, the control device can determine whether the above-mentioned human voice signal is a human voice signal emitted by the user based on the positive consistency coefficient and the negative consistency coefficient, and then determine the response strategy of the human voice signal to reduce the possibility of misrecognition.
[0472] In some embodiments, the human voice signal includes a command word; and the control device determines a response strategy for the human voice signal based on the positive consistency coefficient and the negative consistency coefficient, including:
[0473] The control device obtains the vocabulary confidence and confidence threshold corresponding to the command word;
[0474] The control device adjusts the vocabulary confidence according to the positive consistency coefficient and the negative consistency coefficient, or adjusts the confidence threshold;
[0475] If the adjusted vocabulary confidence is greater than the confidence threshold, or the vocabulary confidence is greater than the adjusted confidence threshold, the control device executes the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
[0476] It should be noted that when the human voice signal includes a command word, the control device can obtain the vocabulary confidence and confidence threshold corresponding to the command word. The vocabulary confidence refers to the degree of credibility of the recognized command word.
[0477] Then, the control device may adjust the vocabulary confidence or the confidence threshold according to the positive consistency coefficient and the negative consistency coefficient.
[0478] When the control device adjusts the lexical confidence of the command word according to the positive consistency coefficient and the negative consistency coefficient, the positive consistency coefficient is positively correlated with the adjusted lexical confidence, and the negative consistency coefficient is negatively correlated with the adjusted lexical confidence.
[0479] That is to say, the larger the positive consistency coefficient is, the more likely it is that the above-mentioned human voice signal is issued by the user, and the control device can appropriately increase the vocabulary confidence; the smaller the positive consistency coefficient is, the less likely it is that the above-mentioned human voice signal is issued by the user, and the control device can appropriately lower the vocabulary confidence.
[0480] The larger the negative consistency coefficient is, the more likely it is that the above-mentioned human voice signal is emitted by the speaker device, and the control device can appropriately lower the vocabulary confidence; the smaller the negative consistency coefficient is, the less likely it is that the above-mentioned human voice signal is emitted by the speaker device, and the control device can appropriately increase the vocabulary confidence.
[0481] When the control device adjusts the confidence threshold of the command word according to the positive consistency coefficient and the negative consistency coefficient, the positive consistency coefficient is negatively correlated with the adjusted confidence threshold, and the negative consistency coefficient is positively correlated with the adjusted confidence threshold.
[0482] That is to say, the larger the positive consistency coefficient is, the smaller the adjusted confidence threshold is; the smaller the positive consistency coefficient is, the larger the adjusted confidence threshold is.
[0483] The larger the negative consistency coefficient, the larger the adjusted confidence threshold; the smaller the negative consistency coefficient, the smaller the adjusted confidence threshold.
[0484] After adjusting the lexical confidence or the confidence threshold of the command word, the control device may compare the lexical confidence of the command word with the confidence threshold.
[0485] When the vocabulary confidence is less than or equal to the confidence threshold, it indicates that the command word recognition fails, and the control device may not respond to the command word.
[0486] When the vocabulary confidence is greater than the confidence threshold, it means that the command word is recognized successfully, and the control device can execute the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word. After receiving the control instruction, the smart home device executes the operation corresponding to the command word.
[0487] In the method described above, the control device updates the vocabulary confidence / confidence threshold of the command word, and performs a secondary judgment on the command word based on the updated vocabulary confidence and confidence threshold. This can improve the accuracy of command word recognition, reduce the possibility of misrecognition of command words due to the sound of electronic devices, and improve the user experience.
[0488] In some other embodiments, the control device determines a response strategy for the human voice signal based on the positive consistency coefficient and the negative consistency coefficient, including:
[0489] When the positive consistency coefficient and the negative consistency coefficient meet the preset correction rules, the control device updates the target sound source direction to the direction where the user is located;
[0490] The control device controls the sound pickup device to perform an audio directional enhancement operation on the updated target sound source direction.
[0491] It should be noted that the preset correction rule is that the positive consistency coefficient is greater than the negative consistency coefficient;
[0492] Alternatively, the preset correction rule is that the positive consistency coefficient is greater than the first deviation threshold, and the negative consistency coefficient is less than the second deviation threshold;
[0493] Alternatively, the preset correction rule is that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than a third deviation threshold.
[0494] When the positive consistency coefficient and the negative consistency coefficient meet the preset calibration rules, it means that the above human voice signal is likely to be emitted by the user. When the positive consistency coefficient and the negative consistency coefficient do not meet the calibration rules, it means that the above human voice signal is likely to be emitted by the speaker device.
[0495] Therefore, when the positive consistency coefficient and the negative consistency coefficient meet the preset correction rules, the control device can update the target sound source position direction to the direction where the user is located. The control device can then control the sound pickup device to perform an audio directional enhancement operation on the corrected target sound source direction, thereby improving the sound pickup effect of the sound pickup device in the direction where the user is located.
[0496] In some other embodiments, the control device determines a response strategy for the human voice signal based on the positive consistency coefficient and the negative consistency coefficient, including:
[0497] The control device adjusts the signal ratio corresponding to the target audio channel according to the positive consistency coefficient and the negative consistency coefficient, where the target audio channel is the audio channel corresponding to the target sound source direction.
[0498] It should be noted that the sound pickup device may be provided with multiple microphones, each corresponding to an audio channel. After the multiple microphones on the sound pickup device collect sound signals, the sound signals collected by each microphone are respectively converged through the corresponding audio channel to form a beam.
[0499] Therefore, after acquiring the target sound source direction, the control device may determine the audio channel of the sound pickup device in the target sound source direction as the target audio channel.
[0500] Then, the control device may adjust the signal ratio of the target audio channel signal in the beam according to the positive consistency coefficient and the negative consistency coefficient.
[0501] The above signal ratio is positively correlated with the positive consistency coefficient, and the above signal ratio is negatively correlated with the negative consistency coefficient.
[0502] That is to say, the larger the positive consistency coefficient, the larger the signal ratio; the smaller the positive consistency coefficient, the smaller the signal ratio. The larger the negative consistency coefficient, the smaller the signal ratio; the smaller the negative consistency coefficient, the larger the signal ratio.
[0503] The control device adjusts the signal ratio corresponding to the target audio channel according to the positive consistency coefficient and the negative consistency coefficient.
[0504] When the human voice signal is likely to be a human voice signal emitted by the user, the control device increases the signal ratio of the target audio channel according to the positive consistency coefficient and the negative consistency coefficient, so that the sound pickup device can more clearly collect the audio signal in the direction of the target sound source.
[0505] When the human voice signal is more likely to be a human voice signal emitted by the speaker device, the control device lowers the signal ratio of the target audio channel through the positive consistency coefficient and the negative consistency coefficient, so that the sound pickup device can reduce interference in the direction of the target sound source.
[0506] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0507] Corresponding to the human voice signal response method described in the above embodiment, Figure 19 A structural block diagram of a human voice signal response device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0508] The device can be used to control equipment, refer to Figure 19 , the device comprises:
[0509] The sound source direction module 191 is used to obtain the target sound source direction of the human voice signal when the sound pickup device detects and collects the human voice signal;
[0510] A location information module 192 is configured to obtain location information of a target object, wherein the target object includes a user and / or a speaker device, wherein the speaker device is an electronic device having a speaker;
[0511] The response strategy module 193 is configured to determine a response strategy for the human voice signal according to the target sound source direction and the position information of the target object.
[0512] Optionally, the response strategy for the human voice signal includes responding to the human voice signal and not responding to the human voice signal.
[0513] Optionally, the human voice signal includes a command word; and the responding includes:
[0514] The control device executes the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
[0515] Optionally, the response strategy module 193 includes:
[0516] a deviation submodule, configured to determine a degree of deviation between the position information of the target object and the direction of the target sound source based on the direction of the target sound source and the position information of the target object;
[0517] The strategy submodule is used to determine a response strategy for the human voice signal according to the degree of deviation.
[0518] Optionally, the strategy submodule includes:
[0519] a response submodule, configured to respond to the human voice signal when the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the user;
[0520] The silent submodule is configured to not respond to the human voice signal when the control device determines, based on the deviation degree, that the human voice signal is a human voice signal emitted by the speaker device.
[0521] Optionally, the target object includes a user and a speaker device; the deviation degree includes a first deviation degree between the position information of the speaker device and the direction of the target sound source, and a second deviation degree between the position information of the user and the direction of the target sound source;
[0522] The strategy submodule includes:
[0523] A negative coefficient submodule, configured to calculate a negative consistency coefficient based on the first deviation degree;
[0524] a forward coefficient submodule, configured to calculate a forward consistency coefficient according to the second deviation degree;
[0525] The comprehensive decision submodule is used to determine a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient.
[0526] Optionally, the human voice signal includes a command word; and the comprehensive decision submodule includes:
[0527] A confidence submodule, configured to obtain a vocabulary confidence and a confidence threshold corresponding to the command word;
[0528] an adjusting submodule, configured to adjust the vocabulary confidence according to the positive consistency coefficient and the negative consistency coefficient, or to adjust the confidence threshold;
[0529] An operation submodule is configured to execute the operation corresponding to the command word, or send a control instruction to the smart home device corresponding to the command word, if the adjusted word confidence is greater than the confidence threshold, or the word confidence is greater than the adjusted confidence threshold.
[0530] Optionally, the comprehensive decision submodule includes:
[0531] a direction updating submodule, configured to update the target sound source direction to the direction of the user when the positive consistency coefficient and the negative consistency coefficient meet a preset correction rule;
[0532] The directional sound pickup submodule is used to control the sound pickup device to perform audio directional enhancement operations on the updated target sound source direction, and the sound pickup device is an electronic device with a microphone.
[0533] Optionally, the preset correction rule is that the positive consistency coefficient is greater than the negative consistency coefficient;
[0534] Alternatively, the preset correction rule is that the positive consistency coefficient is greater than a first deviation threshold, and the negative consistency coefficient is less than a second deviation threshold;
[0535] Alternatively, the preset correction rule is that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than a third deviation threshold.
[0536] Optionally, the comprehensive decision submodule includes:
[0537] A ratio adjustment submodule is used to adjust the signal ratio corresponding to the target audio channel according to the positive consistency coefficient and the negative consistency coefficient, where the target audio channel is the audio channel corresponding to the target sound source direction.
[0538] Optionally, the target object includes a user; and the device further includes:
[0539] The rejection response module is used to not respond to the human voice signal when the control device fails to obtain the user's location information.
[0540] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0541] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0542] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0543] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0544] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0545] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0546] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.
[0547] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for responding to a human voice signal, characterized in that: include: When a sound pickup device collects a human voice signal, the control device obtains a target sound source direction of the human voice signal, wherein the sound pickup device is an electronic device with a microphone; The control device acquires location information of a target object, where the target object includes a user and / or a speaker device, where the speaker device is an electronic device having a speaker; the location information of the speaker device is obtained by the control device through detection of the environment by a detection device; The control device determines a response strategy for the human voice signal according to the target sound source direction and the position information of the target object; The control device determines a response strategy for the human voice signal according to the target sound source direction and the position information of the target object, including: The control device determines, based on the target sound source direction and the position information of the target object, a degree of deviation between the position information of the target object and the target sound source direction; the degree of deviation includes a first degree of deviation between the position information of the speaker device and the target sound source direction, and a second degree of deviation between the position information of the user and the target sound source direction; The control device determines a response strategy for the human voice signal according to the degree of deviation; The target objects include users and speaker devices; The control device determines a response strategy for the human voice signal according to the degree of deviation, including: The control device calculates a negative consistency coefficient according to the first deviation degree; The control device calculates a positive consistency coefficient according to the second deviation degree; The control device determines a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient; The sound pickup device includes multiple microphones; each microphone corresponds to an audio channel, and the sound signals collected by each microphone are respectively converged through the corresponding audio channel to form a beam; the control device determines a response strategy for the human voice signal based on the positive consistency coefficient and the negative consistency coefficient, including: The control device adjusts the signal ratio corresponding to the target audio channel according to the positive consistency coefficient and the negative consistency coefficient, where the target audio channel is the audio channel corresponding to the target sound source direction; the signal ratio is positively correlated with the positive consistency coefficient, and the signal ratio is negatively correlated with the negative consistency coefficient.
2. The method according to claim 1, wherein The response strategy for the human voice signal includes responding to the human voice signal and not responding to the human voice signal.
3. The method according to claim 2, wherein The human voice signal includes a command word; and the responding includes: The control device executes the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
4. The method according to claim 1, wherein The control device determines a response strategy for the human voice signal according to the degree of deviation, including: When the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the user, the control device responds to the human voice signal; When the control device determines, based on the degree of deviation, that the human voice signal is a human voice signal emitted by the speaker device, the control device does not respond to the human voice signal.
5. The method according to claim 1, wherein The human voice signal includes a command word; The control device determines a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient, including: The control device obtains the vocabulary confidence and confidence threshold corresponding to the command word; The control device adjusts the vocabulary confidence according to the positive consistency coefficient and the negative consistency coefficient, or adjusts the confidence threshold; If the adjusted vocabulary confidence is greater than the confidence threshold, or the vocabulary confidence is greater than the adjusted confidence threshold, the control device performs the operation corresponding to the command word, or the control device sends a control instruction to the smart home device corresponding to the command word.
6. The method according to claim 1, wherein The control device determines a response strategy for the human voice signal according to the positive consistency coefficient and the negative consistency coefficient, including: When the positive consistency coefficient and the negative consistency coefficient meet a preset correction rule, the control device updates the target sound source direction to the direction where the user is located; The control device controls the sound pickup device to perform an audio directional enhancement operation on the updated target sound source direction.
7. The method according to claim 6, wherein The preset correction rule is that the positive consistency coefficient is greater than the negative consistency coefficient; Alternatively, the preset correction rule is that the positive consistency coefficient is greater than a first deviation threshold, and the negative consistency coefficient is less than a second deviation threshold; Alternatively, the preset correction rule is that the difference between the positive consistency coefficient and the negative consistency coefficient is greater than a third deviation threshold.
8. The method according to any one of claims 1 to 7, characterized in that The target object includes a user; after the control device obtains the target sound source direction of the human voice signal, the method further includes: When the control device fails to obtain the user's location information, the control device does not respond to the human voice signal.
9. A control device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor is configured to implement the method according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium configured to store a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
11. A computer program product, characterized in that The computer program product is configured to cause the control device to perform the method according to any one of claims 1 to 8 when the computer program product is run on the control device.
12. A chip system, characterized in that: The chip system includes a memory and a processor, and the processor is configured to execute a computer program stored in the memory to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Man-machine interaction method and device, computer equipment and storage medium
CN112581981A