Smart device interaction method, device, storage medium and electronic device
By detecting the environmental volume value and performing lip recognition in a noisy environment, and obtaining auxiliary recognition data to wake up the smart device, the problem of wake-up failure in noisy environments and user accents is solved, improving wake-up accuracy and user experience.
Patent Information
- Application Number
- CN202210178420.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-02-24
AI Technical Summary
In noisy environments or when users have local accents, it is difficult for the prior art to accurately recognize voice wake-up keywords, resulting in the failure of the smart device to wake up.
By detecting the ambient volume value, auxiliary recognition is performed when the threshold is reached, such as lip recognition, to obtain auxiliary recognition data to wake up the device, and if the voice wake-up fails, auxiliary recognition data is used to wake up the device.
Improves the wake-up accuracy of smart devices in noisy environments and user accents, and improves user experience.
Smart Images

Figure CN114678019B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of intelligent recognition technology, in particular to the field of voice recognition and image recognition technology, and specifically to an intelligent device interaction method, device, storage medium and electronic device. Background Art
[0002] Currently, many fields are beginning to adopt voice wake-up to wake up smart devices. However, in some situations, such as in noisy public areas, voice wake-up may not be able to recognize the voice interaction information due to nearby noise, and thus fail to wake up the smart device. In addition, for users with speech impairments, such as those with regional accents, it is also difficult to effectively wake up smart devices.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, storage medium, and electronic device for smart device interaction.
[0005] According to one aspect of the present disclosure, a smart device interaction method is provided, including: detecting an ambient volume value of an environment in which a target device is located; if it is detected that the ambient volume value reaches a preset threshold, performing auxiliary recognition in the process of identifying a wake-up keyword to obtain auxiliary recognition data, wherein the wake-up keyword is used to wake up the target device; if the target device cannot be woken up using the wake-up keyword, waking up the target device according to the auxiliary recognition data.
[0006] According to another aspect of the present disclosure, a smart device interaction device is provided, including: a detection module for detecting an ambient volume value of an environment in which a target device is located; an acquisition module for performing auxiliary recognition in the process of identifying a wake-up keyword if it is detected that the ambient volume value reaches a preset threshold, and obtaining auxiliary recognition data, wherein the wake-up keyword is used to wake up the target device; and a wake-up module for waking up the target device according to the auxiliary recognition data if the target device cannot be woken up using the wake-up keyword.
[0007] According to another aspect of the present disclosure, the above-mentioned auxiliary recognition is provided, which at least includes: lip reading recognition; the above-mentioned acquisition module includes: a first acquisition sub-module, which is used to obtain the above-mentioned wake-up keyword by identifying the voice content output by the wake-up object; a second acquisition sub-module, which is used to identify the facial area of the above-mentioned wake-up object to obtain facial feature information, wherein the above-mentioned facial feature information at least includes: lip shape features when lip sounds are uttered; a third acquisition sub-module, which is used to perform the above-mentioned lip reading recognition according to the above-mentioned lip shape features to obtain the above-mentioned auxiliary recognition data.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above-mentioned smart device interaction methods.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the above-mentioned smart device interaction methods.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements any of the above-mentioned smart device interaction methods when executed by a processor.
[0011] According to another aspect of the present disclosure, a smart device interaction product is provided, including the electronic device as described above.
[0012] The embodiments of the present disclosure can improve the wake-up efficiency of smart devices.
[0013] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a flow chart of a smart device interaction method according to the first embodiment of the present disclosure;
[0015] Figure 2 is a flowchart of an optional smart device interaction method according to the first embodiment of the present disclosure;
[0016] Figure 3 is a flowchart of another optional smart device interaction method according to the first embodiment of the present disclosure;
[0017] Figure 4 is a flowchart of another optional smart device interaction method according to the first embodiment of the present disclosure;
[0018] Figure 5 is a structural diagram of a smart device interaction apparatus according to a second embodiment of the present disclosure;
[0019] Figure 6 It is a block diagram of an electronic device used to implement the smart device interaction method according to the first embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0022] Example 1
[0023] Existing technologies primarily use fixed voice wake-up words to wake up smart devices, such as Xiaodu Xiaodu, Hey Siri, etc. Smart devices can also be woken up by clicking a specific button on the screen. However, when using fixed voice wake-up words to wake up smart devices, the accuracy of recognizing the user's words during the reception process is limited by voice alone. The reception effect also needs to be improved in noisy environments or when users have inaccurate pronunciation (accents). Furthermore, when multiple people are speaking at the same time, the voice of a single user cannot be accurately recognized, resulting in the inability to wake up the smart device.
[0024] According to an embodiment of the present disclosure, an embodiment of a smart device interaction method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0025] Figure 1 is a flow chart of a smart device interaction method according to the first embodiment of the present disclosure. Figure 1 As shown, the method includes the following steps:
[0026] Step S102, detecting the ambient volume value of the environment in which the target device is located;
[0027] Step S104: if it is detected that the ambient volume value reaches a preset threshold, performing auxiliary recognition in the process of identifying the wake-up keyword to obtain auxiliary recognition data, wherein the wake-up keyword is used to wake up the target device;
[0028] Step S106: If the target device cannot be awakened by the awakening keyword, the target device is awakened according to the auxiliary recognition data.
[0029] It can be understood that whether the current target device is in a noisy environment is determined by detecting whether the ambient volume value reaches a preset threshold.
[0030] Optionally, the above-mentioned target device can be, but is not limited to, an intelligent robot (such as indoor and outdoor delivery robots, multi-cabin robots, food delivery robots, etc.), a smart phone, a smart tablet, a smart car, and other smart devices equipped with a voice wake-up device.
[0031] Optionally, the above-mentioned wake-up keywords are used to wake up the above-mentioned target device, such as Xiaodu Xiaodu, Hey S ir i, etc.; the above-mentioned auxiliary recognition can be but is not limited to lip reading recognition; the above-mentioned auxiliary recognition data can be but is not limited to lip reading recognition data.
[0032] It should be noted that when it is detected that the ambient volume value of the environment in which the target device is located reaches a preset threshold, it indicates that the above-mentioned target device is currently in a noisy environment. At this time, the wake-up keyword cannot be accurately identified by voice recognition alone. Therefore, auxiliary recognition is further performed in the process of identifying the wake-up keyword, such as lip reading recognition, to obtain auxiliary recognition data, and wake up the above-mentioned target device according to the above-mentioned auxiliary recognition data, so as to achieve the purpose of accurately identifying the wake-up keyword, improve the wake-up rate of the target device, and thus improve the user experience.
[0033] In the embodiment of the present disclosure, the ambient volume value of the environment in which the target device is located is detected; if it is detected that the above-mentioned ambient volume value reaches a preset threshold, auxiliary recognition is performed in the process of identifying the wake-up keyword to obtain auxiliary recognition data, wherein the above-mentioned wake-up keyword is used to wake up the above-mentioned target device; if the above-mentioned wake-up keyword cannot wake up the above-mentioned target device, the above-mentioned target device is woken up according to the above-mentioned auxiliary recognition data, thereby achieving the purpose of waking up the target device by adopting multiple interactive methods, thereby realizing the technical effect of improving the accuracy of interactive information recognition and improving wake-up efficiency, and further solving the technical problems of low recognition accuracy and poor wake-up effect in the prior art method of waking up smart devices by only using voice interaction.
[0034] As an optional embodiment, Figure 2 is a flow chart of an optional smart device interaction method according to the first embodiment of the present disclosure, such as Figure 2As shown, the auxiliary recognition includes at least: lip reading recognition; the auxiliary recognition is performed in the process of recognizing the wake-up keyword to obtain auxiliary recognition data, including:
[0035] Step S202, obtaining the above-mentioned wake-up keyword by identifying the voice content output by the wake-up object;
[0036] Step S204, identifying the facial area of the awakened object to obtain facial feature information;
[0037] Step S206: performing the lip reading recognition according to the lip shape feature to obtain the auxiliary recognition data.
[0038] Optionally, the above-mentioned facial feature information includes at least: lip shape features when uttering lip sounds. For example, when the wake-up keyword is "xiaodu xiaodu", the user's mouth shape will change from a flat "1" shape to an "o" shape when pronouncing the sound, and repeat it twice, which is used as a feature basis for waking up the target device.
[0039] It should be noted that in the prior art, the target device is awakened only by obtaining the wake-up keyword through voice recognition, and the recognition accuracy of the wake-up keyword is low, which makes it difficult to wake up the target device. The embodiment of the present disclosure obtains the above-mentioned wake-up keyword by recognizing the voice content output by the wake-up object; recognizes the facial area of the above-mentioned wake-up object to obtain facial feature information; and performs the above-mentioned lip reading recognition according to the above-mentioned lip shape features to obtain the above-mentioned auxiliary recognition data, thereby expanding the way to wake up the target device and improving the possibility of waking up the target device.
[0040] In an optional embodiment, waking up the target device according to the auxiliary identification data includes:
[0041] Step S302, determining a wake-up auxiliary word for waking up the target device according to the auxiliary recognition data;
[0042] Step S304: Use the wake-up auxiliary word to wake up the target device.
[0043] It should be noted that in the prior art, the target device is awakened only by obtaining the wake-up keyword through voice recognition, and the recognition accuracy of the wake-up keyword is low, which makes it difficult to wake up the target device. The embodiment of the present disclosure determines the wake-up auxiliary word for waking up the target device based on the auxiliary recognition data; the method of waking up the target device using the wake-up auxiliary word expands the ways to wake up the target device, thereby increasing the possibility of waking up the target device.
[0044] As an optional embodiment, Figure 3 is a flow chart of another optional smart device interaction method according to the first embodiment of the present disclosure, such as Figure 3As shown, after waking up the target device according to the auxiliary identification data, the method further includes:
[0045] Step S402: Using the target device to process the first voice content output by the wake-up object to obtain first received audio data;
[0046] Step S404: identifying a first lip shape feature of the awakened object when outputting the first voice content, and obtaining first lip reading data;
[0047] Step S406, respectively identifying the first text content corresponding to the first sound recording data and the second text content corresponding to the first lip reading data to obtain a first recognition result;
[0048] Step S408: Determine whether to continuously receive the first voice content output by the wake-up object according to the first recognition result.
[0049] Optionally, a directional microphone in the target device is used to collect the first voice content output by the wake-up object to obtain first collected data.
[0050] Optionally, the second text content is text content that matches the first text content and the first lip reading data. After the target object is awakened, the directional microphone in the target device is used to collect the sound of the awakened object, and the lip shape content and the collected sound content are simultaneously detected to see if they are consistent.
[0051] It should be noted that different characters have different lip shape features, and different mouth shapes correspond to different labial sounds and labiodental sounds. In the process of converting speech to text, feature matching is performed on the mouth shapes of the labial sounds and labiodental sounds. Different character pronunciations correspond to different labial sounds and labiodental sounds. For example, the "bo" in "boss" is a labial sound. By double comparing the lip shape features and the speech recognition results, the accuracy of speech recognition can be further improved.
[0052] In an optional embodiment, determining whether to continuously receive the first voice content output by the wake-up object according to the first recognition result includes:
[0053] Step S502: If the first recognition result indicates that the first text content and the second text content are consistent, then determining to continue collecting the first voice content until the voice recognition process ends;
[0054] Step S504: If the recognition result indicates that the first text content and the second text content are inconsistent, it is determined that there is no need to continuously collect the first voice content, and other speaking objects other than the wake-up object are searched within a predetermined range.
[0055] Optionally, if the first recognition result indicates that the first text content and the second text content are consistent, it is determined that the voice recognition result of the wake-up object is accurate, and the sound of the wake-up object is continuously collected.
[0056] Optionally, during the process of continuously collecting the above-mentioned first voice content, lip reading data corresponding to the above-mentioned first voice content can be obtained at the same time, and matching can be continuously performed to achieve the purpose of accurately obtaining the voice interaction signal emitted by the awakened object.
[0057] It should be noted that when the recognition result indicates that the above-mentioned first text content and the above-mentioned second text content are inconsistent, it means that there is a deviation between the first text content obtained by voice recognition and the second text content obtained by lip shape recognition. At this time, it is impossible to accurately know the specific meaning that the current awakening object wants to express. Therefore, there is no need to continuously collect the above-mentioned first voice content. Relevant equipment (such as an image acquisition device) should be used to find other speaking objects other than the above-mentioned awakening object within a predetermined range.
[0058] As an optional embodiment, Figure 4 is a flow chart of another optional smart device interaction method according to the first embodiment of the present disclosure, such as Figure 4 As shown, after searching for other speaking objects other than the awakening object within the predetermined range, the method further includes:
[0059] Step S602: collecting the second speech content output by the other speaker to obtain second collected sound data;
[0060] Step S604, identifying the second lip shape features of the other speaker when outputting the second voice content, and obtaining second lip reading data;
[0061] Step S606, respectively identifying the third text content corresponding to the second collected sound data and the fourth text content corresponding to the second lip reading data to obtain a second recognition result;
[0062] Step S608: If the second recognition result indicates that the third text content is consistent with the fourth text content, then it is determined to continuously collect the second voice content output by the other speaker until the voice recognition process is terminated.
[0063] Optionally, when it is determined that the first text content and the second text content output by the wake-up object are inconsistent, other speaking objects other than the wake-up object are searched within a predetermined range, and the second sound reception data and second lip reading data output by the other speaking objects are obtained. When the text content corresponding to the second sound reception data and the second lip reading data (that is, the third text content and the second lip reading data) are consistent, the second voice content output by the other speaking objects is continuously recorded until the voice recognition process is terminated, so as to achieve the effect of accurately identifying what the speaker wants to express and improving the accuracy of voice interaction.
[0064] In an optional embodiment, the above method further includes:
[0065] Step S702: If it is determined that the first recognition result and the second recognition result are inconsistent after a predetermined period of time, a plurality of text contents to be selected are displayed on the display interface;
[0066] Step S704, in response to the click operation of the awakened object or the other speaking object, selecting the correct expression content from the multiple text contents to be selected;
[0067] Step S706: Control the target device to perform a feedback operation according to the correct expression content.
[0068] Optionally, the above-mentioned text content to be selected includes: the above-mentioned first text content, the above-mentioned second text content, the above-mentioned third text content and the above-mentioned fourth text content.
[0069] Optionally, the target device is controlled to perform the feedback operation of the correct expression content through voice broadcast or text prompt.
[0070] It should be noted that, after a predetermined period of time, if it is determined that the clocks of the above-mentioned first recognition result and the above-mentioned second recognition result cannot indicate the same thing, then multiple text contents to be selected (i.e. the above-mentioned first text content, the above-mentioned second text content, the above-mentioned third text content and the above-mentioned fourth text content) are displayed on the display interface. The user can select the correct expression content from the above-mentioned multiple text contents to be selected by clicking on the above-mentioned display interface. After obtaining the user's input information, the above-mentioned target device feeds back the user's input results through voice broadcast or question prompts, ensuring that the user can get the desired interactive results and improving the user experience.
[0071] It should be noted that the optional or preferred implementations of this embodiment can be found in the relevant descriptions of the vehicle information prompt method embodiment described above, and will not be repeated here. In the technical solution disclosed herein, the acquisition, storage, and application of user personal information involved are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0072] Example 2
[0073] According to an embodiment of the present disclosure, there is also provided an apparatus embodiment for implementing the above-mentioned smart device interaction method. Figure 5 is a structural diagram of a smart device interaction apparatus according to a second embodiment of the present disclosure, such as Figure 5 As shown, the above-mentioned smart device interaction device includes: a detection module 40, an acquisition module 42, and a wake-up module 44, wherein:
[0074] The detection module 40 is used to detect the ambient volume value of the environment in which the target device is located;
[0075] The acquisition module 42 is configured to perform auxiliary recognition during the process of identifying the wake-up keyword to obtain auxiliary recognition data if it is detected that the ambient volume value reaches a preset threshold, wherein the wake-up keyword is used to wake up the target device;
[0076] The wake-up module 44 is configured to wake up the target device according to the auxiliary recognition data if the wake-up keyword fails to wake up the target device.
[0077] In the embodiment of the present disclosure, the above-mentioned detection module 40 is used to detect the ambient volume value of the environment in which the target device is located; the above-mentioned acquisition module 42 is used to perform auxiliary recognition in the process of identifying the wake-up keyword if it is detected that the above-mentioned ambient volume value reaches a preset threshold, and obtain auxiliary recognition data, wherein the above-mentioned wake-up keyword is used to wake up the above-mentioned target device; the above-mentioned wake-up module 44 is used to wake up the above-mentioned target device according to the above-mentioned auxiliary recognition data if the above-mentioned wake-up keyword cannot wake up the above-mentioned target device, thereby achieving the purpose of waking up the target device using multiple interactive methods, thereby achieving the technical effect of improving the accuracy of interactive information recognition and improving wake-up efficiency, and further solving the technical problems of low recognition accuracy and poor wake-up effect in the method of waking up smart devices using only voice interaction in the prior art. It should be noted that each of the above-mentioned modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: each of the above-mentioned modules can be located in the same processor; or each of the above-mentioned modules can be located in different processors in any combination.
[0078] It should be noted that the detection module 40, acquisition module 42, and wake-up module 44 correspond to steps S102 to S106 in Example 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in a computer terminal.
[0079] Optionally, the above-mentioned auxiliary recognition includes at least: lip reading recognition; the above-mentioned acquisition module includes: a first acquisition sub-module, used to obtain the above-mentioned wake-up keyword by identifying the voice content output by the wake-up object; a second acquisition sub-module, used to identify the facial area of the above-mentioned wake-up object to obtain facial feature information, wherein the above-mentioned facial feature information includes at least: lip shape features when lip sounds are uttered; a third acquisition sub-module, used to perform the above-mentioned lip reading recognition according to the above-mentioned lip shape features to obtain the above-mentioned auxiliary recognition data.
[0080] Optionally, the above-mentioned wake-up module includes: a first determination module, used to determine the wake-up auxiliary word used to wake up the above-mentioned target device based on the above-mentioned auxiliary recognition data; and a first wake-up sub-module, used to use the above-mentioned wake-up auxiliary word to wake up the above-mentioned target device.
[0081] Optionally, the above-mentioned device also includes: a fourth acquisition sub-module, which is used to use the above-mentioned target device to collect the first voice content output by the wake-up object to obtain first collection data; a fifth acquisition sub-module, which is used to identify the first lip shape feature of the above-mentioned wake-up object when outputting the above-mentioned first voice content to obtain first lip reading data; a sixth acquisition sub-module, which is used to respectively identify the first text content corresponding to the above-mentioned first collection data and the second text content corresponding to the above-mentioned first lip reading data to obtain a first recognition result; a second determination module, which is used to determine whether to continuously collect the above-mentioned first voice content output by the above-mentioned wake-up object based on the above-mentioned first recognition result.
[0082] Optionally, the second determination module includes: a third determination module, which is used to determine that the first voice content is continuously recorded if the first recognition result indicates that the first text content and the second text content are consistent, until the voice recognition process is terminated; a fourth determination module, which is used to determine that there is no need to continuously record the first voice content if the recognition result indicates that the first text content and the second text content are inconsistent, and to search for other speaking objects other than the above-mentioned wake-up object within a predetermined range.
[0083] Optionally, the above-mentioned device also includes: a seventh acquisition sub-module, which is used to collect the second voice content output by the above-mentioned other speaking object to obtain second collected data; an eighth acquisition sub-module, which is used to identify the second lip shape features of the above-mentioned other speaking object when outputting the above-mentioned second voice content to obtain second lip reading data; a ninth acquisition sub-module, which is used to respectively identify the third text content corresponding to the above-mentioned second collected data and the fourth text content corresponding to the above-mentioned second lip reading data to obtain a second recognition result; and a fifth determination module, which is used to determine to continue collecting the above-mentioned second voice content output by the above-mentioned other speaking object if the above-mentioned second recognition result indicates that the above-mentioned third text content and the above-mentioned fourth text content are consistent, until the voice recognition process is terminated.
[0084] Optionally, the above-mentioned device also includes: a sixth determination module, which is used to display multiple text contents to be selected on the display interface if it is determined that the above-mentioned first recognition result and the above-mentioned second recognition result both indicate inconsistency after a predetermined time period, wherein the above-mentioned text contents to be selected include: the above-mentioned first text content, the above-mentioned second text content, the above-mentioned third text content and the above-mentioned fourth text content; a selection module, which is used to select the correct expression content from the above-mentioned multiple text contents to be selected in response to the click operation of the above-mentioned wake-up object or the above-mentioned other speaking object; and a control module, which is used to control the above-mentioned target device to perform a feedback operation according to the above-mentioned correct expression content.
[0085] It should be noted that the optional or preferred implementation of this embodiment can be found in the relevant description of Example 1 and will not be repeated here. In the technical solution disclosed herein, the acquisition, storage, and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0086] Example 3
[0087] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, a computer program product, and a smart device interactive product.
[0088] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0089] like Figure 6 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0090] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0091] The computing unit 801 can be a variety of general and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the method for detecting the ambient volume value of the environment in which the target device is located. For example, in some embodiments, the method for detecting the ambient volume value of the environment in which the target device is located can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method for detecting the ambient volume value of the environment in which the target device is located described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured in any other appropriate manner (eg, by means of firmware) to execute the method to detect the ambient volume value of the environment in which the target device is located.
[0092] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0093] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0094] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0096] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0097] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0098] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0099] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A smart device interaction method, comprising: Detect the ambient volume value of the target device's environment; If it is detected that the ambient volume value reaches a preset threshold, performing auxiliary recognition in the process of identifying the wake-up keyword to obtain auxiliary recognition data, wherein the wake-up keyword is used to wake up the target device; If the target device cannot be awakened by using the awakening keyword, awakening the target device according to the auxiliary identification data; After waking up the target device according to the auxiliary recognition data, the method further includes: using the target device to collect the first voice content output by the wake-up object to obtain first collected data; identifying the first lip shape feature of the wake-up object when outputting the first voice content to obtain first lip reading data; respectively identifying the first text content corresponding to the first collected data and the second text content corresponding to the first lip reading data to obtain a first recognition result; and determining whether to continuously collect the first voice content output by the wake-up object based on the first recognition result.
2. The method according to claim 1, wherein The auxiliary recognition includes at least: lip reading recognition; the auxiliary recognition is performed during the process of recognizing the wake-up keyword to obtain auxiliary recognition data, including: Obtaining the wake-up keyword by identifying the voice content output by the wake-up object; Identify the facial area of the awakened object to obtain facial feature information, wherein the facial feature information at least includes: lip shape features when uttering lip sounds; The lip reading recognition is performed according to the lip shape feature to obtain the auxiliary recognition data.
3. The method according to claim 1, wherein The waking up the target device according to the auxiliary identification data includes: determining a wake-up auxiliary word for waking up the target device according to the auxiliary recognition data; The target device is woken up using the wake-up auxiliary word.
4. The method according to claim 1, wherein The determining, according to the first recognition result, whether to continuously receive the first voice content output by the wake-up object includes: If the first recognition result indicates that the first text content and the second text content are consistent, determining to continue collecting the first voice content until the voice recognition process ends; If the recognition result indicates that the first text content and the second text content are inconsistent, it is determined that there is no need to continuously collect the first voice content, and other speaking objects other than the wake-up object are searched within a predetermined range.
5. The method according to claim 4, wherein After searching for other speaking objects other than the wake-up object within a predetermined range, the method further includes: Performing sound collection processing on the second speech content output by the other speaker to obtain second sound collection data; Identifying second lip shape features of the other speaker when outputting the second voice content to obtain second lip reading data; Respectively identifying a third text content corresponding to the second collected sound data and a fourth text content corresponding to the second lip reading data to obtain a second recognition result; If the second recognition result indicates that the third text content is consistent with the fourth text content, it is determined to continuously collect the second voice content output by the other speaker until the voice recognition process is terminated.
6. The method according to claim 5, wherein: The method further comprises: If, after a predetermined period of time, it is determined that both the first recognition result and the second recognition result indicate inconsistency, a plurality of text contents to be selected are displayed on a display interface, wherein the text contents to be selected include: the first text content, the second text content, the third text content, and the fourth text content; In response to a click operation of the awakened object or the other speaking object, selecting a correct expression content from the multiple text contents to be selected; The target device is controlled to perform a feedback operation according to the correct expression content.
7. A smart device interaction device, comprising: A detection module, used to detect the ambient volume value of the environment where the target device is located; an acquisition module, configured to perform auxiliary recognition in the process of identifying a wake-up keyword to obtain auxiliary recognition data if it is detected that the ambient volume value reaches a preset threshold, wherein the wake-up keyword is used to wake up the target device; a wake-up module, configured to wake up the target device according to the auxiliary identification data if the wake-up keyword fails to wake up the target device; The wake-up module is also used to use the target device to collect the first voice content output by the wake-up object to obtain first collected data; identify the first lip shape feature of the wake-up object when outputting the first voice content to obtain first lip reading data; respectively identify the first text content corresponding to the first collected data and the second text content corresponding to the first lip reading data to obtain a first recognition result; and determine whether to continuously collect the first voice content output by the wake-up object based on the first recognition result.
8. The device according to claim 7, wherein The auxiliary recognition includes at least: lip reading recognition; the acquisition module includes: A first acquisition submodule is configured to obtain the wake-up keyword by identifying the voice content output by the wake-up object; The second acquisition submodule is configured to identify the facial region of the awakened object to obtain facial feature information, wherein the facial feature information at least includes: lip shape features when uttering lip sounds; The third acquisition submodule is configured to perform the lip reading recognition according to the lip shape feature to obtain the auxiliary recognition data.
9. The device according to claim 7, wherein The wake-up module includes: A first determining module, configured to determine a wake-up auxiliary word for waking up the target device according to the auxiliary recognition data; The first wake-up submodule is configured to wake up the target device using the wake-up auxiliary word.
10. The device according to claim 7, wherein The device further comprises: A fourth acquisition submodule is configured to use the target device to perform sound reception processing on the first voice content output by the wake-up object to obtain first sound reception data; a fifth acquisition submodule, configured to identify a first lip shape feature of the awakened object when outputting the first voice content, and obtain first lip reading data; a sixth acquisition submodule, configured to respectively identify first text content corresponding to the first collected sound data and second text content corresponding to the first lip reading data, to obtain a first recognition result; The second determination module is used to determine whether to continuously collect the first voice content output by the wake-up object according to the first recognition result.
11. The device according to claim 10, wherein The second determining module includes: a third determining module, configured to determine to continue collecting the first voice content until the voice recognition process ends if the first recognition result indicates that the first text content and the second text content are consistent; The fourth determination module is configured to determine that there is no need to continuously collect the first voice content if the recognition result indicates that the first text content and the second text content are inconsistent, and to search for other speaking objects other than the wake-up object within a predetermined range.
12. The device according to claim 11, wherein The device further comprises: a seventh acquisition submodule, configured to perform sound collection processing on the second speech content output by the other speaker to obtain second sound collection data; an eighth acquisition submodule, configured to identify second lip shape features of the other speaker when outputting the second voice content, and obtain second lip reading data; a ninth acquisition submodule, configured to respectively identify third text content corresponding to the second collected sound data and fourth text content corresponding to the second lip reading data, to obtain a second recognition result; The fifth determining module is configured to determine to continuously collect the second voice content output by the other speaker until the voice recognition process is terminated if the second recognition result indicates that the third text content is consistent with the fourth text content.
13. The device according to claim 12, wherein The device further comprises: a sixth determining module, configured to display a plurality of text contents to be selected on a display interface if it is determined that both the first recognition result and the second recognition result indicate inconsistency after a predetermined period of time, wherein the text contents to be selected include: the first text content, the second text content, the third text content, and the fourth text content; a selection module, configured to select a correct expression from the plurality of text contents to be selected in response to a click operation of the awakened object or the other speaking object; A control module is used to control the target device to perform a feedback operation according to the correct expression content.
14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the smart device interaction method according to any one of claims 1 to 6.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the smart device interaction method according to any one of claims 1 to 6. 16 . A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the smart device interaction method according to claim 1 .
17. A smart device interactive product, comprising the electronic device according to claim 14.
Citation Information
Patent Citations
Voice recognition method, mobile terminal and computer readable storage medium
CN107799125A
Sound awakening method and device, storage medium and electrical equipment
CN111651135A