Method and system for monitoring voice wake-up, electronic device and readable storage medium

CN117116257BActive Publication Date: 2026-09-04PATEO CONNECT (NANJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210534522.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2026-09-04
Estimated Expiration
2042-05-17

Smart Images

  • Figure CN117116257B_ABST
    Figure CN117116257B_ABST
Patent Text Reader

Abstract

This application discloses a method, system, electronic device, and readable storage medium for monitoring voice wake-up, relating to the field of automotive technology. The method includes: in response to determining that a user's real-time lip movement information includes preset lip movement information corresponding to a preset wake-up word, determining time information corresponding to the real-time lip movement information and the preset lip movement information; in response to determining that the user's real-time voice information does not include the preset voice information corresponding to the preset wake-up word, extracting a voice segment from the real-time voice information that failed to wake up based on the time information. This application, by determining the real-time lip movement information and the time information corresponding to the preset lip movement information when the user speaks the preset wake-up word, can extract the voice segment that failed to wake up based on the time information when wake-up fails. Since the above voice segment is directly extracted from real-time voice information, it is authentic and reliable, and has significant reference value for subsequent improvements in wake-up rates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automotive technology, and more particularly to methods and systems for monitoring voice wake-up, electronic devices, and readable storage media. Background Technology

[0002] The Internet of Vehicles (IOV) is a network that connects cars. In the IOV, cars form a vehicle network, which is connected to the internet. Based on a unified protocol, these three entities enable data exchange between people, vehicles, roads, and the network, ultimately achieving functions such as intelligent transportation, intelligent vehicles, and intelligent driving.

[0003] Taking intelligent driving as an example, to avoid affecting the driver's control of the steering wheel during driving, the driver can activate the central control voice system by speaking a specific wake-up word to achieve voice recognition interaction and control programs such as broadcasting and navigation. However, in some cases, the wake-up word may fail to activate the central control voice system, thus preventing voice recognition interaction. To obtain voice data of wake-up failures, related technologies usually need to simulate wake-up failure scenarios afterward. However, the voice data obtained in this way differs significantly from the actual wake-up failure voice data, offering little reference value for subsequent improvements in wake-up rates.

[0004] Application content

[0005] The method, system, electronic device, and readable storage medium for monitoring voice wake-up provided in this application can solve or partially solve the deficiencies in the prior art or other deficiencies in the prior art.

[0006] The method for monitoring voice wake-up provided in the first aspect of this application may include:

[0007] In response to determining that the user's real-time lip movement information includes preset lip movement information corresponding to a preset wake word, the time information corresponding to the real-time lip movement information and the preset lip movement information is determined; and

[0008] In response to determining that the user's real-time voice information does not include preset voice information corresponding to the preset wake-up word, the voice segment that failed to wake up is extracted from the real-time voice information based on the time information.

[0009] The system for monitoring voice wake-up provided according to the second aspect of this application may include:

[0010] The voice acquisition unit collects real-time voice information.

[0011] The image acquisition unit collects real-time lip movement information;

[0012] The wake-up unit is configured to perform the method for monitoring voice wake-up described in the first aspect above.

[0013] The electronic device provided according to the third aspect of this application may include:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the voice wake-up monitoring method provided in the first aspect of this application.

[0016] The computer-readable storage medium provided in the fourth aspect of this application stores a computer program, which, when executed by a processor, implements the method for monitoring voice wake-up provided in the first aspect of this application.

[0017] According to the method, system, electronic device, and readable storage medium for monitoring voice wake-up provided in this application, by determining the real-time lip movement information and the corresponding time information when the user speaks a preset wake-up word, the speech segment that failed to wake up can be extracted from the real-time speech information based on the time information when wake-up fails. Since the aforementioned speech segment is directly extracted from the real-time speech information, it is authentic and reliable, and has great reference value for subsequent improvement of the wake-up rate.

[0018] It should be understood that the description in this section is not intended to identify key or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0019] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this application. Wherein:

[0020] Figure 1 This is an exemplary system architecture diagram in which embodiments of this application can be applied;

[0021] Figure 2 This is a flowchart illustrating a method for monitoring voice wake-up according to an embodiment of this application;

[0022] Figure 3 This is a flowchart illustrating step S100 of the method for monitoring voice wake-up according to an embodiment of this application;

[0023] Figure 4 This is another flowchart illustrating step S100 in the method for monitoring voice wake-up according to an embodiment of this application.

[0024] Figure 5 This is a flowchart illustrating step S200 of the method for monitoring voice wake-up according to an embodiment of this application;

[0025] Figure 6 This is another flowchart illustrating a method for monitoring voice wake-up according to an embodiment of this application;

[0026] Figure 7 This is a block diagram of an electronic device used to implement the voice wake-up monitoring method of the embodiments of this application.

[0027] Figure label:

[0028] 100. System Architecture; 101. Image Acquisition Unit; 102. Voice Acquisition Unit;

[0029] 103. Wake-up unit; 104. Cloud server; 105. Network; 200. Electronic device;

[0030] 201. Computing unit; 202. Memory (ROM); 203. Memory (RAM);

[0031] 204. Bus; 205. I / O interface; 206. Input unit; 207. Output unit;

[0032] 208. Storage unit; 209. Communication unit. Detailed Implementation

[0033] The exemplary embodiments of this application are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0034] It should be noted that, where there is no conflict, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0035] Figure 1 An exemplary system architecture 100 is shown, in which the method for monitoring voice wake-up of this application can be applied.

[0036] like Figure 1As shown, the system architecture 100 may include an image acquisition unit 101, a voice acquisition unit 102, a wake-up unit 103, a cloud server 104, and a network 105. The network 105 is a medium for providing communication links between the image acquisition unit 101 and the wake-up unit 103, between the voice acquisition unit 102 and the wake-up unit 103, and between the wake-up unit 103 and the cloud server 104. The network 105 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0037] It should be noted that the cloud server 104 can be hardware, software, or a combination of both. When the cloud server 104 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the cloud server 104 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module. No specific limitations are made here.

[0038] It should be noted that the method for monitoring voice wake-up provided in this application is generally executed by the wake-up unit 103. Furthermore, it should be understood that... Figure 1 The number of image acquisition unit 101, voice acquisition unit 102, wake-up unit 103, cloud server 104, and network 105 is merely illustrative. Depending on implementation needs, any number of these units can be used. Furthermore, although the above units are schematically divided according to function in the figure, in practical applications, some units can be further divided or merged as needed. For example, in one embodiment, image acquisition unit 101 and voice acquisition unit 102 can be merged into a single unit, i.e., this unit can acquire both images and voice. In an example where image acquisition unit 101 and voice acquisition unit 102 are separately configured, image acquisition unit 101 could be, for example, a high-definition front-facing camera installed in the vehicle, and voice acquisition unit 102 could be, for example, a vehicle microphone.

[0039] In daily life, when driving, people expect to be able to immediately activate the central control voice system for voice recognition and interaction after saying a wake-up word. However, due to noise, the user's accent, and other factors while driving, the system may not wake up after the user says the wake-up word, thus affecting the user experience. To improve the wake-up rate, it is necessary to collect voice data from failed wake-up attempts, and this voice data must be authentic and reliable.

[0040] Based on this, the embodiments of this application provide a method for monitoring voice wake-up. Figure 2An exemplary flowchart of a method for monitoring voice wake-up according to an embodiment of this application is shown. The method for monitoring voice wake-up includes the following steps:

[0041] S100, in response to determining that the user’s real-time lip movement information includes preset lip movement information corresponding to a preset wake word, determine the time information corresponding to the real-time lip movement information and the preset lip movement information;

[0042] S200: In response to determining that the user's real-time voice information does not include preset voice information corresponding to the preset wake-up word, extract the voice segment that failed to wake up from the real-time voice information based on the time information.

[0043] When the vehicle starts, its image acquisition unit 101, voice acquisition unit 102, and wake-up unit 103 start simultaneously. The image acquisition unit 101 continuously acquires real-time lip movement information of occupants, such as the driver, while the voice acquisition unit 102 simultaneously and continuously acquires real-time voice information of occupants. If the driver utters a preset wake-up word, the wake-up unit 103 determines the time information corresponding to the driver's real-time lip movement information and the preset lip movement information. Simultaneously, if the driver utters the preset wake-up word but fails to wake up the vehicle's central control voice system, it indicates a wake-up failure. The wake-up unit 103 can then extract the failed wake-up voice segment from the real-time voice information based on the aforementioned time information.

[0044] As can be seen, the voice wake-up monitoring method of this application, by determining the real-time lip movement information and the corresponding time information when the user speaks a preset wake-up word, can extract the speech segment of the wake-up failure from the real-time speech information based on the time information when wake-up fails. Since the above speech segment is directly extracted from the real-time speech information, it is authentic and reliable, and has great reference value for subsequent improvement of wake-up rate.

[0045] It should be noted that the aforementioned time information may include, but is not limited to, the start and end times corresponding to the real-time lip movement information and the preset lip movement information. For example, the aforementioned time information may also include the duration of the lip movement corresponding to the real-time lip movement information and the preset lip movement information.

[0046] The following is a detailed description of each step of the method for monitoring voice wake-up in the embodiments of this application.

[0047] In step S100, in response to determining that the user's real-time lip movement information includes preset lip movement information corresponding to a preset wake-up word, the time information corresponding to the real-time lip movement information and the preset lip movement information is determined. In some examples, before performing step S100, the method further includes: step S000, acquiring real-time lip movement information and real-time voice information. The real-time lip movement information can be acquired through an image acquisition unit, such as a high-definition front-facing camera installed in the vehicle, and the real-time voice information can be acquired through a voice acquisition unit, such as a vehicle microphone.

[0048] It should be noted that whether the real-time lip movement information includes the preset lip movement information can be determined by comparing the lip movement similarity between the real-time lip movement information and the preset lip movement information. The lip movement similarity comparison can be performed locally, such as in the vehicle system or wake-up unit 103, or it can be sent to the cloud server 104 via the network 105.

[0049] If the lip movement similarity comparison is performed locally, then, as Figure 3 As shown, step S100 according to an embodiment of this application may include:

[0050] S110. Determine the initial time for acquiring real-time lip movement information as the start time;

[0051] S120. Starting from the beginning, compare the real-time lip movement information with the preset lip movement information to obtain the lip movement similarity ratio.

[0052] S130. In response to determining that the lip movement similarity ratio is not less than the first preset threshold, the moment when the lip movement similarity ratio is not less than the first preset threshold is determined as the termination moment.

[0053] S140. In response to determining that the lip movement similarity ratio is less than a first preset threshold, the determined start time is updated to the time when the lip movement similarity ratio is less than the first preset threshold.

[0054] For example, the preset wake-up word is "Hello Xiaoling," and the real-time voice message is "Hello Xiaoling, please open the navigation map." Since the real-time voice message and real-time lip movement information are acquired synchronously, the moment the user says "you," i.e., the initial moment when the user's real-time lip movement information is acquired, can be defined as the start time t1. Starting from the start time t1, the lip movement similarity between the real-time lip movement information and the preset lip movement information is compared. When lip movement information corresponding to "Ling" is detected, the lip movement similarity ratio obtained is exactly not less than a first preset threshold, such as 90%, then this moment can be defined as the end time t2. Similarly, if the preset wake-up word is still "Hello Xiaoling," and the real-time voice message is "Hello, please open the navigation map," then the moment the user says "you," i.e., the initial moment when the user's real-time lip movement information is acquired, can be defined as the start time t1. Starting from the beginning, the lip movement information in real time is compared with the preset lip movement information. If the lip movement similarity ratio is still less than the first preset threshold until the lip movement information corresponding to the "image" is detected, it means that the user has not said the preset wake word, and the beginning time t1 can be updated to this time.

[0055] If the lip movement similarity comparison is performed on cloud server 104, then, as Figure 4 As shown, step S100 according to the embodiments of this application may include:

[0056] S110. Determine the initial time for acquiring real-time lip movement information as the start time;

[0057] S120. Starting from the beginning, the real-time lip movement information is uploaded to the cloud server in real time, so that the cloud server can compare the real-time lip movement information with the preset lip movement information to obtain the lip movement similarity ratio.

[0058] S130. In response to determining that the lip movement similarity ratio received from the cloud server is not less than a first preset threshold, the moment when the lip movement similarity ratio is not less than the first preset threshold is determined as the termination moment.

[0059] S140. In response to the lip movement similarity ratio received from the cloud server being less than a first preset threshold, the determined start time is updated to the time when the lip movement similarity ratio is less than the first preset threshold.

[0060] In some embodiments, the first preset threshold may be, but is not limited to, 80% to 100%.

[0061] The following example uses the preset wake-up word "Hello Xiaoling" and the real-time voice message "Hello Xiaoling, please open the navigation map." Since the real-time voice message and real-time lip movement message are acquired synchronously, the moment the user says "you," i.e., the initial moment when the user's real-time lip movement message is acquired, can be defined as the start time t1. From start time t1, the acquired real-time lip movement message is uploaded to the cloud server 104 in real time. After receiving the real-time lip movement message, the cloud server 104 compares the real-time lip movement message with the preset lip movement message for lip movement similarity and sends the comparison result, i.e., the lip movement similarity ratio, back to the wake-up unit. When the cloud server 104 receives the lip movement message corresponding to "Ling," the lip movement similarity ratio obtained is exactly not less than the first preset threshold. Upon receiving this lip movement similarity ratio, the wake-up unit can define this moment as the end time t2. Similarly, if the preset wake-up word is still "Hello Xiaoling" and the real-time voice message is "Hello, please open the navigation map," the moment the user says "you," i.e., the initial moment when the user's real-time lip movement message is acquired, can be defined as the start time t1. If the lip movement similarity ratio is still less than the first preset threshold until the cloud server 104 receives the lip movement information corresponding to the "image", it means that the user has not said the preset wake-up word. The wake-up unit can then update the start time t1 to the preset time based on the lip movement similarity ratio.

[0062] As can be seen from the above, lip movement similarity comparison can be performed either locally or on a cloud server. However, the latter is more accurate than the former. But because the latter requires transmitting real-time lip movement information and other data between the local and cloud servers, it takes longer.

[0063] In step S200, whether the real-time voice information includes the preset voice information can be determined by comparing the voice similarity between the real-time voice information and the preset voice information. The voice similarity comparison can be performed locally or sent to a cloud server.

[0064] Taking local speech similarity comparison as an example, the comparison is determined based on the current local wake-up model. Specifically: the speech similarity between real-time speech information and a preset wake-up word is determined based on the current wake-up model to obtain a speech similarity ratio; if the speech similarity ratio is less than a second preset threshold, it is determined that the real-time speech information does not include the preset speech information corresponding to the preset wake-up word. Conversely, if the speech similarity ratio is not less than the second preset threshold, it is determined that the real-time speech information includes the preset speech information corresponding to the preset wake-up word. The wake-up model is a recognized and applicable model, such as a conventional convolutional neural network (CNN) or other deep learning models. These models can be trained locally or sent to a cloud server for remote training, which will be further described later.

[0065] In some embodiments, the second preset threshold may be, but is not limited to, 80% to 100%.

[0066] like Figure 5 As shown, when the time information includes the start and end times corresponding to real-time lip movement information and preset lip movement information, step S200 may include:

[0067] S210. Calculate the difference between the termination time and the start time to obtain the backtracking duration of the speech segment;

[0068] S220. Determine the data length of the speech segment based on the backtracking duration;

[0069] S230. Extract the speech segment of data length backward from the real-time speech information starting from the termination time.

[0070] The following example uses the preset wake-up word "Hello Xiaoling" and the real-time voice message "Hello Xiaoling, please open the navigation map." Assuming that due to noise in the car or surrounding environment, or the user's accent, the user fails to wake up the central control voice system after saying "Hello Xiaoling, please open the navigation map," meaning the real-time voice message is considered to exclude the preset voice message (which corresponds to the preset wake-up word "Hello Xiaoling"), the backtracking duration Δt can be determined based on the start time t1 and end time t2, where Δt = t2 - t1. The data length of the voice segment that failed to wake up can be determined based on the backtracking duration Δt. For example, if the length of the real-time voice message "Hello Xiaoling, please open the navigation map" is 1000k, the data length of the voice segment that failed to wake up can be 500k. Since, as mentioned above, the end time t2 is the time when lip movement information corresponding to "Ling" is detected, the voice data "Hello Xiaoling" that failed to wake up can be obtained by backtracking from the real-time voice message by the above-mentioned data length.

[0071] Furthermore, in order to improve the wake-up rate by utilizing the extracted speech segments, such as Figure 6 As shown, the method for monitoring voice wake-up also includes the following steps:

[0072] S300, Obtain the wake-up model trained based on the speech segment;

[0073] S400: Update the current wake-up model to the trained wake-up model.

[0074] It should be noted that the wake-up model can be trained locally or sent to the cloud server 104 for remote training. Taking remote training as an example, step S300 specifically includes:

[0075] S310. Upload the voice segment to the cloud server 104 so that the cloud server 104 can train the previously stored current wake-up model based on the voice segment to obtain the trained wake-up model.

[0076] S320: Obtain the trained wake-up model from cloud server 104.

[0077] For example, the wake-up unit 103 can upload the extracted wake-up failure speech segments to the cloud server 104. The cloud server 104 then trains the currently stored wake-up model based on the received speech segments. After training, the wake-up rate of the current wake-up model improves. The cloud server 104 then sends the trained wake-up model back to the wake-up unit 103, which can then use this trained wake-up model to wake up the central control voice system. That is, it determines whether the next received real-time voice information includes preset voice information corresponding to the preset wake-up word based on the trained wake-up model. At the same time, the cloud server 104 stores the trained wake-up model as the current wake-up model for the next model training. Thus, each time the cloud server 104 receives a speech segment, it retrains the previously trained wake-up model. Each training iteration improves the wake-up rate of the wake-up model, resulting in a significant increase in the wake-up rate after repeated training by the cloud server 104.

[0078] Additionally, considering that the failure to wake up the central control voice system may be due to the user's voice being too soft, there is no need to extract the voice segment that caused the wake-up failure in this case. Therefore, the following steps are included before extracting the voice segment that caused the wake-up failure:

[0079] S201. Compare the volume of real-time voice information with the preset volume starting from the start time;

[0080] S202. In response to the determination that the volume of the real-time voice information from the start time to the end time is always less than the preset volume, the real-time voice information and real-time lip movement information are reacquired.

[0081] Assuming the start time determined by step S100 is t1 and the end time is t2, then the relative volume of the real-time voice information can be continuously compared with the preset volume starting from the start time t1. If the volume of the real-time voice information is lower than the preset volume from the start time t1 to the end time t2, it means that the user's voice is very low and almost indistinguishable, thus failing to wake up the central control voice system. Therefore, in this case, there is no need to extract the voice segment that failed to wake up, and we can continue to wait for the user to issue a command again, that is, to reacquire the real-time voice information and real-time lip movement information.

[0082] It should be noted that, according to one or more embodiments of this application, after acquiring real-time voice information, the acquired real-time voice information can be cached; in response to the failure to acquire real-time voice information after a preset time, indicating that the user has not spoken for a long time, the cached voice segment can be uploaded to the cloud server 104, so that the cloud server 104 can train the previously stored current wake-up model based on the voice segment. Of course, to avoid occupying local data, voice segments that have already been uploaded to the cloud server 104 can be deleted.

[0083] In addition, this application also provides a system for monitoring voice wake-up, which includes a voice acquisition unit, an image acquisition unit, and a wake-up unit 103. The voice acquisition unit is used to acquire real-time voice information, the image acquisition unit is used to acquire real-time lip movement information, and the wake-up unit 103 is configured to execute the above-described method for monitoring voice wake-up.

[0084] In some embodiments, the system for monitoring voice wake-up also includes a cloud server 104, which is communicatively connected to the wake-up unit 103 and stores the current wake-up model; wherein the cloud server 104 is configured to receive voice segments from the wake-up unit 103 and train the current wake-up model based on the voice segments.

[0085] In some embodiments, the system for monitoring voice wake-up also includes a cloud server 104, which is communicatively connected to the wake-up unit 103. The cloud server 104 is configured to receive real-time lip movement information from the wake-up unit 103 from the start time, and compare the real-time lip movement information with preset lip movement information to obtain a lip movement similarity ratio.

[0086] In addition, this application also provides an electronic device, which includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-described method for monitoring voice wake-up.

[0087] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for monitoring voice wake-up.

[0088] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0089] like Figure 7 As shown, the electronic device 200 includes a computing unit 201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 202 or a computer program loaded from a storage unit 208 into a random access memory (RAM) 203. The RAM 203 may also store various programs and data required for the operation of the electronic device 200. The computing unit 201, ROM 202, and RAM 203 are interconnected via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.

[0090] Multiple components in electronic device 200 are connected to I / O interface 205, including: input unit 206, such as keyboard, mouse, etc.; output unit 207, such as various types of displays, speakers, etc.; storage unit 208, such as disk, optical disk, etc.; and communication unit 209, such as network card, modem, wireless transceiver, etc. Communication unit 209 allows electronic device 200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0091] The computing unit 201 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 201 performs the various methods and processes described above, such as the method of monitoring voice wake-up. For example, in some embodiments, the method of monitoring voice wake-up can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 208. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 200 via ROM 202 and / or communication unit 209. When the computer program is loaded into RAM 203 and executed by the computing unit 201, one or more steps of the method of monitoring voice wake-up described above can be performed. Alternatively, in other embodiments, the computing unit 201 can be configured to perform the method of monitoring voice wake-up by any other suitable means (e.g., by means of firmware).

[0092] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0093] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0094] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0096] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0097] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0098] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0099] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for monitoring voice wake-up, characterized in that, include: In response to determining that the user’s real-time lip movement information includes preset lip movement information corresponding to a preset wake word, the time information corresponding to the real-time lip movement information and the preset lip movement information is determined; as well as In response to determining that the user’s real-time voice information does not include preset voice information corresponding to the preset wake-up word, the voice segment that failed to wake up is extracted from the real-time voice information according to the time information; in, The time information for determining the real-time lip movement information and the preset lip movement information includes: The initial time at which the real-time lip movement information is acquired is determined as the start time; Starting from the aforementioned start time, the real-time lip movement information and the preset lip movement information are compared to obtain a lip movement similarity ratio. In response to determining that the lip movement similarity ratio is not less than a first preset threshold, the moment when the lip movement similarity ratio is not less than the first preset threshold is determined as the termination moment; and In response to determining that the lip movement similarity ratio is less than the first preset threshold, the determined start time is updated to the time when the lip movement similarity ratio is less than the first preset threshold.

2. The method for monitoring voice wake-up according to claim 1, wherein, The real-time voice information does not include the preset voice information corresponding to the preset wake-up word, which is determined based on the current wake-up model; The method for monitoring voice wake-up also includes: Obtain the wake-up model trained based on the speech segment; and Update the current wake-up model to the trained wake-up model.

3. The method for monitoring voice wake-up according to claim 2, wherein, Obtaining the wake-up model trained based on the speech segment includes: The speech segment is uploaded to a cloud server, so that the cloud server can train the previously stored current wake-up model based on the speech segment to obtain the trained wake-up model; and The trained wake-up model is obtained from the cloud server.

4. The method for monitoring voice wake-up according to claim 2, wherein, Determining that the user's real-time voice information does not include the preset voice information corresponding to the preset wake word includes: Based on the current wake-up model, the speech similarity between the real-time speech information and the preset wake-up word is determined to obtain the speech similarity ratio; and In response to the voice similarity ratio being less than a second preset threshold, it is determined that the real-time voice information does not include preset voice information corresponding to the preset wake-up word.

5. The method for monitoring voice wake-up according to claim 1, wherein, The method further includes: The system acquires the real-time voice information in real time and caches the real-time voice information; and If the real-time voice information is not obtained after a preset time period, the cached voice segment is uploaded to the cloud server.

6. The method for monitoring voice wake-up according to any one of claims 1 to 5, wherein, The time information includes the start and end times corresponding to the real-time lip movement information and the preset lip movement information.

7. The method for monitoring voice wake-up according to claim 6, wherein, Extracting the wake-up failure audio segment from the real-time voice information based on the time information includes: Calculate the difference between the termination time and the start time to obtain the backtracking duration of the speech segment; The data length of the speech segment is determined based on the backtracking duration; and Starting from the termination time, the audio segment of the specified data length is extracted by backtracking from the real-time audio information.

8. The method for monitoring voice wake-up according to claim 6, wherein, Before determining the time information corresponding to the real-time lip movement information and the preset lip movement information, the method further includes: The real-time lip movement information and the real-time voice information are acquired in real time.

9. The method for monitoring voice wake-up according to claim 8, wherein, The time information for determining the real-time lip movement information and the preset lip movement information includes: The initial time at which the real-time lip movement information is acquired is determined as the starting time. Starting from the aforementioned start time, the real-time lip movement information is uploaded to the cloud server in real time, so that the cloud server compares the real-time lip movement information with the preset lip movement information to obtain a lip movement similarity ratio. In response to determining that the lip movement similarity ratio received from the cloud server is not less than a first preset threshold, the moment when the lip movement similarity ratio is not less than the first preset threshold is determined as the termination moment; and In response to the lip movement similarity ratio received from the cloud server being less than the first preset threshold, the determined start time is updated to the time when the lip movement similarity ratio is less than the first preset threshold.

10. The method for monitoring voice wake-up according to claim 8, wherein, Before extracting the wake-up-failed speech segment from the real-time speech information based on the time information, the method further includes: Compare the volume of the real-time voice information with the preset volume starting from the stated start time; and In response to determining that the volume of the real-time voice information from the start time to the end time is always less than the preset volume, the real-time voice information and real-time lip movement information are reacquired.

11. A system for monitoring voice wake-up, characterized in that, include: The voice acquisition unit collects real-time voice information. The image acquisition unit collects real-time lip movement information; The wake-up unit is configured to perform the method of monitoring voice wake-up as described in any one of claims 1 to 10.

12. The system for monitoring voice wake-up according to claim 11, wherein, The system also includes: A cloud server is communicatively connected to the wake-up unit and stores the current wake-up model; wherein the cloud server is configured to receive the voice segment from the wake-up unit and train the current wake-up model based on the voice segment.

13. The system for monitoring voice wake-up according to claim 11, wherein, The system also includes: A cloud server is communicatively connected to the wake-up unit; wherein the cloud server is configured to receive real-time lip movement information from the wake-up unit from the start time, and compare the real-time lip movement information with preset lip movement information to obtain a lip movement similarity ratio.

14. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the monitoring voice wake-up method according to any one of claims 1 to 10.

15. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for monitoring voice wake-up as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method for training wake-up model and device thereof

    CN111667818A

  • Man-machine conversation method, electronic equipment and computer readable storage medium

    CN112634911A