Display device and voice interaction method

By integrating a camera into the display device for face and lip movement detection, combined with audio signal processing, the problem of noise interference in far-field voice interaction is solved, improving the accuracy of voice recognition and user experience.

CN114299940BActive Publication Date: 2025-12-05HISENSE VISUAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110577525.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-26
Publication Date
2025-12-05
Estimated Expiration
2041-05-26

AI Technical Summary

Technical Problem

In far-field voice interaction scenarios, noise interference can lead to a decrease in wake-up rate and speech recognition accuracy, resulting in a poor user experience.

Method used

By integrating a camera into the display device, images are captured and facial information and lip movement are detected. Combined with audio signal processing, sound source localization and voiceprint recognition are performed, eliminating interference from non-target people and responding only to the real-time voice commands of the target person.

Benefits of technology

It improves speech recognition accuracy, reduces the probability of noise sampling, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299940B_ABST
    Figure CN114299940B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a display device and a voice interaction method. The display device includes a display configured to present a user interface; and a controller connected to the display. The controller is configured to: acquire user identity information of a target person, and collect a voice real-time instruction, the target person including a person who issues the wake-up instruction or a registered user; detect face information in an image collected by a camera; if the face information is face information of the target person, perform face tracking and lip movement detection on the target person; if the face of the target person has lip movement, and the voice real-time instruction includes a voice of the target person, respond to the voice real-time instruction; and if the face of the target person does not have lip movement, or the voice real-time instruction does not include the voice of the target person, do not respond to the voice real-time instruction. The present application solves the technical problem of poor voice interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice interaction, and in particular to a display device and a voice interaction method. BACKGROUND

[0002] With the rise of smart home, controlling smart TV and other home devices through voice interaction has become a more and more popular control method. The wake-up rate and the voice recognition accuracy are two important indicators affecting the user experience of voice interaction. In the early stage of the development of voice interaction technology, voice interaction is usually near-field interaction. In the near-field voice interaction scene, the distance between the human and the machine is small, the influence of noise interference is small, and the wake-up rate and the voice recognition accuracy are high. However, when people watch TV, they are usually far away from the TV, and the near-field interaction cannot meet the needs of people. In order to improve the convenience of voice interaction, far-field voice interaction technology emerges as the times require. In the far-field voice interaction scene, the distance between the human and the machine is large, and the influence of noise interference is also large, which leads to the decrease of the wake-up rate and the voice recognition accuracy, resulting in poor voice interaction experience. SUMMARY

[0003] To solve the technical problem of poor voice interaction experience, the present application provides a display device and a voice interaction method.

[0004] In a first aspect, the present application provides a display device, which comprises:

[0005] a display for presenting a user interface;

[0006] a camera for collecting images;

[0007] a controller connected to the display, the controller being configured to:

[0008] collect a voice wake-up instruction;

[0009] in response to the voice wake-up instruction, obtain user identity information of a target person, and collect a voice real-time instruction, the target person including a person issuing the wake-up instruction or a registered user;

[0010] detect face information in the images collected by the camera;

[0011] if the face information of the target person is detected, perform face tracking and lip movement detection on the target person, if the face of the target person has lip movement, and the voice real-time instruction includes the voice of the target person, respond to the voice real-time instruction;

[0012] if the face of the target person does not have lip movement, or the voice real-time instruction does not include the voice of the target person, do not respond to the voice real-time instruction.

[0013] In some embodiments, the face information in the image captured by the camera is detected, including:

[0014] The voice wake-up instruction is subjected to sound source positioning to obtain a wake-up sound source position;

[0015] The camera is rotated towards the wake-up sound source position, and in the rotation process, the face information of the target person is detected in the image captured by the camera, and if the face information of the target person is detected, the rotation of the camera is stopped.

[0016] In some embodiments,

[0017] The face tracking and lip movement detection of the target person are performed, including:

[0018] The real-time coordinate range of the target face in the image captured by the camera is obtained;

[0019] According to the change trend of the real-time coordinate range, the camera is controlled to rotate so that the face of the target person is located in a preset area in the image captured by the camera;

[0020] The image of the target face is subjected to lip movement detection.

[0021] In a second aspect, the present application provides a voice interaction method, including:

[0022] A voice wake-up instruction is collected;

[0023] In response to the voice wake-up instruction, user identity information of a target person is obtained, and a voice real-time instruction is collected, the target person including a person who issues the wake-up instruction or a registered user;

[0024] Face information in an image captured by a camera is detected;

[0025] If the face information of the target person is detected, face tracking and lip movement detection of the target person are performed, if the face of the target person has lip movement, and the voice real-time instruction includes the voice of the target person, the voice real-time instruction is responded to;

[0026] If the face of the target person does not have lip movement, or the voice real-time instruction does not include the voice of the target person, the voice real-time instruction is not responded to.

[0027] The display device and the voice interaction method provided by the present application have the following beneficial effects:

[0028] The display device provided in the application can exclude the interference of non-target persons by tracking the face of the target person corresponding to the user identity information after receiving the voice wake-up instruction; the display device can respond to the voice real-time instruction again after tracking the face of the target person, and the target person has lip movement, and the received voice real-time instruction includes the voice of the target person, and the display device does not respond to the voice real-time instruction when the target person does not have lip movement, or the voice real-time instruction does not include the voice of the target person, thereby reducing the probability of noise collection, improving the voice recognition accuracy, and improving the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the application, the drawings needed in the embodiments will be briefly introduced below. Obviously, other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0030] Figure 1 Fig. 1 exemplarily shows a schematic diagram of an operation scenario between a display device and a control device according to some embodiments;

[0031] Figure 2 Fig. 2 exemplarily shows a hardware configuration block diagram of the control device 100 according to some embodiments;

[0032] Figure 3 Fig. 3 exemplarily shows a hardware configuration block diagram of the display device 200 according to some embodiments;

[0033] Figure 4 Fig. 4 exemplarily shows a software configuration schematic diagram in the display device 200 according to some embodiments;

[0034] Figure 5 Fig. 5 exemplarily shows a schematic diagram of a voice interaction principle according to some embodiments;

[0035] Figure 6 Fig. 6 exemplarily shows a scene schematic diagram of voice interaction according to some embodiments;

[0036] Figure 7 Fig. 7 exemplarily shows a signal processing schematic diagram of voice interaction according to some embodiments;

[0037] Figure 8 Fig. 8 exemplarily shows a timing schematic diagram of voice interaction according to some embodiments;

[0038] Figure 9 Fig. 9 exemplarily shows a whole flow schematic diagram of a voice interaction method according to some embodiments;

[0039] Figure 10Fig. 2 shows a flow diagram of a method for processing voice wake-up instructions according to some embodiments;

[0040] Figure 11 Fig. 3 shows a flow diagram of a method for processing when a single person face is detected during a voice interaction according to some embodiments;

[0041] Figure 12 Fig. 4 shows a flow diagram of a method for processing when multiple person faces are detected during a voice interaction according to some embodiments. DETAILED DESCRIPTION

[0042] For the purpose of clarity and enabling embodiments of the present application to be more clearly described, the following description of the example embodiments of the present application will be made with reference to the accompanying drawings. It is apparent that the example embodiments described are only a part of the embodiments of the present application, and not all of the embodiments of the present application.

[0043] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the embodiments described next, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.

[0044] The terms "first", "second", "third", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances.

[0045] The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not necessarily limit to all components clearly listed, but can include other components not clearly listed or inherent to these products or devices.

[0046] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code capable of performing a function associated with the element.

[0047] Figure 1 Fig. 1 shows a schematic diagram of an operating scenario between a display device and a control device according to an embodiment. As shown in Fig. 1, a user can operate the display device 200 through the smart device 300 or the control device 100. Figure 1

[0048] ​In some embodiments, the control device 100 can be a remote controller, and the communication between the remote controller and the display device can be infrared protocol communication or Bluetooth protocol communication, or other short-distance communication, and the display device 200 can be controlled by the remote controller through wireless or wired communication. The user can input user instructions through the buttons on the remote controller, voice input, control panel input, etc., to control the display device 200.

[0049] In some embodiments, the smart device 300 (such as a mobile terminal, a tablet computer, a computer, a notebook computer, etc.) can also be used to control the display device 200. For example, the display device 200 can be controlled by using an application running on the smart device.

[0050] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, the display device 200 can directly receive voice instructions from the user through a voice instruction acquisition module configured inside the display device 200, or the display device 200 can receive voice instructions from the user through a voice control device configured outside the display device 200.

[0051] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 can be connected to the server 400 through a local area network (LAN), a wireless local area network (WLAN), or other networks. The server 400 can provide various content and interactions to the display device 200. The server 400 can be a cluster or multiple clusters, and can include one or more types of servers.

[0052] Figure 2 An exemplary configuration block diagram of the control device 100 according to an exemplary embodiment is shown. As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, a power supply. The control device 100 can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, acting as an intermediary between the user and the display device 200. Figure 2

[0053] An exemplary hardware configuration block diagram of the display device 200 according to an exemplary embodiment is shown. Figure 3 In some embodiments, the display device 200 includes at least one of a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.

[0054]

[0055] ​In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphic processor, a RAM, a ROM, a first interface to an n-th interface for input / output.

[0056] In some embodiments, the display 260 includes a display screen component for presenting a picture, and a driving component for driving the image display, a component for receiving the image signal originated from the output of the controller, and a component for displaying the video content, the image content, and the menu operation interface, and a user operation UI interface.

[0057] In some embodiments, the display 260 can be a liquid crystal display, an OLED display, and a projection display, and can also be a projection device and a projection screen.

[0058] In some embodiments, the communicator 220 is a component for communicating with external devices or servers according to various communication protocol types. For example, the communicator can include at least one of a Wifi module, a Bluetooth module, a wired Ethernet module, and other network communication protocol chips or near field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.

[0059] In some embodiments, the user interface can be used to receive the control signal of the control device 100 (such as an infrared remote controller, etc.).

[0060] In some embodiments, the detector 230 is used to collect signals of the external environment or interaction with the outside. For example, the detector 230 includes a light receiver for collecting ambient light intensity, or the detector 230 includes an image collector such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures, or the detector 230 includes a sound collector such as a microphone, etc., for receiving external sounds.

[0061] In some embodiments, the external device interface 240 can include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It can also be a composite input / output interface formed by multiple interfaces described above.

[0062] In some embodiments, the tuner demodulator 210 receives broadcast television signals through wired or wireless reception, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.

[0063] In some embodiments, the controller 250 and the tuner-demodulator 210 can be located in different separate devices, i.e. the tuner-demodulator 210 can also be located in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.

[0064] In some embodiments, the controller 250 controls the operation of the display device and the response to the user's operation by storing various software control programs in the memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command for selecting a UI object displayed on the display 260, the controller 250 can perform an operation related to the object selected by the user command.

[0065] In some embodiments, the object can be any one of the selectable objects, such as a hyperlink, an icon, or other operable control. The operation related to the selected object can be an operation of displaying a page connected to the hyperlink, a document, an image, etc., or an operation of executing a program corresponding to the icon.

[0066] In some embodiments, the controller includes at least one of a Central Processing Unit (CPU), a video processor, an audio processor, a Graphics Processing Unit (GPU), a RAM (Random Access Memory), a ROM (Read-Only Memory), a first interface to an n-th interface for input / output, a communication bus, etc.

[0067] The CPU processor is used to execute the operating system and application program instructions stored in the memory, and to execute various application programs, data and content according to various interactive instructions received from external input, so as to finally display and play various audio and video content. The CPU processor can include multiple processors. For example, it can include a main processor and one or more sub-processors.

[0068] In some embodiments, the graphics processor is used to generate various graphical objects, such as icons, operation menus, and user input instruction display graphics, etc. The graphics processor includes an operator that performs operations by receiving various interactive instructions from the user input, and displays various objects according to display attributes; and a renderer that renders various objects obtained based on the operator, and the rendered objects are used for display on the display.

[0069] In some embodiments, the video processor is configured to receive an external video signal, and perform video processing according to a standard codec protocol of the input signal, such as decompression, decoding, scaling, noise reduction, frame rate conversion, resolution conversion, image composition, etc., to obtain a signal that can be directly displayed on the display device 200.

[0070] In some embodiments, the video processor includes a demultiplexing module, a video decoding module, an image composition module, a frame rate conversion module, a display formatting module, etc. The demultiplexing module is configured to perform demultiplexing processing on the input audio / video data stream. The video decoding module is configured to process the demultiplexed video signal, including decoding and scaling processing, etc. The image composition module, such as an image compositor, is configured to perform superimposition and mixing processing on the video image after scaling processing and the GUI signal generated by the graphics generator according to user input or self-generation, to generate an image signal that can be displayed. The frame rate conversion module is configured to convert the input video frame rate. The display formatting module is configured to change the output signal after frame rate conversion to a signal that conforms to the display format, such as an output RGB data signal.

[0071] In some embodiments, the audio processor is configured to receive an external audio signal, and perform processing such as decompression and decoding, noise reduction, digital-to-analog conversion, and amplification processing, etc., according to a standard codec protocol of the input signal, to obtain a sound signal that can be played on a loudspeaker.

[0072] In some embodiments, the user can input a user command through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user can input a user command by inputting a specific sound or gesture, and the user input interface receives the user input command by recognizing the sound or gesture through a sensor.

[0073] In some embodiments, the "user interface" is a medium interface for interaction and information exchange between an application program or an operating system and a user, which realizes the conversion between the internal form of information and the form that the user can accept. The commonly used form of the user interface is a graphical user interface (GUI), which refers to a user interface related to computer operation displayed in a graphical manner. It can be an icon, window, control, etc. interface element displayed in the display screen of an electronic device, wherein the control can include an icon, button, menu, tab, text box, dialog box, status bar, navigation bar, Widget, etc. visual interface element.

[0074] In some embodiments, the system of the display device can include a kernel, a shell, a file system and an application. The kernel, the shell and the file system together form a basic operating system structure, which allows a user to manage files, run programs and use the system. After power on, the kernel starts, activates the kernel space, abstracts hardware, initializes hardware parameters, etc., runs and maintains virtual memory, scheduler, signals and inter-process communication (IPC). After the kernel starts, the shell and user applications are loaded. The application is compiled into machine code after starting, forming a process.

[0075] The system of the display device can include a kernel, a shell, a file system and an application. The kernel, the shell and the file system together form a basic operating system structure, which allows a user to manage files, run programs and use the system. After power on, the kernel starts, activates the kernel space, abstracts hardware, initializes hardware parameters, etc., runs and maintains virtual memory, scheduler, signals and inter-process communication (IPC). After the kernel starts, the shell and user applications are loaded. The application is compiled into machine code after starting, forming a process.

[0076] Referring to Figure 4 In some embodiments, the system is divided into four layers, from top to bottom, the application layer (referred to as "application layer" for short), the application framework layer (referred to as "framework layer" for short), the Android runtime and system library layer (referred to as "system runtime library layer" for short), and the kernel layer.

[0077] In some embodiments, at least one application program is running in the application layer, which can be a window program, a system setting program or a clock program provided by the operating system, etc.; or an application program developed by a third-party developer. In specific implementation, the application program package in the application layer is not limited to the above examples.

[0078] The framework layer provides application programming interface (API) and programming framework for the application. The application framework layer includes some pre-defined functions. The application framework layer is equivalent to a processing center, which decides the action of the application in the application layer. The application can access the resources in the system and obtain the services of the system through the API interface in the execution.

[0079] As Figure 4As shown, the application framework layer in the embodiments of the present application includes managers (Managers) and content providers (Content Provider), wherein the managers include at least one of the following modules: an activity manager (ActivityManager) for interacting with all activities running in the system; a location manager (Location Manager) for providing system location services to system services or applications; a package manager (Package Manager) for retrieving various information related to application packages currently installed on the device; a notification manager (NotificationManager) for controlling the display and clearing of notification messages; and a window manager (Window Manager) for managing icons, windows, toolbars, wallpapers and desktop components on the user interface.

[0080] In some embodiments, the activity manager is used to manage the life cycle of each application and the general navigation back function, such as controlling the exit, opening, back, etc. of the application. The window manager is used to manage all window programs, such as obtaining the size of the display screen, determining whether there is a status bar, locking the screen, intercepting the screen, controlling the display window change (such as reducing the display window, shaking the display, twisting the display, etc.) and the like.

[0081] In some embodiments, the system runtime library layer provides support for the upper layer, i.e. the framework layer. When the framework layer is used, the Android operating system runs the C / C++ library contained in the system runtime library layer to realize the functions of the framework layer.

[0082] In some embodiments, the kernel layer is the layer between hardware and software. As Figure 4 As shown, the kernel layer includes at least one of the following drivers: an audio driver, a display driver, a Bluetooth driver, a camera driver, a WIFI driver, a USB driver, an HDMI driver, a sensor driver (such as a fingerprint sensor, a temperature sensor, a pressure sensor, etc.), and a power supply driver, etc.

[0083] In some embodiments, the hardware or software architecture can be based on the above-mentioned embodiments, and in some embodiments, it can be based on other similar hardware or software architectures, as long as it can realize the technical solutions of the present application.

[0084] In order to clearly illustrate the embodiments of the present application, the following will be combined with Figure 5 A voice recognition network architecture provided by the embodiments of the present application is described.

[0085] Referring to Figure 5 , Figure 5 A voice recognition network architecture provided by the embodiments of the present application is described. Figure 5In some embodiments, the smart device is configured to receive input information and output processing results of the information. The voice recognition service device is an electronic device deployed with a voice recognition service, the semantic service device is an electronic device deployed with a semantic service, and the business service device is an electronic device deployed with a business service. The electronic device can include a server, a computer, etc. The voice recognition service, the semantic service (also referred to as a semantic engine), and the business service are web services that can be deployed on the electronic device. The voice recognition service is configured to recognize audio as text, the semantic service is configured to perform semantic analysis on the text, and the business service is configured to provide a specific service, such as a weather query service of Ink Weather, a music query service of QQ Music, etc. In one embodiment, Figure 5 In the illustrated architecture, there can be multiple entity service devices deployed with different business services, and one or more entity service devices can be grouped with one or more function services.

[0086] In some embodiments, the following is based on the architecture shown in FIG. 1. Figure 5 The process of processing input information of the smart device in the illustrated architecture is described by way of example, taking a query sentence input by voice as an example. The process can include the following three processes:

[0087] [Voice Recognition]

[0088] After receiving the query sentence input by voice, the smart device can upload audio of the query sentence to the voice recognition service device, so that the voice recognition service device recognizes the audio as text by the voice recognition service and returns the text to the smart device. In one embodiment, before uploading the audio of the query sentence to the voice recognition service device, the smart device can perform noise removal processing on the audio of the query sentence. The noise removal processing can include steps such as removing echo and environmental noise.

[0089] [Semantic Understanding]

[0090] The smart device uploads the text of the query sentence recognized by the voice recognition service to the semantic service device, so that the semantic service device performs semantic analysis on the text by the semantic service to obtain a business domain, an intent, etc. of the text.

[0091] [Semantic Response]

[0092] According to the semantic analysis result of the text of the query sentence, the semantic service device issues a query instruction to the corresponding business service device to obtain a query result given by the business service. The smart device can obtain the query result from the semantic service device and output the query result. As an example, the semantic service device can also send the semantic analysis result of the query sentence to the smart device, so that the smart device outputs a feedback sentence in the semantic analysis result.

[0093] It should be noted that, Figure 5The illustrated architecture is only an example and does not limit the scope of protection of the present application. In the embodiments of the present application, other architectures can also be used to implement similar functions, for example, all or part of the three processes can be completed by the intelligent terminal, which will not be described here.

[0094] In some embodiments, Figure 5 The illustrated intelligent device can be a display device, such as a smart television. The functions of the voice recognition service device can be implemented by the sound collector and the controller provided on the display device. The functions of the semantic service device and the service service device can be implemented by the controller of the display device or by the server of the display device.

[0095] In some embodiments, the query statement or other interactive statement input by the user to the display device through voice can be referred to as a voice instruction.

[0096] In some embodiments, the display device obtains the query result given by the service service device from the semantic service device. The display device can analyze the query result, generate response data of the voice instruction, and then control the display device to perform corresponding actions according to the response data.

[0097] In some embodiments, the display device obtains the semantic analysis result of the voice instruction from the semantic service device. The display device can analyze the semantic analysis result, generate response data, and then control the display device to perform corresponding actions according to the response data.

[0098] In some embodiments, the remote controller of the display device can be provided with a voice control button. After the user presses the voice control button on the remote controller, the controller of the display device can control the display of the display device to display a voice interaction interface and control the sound collector, such as a microphone, to collect the sound around the display device. At this time, the user can input a voice instruction to the display device.

[0099] In some embodiments, the display device can support voice wake-up function. The sound collector of the display device can be in a state of continuously collecting sound. After the user speaks the wake-up word, the display device performs voice recognition on the voice instruction input by the user. After recognizing that the voice instruction is the wake-up word, the display device can control the display of the display device to display a voice interaction interface. At this time, the user can continue to input a voice instruction to the display device. The wake-up word can be referred to as a voice wake-up instruction, and the voice instruction continuously input by the user can be referred to as a voice real-time instruction.

[0100] In some embodiments, after a user inputs a voice instruction, the sound collector of the display device can keep the state of sound collection during the display device acquires response data of the voice instruction or the display device responds according to the response data. The user can press the voice control button on the remote control to re-input the voice instruction at any time, or speak the wake-up word. At this time, the display device can end the previous voice interaction process, start a new voice interaction process according to the newly input voice instruction of the user, thereby ensuring the real-time performance of voice interaction.

[0101] In some embodiments, when the current interface of the display device is a voice interaction interface, after the display device performs voice recognition on the voice instruction input by the user, the display device obtains the text corresponding to the voice instruction. The display device or a server of the display device performs semantic understanding on the text to obtain a user intent, processes the user intent to obtain a semantic analysis result, and generates response data according to the semantic analysis result.

[0102] For example, in the voice interaction mode in which the display device starts a voice dialogue according to the received voice wake-up instruction, the user who issues the voice wake-up instruction can be referred to as a target person.

[0103] In some embodiments, the target person can also be a registered user of the display device. When the user registers on the display device, the user can input voiceprint information and a face image on the display device.

[0104] In some voice interaction scenarios, due to environmental noise interference and voice interference of non-target persons, the display device can be mistakenly woken up, or cannot accurately identify the intent of the target person according to the collected audio, which will seriously affect the voice interaction experience of the user.

[0105] To solve the above problems, embodiments of the present application show a voice interaction scheme, which can effectively improve the voice interaction experience in a complex voice interaction scenario by combining audio signal processing and video signal processing.

[0106] In some embodiments, to collect video signals required in the voice interaction process, the display device can be provided with a camera, or be connected with an external camera.

[0107] Taking the display device provided with a camera as an example, refer to Figure 6 FIG. 1 is a schematic diagram of a voice interaction scenario according to some embodiments. As shown in Figure 6 In some embodiments, the display device 200 can be provided with a camera 201, and the camera 201 can capture images. If the camera is fixed on the display device 200, it can only capture images within a certain field of view, as shown in Figure 6The image in the A region has a field of view angle of a, which is less than 180 degrees. When the user stands in the A region, the camera 201 can capture the user, and when the user stands in the B region and the C region, the camera 201 cannot capture the user, where the B region is a left region of the A region, and the C region is a right region of the A region.

[0108] In some embodiments, to expand the field of view of the camera 201, the camera 201 can be provided with a holder or other structure that can adjust the field of view of the camera. The holder can adjust the field of view of the camera 201, and the controller of the display device can be connected to the camera to control the dynamic change of the field of view of the camera through the holder, so that the field of view of the camera 201 after rotation can reach 0-180 degrees, so that the user in the range of 0-180 degrees can be captured, where 0 degrees means that the user stands on the left side of the display device 200 and is in the same plane as the display device 200, and 180 degrees means that the user stands on the right side of the display device 200 and is in the same plane as the display device 200. It can be seen that by rotating the camera, the camera can capture the user in the B region and the C region, achieving the effect that as long as the user stands in front of the display device 200, the user can be captured by rotating the camera 201.

[0109] In some embodiments, the display device captures images through an external camera, which can be a camera provided with a holder, so as to realize dynamic field of view image capture.

[0110] In some embodiments, the camera 201 itself does not have a holder, and can be installed on a holder in communication connection with the display device, and through the control of the holder by the display device, dynamic field of view image capture can also be realized.

[0111] In some embodiments, the method of combining audio signal processing and video signal processing of the display device can be seen in Figure 7 The signal processing schematic diagram of the voice interaction process according to some embodiments.

[0112] As shown in Figure 7 In some embodiments, the display device obtains an audio signal through the audio collected by the microphone, and the processing of the audio signal includes sound source positioning, user attribute recognition, voiceprint recognition, speech recognition, and noise reduction enhancement.

[0113] The sound source positioning can include determining the angle between the audio source and the display device, as shown in Figure 6If α is 90 degrees, then when the user is in area B, the wake-up angle is between 0 and 45 degrees; when the user is in area A, the wake-up angle is between 45 and 135 degrees; and when the user is in area C, the wake-up angle is between 135 and 180 degrees. Sound source localization can be achieved through various algorithms, such as the time difference of arrival method and beamforming. To achieve sound source localization, the display device can be equipped with a microphone array. The microphone array includes multiple microphones positioned at different locations on the display device, each connected to the display device's controller. The display device collects multiple audio signals through the microphone array and obtains the wake-up angle by comprehensively analyzing these signals. For example, in the time difference of arrival method, the display device can calculate the position of the audio source relative to the display device based on the time difference between the audio signals received by multiple microphones and the relative positions of the microphones. In beamforming, the display device can filter and weight the audio signals collected by each microphone to form a sound pressure distribution beam. Based on the distribution characteristics of the sound pressure beam, the position of the audio source relative to the display device can be obtained. The angle between the audio source and the display device is obtained based on the position of the audio source relative to the display device.

[0114] User attribute identification can determine user attributes such as gender and age, where age can be a range, such as 1-10 years old, 11-20 years old, 20-40 years old, 40-60 years old, etc. User attribute identification can be achieved based on a pre-trained model. By collecting a large number of audio samples with different user attributes, a model capable of predicting user attributes can be trained based on a neural network. After inputting the audio signal into the model, the user attributes can be obtained.

[0115] Noise reduction enhancement can include enhancing the speech of the target person in the audio signal and reducing noise in the audio of non-target persons. Noise reduction enhancement can be achieved through speech-oriented enhancement technology. By dynamically adjusting the speech enhancement beam to be centered on the target person, the speech of the target person is enhanced, while sounds outside the beam are suppressed.

[0116] As can be seen, the processing of audio signals enables the determination of user location, user identity, and audio content. Once the user location is determined, the camera can be rotated to quickly locate the target person. Once the user identity is determined, the system can distinguish between the target person engaging in voice interaction and other individuals. Once the audio content is determined, the user's intent can be obtained.

[0117] like Figure 7 As shown, in some embodiments, the display device obtains video signals from images captured by a camera, and the processing of the video signals includes face detection and tracking, face recognition, lip movement detection, and lip reading recognition.

[0118] During face detection and tracking, the camera can be rotated to ensure that the target person remains within the camera's field of view.

[0119] Lip movement detection can determine whether there are changes in the lips of a person captured by a camera. If the lips change, it can be determined that the person is speaking; if the lips do not change, it can be determined that the person is not speaking. If someone is speaking, it can be combined with facial recognition technology to determine whether the person speaking is the target person. If the person speaking is the target person, it can be determined that the received audio signal contains the target person's voice; if the person speaking is not the target person, it can be determined that the received audio signal does not contain the target person's voice.

[0120] Lip reading can identify the content of a speaker's speech, which can then be compared and analyzed with audio content obtained from speech recognition of the audio signal. If they match or are largely consistent, the audio content is considered to originate from a person in the image captured by the camera. Conversely, if there is a significant difference, the audio content is considered to originate from a person or environment outside the image captured by the camera. In some embodiments, lip reading can be omitted, and lip movement detection alone can be used to determine whether the audio content originates from a person in the image captured by the camera. This reduces the resource consumption of video signal processing and improves processing efficiency.

[0121] It is evident that the processing of video signals enables the tracking of the user's location and the determination of whether the user is speaking.

[0122] like Figure 7 As shown, in some embodiments, after obtaining the processing results of the audio signal and the video signal, the display device can also obtain application scenario information. The application scenario information may include preset interactive control information of the foreground application. This interactive control information may include audio acquisition control parameters and video acquisition control parameters. For example, an audio acquisition control parameter value of 1 indicates that audio data acquisition is currently possible, and a value of 0 indicates that audio data acquisition is currently not possible. Similarly, a video acquisition control parameter value of 1 indicates that video data acquisition is currently possible, and a value of 0 indicates that video data acquisition is currently not possible.

[0123] For example, when the front-end application is a video chat application, both the audio acquisition control parameter and the video acquisition control parameter are 1. The fusion decision engine can comprehensively analyze the processing results of the audio signal and the video signal based on the fact that both the audio acquisition control parameter and the video acquisition control parameter are 1, and obtain the recognition result.

[0124] When the foreground application is an online teaching application, if the display device is a student's terminal, there may be times when students are not allowed to speak. In this case, the audio acquisition control parameter can be 0. The fusion decision engine can determine whether to adopt the processing results of the audio signal and the video signal to obtain the recognition result, or determine whether to control the display device to acquire audio and video data, based on the audio acquisition control parameter and the video acquisition control parameter.

[0125] In some embodiments, the processing results of audio signals, video signals, and application scenario information are input into a feature fusion decision engine, which can then output multimodal recognition results. These multimodal recognition results may include the speech content of the audio signal, the speaker, the non-speaking speaker, and the correspondence between the speaker and the speech content. Based on this correspondence, it can be determined whether the target person has spoken, and if so, the content of their speech, thereby improving the accuracy of speech recognition and reducing the probability of false wake-up.

[0126] In some embodiments, Figure 7 The processing of audio and video signals can be performed locally by the display device, or the display device can send the audio and video signals to the server, which will process them in the cloud and then return the processing results to the display device. Alternatively, some functions can be implemented locally by the display device and some functions can be implemented by the server.

[0127] In some embodiments, the user's voiceprint and facial information are private information. The display device can be configured to display option controls for privacy functions such as voiceprint recognition and facial recognition during startup navigation and when the user uses the voice assistant function for the first time, and display a prompt to confirm the use of the camera and far-field voice. After viewing the prompt, the user can click the option control to enable the above-mentioned privacy functions to improve the voice interaction effect, or choose not to trigger the option control, thereby improving privacy and security.

[0128] To Figure 7 The signal processing procedure of the voice interaction process shown in the figure will be further described. Figure 8 A timing diagram of a voice interaction process according to some embodiments is shown. It should be noted that this timing diagram is only for... Figure 7 The illustrated timing diagram represents an exemplary signal processing procedure. In a practical embodiment, Figure 7 The signal processing procedure shown may also include other timing sequences.

[0129] See Figure 8 During voice interaction, the display device can interact with the user and the server respectively, providing the user with voice control services.

[0130] Figure 8 In some embodiments, the user can be a target person who wants to control the display device, and in addition to the target person, the microphone of the display device can also collect the voice of other users or environmental noise.

[0131] In some embodiments, the wake-up word input by the user to the display device can be a voice wake-up instruction. After receiving the voice wake-up instruction, the display device can perform sound source positioning according to the voice wake-up instruction, calculate the wake-up sound source position, and obtain wake-up sound source position information, which can include a wake-up angle.

[0132] In some embodiments, after calculating the wake-up angle, the display device can turn on the camera and turn the camera towards the wake-up sound source according to the wake-up angle. During the turning process, continuous shooting can be performed to obtain images in the dynamic field of view. The wake-up sound source is the target person. Of course, if there is no user inputting the wake-up word to the display device, i.e., the display device is mistakenly woken up, the wake-up sound source can not actually exist.

[0133] In some embodiments, after collecting the voice wake-up instruction, to avoid that the current voice wake-up instruction is a mistakenly woken-up voice instruction, the display device can obtain user identity information of the voice wake-up instruction, and then detect the target person in the collected images in the dynamic field of view. If the target person can be located in the collected images, it can be determined that the current wake-up is not a mistaken wake-up. If the target person cannot be located, it can be determined that the current wake-up is a mistaken wake-up.

[0134] In some embodiments, to obtain the user identity information, the display device can generate a wake-up word recognition request containing the wake-up word voice, and send the wake-up word recognition request to the server. After receiving the wake-up word recognition request, the server extracts the wake-up word voice from the request, performs user attribute recognition and voiceprint recognition on the wake-up word voice, and obtains the user identity information of the target person.

[0135] In some embodiments, when performing user attribute recognition and voiceprint recognition, the server can obtain a voiceprint feature, and then match the voiceprint feature with voiceprint features in a database, and obtain the user identity information of the voice wake-up instruction according to the matching result. The voiceprint features in the database can be obtained according to audio data pre-recorded by the user. The user identity information can include a voiceprint mark U1, a gender US1, and an age UA1 of the target person. The gender US1 and the age UA1 can be user attributes, and the age UA1 can represent an age range, such as 1-10 years old, 11-20 years old, 20-40 years old, 40-60 years old, and so on.

[0136] In some embodiments, the user attribute recognition and the voiceprint recognition can also be implemented locally by the display device. The user can pre-record audio on the display device. The display device can compare the voice wake-up instruction with the pre-recorded audio of the user through voiceprint recognition to obtain the user identity information corresponding to the voice wake-up instruction.

[0137] In some embodiments, the user identity information can also include a face image of the target person, facilitating subsequent face recognition. The head portrait can be stored in the server and / or the display device. When the display device obtains the user identity information corresponding to the voice wake-up instruction, the display device can obtain the face image.

[0138] In some embodiments, during the process of requesting the user identity information from the server, the display device can not adjust the field of view of the camera first, but directly control the camera to capture images and perform face recognition on the images captured by the camera locally. At this time, if the target person is just standing in the current field of view of the camera, the target person can be photographed. If the user is outside the current field of view of the camera, the target person cannot be photographed. In order to reduce the resource consumption of the display device, only the age feature and the gender feature of the face can be detected during face recognition. After receiving the user identity information from the server, the gender US1 and the age UA1 in the user identity information can be extracted and compared with the result of face recognition. If no face is detected or the age or gender of the detected face does not match the user identity information, the field of view of the camera is adjusted again and the image is recaptured for face detection. Of course, if the user identity information includes a face image, during face recognition, it can also be determined whether the face in the image captured by the camera is consistent with the face image in the user identity information. Compared with only recognizing the age and gender, the accuracy of target person recognition can be improved, but the speed of target person recognition can be reduced.

[0139] In some embodiments, if the display device cannot recognize the target person from the image captured by the camera in the current field of view, after obtaining the wake-up angle, the display device can adjust the field of view of the camera according to the wake-up angle to capture the target person.

[0140] In some embodiments, the display device can adjust the field of view of the camera to cover the wake-up angle. For example, the wake-up angle is 30 degrees, and the current field of view of the display device is 40-140 degrees. The camera can be rotated more than 10 degrees to the left, so that the target person can be located in the captured image. During the rotation, the camera can continuously take pictures and perform face recognition. If a face matching the user identity information is recognized, the positioning of the target person is achieved. At this time, the camera can stop rotating.

[0141] In some embodiments, after the display device adjusts the field of view of the camera according to the wake-up angle, if the face matching the user identity information still cannot be recognized from the photographed image, it can be determined that the wake-up angle calculation is wrong or a false wake-up occurs. To solve the problem of wake-up angle calculation error, the display device can control the camera to rotate to cover the maximum field of view, i.e. 0-180 degrees of field of view, and if the target person still cannot be recognized within the 0-180 degree field of view, it is confirmed that a false wake-up occurs when the voice wake-up instruction. Alternatively, the display device can re-calculate the angle of the target person relative to the display device according to the real-time voice instruction received after the voice wake-up instruction, and then adjust the field of view of the camera according to the recalculated angle. If the target person still cannot be recognized after adjusting the field of view, it can be determined that a false wake-up has occurred. At this time, the display device can exit the current voice interaction process.

[0142] In some embodiments, after the display device identifies the target person in the image photographed by the camera, it can perform face tracking on the target person. Since the target person may move back and forth, the field of view of the camera is dynamically adjusted through face tracking, so that the target person is always located within the field of view of the camera. During face tracking, a preset area can be determined in the image photographed by the camera, and then the real-time coordinate range of the target face in the image photographed by the camera is obtained, and the camera is controlled to rotate according to the change trend of the real-time coordinate range. If the real-time coordinate range is within the preset area, the camera can not be rotated, and if one boundary of the real-time coordinate range coincides with the boundary of the preset area, the camera can be rotated according to the change trend of the real-time coordinate range, so that the face is located within the preset area again. For example, if the change trend of the real-time coordinate range is to translate to the left, the camera is controlled to rotate to the left.

[0143] In some embodiments, the real-time voice instruction received by the display device after the voice wake-up instruction can include any one or more of the target person's voice, other people's voice, and environmental noise. In the image photographed by the camera, by performing lip movement detection on the target person, it can be determined whether the target person is speaking. If the target person is speaking, the real-time voice instruction can be subjected to semantic recognition, and if the target person is not speaking, the real-time voice instruction can not be subjected to semantic recognition.

[0144] In some embodiments, the display device can generate a voice recognition request containing real-time voice and send the request to the server. After receiving the voice recognition request, the server can perform semantic recognition on the real-time voice in the voice recognition request and return the semantic recognition result to the display device.

[0145] In some embodiments, the server can first perform voiceprint recognition on the real-time voice, and if the voiceprint recognition result determines that the real-time voice contains the voice of the target person, then perform semantic recognition on the real-time voice, and return the semantic recognition result to the display device, or the server can simultaneously perform voiceprint recognition and semantic recognition, and return the voiceprint recognition result and the semantic recognition result to the display device.

[0146] In some embodiments, to improve the accuracy of semantic recognition, the server can perform noise reduction processing on the voice real-time instruction when performing semantic recognition.

[0147] In some embodiments, if there are other people in the image captured by the camera in addition to the target person, the noise reduction processing can be achieved through a human voice separation mechanism. Based on mixed human voice, a single-channel human voice is separated, and voiceprint recognition is performed on the single-channel audio to determine whether the single-channel audio belongs to the voice of the target speaker, and then semantic recognition is performed on the voice of the target speaker, and the voice of the non-target speaker is discarded. The human voice separation mechanism uses the voiceprint features in the identity information of the target person, and models based on noise sources to enhance the voice of the target person, suppress the sound other than the voice of the target speaker, and optimize the recognition effect of the voice of the target person in this scenario.

[0148] In some embodiments, after receiving the recognition result from the server, the display device can respond according to the recognition result. For example, the recognition result can include: R = {U = U1, V = increase volume}, where R represents the recognition result, U represents the voiceprint, U1 represents the voiceprint of the target person, and the voice content of the target person is "increase volume". According to the recognition result, the display device can increase the volume of the display device.

[0149] In some embodiments, the specific process of the voice recognition method of the present application can be seen in Figures 9-12 , and the technical solutions of the present application will be introduced below in combination with the specific process of the voice recognition method.

[0150] See Figure 9 , the overall flowchart of the voice interaction method according to some embodiments. After the user inputs the voice wake-up instruction to the display device, as shown in Figure 9 , the display device can collect the voice wake-up instruction of the user, calculate the wake-up angle according to the voice wake-up instruction, and then rotate the camera according to the wake-up angle so that the camera is directed towards the wake-up angle, thereby making the capture range of the camera cover the wake-up angle.

[0151] After the display device controls the camera to be directed towards the wake-up angle, the display device can perform face detection and face recognition on the image captured by the camera, so as to determine whether the target person is detected in the image captured by the camera.

[0152] If the target person cannot be detected in the image captured by the camera, the wake-up angle is recalculated according to the real-time audio collected, real-time sound source positioning is realized, and then the camera is rotated to face the recalculated wake-up angle according to the result of the sound source positioning, the face is detected from the image captured by the camera, and face recognition is performed on the detected face to determine whether it is the target person.

[0153] If the target person is detected in the image captured by the camera, the display device controls the camera to track the face of the target person and performs lip movement detection and recognition on the image captured by the camera. Among them, the image captured by the camera may contain the face of the target person and the face of the non-target person, or only contain the face of the target person.

[0154] If it is detected that the person who performs the lip movement in the image captured by the camera is the target person, the real-time recording, i.e. the real-time voice instruction collected by the audio input device of the display device, can be obtained. The display device can use beamforming to process the beam of the real-time recording, and the processing content can include enhancing the voice beam of the target person and suppressing other voice beams, wherein the voice beam of the target person can be determined according to the position of the target person in the image.

[0155] If it is detected that the person who performs the lip movement in the image captured by the camera is the non-target person, the recording may include one or more of noise and non-target person sound, and the display device can perform noise suppression and non-target person sound suppression on the real-time recording in order to improve the recognition accuracy of the next segment of real-time recording. Of course, the display device can also directly delete the noise or non-target person sound.

[0156] Among them, before face detection and recognition, the display device can determine the target person in advance according to the voice wake-up instruction, and the target person can be the user who issues the voice wake-up instruction. Referring to Figure 10 The flowchart of the processing method of the voice wake-up instruction according to some embodiments is shown. As Figure 10 shown, when the display device collects the wake-up audio, it determines that the wake-up audio includes the voice wake-up instruction through the wake-up word detection, can calculate the wake-up angle according to the wake-up audio, and save the wake-up audio to a preset path, then upload the wake-up audio of the preset path to the cloud, and make the server of the cloud process the wake-up audio. The processing of the server on the wake-up audio can include voiceprint recognition and user attribute recognition, through which the voiceprint of the user who issues the wake-up audio can be determined, and the gender and age range of the user corresponding to the wake-up audio can be obtained. Finally, the wake-up angle calculated locally by the display device, and the voiceprint and user attribute recognized by the server of the cloud, are determined as the recognition result of the wake-up audio.

[0157] by Figure 10It can be seen that the display device collects the voice wake-up instruction, calculates the wake-up angle according to the voice wake-up instruction, and controls the camera to rotate towards the wake-up angle, which can increase the probability of detecting the target person in the captured image.

[0158] In some embodiments, the image captured by the camera can include a single-person face and a multi-person face, and the two cases are analyzed in detail, Figure 11 A flowchart of a processing method when a single-person face is detected in a voice interaction process according to some embodiments is shown, Figure 12 A flowchart of a processing method when a multi-person face is detected in a voice interaction process is shown.

[0159] Referring to Figure 11 If the display device detects a single-person face within the wake-up angle, it can collect real-time recording and perform lip movement detection on the detected face. The processing of the real-time recording includes real-time sound source positioning and cloud recognition. Real-time sound source positioning includes calculating the real-time sound source angle of the real-time recording. Cloud recognition includes voice recognition and voiceprint recognition. When the display device requests the server in the cloud to perform voiceprint recognition, it can upload the voiceprint of the voice wake-up instruction and the real-time recording to the server, so that the server can confirm whether the voiceprint is the voiceprint of the same user as the voice wake-up instruction after obtaining the voiceprint of the real-time recording, and include the confirmation result in the voiceprint recognition result. According to the voice recognition result and the voiceprint recognition result returned by the server, and the result of real-time sound source positioning obtained locally by the display device, the display device can obtain the real-time sound source angle, voiceprint, attribute, voice content, etc. information corresponding to the real-time recording.

[0160] The display device can correct whether the real-time recording is the voice of the target speaker according to the voiceprint recognition result. If the real-time recording is not the voice of the target speaker, the camera is rotated again based on the real-time sound source angle determined by the real-time sound source positioning, so that the camera is directed towards the real-time sound source angle, and then face detection is performed in the image captured by the camera. If the real-time recording is the voice of the target speaker, face tracking is performed on the face detected by the camera, and then the lip movement detection result of the face is obtained. If the lip movement detection result is that lip movement has occurred, voice interaction is performed according to the real-time recording.

[0161] Referring to Figure 12If the display device detects multiple human faces within the wake-up angle, real-time recording can be collected, and lip movement detection can be performed on the detected human faces. The processing of the real-time recording includes real-time sound source positioning and cloud recognition. Real-time sound source positioning includes calculating the real-time sound source angle of the real-time recording. Cloud recognition includes speech recognition and voiceprint recognition. When the display device requests the server in the cloud to perform voiceprint recognition, the voiceprint of the voice wake-up instruction and the real-time recording can be uploaded to the server. After obtaining the voiceprint of the real-time recording, the server can confirm whether the real-time recording contains the voiceprint of the same user as the voiceprint of the voice wake-up instruction, that is, whether it contains the voiceprint of the target person, and include the confirmation result in the voiceprint recognition result. According to the speech recognition result and the voiceprint recognition result returned by the server, and the result of real-time sound source positioning obtained locally by the display device, the display device can obtain the real-time sound source angle, voiceprint, attribute, speech content and other information corresponding to the real-time recording.

[0162] The display device can correct whether the real-time recording is the voice of the target speaker according to the voiceprint recognition result. If the real-time recording is not the voice of the target speaker, the camera is rotated again based on the real-time sound source angle determined by the real-time sound source positioning, so that the camera is directed towards the real-time sound source angle, and then human face detection is performed in the image captured by the camera. If the real-time recording contains the voice of the target speaker, the target person is located from the human face detected by the camera, and then human face tracking is performed on the target person, and then the lip movement detection result of the human face of the target person is obtained. If the lip movement detection result indicates that lip movement has occurred, speech interaction is performed according to the real-time recording.

[0163] In the above embodiments, the camera provided by the display device or the camera connected to the display device can obtain a larger field of view angle by rotating, so as to realize human face detection, tracking and lip movement detection in a larger field of view range, improve the success probability of target person positioning in speech interaction, and further improve the wake-up rate and speech recognition accuracy of speech interaction. However, in some embodiments, the camera provided by the display device or the camera connected to the display device is not provided with a pan-tilt head, and the camera cannot be rotated. In this case, human face detection, tracking and lip movement detection can still be performed in the image within the fixed field of view range captured by the camera. At this time, human face tracking can be used to obtain the real-time coordinate range of the target human face, which can be the coordinate range of a rectangular area containing the human face of the target person, and lip movement detection is performed on the image within the real-time coordinate range. Through human face tracking, the range of lip movement detection can be reduced, and the response speed of speech interaction can be improved.

[0164] As can be seen from the above embodiments, the present application can effectively improve the speech interaction experience in complex speech interaction scenarios by comprehensively using sound source positioning, human face tracking, voiceprint recognition, lip movement detection and noise reduction processing technologies.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

[0166] The foregoing description has been presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the above teachings or can be acquired from practice of the embodiments. The described embodiments were chosen and described in order to best explain the principles of the embodiments and its practical application. This enables others skilled in the art to best use the embodiments in various embodiments and with various modifications as are suited to the particular use contemplated.

Claims

1. A display device, characterized by comprising: The application relates to a voice wake-up device, comprising: a display for presenting a user interface; a camera for collecting images to obtain a video signal; a controller connected with the display, the controller being configured to: collect a voice wake-up instruction; in response to the voice wake-up instruction, perform voiceprint recognition and user attribute recognition on the voice wake-up instruction to obtain a voiceprint feature, match the voiceprint feature with voiceprint features in a database to obtain user identity information of a target person, the user identity information comprising a voiceprint feature, gender and age of the target person, and collect a voice real-time instruction, the target person comprising a person who issues the wake-up instruction; control the camera to collect images, detect age features and gender features of a face in the images collected by the camera; compare the gender with the gender features and the age with the age features, and if both are matched, determine that face information of the target person is detected, perform face tracking, lip movement detection and lip language recognition on the target person according to the video signal, if the face of the target person has lip movement and the speech content obtained after the lip language recognition matches the audio content obtained by performing voice recognition on the voice real-time instruction, determine that the person who is speaking comprises the target person, and the voice real-time instruction is from the person in the video signal, if there are other people in addition to the target person in the images captured by the camera, separate the single-channel audio of the target person from the voice real-time instruction based on the voiceprint feature, and respond to the voice real-time instruction according to the single-channel audio; if the face of the target person has no lip movement or the voice real-time instruction does not comprise the voice of the target person, do not respond to the voice real-time instruction.

2. The display device of claim 1, wherein, detecting face information in the images collected by the camera comprises: performing sound source positioning on the voice wake-up instruction to obtain a wake-up sound source position; turning the camera towards the wake-up sound source position, and detecting face information in the images collected by the camera during the turning process, if the face information of the target person is detected, controlling the camera to stop turning.

3. The display device of claim 1, wherein, performing face tracking and lip movement detection on the target person comprises: obtaining a real-time coordinate range of a target face in the images captured by the camera; performing lip movement detection on the images in the real-time coordinate range.

4. The display device of claim 1, wherein, performing face tracking and lip movement detection on the target person comprises: obtaining a real-time coordinate range of a target face in the images captured by the camera; controlling the camera to turn according to the change trend of the real-time coordinate range, so that the face of the target person is located in a preset area in the images collected by the camera; performing lip movement detection on the images of the target face.

5. The display device of claim 1, wherein, The controller is further configured to: if the face of the target person is not detected, performing sound source positioning on the voice real-time instruction to obtain a real-time sound source position; turning the camera towards the real-time sound source position, and detecting the face of the target person corresponding to the user identity information in the images collected by the camera.

6. The display device of claim 1, wherein, obtaining user identity information corresponding to the voice wake-up instruction comprises: The voice wake-up instruction is subjected to voiceprint recognition to obtain user identity information, which includes voiceprint information of the target person.

7. The display device of claim 1, wherein, The voice real-time instruction is responded to, including: If the image captured by the camera only includes a single face of the target person, the voice corresponding to the wake-up sound source position is subjected to directional voice enhancement; The voice real-time instruction subjected to directional voice enhancement is responded to.

8. The display device of claim 1, wherein, The controller is further configured to: After the voice real-time instruction is collected, the voice real-time instruction is sent to a server, so that the server performs voiceprint recognition, speech recognition and semantic recognition on the voice real-time instruction to obtain a recognition result; The recognition result of the voice real-time instruction by the server is received.

9. A voice interaction method, characterized in that, It includes: Collecting a voice wake-up instruction; In response to the voice wake-up instruction, the voice wake-up instruction is subjected to voiceprint recognition and user attribute recognition to obtain voiceprint features, and the voiceprint features are matched with voiceprint features in a database to obtain user identity information of a target person, which includes voiceprint features, gender and age of the target person, and a voice real-time instruction is collected, the target person including a person who issues the wake-up instruction; A camera is controlled to collect an image, and age features and gender features of a face are detected in the image collected by the camera, the camera being used to collect an image to obtain a video signal; The gender is compared with the gender features, and the age is compared with the age features, and if both are matched, it is determined that the face information of the target person is detected, and the target person is subjected to face tracking, lip movement detection and lip reading according to the video signal, if the face of the target person has lip movement, and the speech content obtained after the lip reading matches the audio content obtained by performing speech recognition on the voice real-time instruction, it is determined that the person speaking includes the target person, and the voice real-time instruction is derived from the person in the video signal, if the image captured by the camera has the target person and other persons, a single-channel audio of the target person is separated from the voice real-time instruction based on the voiceprint features, and the voice real-time instruction is responded to according to the single-channel audio; If the face of the target person does not have lip movement, or the voice real-time instruction does not include the voice of the target person, the voice real-time instruction is not responded to.

Citation Information

Patent Citations

  • Human-machine interaction method and device, storage medium, and smart terminal

    CN108766438A

  • Voice conversion method and device and electronic equipment

    CN111883135A

  • Voice interaction method and device for massage chair

    CN112558911A

  • Man-machine conversation method, electronic equipment and computer readable storage medium

    CN112634911A