A projection device and a far-field speech recognition method

By transmitting media audio data back to the projector via the audio output device on the projection screen, signal alignment and echo cancellation are achieved, solving the problem of media audio interference with far-field speech recognition and improving the accuracy of far-field speech recognition.

CN119865648BActive Publication Date: 2025-10-31HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411942776.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-10-31
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The media sound output from the audio output device in a traditional projection screen can easily interfere with far-field speech recognition, leading to a decrease in recognition accuracy.

Method used

When the audio output device in the projection screen outputs media sound, it sends back the corresponding echo reference data to the projector host. The data is then transmitted to the controller via a wireless transmission module for signal alignment and echo cancellation, thus eliminating interference from the media sound to far-field speech.

Benefits of technology

By using signal alignment and echo cancellation, the accuracy of far-field speech recognition is improved, and the interference of media sound on far-field speech recognition is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119865648B_ABST
    Figure CN119865648B_ABST
Patent Text Reader

Abstract

This application relates to a projection device and a far-field speech recognition method. The projection device includes a first sound acquisition unit, a wireless transmission module, a controller, a signal processing module, and an audio output device. The audio output device is configured to output media sound and send back-collected reference data corresponding to the media sound to the signal processing module. The signal processing module is configured to send the back-collected reference data to the wireless transmission module. The wireless transmission module is configured to send the back-collected reference data to the controller. The first sound acquisition unit is configured to acquire first external sound data. The controller is configured to: perform signal alignment on the back-collected reference data and the first external sound data; perform echo cancellation on the media sound in the signal-aligned first external sound data based on the signal-aligned back-collected reference data to obtain far-field speech data; and perform far-field speech recognition on the far-field speech data. Using this projection device can improve the accuracy of far-field speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of projection equipment technology, and in particular to a projection device and a far-field speech recognition method. Background Technology

[0002] With the rapid development of smart devices and IoT technology, more and more laser TVs are being equipped with far-field voice control, which greatly improves the ease of operation of laser TVs. Furthermore, many high-end laser TVs now incorporate audio output devices within the projection screen to enhance the audio-visual experience.

[0003] In traditional technology, screen-sound laser TVs collect far-field speech and perform far-field speech recognition through a sound acquisition device in the projector host.

[0004] However, when performing far-field speech recognition, the media sound output by the audio output device on the projection screen can easily interfere with or even mask the user's far-field speech, leading to a decrease in the accuracy of far-field speech recognition. Summary of the Invention

[0005] This application provides a projection device and a far-field speech recognition method to solve the problem of low accuracy in far-field speech recognition.

[0006] In a first aspect, some embodiments provide a projection device, including: a projection screen and a projection host, the projection host being configured to project a display image onto the projection screen;

[0007] The projector unit includes a first sound acquisition unit, a wireless transmission module, and a controller; the projection screen includes a signal processing module and an audio output device, wherein:

[0008] The projection screen is configured to display the user interface;

[0009] The audio output device is configured to output media sound and, when outputting media sound, sends the corresponding retrieval reference data of the media sound to the signal processing module;

[0010] The signal processing module is configured to send the acquired reference data to the wireless transmission module;

[0011] The wireless transmission module is configured to receive the back-collected reference data and send the back-collected reference data to the controller;

[0012] The first sound acquisition device is configured to acquire first external sound data and send the first external sound data to the controller;

[0013] The controller is configured as follows:

[0014] Upon receiving the acquired reference data and the first external sound data, signal alignment is performed on the acquired reference data and the first external sound data;

[0015] Based on the back-sampling reference data after signal alignment, echo cancellation is performed on the media sound in the first external sound data after signal alignment to obtain far-field speech data.

[0016] Perform far-field speech recognition on far-field speech data.

[0017] Technical effects: On the one hand, the audio output device in the projection screen transmits back the corresponding echo reference data while outputting media sound, and transmits it to the projector wirelessly. On the other hand, the media sound output by the audio output device in the projection screen can be captured by the first sound acquisition unit in the projector. In this way, the projector can obtain the echo reference data corresponding to the media sound, as well as the first external sound data mixed with the media sound and far-field speech. Then, by signal alignment, the phase difference caused by the different transmission times of the two can be eliminated, improving the echo cancellation effect. Then, the far-field speech with the media sound eliminated can be used for far-field speech recognition, which can effectively reduce the interference of media sound on far-field speech recognition and improve the accuracy of far-field speech recognition.

[0018] Secondly, some embodiments also provide a far-field speech recognition method, which is applied to a projector in a projection device, the projection device further including a projection screen, and the projector being configured to project a display image onto the projection screen; the method includes:

[0019] The system collects the first external sound data and receives the back-collected reference data. The back-collected reference data is wirelessly transmitted from the projection screen to the projection host when the media sound is output.

[0020] Signal alignment is performed between the re-acquired reference data and the first external sound data;

[0021] Based on the back-sampling reference data after signal alignment, echo cancellation is performed on the media sound in the first external sound data after signal alignment to obtain far-field speech data.

[0022] Perform far-field speech recognition on far-field speech data.

[0023] Technical effects: On the one hand, the audio output device in the projection screen transmits back the corresponding echo reference data while outputting media sound, and transmits it to the projector wirelessly. On the other hand, the media sound output by the audio output device in the projection screen can be captured by the first sound acquisition unit in the projector. In this way, the projector can obtain the echo reference data corresponding to the media sound, as well as the first external sound data mixed with the media sound and far-field speech. Then, by signal alignment, the phase difference caused by the different transmission times of the two can be eliminated, improving the echo cancellation effect. Then, the far-field speech with the media sound eliminated can be used for far-field speech recognition, which can effectively reduce the interference of media sound on far-field speech recognition and improve the accuracy of far-field speech recognition.

[0024] Thirdly, some embodiments also provide a far-field speech recognition method, which is applied to a projection screen in a projection device, the projection device further including a projection host configured to project a display image onto the projection screen; the method includes:

[0025] The system outputs media audio and wirelessly transmits the corresponding backtracking reference data to the projector host while outputting the media audio. This allows the projector host to perform signal alignment between the backtracking reference data and the acquired first external audio data. Based on the backtracking reference data after signal alignment, the system performs echo cancellation on the media audio in the first external audio data after signal alignment to obtain far-field speech data. The system then performs far-field speech recognition on the far-field speech data.

[0026] Technical effects: On the one hand, the audio output device in the projection screen transmits back the corresponding echo reference data while outputting media sound, and transmits it to the projector wirelessly. On the other hand, the media sound output by the audio output device in the projection screen can be captured by the first sound acquisition unit in the projector. In this way, the projector can obtain the echo reference data corresponding to the media sound, as well as the first external sound data mixed with the media sound and far-field speech. Then, by signal alignment, the phase difference caused by the different transmission times of the two can be eliminated, improving the echo cancellation effect. Then, the far-field speech with the media sound eliminated can be used for far-field speech recognition, which can effectively reduce the interference of media sound on far-field speech recognition and improve the accuracy of far-field speech recognition. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a schematic diagram illustrating the operational scenarios between a projection device and a control device provided in some embodiments of this application;

[0029] Figure 2 This is a schematic diagram of the hardware configuration of a projection device provided in some embodiments of this application;

[0030] Figure 3 This is a schematic diagram of the hardware configuration of the control device provided in some embodiments of this application;

[0031] Figure 4 This is a schematic diagram illustrating the software configuration of a projection device provided in some embodiments of this application;

[0032] Figure 5 This application provides schematic diagrams of the structure of a projection device according to some embodiments;

[0033] Figure 6 Timing diagrams for implementing far-field speech recognition via a projection device in some embodiments of this application;

[0034] Figure 7 A schematic diagram of the audio signal of the speech data to be tested provided in some embodiments of this application;

[0035] Figure 8 A schematic flowchart of far-field speech detection provided for some embodiments of this application;

[0036] Figure 9 Schematic diagrams of audio and reverse audio provided for other embodiments of this application;

[0037] Figure 10 A schematic diagram illustrating the decoding speed feedback adjustment process provided for other embodiments of this application;

[0038] Figure 11 This is a flowchart illustrating the implementation of a far-field speech recognition method provided in some embodiments of this application. Detailed Implementation

[0039] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0040] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0041] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0042] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0043] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0044] In this application embodiment, projection device 200 generally refers to a device with image display and data processing capabilities. For example, projection device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc., that use laser light sources.

[0045] Figure 1 This is a schematic diagram illustrating an operational scenario between a projection device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, users can operate the projection device 200 via touch operation, mobile terminal 300, and control device 100. For example, control device 100 can be a remote control, stylus, handle, etc.

[0046] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the projection device 200. It can also function as a communication device for establishing a communication connection with the projection device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the projection device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the projection device 200 for synchronized display.

[0047] like Figure 1 The diagram also shows that the projection device 200 communicates with the server 400 via various communication methods. The projection device 200 can communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0048] The projection device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0049] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the projection device 200.

[0050] In some embodiments, the projection device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a projection screen 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0051] In some embodiments, the projection device includes a projection host and a projection screen, the projection host being configured to project display images onto the projection screen. The projection screen includes at least a signal processing module and an audio output device, and the projection host includes at least a first sound acquisition unit, a wireless transmission module, and a controller. The projection host may also include at least one of a tuner / demodulator, a communication device, a detector, a device interface, a controller, a projection screen, an audio output device, a memory, a power supply, and a user input interface.

[0052] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0053] In some embodiments, the projection screen 260 includes display function components for presenting images and driving components for driving image display. The projection screen 260 is used to receive and display image signals output from the controller 250. For example, the projection screen 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0054] In some embodiments, the communication device 220 is a component used to communicate with a first external device or server 400 according to various communication protocol types. The projection device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the projection device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the projection device 200 supports Bluetooth connection communication, it needs to have a communication device 220 with Bluetooth functionality.

[0055] The communication device 220 enables the projection device 200 to communicate with the first external device or server 400 via wireless or wired connection. Wired connection uses data cables, interfaces, and other components to display personalized recommendations. Wireless connection uses wireless signals or wireless networks to display personalized recommendations. The projection device 200 can directly connect to the first external device or indirectly through gateways, routers, or other connection devices.

[0056] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the projection device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the projection device 200.

[0057] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0058] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a projection screen 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0059] In some embodiments, the audio output device 270 can be the built-in speaker of the projection device 200 or an external audio output device connected to the projection device 200. For the external audio output device connected to the projection device 200, the projection device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the projection device 200 to output sound from the projection device 200.

[0060] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0061] Figure 3 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the central control device. (Example) Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0062] The control device 100 is configured to control the projection device 200, and to receive user input operation commands and convert the operation commands into commands that the projection device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the projection device 200.

[0063] In some embodiments, the control device 100 may be a smart device. For example, the control device 100 may be equipped with various applications for controlling the projection device 200 according to user needs.

[0064] In some embodiments, such as Figure 1 As shown, a mobile terminal 300 or other smart electronic device can perform similar functions to control device 100 after installing an application to control the projection device 200.

[0065] The controller 110 includes a processor 112, RAM 113, ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation of the control device 100, as well as the communication and cooperation between internal components and the external and internal data processing functions.

[0066] Under the control of the controller 110, the communication interface 130 enables communication of control signals and data signals with the projection device 200. The communication interface 130 may include at least one of other near-field communication modules such as WiFi chip 131, Bluetooth module 132, and NFC module 133.

[0067] User input / output interface 140, wherein the input interface includes at least one of other input interfaces such as microphone 141, touchpad 142, sensor 143, and button 144.

[0068] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, which can encode user input commands via WiFi, Bluetooth, or NFC protocols and send them to the projection device 200.

[0069] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can also store various control signal instructions input by the user.

[0070] The power supply 180 is used to provide operating power support for the various components of the control device 100 under the control of the controller.

[0071] In some embodiments, the projection device 200 may run an operating system to enable user interaction. An operating system is a computer program that manages and controls the hardware and software resources of the projection device 200. The operating system can (control the projection device) provide a user interface, allowing users to interact with the projection device 200 and supporting the running of various applications.

[0072] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for projection devices.

[0073] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer.

[0074] In some embodiments, the application layer provides services and interfaces for applications, enabling the projection device 200 to run applications and interact with users based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0075] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0076] like Figure 4As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0077] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0078] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0079] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the projection device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 4 As shown, hardware drivers can be configured in the kernel layer. The kernel layer can contain at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0080] It should be noted that the above examples are merely a simple division of the functions of the operating system and do not limit the specific form of the operating system of the projection device 200 in this application embodiment. Depending on the functions of the projection device, the type of the operating system, and other factors, the number of layers and the specific type of the operating system may be expressed in other forms.

[0081] To enhance the audio-visual experience of laser TVs, a special type of screen-sound laser TV has been developed. This screen-sound technology uses screen vibration to generate sound by pushing air in front of and behind the screen, utilizing the entire screen to produce sound and improving sound quality. Screen-sound technology provides users with a "sound and picture in perfect harmony" experience, allowing the image and sound to blend seamlessly in three-dimensional space, similar to a top-tier cinema experience. Screen-sound laser TVs incorporate an audio output device within the projection screen. To achieve a cleaner and more aesthetically pleasing design for both the projector and screen, the connection cable between the projector and screen has been eliminated. The projector drives the audio output device in the projection screen via wireless communication. With the rapid development of smart devices and IoT technology, more and more laser TVs are incorporating far-field voice functionality, which greatly improves the ease of operation. In traditional technology, screen-sound laser TVs collect far-field voice data through a sound acquisition unit in the projector and perform far-field voice recognition.

[0082] However, echo cancellation usually requires a certain amount of computing power, so it is usually implemented by the projector host. But during the playback of multimedia content on a screen-sound laser TV, the external sound collected by the sound acquisition device in the projector host includes media sound output by the audio output device in the projection screen. However, the echo acquisition signal of this media sound cannot be sent back to the projector host in time, making it difficult for the projector host to effectively eliminate the media sound in the external sound. As a result, far-field speech is easily affected by the media sound, or even masked, thus reducing the accuracy of far-field speech recognition.

[0083] In some embodiments, a projection device is provided, the projection device including a projection screen and a projection host, the projection host being configured to project a display image onto the projection screen;

[0084] The projector unit includes a first sound acquisition unit, a wireless transmission module, and a controller; the projection screen includes a signal processing module and an audio output device, wherein:

[0085] The projection screen is configured to display the user interface;

[0086] The audio output device is configured to output media sound and, when outputting media sound, sends the corresponding retrieval reference data of the media sound to the signal processing module;

[0087] The signal processing module is configured to send the acquired reference data to the wireless transmission module;

[0088] The wireless transmission module is configured to receive the back-collected reference data and send the back-collected reference data to the controller;

[0089] The first sound acquisition device is configured to acquire first external sound data and send the first external sound data to the controller;

[0090] The controller is configured as follows:

[0091] Upon receiving the acquired reference data and the first external sound data, signal alignment is performed on the acquired reference data and the first external sound data;

[0092] Based on the back-sampling reference data after signal alignment, echo cancellation is performed on the media sound in the first external sound data after signal alignment to obtain far-field speech data.

[0093] Perform far-field speech recognition on far-field speech data.

[0094] In this context, "projection equipment" can refer to screen-sound laser projection equipment, such as screen-sound laser TVs. Screen-sound laser projection equipment can also refer to projection equipment that combines laser display technology and screen-sound technology.

[0095] In some embodiments, such as Figure 5 As shown, the projection device includes at least a projection screen 510 and a projection host 520. The projection host 520 is configured to project display images onto the projection screen 510. The projection screen 510 includes at least a signal processing module 511 and an audio output device 512. The projection host 520 includes at least a first sound acquisition unit 521, a wireless transmission module 523, and a controller 522. The projection host 520 may also include at least one of the following: a tuner / demodulator, a communication device, a detector, a device interface, a controller, a display, an audio output device, a memory, a power supply, and a user input interface. It is understood that the projection host can be equipped with a more powerful processor and functional modules, possessing stronger computing power and heat dissipation capabilities, thereby performing more complex processing tasks. The projection screen, in order to maintain a thin and aesthetically pleasing design, has relatively limited hardware and computing resources, and its heat dissipation capabilities are also relatively limited. Therefore, it is mainly used for display and performing simple processing tasks such as audio driving and local image enhancement.

[0096] Media sound can refer to the sound of multimedia content. When a projection device plays multimedia content, it can simultaneously display the image of the multimedia content on the projection screen and display the sound of the multimedia content through an audio output device.

[0097] Reference data can refer to a copy of the media sound data output by an audio output device.

[0098] The first external sound data can refer to the sounds in the environment that can be collected by the first sound collector. It can include media sounds output by the audio output device, far-field speech, and other sounds that may exist in the environment, such as rain, wind, knocking, footsteps, animal calls, etc.

[0099] Far-field speech can refer to voice commands issued by the user that can be captured by the sound acquisition device of the projection equipment.

[0100] Far-field speech data can refer to the audio data after eliminating media sounds from the first external sound data. It can be understood that far-field speech data may contain not only far-field speech but also other interfering sounds in the environment. Compared to media sounds, other interfering sounds have a smaller impact on the accuracy of far-field speech recognition. Media sounds may contain human voices, and media sounds are emitted from the projection screen. The distance between the projection screen and the projection host is usually smaller than the distance between the user and the projection host. Therefore, whether from the perspective of volume or sound type, media sounds have a greater impact on the accuracy of far-field speech recognition. Eliminating media sounds can effectively improve the accuracy of far-field speech recognition.

[0101] Acoustic echo refers to the sound emitted by an audio output device that has traveled through multiple paths and been picked up by a sound acquisition device. Because the sound is reflected along multiple paths, it produces echoes with varying delays. Acoustic echoes can include direct echoes and indirect echoes. Direct echoes refer to the sound emitted by the audio output device that enters the sound acquisition device directly without any reflection. This type of echo has the shortest delay and is directly related to factors such as the distance and angle between the audio output device and the sound acquisition device. Indirect echoes refer to the collection of sound emitted by the audio output device that enters the sound acquisition device after one or more reflections along different paths. Any movement of any object in space will change the echo path; therefore, indirect echoes are characterized by being multi-path and time-varying. The echo cancellation in this embodiment can include at least one of direct echo cancellation and indirect echo cancellation. Regardless of whether it is direct or indirect echo cancellation, signal alignment of the echo-collected reference data and the first external sound data is required to accurately cancel out the media sound from the first external sound data, thus achieving echo cancellation.

[0102] In some feasible implementations, signal alignment methods may include fixed delay estimation, adaptive filtering, frequency domain alignment, feature point matching alignment, machine learning or deep learning, etc. The specific method can be selected according to the actual situation, and this embodiment does not limit it.

[0103] For example, such as Figure 6As shown, during the playback of multimedia content, the projection device displays the media image of the multimedia content on the projection screen and simultaneously outputs the media sound of the multimedia content through the audio output device. The output media sound can travel through the air to the user's ear, or it can be collected by the first sound acquisition device after traveling through the air. The first sound acquisition device can collect various sounds in the environment, including media sound and far-field speech, generate first external sound data, and transmit the generated first external sound data to the controller. While outputting the media sound, the audio output device can also back up the media sound data and use the backed-up copy as a reference for re-acquisition. The data is sent to the signal processing module, which then wirelessly transmits the acquired reference data to the wireless transmission module in the projector. After receiving the acquired reference data, the wireless transmission module transmits the acquired reference data to the controller. After receiving the first external sound data and the acquired reference data, the controller can first perform signal alignment on the first external sound data and the acquired reference data, and then use the aligned acquired reference data to perform echo cancellation on the signal-aligned first external sound data to eliminate the interference of media sound in the first external sound data on far-field speech, thereby obtaining far-field speech data, and then performing far-field speech recognition based on the far-field speech data.

[0104] In some feasible implementations, the process of an audio output device outputting the media sound of multimedia content may include: receiving media sound data sent by a projector host, performing sound effect processing, enhancement processing, etc. on the received media sound data, and then outputting the processed media sound data.

[0105] In some feasible implementations, the process of performing far-field speech recognition on far-field speech data may include: extracting features from far-field speech data using a preset acoustic model to obtain far-field speech features, performing similarity detection between the far-field speech features and preset wake-up word features, determining to trigger the far-field speech wake-up function if the similarity is higher than a preset similarity threshold, establishing a communication connection with the cloud, and transmitting audio streams to the cloud in real time so that the cloud can perform speech recognition through a speech recognition model and send the speech recognition results to a projection device. The projection device can display the speech recognition results through a projection screen or an audio output device.

[0106] In this embodiment, on the one hand, the audio output device in the projection screen outputs media sound while simultaneously transmitting back the corresponding echo reference data, and transmits it to the projection host via wireless transmission. On the other hand, the media sound output by the audio output device in the projection screen can be captured by the first sound acquisition unit in the projection host. In this way, the projection host can obtain the echo reference data corresponding to the media sound, as well as the first external sound data mixed with the media sound and far-field speech. Then, by signal alignment, the phase difference caused by the different transmission times of the two can be eliminated, improving the echo cancellation effect. Furthermore, the far-field speech with the media sound eliminated can be used for far-field speech recognition, which can effectively reduce the interference of the media sound on far-field speech recognition and improve the accuracy of far-field speech recognition.

[0107] In some embodiments, during the process of sending the acquired reference data to the wireless transmission module, the signal processing module is further configured to:

[0108] When the far-field voice function is triggered, the sampled reference data is sent to the wireless transmission module.

[0109] It should be noted that since the backsampling reference data needs to be transmitted in real time, transmitting it wirelessly from the projection screen to the projector increases the burden on wireless signal transmission and processing, and consumes more energy. Furthermore, the projector will continuously perform echo cancellation processing due to the transmission of backsampling reference data, consuming its computing resources. However, echo cancellation is typically only needed when the user is using far-field voice functionality. When far-field voice functionality is not used, there is no need for echo cancellation, and it would only result in unnecessary resource waste.

[0110] Among them, far-field voice function refers to the ability of a projection device to effectively receive and process user voice input at a greater distance. It allows users to interact with the projection device through natural language without having to be close to or very close to the microphone.

[0111] For example, the audio output device can continuously send back the corresponding back-acquisition reference data to the signal processing module when outputting media sound. After receiving the back-acquisition reference data, the signal processing module can perform operations such as storage, processing, or transmission. During the operation of the signal processing module, the signal processing module can detect whether the far-field voice function has been triggered. Before the far-field voice function is detected to be triggered, the signal processing module can perform operations such as storage or processing after receiving the back-acquisition reference data, but will not send the back-acquisition reference data to the projector. After the far-field voice function is not detected to be triggered, the signal processing module can perform operations such as storage or processing after receiving the back-acquisition reference data, and can also send the back-acquisition reference data to the projector.

[0112] In some feasible implementations, the method for detecting whether the far-field voice function has been triggered may include: the projection screen determining whether the far-field voice function has been triggered automatically, or the projection host determining whether the far-field voice function has been triggered, and after the projection host determines that the far-field voice function has been triggered, sending a far-field voice function activation command to the projection screen, and the projection screen determining that the far-field voice function has been triggered after receiving the far-field voice function activation command sent by the projection host. The method for determining whether the far-field voice function has been triggered may include voice recognition, voice wake-up, human voice recognition, button triggering, etc., and the specific method can be determined according to the actual situation. This embodiment does not limit this.

[0113] In this embodiment, the signal processing module on the projection screen determines whether the far-field voice function is triggered. Only when the far-field voice function is triggered is the back-collection reference data sent to the projection host. This can effectively reduce unnecessary transmission of back-collection reference data and echo cancellation processing, thereby effectively saving communication and computing resources and ensuring the normal operation of other functions of the projection device.

[0114] In some embodiments, the projection screen includes a second sound acquisition unit, configured to acquire second external sound data; and during the process of sending the acquired reference data to the wireless transmission module when the far-field voice function is triggered, the signal processing module is further configured to:

[0115] Acquire second external sound data and re-collected reference data;

[0116] Based on the re-acquired reference data, echo cancellation is performed on the second external sound data to obtain the speech data to be tested.

[0117] If far-field voice data is present in the voice data to be tested, the far-field voice function is determined to be triggered, and the back-collected reference data is sent to the wireless transmission module.

[0118] It should be noted that compared to the projection screen, the projector host has higher computing power. Therefore, the projector host has more options for detecting whether the far-field voice function has been triggered. However, after the projector host detects that the far-field voice function has been triggered, it still needs to communicate with the projection screen, which not only consumes communication resources but also causes a delay in the start-up time of echo cancellation, which may affect the effectiveness and accuracy of the far-field voice function.

[0119] Detecting whether the far-field voice function has been triggered can include detecting whether human voice is present in the second external sound data. However, media audio often contains human voice as well. In cases where the second external sound data contains both far-field voice and human voice from the media audio, detecting human voice may also lead to false triggering of the far-field voice function.

[0120] The second sound acquisition device can refer to a sound acquisition device deployed in the projection screen. In some feasible implementations, compared to the projection host, the projection screen has greater space and resource constraints due to aesthetic and thinness requirements. Therefore, the first sound acquisition device set in the projection host can be a microphone array to ensure the comprehensiveness and accuracy of sound acquisition, while the second sound acquisition device set in the projection screen can be a single microphone, effectively balancing the actual needs of the projection screen and the actual needs of sound acquisition.

[0121] For example, during the operation of the projection screen, the second sound acquisition device can continuously acquire second external sound data and transmit the second external sound data to the signal processing module. Simultaneously, the audio output device can transmit the corresponding back-collection reference data to the signal processing module while outputting media sound. Since neither the second external sound data acquired by the second sound acquisition device nor the back-collection reference data has undergone wireless transmission before comparison, there is no audio delay issue. After receiving the second external sound data and the back-collection reference data, the signal processing module can perform echo cancellation on the second external sound data based on the back-collection reference data to obtain the speech data to be tested, and detect whether there is far-field speech data in the speech data to be tested. If there is far-field speech data in the speech data to be tested, it can be determined that the far-field speech function has been triggered, and then the back-collection reference data can be sent to the wireless transmission module. If no far-field speech data is detected in the speech data to be tested, there is no need to send back-collection reference data to the wireless transmission module, and the detection of whether there is far-field speech data in the speech data to be tested can continue.

[0122] In this embodiment, by setting a second sound acquisition device on the projection screen, the projection screen can accurately determine whether the far-field voice function has been triggered. This can effectively reduce the communication resources required to determine whether the far-field voice function has been triggered, improve the timeliness of echo cancellation, and thus improve the effectiveness and accuracy of the far-field voice function.

[0123] In some embodiments, when far-field speech data exists in the speech data to be tested, during the process of determining that the far-field speech function is triggered, the signal processing module is further configured to:

[0124] If the energy amplitude of the voice data to be tested exceeds the preset energy amplitude threshold, it is determined that there is far-field voice data in the voice data to be tested, and the far-field voice function is triggered.

[0125] It should be noted that the external sounds collected by the sound acquisition device may include not only media sounds and far-field speech, but also environmental interference sounds such as footsteps, rain, and wind, as well as noise. If the external sounds do not include far-field speech, after echo cancellation using the back-collected reference data, the test speech data may still contain environmental interference sounds and noise that are not far-field speech. The presence of human voices in the test speech data can be determined by timbre recognition or human voice recognition, thereby determining whether the far-field speech function has been triggered.

[0126] However, voice recognition or human voice recognition methods require high computing power, and the frequency of detecting whether there is far-field voice data in the voice data to be tested is high. Frequent voice recognition or human voice recognition may put a heavy burden on the projection screen and affect the normal use of other functions of the projection screen.

[0127] Energy amplitude refers to the instantaneous or average intensity level of the audio signal corresponding to the speech data being tested. Amplitude is an important parameter describing the characteristics of an audio waveform, indicating the degree to which the audio signal deviates from zero at a certain moment. The energy amplitude of environmental interference sounds and noise is usually low, while the energy amplitude of far-field speech is usually high. Therefore, by limiting the energy amplitude threshold, far-field speech can be effectively distinguished from environmental interference sounds and noise, thus enabling the detection of far-field speech data.

[0128] For example, after obtaining the voice data to be tested through echo cancellation, the energy amplitude of the audio signal corresponding to the voice data can be detected to determine whether the energy amplitude exceeds a preset energy amplitude threshold. If the energy amplitude exceeds the preset energy threshold, it is determined that far-field voice data exists in the voice data to be tested. From the moment the energy amplitude exceeds the preset energy threshold, it is determined that the far-field voice function is triggered, and the back-collection reference data begins to be transmitted from that moment. For example, in a living room scenario, the volume of the user's voice reaching the microphone is approximately 65 dB, corresponding to an amplitude of approximately 1000 smpl. Therefore, 1000 smpl can be set as the energy amplitude threshold. When the energy amplitude is detected to reach the energy amplitude threshold, the back-collection reference data begins to be transmitted.

[0129] In some feasible implementations, such as Figure 7 As shown, the energy amplitude of the audio signal of the speech data to be tested reaches the preset energy amplitude threshold at time t1, so the reference data can be transmitted back from time t1.

[0130] In this embodiment, by limiting the energy amplitude threshold, far-field speech can be effectively distinguished from environmental interference sounds and noise, thus realizing the detection of far-field speech data and improving the accuracy of far-field speech function triggering. Reference data is only transmitted back after far-field speech data is detected, which can reduce unnecessary data transmission and data processing and improve resource utilization.

[0131] In some embodiments, the signal processing module includes a bandpass filter; in the process of determining that far-field speech data exists in the speech data under test when the energy amplitude of the speech data under test exceeds a preset energy amplitude threshold, the signal processing module is further configured to:

[0132] By using a bandpass filter, the speech data to be tested is filtered to obtain the target human voice speech data corresponding to the preset human voice signal frequency band.

[0133] If the energy amplitude of the target human voice data exceeds the preset energy amplitude threshold, it is determined that there is far-field voice data in the voice data to be tested.

[0134] It should be noted that for loud environmental interference sounds and noises, such as the sound of musical instruments or objects falling, it is difficult to distinguish them by limiting the energy amplitude threshold, which may still lead to false triggering of far-field voice functions.

[0135] A bandpass filter is an electronic filter that allows signals within a specific frequency range to pass through while attenuating frequency components outside that range. In this embodiment, the bandpass filter can only allow signals within a preset human voice frequency band to pass through, attenuating signals in other frequency bands, thereby filtering out sounds other than human voice.

[0136] For example, after obtaining the speech data to be tested through echo cancellation, the signals of other signal bands in the speech data to be tested, excluding the preset human voice signal frequency band, can be first passed through a bandpass filter to obtain the target human voice speech data corresponding to the preset human voice signal frequency band; then the energy amplitude of the audio signal corresponding to the target human voice speech data is detected, and it is determined whether the energy amplitude exceeds the preset energy amplitude threshold. If the energy amplitude exceeds the preset energy threshold, it is determined that there is far-field speech data in the target human voice speech data. From the moment the energy amplitude exceeds the preset energy threshold, it is determined that the far-field speech function is triggered, and the back-collected reference data is transmitted from that moment.

[0137] In some feasible implementations, such as Figure 8As shown, the second sound acquisition device acquires second external sound data and transmits it to the signal processing module. Simultaneously, the audio output device sends back the corresponding back-collected reference data of the output media sound to the signal processing module. The signal processing module receives the second external sound data and the back-collected reference data, denotes the second external sound data as audio A, and the back-collected reference data as audio B. Audio A is subtracted from audio B to obtain audio C, which still contains far-field speech and environmental noise. The process of subtracting audio B from audio A includes: Figure 9 As shown, first multiply each frequency point of audio B by -1 to obtain the inverse audio B*; then add audio A and the inverse audio B* to cancel out the energy of the same frequency points in audio A as audio B, and obtain audio C.

[0138] Human voices typically generate energy within the 500Hz to 4kHz frequency band, while other sounds such as noise and musical instruments are usually distributed in different frequency bands. Therefore, a bandpass filter can be designed to filter out non-human voice signal frequency bands below 500Hz and above 4kHz. According to the cutoff frequency formula:

[0139]

[0140] Where F is the cutoff frequency; R is the resistance value of the resistor in the bandpass filter; C is the capacitance of the capacitor in the bandpass filter; 2π refers to twice the mathematical constant π, used to convert the time constant into frequency.

[0141] The high-pass filter filters audio frequencies below 500Hz. Let F = 500Hz, and select appropriate resistor values ​​and capacitor values. The low-pass filter filters audio frequencies above 4kHz. Let F = 4kHz, and select appropriate resistor values ​​and capacitor values. After filtering, audio C becomes audio D. Audio D is then analyzed for energy. When the energy amplitude of audio D reaches a threshold, reference data transmission begins.

[0142] In this embodiment, by filtering out sound from other signal frequency bands besides the human voice signal frequency band using a bandpass filter, the interference of environmental noise and other sounds can be effectively eliminated, improving the detection accuracy of far-field speech and thus improving the accuracy of far-field speech function triggering.

[0143] In some embodiments, during the signal alignment of the re-acquired reference data and the first external sound data, the controller is further configured to:

[0144] The back-collected reference data is decoded to obtain the back-collected reference signal, and the first external sound data is decoded to obtain the first external sound signal;

[0145] Acquire a first timestamp in the back-collection reference signal and a second timestamp in the first external sound signal, wherein the first timestamp is used to characterize the time when the back-collection reference signal is emitted from the audio output device, and the second timestamp is used to characterize the time when the first sound collector collects the first external sound signal.

[0146] Based on the preset sound propagation duration range and the difference between the second and first timestamps, the re-acquired reference data and the first external sound data are signal aligned.

[0147] It should be noted that the echo cancellation effect is affected by signal delay during wireless transmission of the echo-recorded reference data. This delay can cause the first external sound data collected by the first sound acquisition device to be difficult to align with the echo-recorded reference data in the time domain.

[0148] The first timestamp represents the time when the echo-collected reference signal is emitted from the audio output device. The second timestamp represents the time when the first sound acquisition device collects the first external sound signal. In the absence of signal delay, the first and second timestamps can be aligned based on the sound propagation time, making it difficult to align the first external sound data and the echo-collected reference data in the time domain. However, during wireless transmission, the echo-collected reference data may experience signal delay. This means the time difference between the first external sound data and the echo-collected reference data is not solely due to sound propagation; the signal delay caused by wireless transmission is not stable and varies with the actual situation, resulting in poor echo cancellation effects from wirelessly transmitted echo-collected reference data.

[0149] The range of sound propagation duration can be determined based on the sound propagation duration itself. It can be equal to the sound propagation duration, or it can be the sound propagation duration plus or minus a preset error value to determine the sound propagation market range. For example, if the sound propagation duration is t, the sound propagation duration range can be t, or it can be t ± t', where t' can be the preset error value.

[0150] For example, after receiving the re-acquisition reference data and the first external sound data, the controller can first decode the re-acquisition reference data into a re-acquisition reference signal and decode the first external sound data into a first external sound signal. It can then extract a first timestamp from the re-acquisition reference signal and a second timestamp from the first external sound signal. Furthermore, it can calculate the difference between the first and second timestamps and detect whether the difference falls within a preset sound propagation time range. If the difference falls within the preset sound propagation time range, it indicates that the re-acquisition reference data and the first external sound data are signal-aligned, and echo cancellation can be performed. If the difference does not fall within the preset sound propagation time range, it indicates that the re-acquisition reference data is not signal-aligned. If the reference data and the first external sound data are not signal-aligned, the echo cancellation effect will be poor. In this case, you can wait until the back-acquired reference data or the first external sound data that is signal-aligned is decoded, and then use the signal-aligned back-acquired reference signal and the first external sound signal for echo cancellation. Alternatively, you can adjust the decoding speed of the back-acquired reference data and the first external sound data to ensure that the subsequent back-acquired reference data and the first external sound data signals are aligned, and continue to monitor whether the back-acquired reference data and the first external sound data are signal-aligned until they are aligned, and then use the signal-aligned back-acquired reference data and the first external sound data for echo cancellation.

[0151] In this embodiment, based on the preset sound propagation duration range and the difference between the second timestamp and the first timestamp, it is possible to accurately determine whether there is a signal delay in the wireless transmission of the re-collected reference data, and guide signal alignment, thereby improving the accuracy of signal alignment and thus improving the echo cancellation effect.

[0152] In some embodiments, during the process of signal alignment of the re-acquired reference data and the first external sound data based on a preset sound propagation duration range and the difference between the second timestamp and the first timestamp, the controller is further configured to perform at least one of the following:

[0153] If the difference between the second timestamp and the first timestamp is greater than the preset sound propagation duration range, the decoding speed of the first external sound data is slowed down according to the difference value.

[0154] If the difference between the second timestamp and the first timestamp is less than the preset sound propagation duration range, the decoding speed of the re-collected reference data is slowed down according to the difference value.

[0155] It should be noted that while aligning the decoded re-acquired reference signal and the first external sound signal before echo cancellation can achieve signal alignment, it will cause the decoded re-acquired reference signal and the first external sound signal to accumulate, occupying storage resources and reducing resource utilization.

[0156] For example, such as Figure 10 As shown, let the first timestamp be T0 and the second timestamp be Tn. If the difference between the first and second timestamps does not fall within the preset sound propagation time range, it indicates that the re-acquired reference data and the first external sound data are not signal-aligned, and the echo cancellation effect is poor. In this case, the order of the first and second timestamps can be used to determine whether the re-acquired reference data is delayed or transmitted too quickly, thus guiding the feedback adjustment of the decoding speed. If the difference between the second and first timestamps is greater than the preset sound propagation time range, it indicates that the re-acquired reference data is delayed, so the decoding speed of the first external sound data can be slowed down based on the difference. If the difference between the second and first timestamps is less than the preset sound propagation time range, it indicates that the re-acquired reference data is transmitted too quickly, so the decoding speed of the re-acquired reference data can be slowed down based on the difference. This continues until the difference between the first and second timestamps falls within the preset sound propagation time range, indicating that the re-acquired reference data and the first external sound data are signal-aligned, and echo cancellation can be performed.

[0157] In this embodiment, by adjusting the decoding speed with feedback, signal alignment can be effectively achieved without requiring additional resources.

[0158] In some embodiments, the audio output device is configured to output a test sound; the first sound acquisition unit is further configured to acquire third external sound data corresponding to the test sound; the projection screen includes a second sound acquisition unit, which is configured to acquire fourth external sound data corresponding to the test sound; before aligning the back-collected reference data and the first external sound data according to a preset sound propagation duration range and the difference between the second timestamp and the first timestamp, the controller is further configured to:

[0159] The third timestamp of the target signal is detected from the third external sound data, and the fourth timestamp of the target signal is detected from the fourth external sound data;

[0160] Calculate the time difference between the third and fourth timestamps, and calculate the sound propagation duration based on the distance between the audio output device and the second sound acquisition device;

[0161] The range of sound propagation duration is determined by summing the time difference with the sound propagation duration.

[0162] It should be noted that since the distance between the projector and the projection screen may vary from home to home, the estimation of the distance between the media sound from the audio output device and the first sound acquisition device may be inaccurate, and the accuracy of the estimation of the sound propagation time range may be reduced. This would lead to a decrease in the accuracy of signal alignment detection and a deterioration in echo cancellation.

[0163] For example, after placing the projector and projection screen, the user can first test the sound propagation time range. Specifically, a frequency sweep wave can be output through the audio output device, such as a frequency sweep wave of 100Hz-8kHz. Each frequency point of the frequency sweep wave appears only once during playback, making it unique. After the frequency sweep wave propagates through the air, it can be collected by the first and second sound collectors. The data collected by the first sound collector is recorded as the third external sound data, and the data collected by the second sound collector is recorded as the fourth external sound data. Furthermore, the third timestamp of the target signal at any target frequency in the third external sound data and the fourth timestamp of the target signal in the fourth external sound data can be detected. Then, the time difference between the third and fourth timestamps can be calculated. Since the distance between the second sound collector and the audio output device on the projection screen is fixed, the sound propagation time of the target signal from the audio output device to the second sound collector can be calculated. By adding the sound propagation time to the time difference between the third and fourth timestamps, the sound propagation time range can be determined.

[0164] In this embodiment, by outputting test sounds and utilizing the fixed distance between the second sound collector in the projection screen and the audio output device, the range of sound propagation time can be accurately determined, thereby improving the detection accuracy of signal alignment and enhancing the echo cancellation effect.

[0165] In some embodiments, during the process of echo cancellation of media audio in the first external audio data after signal alignment, based on the signal-aligned re-acquisition reference data, to obtain far-field speech data, the controller is further configured to:

[0166] By using a pre-set acoustic model, echo prediction is performed based on the echo-acquired reference data after signal alignment, and echo estimation data is obtained.

[0167] Echo cancellation is performed on the media sound in the first external sound data after signal alignment based on the echo prediction data to obtain far-field speech data.

[0168] It should be noted that indirect echo can refer to the collection of sound emitted by an audio output device that enters the sound acquisition device after one or more reflections along different paths. Any movement of any object in the space will change the echo path; therefore, indirect echo is characterized by being multipath-dependent and time-varying.

[0169] The acoustic model refers to a mathematical model used to simulate and predict the reflection path of indirect echoes, enabling the subtraction or suppression of indirect echo components from received external sound data. Echo estimation data refers to the sound data predicted by the acoustic model based on echo sampling reference data, which is then collected by the first sound acquisition device after spatial reflection of the media sound. After the audio output device plays the media sound, the first sound acquisition device can simultaneously collect both the far-field sound and the media sound played by the audio output device. Since the media sound played by the audio output device undergoes multi-path transmission, the characteristic parameters of the echo path can be extracted by utilizing the correlation between the speaker signal and the multi-path echoes generated by the media sound, thereby establishing a corresponding acoustic model. By estimating the echo using the acoustic model and continuously modifying the filter coefficients to make the echo estimation data more closely approximate the real echo, and then subtracting the echo estimation data from the first external sound data, echo cancellation can be achieved.

[0170] For example, after signal alignment, the back-acquired reference data after signal alignment can be input into the acoustic model. The acoustic model predicts the echo estimation data corresponding to the back-acquired reference data. Then, based on the echo estimation data, echo cancellation is performed on the media sound in the first external sound data after signal alignment to obtain far-field speech data.

[0171] In this embodiment, indirect echo cancellation using an acoustic model can further improve the echo cancellation effect, thereby improving the accuracy of far-field speech recognition.

[0172] In some embodiments, a far-field speech recognition method is provided, such as Figure 11 As shown, this method is applied to a projector in a projection device, which also includes a projection screen. The projector is configured to project a display image onto the projection screen. The method includes:

[0173] Step 1102: Collect the first external sound data and receive the back-collection reference data. The back-collection reference data is wirelessly transmitted from the projection screen to the projection host when the media sound is output.

[0174] Step 1104: Align the acquired reference data and the first external sound data.

[0175] Step 1106: Based on the back-sampling reference data after signal alignment, perform echo cancellation on the media sound in the first external sound data after signal alignment to obtain far-field speech data.

[0176] Step 1108: Perform far-field speech recognition on the far-field speech data.

[0177] In some embodiments, signal alignment is performed on the re-acquired reference data and the first external sound data, including:

[0178] The back-collected reference data is decoded to obtain the back-collected reference signal, and the first external sound data is decoded to obtain the first external sound signal;

[0179] Acquire a first timestamp in the back-collection reference signal and a second timestamp in the first external sound signal, wherein the first timestamp is used to characterize the time when the back-collection reference signal is emitted from the audio output device, and the second timestamp is used to characterize the time when the first sound collector collects the first external sound signal.

[0180] Based on the preset sound propagation duration range and the difference between the second and first timestamps, the re-acquired reference data and the first external sound data are signal aligned.

[0181] In some embodiments, signal alignment is performed on the re-acquired reference data and the first external sound data based on a preset sound propagation duration range and the difference between the second timestamp and the first timestamp, including:

[0182] If the difference between the second timestamp and the first timestamp is greater than the preset sound propagation duration range, the decoding speed of the first external sound data is slowed down according to the difference value.

[0183] If the difference between the second timestamp and the first timestamp is less than the preset sound propagation duration range, the decoding speed of the re-collected reference data is slowed down according to the difference value.

[0184] In some embodiments, before aligning the re-acquired reference data and the first external sound data according to a preset sound propagation duration range and the difference between the second timestamp and the first timestamp, the method further includes:

[0185] Collect third external sound data corresponding to the test sound, wherein the test sound is output by an audio output device;

[0186] The third timestamp of the target signal is detected from the third external sound data, and the fourth timestamp of the target signal is obtained from the fourth external sound data. The fourth external sound data is obtained by the test sound collected by the second sound acquisition device in the projection screen.

[0187] Calculate the time difference between the third and fourth timestamps, and calculate the sound propagation duration based on the distance between the audio output device and the second sound acquisition device;

[0188] The range of sound propagation duration is determined by summing the time difference with the sound propagation duration.

[0189] In some embodiments, echo cancellation is performed on the media audio in the first external audio data after signal alignment, based on the signal-aligned re-acquisition reference data, to obtain far-field speech data, including:

[0190] By using a pre-set acoustic model, echo prediction is performed based on the echo-acquired reference data after signal alignment, and echo estimation data is obtained.

[0191] Echo cancellation is performed on the media sound in the first external sound data after signal alignment based on the echo prediction data to obtain far-field speech data.

[0192] In some embodiments, a far-field speech recognition method is provided, such as Figure 11 As shown, the method is applied to a projection screen in a projection device, which also includes a projection host configured to project a display image onto the projection screen. The method includes:

[0193] The system outputs media audio and wirelessly transmits the corresponding backtracking reference data to the projector host while outputting the media audio. This allows the projector host to perform signal alignment between the backtracking reference data and the acquired first external audio data. Based on the backtracking reference data after signal alignment, the system performs echo cancellation on the media audio in the first external audio data after signal alignment to obtain far-field speech data. The system then performs far-field speech recognition on the far-field speech data.

[0194] In some embodiments, wirelessly transmitting back-sampling reference data corresponding to the media audio to the projector host includes:

[0195] When the far-field voice function is triggered, the corresponding back-collection reference data of the media sound is wirelessly transmitted to the projector.

[0196] In some embodiments, the projection screen includes a second sound acquisition device configured to acquire second external sound data; when the far-field voice function is triggered, it wirelessly transmits back-collected reference data corresponding to the media sound to the projection host, including:

[0197] Acquire the second external sound data and the corresponding re-collected reference data of the media sound;

[0198] Based on the re-acquired reference data, echo cancellation is performed on the second external sound data to obtain the speech data to be tested.

[0199] If far-field voice data is present in the voice data to be tested, the far-field voice function is determined to be triggered, and the back-collected reference data is wirelessly transmitted to the projection host.

[0200] In some embodiments, when far-field speech data exists in the speech data to be tested, it is determined that the far-field speech function has been triggered, and the back-collected reference data is wirelessly transmitted to the projection host, including:

[0201] If the energy amplitude of the voice data to be tested exceeds the preset energy amplitude threshold, it is determined that there is far-field voice data in the voice data to be tested, and the far-field voice function is triggered.

[0202] In some embodiments, the projection screen includes a bandpass filter; when the energy amplitude of the voice data to be tested exceeds a preset energy amplitude threshold, it is determined that far-field voice data exists in the voice data to be tested, and it is determined that the far-field voice function is triggered, including:

[0203] By using a bandpass filter, the speech data to be tested is filtered to obtain the target human voice speech data corresponding to the preset human voice signal frequency band.

[0204] If the energy amplitude of the target human voice data exceeds the preset energy amplitude threshold, it is determined that there is far-field voice data in the voice data to be tested.

[0205] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods of the above embodiments.

[0206] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods of the above embodiments.

[0207] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods of the above embodiments.

[0208] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0209] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0210] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0211] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A projection device, characterized in that, Includes a projection screen and a projection host, the projection host being configured to project display images onto the projection screen; The projection host includes a first sound acquisition unit, a wireless transmission module, and a controller; the projection screen includes a signal processing module and an audio output device, wherein: The projection screen is configured to display a user interface; The audio output device is configured to output media sound, and when outputting the media sound, it sends the retrieval reference data corresponding to the media sound to the signal processing module; The signal processing module is configured to send the acquired reference data to the wireless transmission module; The wireless transmission module is configured to receive the back-collected reference data and send the back-collected reference data to the controller; The first sound collector is configured to collect first external sound data and send the first external sound data to the controller; The controller is configured as follows: Upon receiving the acquired reference data and the first external sound data, the acquired reference data is decoded to obtain the acquired reference signal, and the first external sound data is decoded to obtain the first external sound signal. Obtain a first timestamp in the back-collection reference signal and a second timestamp in the first external sound signal, wherein the first timestamp is used to characterize the time when the back-collection reference signal is emitted from the audio output device, and the second timestamp is used to characterize the time when the first sound collector collects the first external sound signal; Based on the preset sound propagation duration range and the difference between the second timestamp and the first timestamp, the back-collected reference data and the first external sound data are signal aligned; Based on the back-sampling reference data after signal alignment, echo cancellation is performed on the media sound in the first external sound data after signal alignment to obtain far-field speech data; Perform far-field speech recognition on the far-field speech data.

2. The projection device according to claim 1, characterized in that, During the process of sending the acquired reference data to the wireless transmission module, the signal processing module is further configured to: When the far-field voice function is triggered, the acquired reference data is sent to the wireless transmission module.

3. The projection device according to claim 2, characterized in that, The projection screen includes a second sound collector, which is configured to collect second external sound data. During the process of sending the acquired reference data to the wireless transmission module when the far-field voice function is triggered, the signal processing module is further configured to: Acquire the second external sound data and the re-collected reference data; Based on the acquired reference data, echo cancellation is performed on the second external sound data to obtain the speech data to be tested. If far-field voice data exists in the voice data to be tested, the far-field voice function is determined to be triggered, and the acquired reference data is sent to the wireless transmission module.

4. The projection device according to claim 3, characterized in that, In the process of determining that the far-field speech function is triggered when far-field speech data exists in the speech data to be tested, the signal processing module is further configured to: If the energy amplitude of the voice data to be tested exceeds a preset energy amplitude threshold, it is determined that there is far-field voice data in the voice data to be tested, and the far-field voice function is triggered.

5. The projection device according to claim 4, characterized in that, The signal processing module includes a bandpass filter; in the process of determining that far-field speech data exists in the speech data under test when the energy amplitude of the speech data under test exceeds a preset energy amplitude threshold, the signal processing module is further configured to: The speech data to be tested is filtered by a bandpass filter to obtain the target human voice speech data corresponding to the preset human voice signal frequency band. If the energy amplitude of the target human voice data exceeds a preset energy amplitude threshold, it is determined that there is far-field voice data in the voice data to be tested.

6. The projection device according to claim 1, characterized in that, During the process of aligning the acquired reference data and the first external sound data according to a preset sound propagation duration range and the difference between the second timestamp and the first timestamp, the controller is further configured to perform at least one of the following: If the difference between the second timestamp and the first timestamp is greater than a preset sound propagation duration range, the decoding speed of the first external sound data is slowed down according to the difference value. If the difference between the second timestamp and the first timestamp is less than a preset sound propagation duration range, the decoding speed of the re-acquired reference data is slowed down according to the difference value.

7. The projection device according to claim 1, characterized in that, The audio output device is configured to output test sounds; The first sound acquisition device is further configured to acquire third external sound data corresponding to the test sound; The projection screen includes a second sound collector, which is configured to collect fourth external sound data corresponding to the test sound. Before performing signal alignment on the re-acquired reference data and the first external sound data based on a preset sound propagation duration range and the difference between the second timestamp and the first timestamp, the controller is further configured to: Detect the third timestamp of the target signal from the third external sound data, and detect the fourth timestamp of the target signal from the fourth external sound data; Calculate the time difference between the third timestamp and the fourth timestamp, and calculate the sound propagation duration based on the distance between the audio output device and the second sound acquisition device; The range of sound propagation duration is determined based on the sum of the time difference and the sound propagation duration.

8. A far-field speech recognition method, characterized in that, The method is applied to a projection host in the projection device of claim 1, wherein the projection device further includes a projection screen, and the projection host is configured to project a display image onto the projection screen; the method includes: The system collects first external sound data and receives back-collected reference data, which is wirelessly transmitted from the projection screen to the projection host when outputting media sound. The acquired reference data and the first external sound data are signal aligned; Based on the back-sampling reference data after signal alignment, echo cancellation is performed on the media sound in the first external sound data after signal alignment to obtain far-field speech data. Perform far-field speech recognition on the far-field speech data.

9. A far-field speech recognition method, characterized in that, The method is applied to a projection screen in the projection device of claim 1, wherein the projection device further includes a projection host configured to project a display image onto the projection screen; the method includes: The system outputs media audio and, while outputting the media audio, wirelessly transmits the back-collection reference data corresponding to the media audio to the projector host, so that the projector host performs signal alignment between the back-collection reference data and the first external audio data acquired. Based on the signal-aligned back-collection reference data, the system performs echo cancellation on the media audio in the signal-aligned first external audio data to obtain far-field speech data, and performs far-field speech recognition on the far-field speech data.

Citation Information

Patent Citations

  • Audio processing method and device, electronic equipment and storage medium

    CN111583952A

  • Speech recognition method, device, apparatus and computer-readable storage medium

    US20190325888A1