Display device, server and audio noise reduction and model training method thereof

By acquiring the audio interference characteristics of the WiFi module in the display device and performing software noise reduction, the problem of WiFi module interference with the microphone was solved, thereby improving the clarity of audio data and reducing costs.

CN119865647BActive Publication Date: 2026-01-02HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411907621.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2026-01-02
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

In traditional technologies, WiFi modules interfere with the microphone's voice capture function, and existing solutions increase cost or operational complexity.

Method used

The controller acquires audio interference characteristics of the WiFi module in different operating modes, and uses software to perform noise reduction processing to eliminate interference characteristics and improve audio data clarity, avoiding hardware adjustments and complicated analysis.

Benefits of technology

It effectively eliminates WiFi module interference without increasing hardware costs and complexity, thus improving the clarity and operability of audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119865647B_ABST
    Figure CN119865647B_ABST
Patent Text Reader

Abstract

The application relates to a display device, a server and an audio noise reduction and model training method thereof, and relates to the technical field of intelligent voice. The display device comprises a WiFi module configured to perform network search in different working modes; a microphone arranged on the same circuit board as the WiFi module and configured to collect audio data; and at least one controller connected with the microphone and the WiFi module and configured to: control the microphone to collect first audio data; the first audio data at least comprises interference audio of the WiFi module in a current working mode; take audio interference features of the WiFi module in the current working mode as target interference features; wherein the feature frequency or the feature strength of the audio interference features in different working modes are different; and perform noise reduction processing on the first audio data according to the feature frequency and the feature strength of the target interference features, to obtain second audio data, thereby improving the intelligibility of the second audio data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent voice, in particular to a display device, a server and an audio noise reduction and model training method thereof. BACKGROUND

[0002] With the rapid development of intelligent voice technology, products with intelligent voice function, such as display devices, are increasing. In order to save product cost and product internal space, it has become a new development trend to replace multiple single hardware modules with a multi-in-one module. For example, in a display device, a WiFi (Wireless Fidelity) module, a light sensing module, a Bluetooth module and a voice module are arranged on one circuit board.

[0003] However, since the WiFi module is close to the microphone, during the working process of the WiFi module, the scanning signals emitted by the WiFi module are reflected to the microphone circuit through the internal structure, which will interfere with the data collected by the microphone.

[0004] In the traditional technology, in order to overcome the above problems, a shielding cloth is usually added around the microphone circuit to shield the radiation signals, or the relative position relationship between the WiFi module and the microphone in the circuit board is adjusted to avoid the radiation signals. However, the addition of the shielding cloth will greatly increase the cost, and the adjustment of the distribution of the modules in the circuit board requires a large amount of manpower and time cost to analyze the radiation path of multiple antennas in the WiFi, which is also relatively complex and has poor operability. SUMMARY

[0005] The present application provides a display device, a server and an audio noise reduction and model training method thereof to solve the influence of the WiFi module on the voice collection function of the microphone in the traditional technology, while taking into account the cost investment and operability.

[0006] In a first aspect, some embodiments provide a display device, comprising:

[0007] a WiFi module configured to perform network search in different working modes;

[0008] a microphone arranged on the same circuit board as the WiFi module and configured to collect audio data;

[0009] at least one controller connected to the microphone and the WiFi module and configured to:

[0010] control the microphone to collect first audio data; wherein the first audio data at least includes interference audio of the WiFi module in the current working mode;

[0011] The audio interference feature of the WiFi module in the current working mode is taken as a target interference feature, wherein the feature frequency or the feature strength of the audio interference feature in different working modes is different;

[0012] The first audio data is denoised according to the feature frequency and the feature strength of the target interference feature, to obtain second audio data.

[0013] In some embodiments, the target interference feature of the WiFi module in the current working mode for network searching is obtained by a controller in the display device, and the first audio data collected by the microphone is denoised according to the feature frequency and the feature strength of the target interference feature, to obtain second audio data, so that the obtained second audio data can deduct the radiation interference generated by the WiFi module, and the clarity of the second audio data is improved. Since the above denoising process adopts a pure software manner, no other hardware structures or materials such as shielding cloth need to be arranged, the hardware cost is reduced, and no relative position relationship between the modules arranged in the circuit board needs to be adjusted, so no complicated radiation path analysis is needed, the labor and time costs are reduced, and the operability and universality are better. In addition, since the scanning frequency and / or scanning strength of the WiFi module in different working modes are different, the feature frequency or the feature strength of the audio interference feature in different working modes also has certain differences, so the audio interference feature of the WiFi module in the current working mode is taken as the target interference feature, instead of using a unified audio interference feature, which can improve the accuracy of the selected target interference feature, thereby avoiding the situation that the key information is lost or the noise is not completely deducted, resulting in poor accuracy of the second audio data, due to the selection of the wrong target interference feature.

[0014] In a second aspect, some embodiments further provide a server, comprising:

[0015] A communication device is in communication connection with the sample display device;

[0016] At least one processor is connected with the communication device and is configured to:

[0017] Obtain a first audio sample generated during the operation of the sample display device when the indoor noise intensity is lower than a first intensity threshold, and a second audio sample generated during the operation of the sample display device after the WiFi driver is unloaded;

[0018] Determine a first difference feature between the first audio sample and the second audio sample;

[0019] Take the first audio sample as input data and take the first difference feature as a label to train the feature extraction network to be trained until the difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition.

[0020] The feature extraction network is configured to extract an audio interference feature of the audio generated by the WiFi module of the display device in a network search condition.

[0021] The standard feature extraction network is trained in the following manner:

[0022] The third audio sample generated by the sample display device in a soundproof environment and the fourth audio sample generated by the sample display device after uninstalling the WiFi driver are obtained.

[0023] The second difference feature between the third audio sample and the fourth audio sample is determined.

[0024] The third audio sample is taken as input data, and the second difference feature is taken as a label to train the standard feature extraction network.

[0025] In some embodiments, the first audio sample generated by the sample display device in a soundproof environment and the second audio sample generated by the sample display device after uninstalling the WiFi driver are obtained by the server when the indoor noise intensity is lower than a first intensity threshold. The first audio sample is taken as input data, and the first difference feature between the first audio sample and the second audio sample is taken as a label to train the feature extraction network until the difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition. Since the standard feature extraction network is trained by taking the third audio sample generated by the sample display device in a soundproof environment as input data and taking the second difference feature between the third audio sample and the fourth audio sample generated by the sample display device in a soundproof environment after uninstalling the WiFi driver as a label, the feature extraction capability of the standard feature extraction network is good. Since the feature extraction capability of the standard feature extraction network in a soundproof room environment is good, it is not suitable for a regular indoor environment, and therefore the standard feature extraction network is used as a reference for the feature extraction network training process applied to a regular indoor environment. When the difference between the trained feature extraction network and the standard feature extraction network meets the preset difference condition, the training of the feature extraction network is terminated, and the feature extraction capability of the feature extraction network is ensured to be equivalent to that of the standard feature extraction network.

[0026] In a third aspect, some embodiments further provide an audio noise reduction method applied to a display device, comprising:

[0027] obtaining first audio data collected by a microphone of the display device; wherein the first audio data at least includes interference audio of the WiFi module in a current working mode; and

[0028] The WiFi module of the display device acquires an audio interference feature in a current working mode under network search as a target interference feature; wherein the feature frequency or feature strength of the audio interference feature under different working modes is different;

[0029] According to the feature frequency and feature strength of the target interference feature, the first audio data is denoised to obtain second audio data.

[0030] In some embodiments, by acquiring the target interference feature of the WiFi module under the current working mode for network search, and according to the feature frequency and feature strength of the target interference feature, the first audio data collected by the microphone is denoised to obtain the second audio data, so that the obtained second audio data can deduct the radiation interference generated by the WiFi module, and the clarity of the second audio data is improved. Since the above denoising process adopts a pure software manner, it does not need to set other hardware structures or materials such as shielding cloth, reduces the hardware cost, and also does not need to adjust the relative position relationship between the modules set in the circuit board, so it does not need to perform complex radiation path analysis, reduces the labor and time cost, and has better operability and universality. In addition, since the scanning frequency and / or scanning strength of the WiFi module under different working modes are different, the feature frequency or feature strength of the audio interference feature under different working modes also has certain differences, so the audio interference feature of the WiFi module under the current working mode is taken as the target interference feature, rather than using a unified audio interference feature, which can improve the accuracy of the selected target interference feature, thereby avoiding the situation that the key information is lost or the noise is not completely deducted due to the selection of the wrong target interference feature, resulting in poor accuracy of the second audio data.

[0031] In a fourth aspect, some embodiments further provide a model training method applied to a server, comprising:

[0032] Acquire a first audio sample generated during the operation of a sample display device under the condition that the indoor environment meets the preset quiet condition, and a second audio sample generated during the operation of the sample display device after uninstalling the WiFi driver;

[0033] Determine a first difference feature between the first audio sample and the second audio sample;

[0034] Take the first audio sample as input data and the first difference feature as a label to train the feature extraction network to be trained until the difference between the trained feature extraction network and the standard feature extraction network meets the preset difference condition;

[0035] The feature extraction network is used to extract the audio interference feature of the audio generated by the WiFi module of the display device under network search;

[0036] The standard feature extraction network is trained in the following manner:

[0037] The third audio sample generated during the operation of the sample display device in a soundproof environment and the fourth audio sample generated during the operation of the sample display device after uninstalling the WiFi driver are obtained.

[0038] The second difference feature between the third audio sample and the fourth audio sample is determined.

[0039] The third audio sample is taken as input data, and the second difference feature is taken as a label to train the standard feature extraction network.

[0040] In some embodiments, the server obtains the first audio sample generated during the operation of the sample display device when the indoor noise intensity is lower than a first intensity threshold and the second audio sample generated during the operation of the sample display device after uninstalling the WiFi driver. The first audio sample is taken as input data, and the first difference feature between the first audio sample and the second audio sample is taken as a label to train the feature extraction network until the difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition. Since the standard feature extraction network is trained by taking the third audio sample generated during the operation of the sample display device in a soundproof environment as input data and taking the second difference feature between the third audio sample and the fourth audio sample generated during the operation of the sample display device in the soundproof environment after uninstalling the WiFi driver as a label, the feature extraction capability of the standard feature extraction network is good. Since the feature extraction capability of the standard feature extraction network in the soundproof room environment is good, it is not suitable for the conventional indoor environment, so the standard feature extraction network is used as a reference for the feature extraction network training process applied to the conventional indoor environment. When the difference between the trained feature extraction network and the standard feature extraction network meets the preset difference condition, the training of the feature extraction network is terminated, which ensures that the feature extraction capability of the feature extraction network is equivalent to that of the standard feature extraction network. Accordingly, the audio interference feature of the audio generated by the WiFi module in the display device under network search is extracted based on the trained feature extraction network, which improves the accuracy of the extracted audio interference feature.

[0041] In a fifth aspect, some embodiments further provide an audio noise reduction device configured in a display device, comprising:

[0042] A first obtaining module is configured to obtain first audio data collected by a microphone of the display device, wherein the first audio data at least includes interference audio of the WiFi module under a current working mode.

[0043] The first obtaining module is further configured to obtain an audio interference feature of the WiFi module of the display device in a current working mode for network searching as a target interference feature; wherein the feature frequency or the feature strength of the audio interference feature in different working modes is different;

[0044] The noise reduction module is configured to perform noise reduction processing on the first audio data according to the feature frequency and the feature strength of the target interference feature, to obtain second audio data.

[0045] In some embodiments, the first obtaining module obtains the target interference feature of the WiFi module in the current working mode for network searching, and the noise reduction module performs noise reduction processing on the first audio data collected by the microphone according to the feature frequency and the feature strength of the target interference feature, to obtain second audio data, so that the obtained second audio data can deduct the radiation interference generated by the WiFi module, and the clarity of the second audio data is improved. Since the above noise reduction process adopts a pure software manner, no other hardware structures or materials such as shielding cloth need to be arranged, the hardware cost is reduced, and no relative position relationship between the modules arranged in the circuit board needs to be adjusted, so no complex radiation path analysis is needed, the labor and time costs are reduced, and the operability and universality are better. In addition, since the scanning frequency and / or scanning strength of the WiFi module in different working modes are different, the feature frequency or the feature strength of the audio interference feature in different working modes also has certain differences, so the audio interference feature of the WiFi module in the current working mode is taken as the target interference feature, instead of using a unified audio interference feature, which can improve the accuracy of the selected target interference feature, thereby avoiding the situation that the key information is lost or the noise is not completely deducted, resulting in poor accuracy of the second audio data, due to the selection of the wrong target interference feature.

[0046] In a sixth aspect, some embodiments further provide a model training device configured in a server, comprising:

[0047] The second obtaining module is configured to obtain a first audio sample generated by a sample display device during operation when the indoor noise intensity is lower than a first intensity threshold, and a second audio sample generated by the sample display device during operation after uninstalling a WiFi driver;

[0048] The determining module is configured to determine a first difference feature between the first audio sample and the second audio sample;

[0049] The training module is configured to take the first audio sample as input data, take the first difference feature as a label, train the feature extraction network to be trained, until the difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition.

[0050] The feature extraction network is configured to extract an audio interference feature of the audio generated by the WiFi module of the display device in a network search condition.

[0051] The standard feature extraction network is trained in the following manner:

[0052] The third audio sample generated by the sample display device in a soundproof environment and the fourth audio sample generated by the sample display device after uninstalling the WiFi driver are obtained.

[0053] The second difference feature between the third audio sample and the fourth audio sample is determined.

[0054] The third audio sample is taken as input data, and the second difference feature is taken as a label to train the standard feature extraction network.

[0055] In some embodiments, the first audio sample generated by the sample display device in a soundproof environment and the second audio sample generated by the sample display device after uninstalling the WiFi driver are obtained by setting the second acquisition module when the indoor noise intensity is lower than the first intensity threshold. The first difference feature between the first audio sample and the second audio sample is determined by setting the determination module. The feature extraction network is trained by setting the training module, taking the first audio sample as input data, and taking the first difference feature as a label until the difference between the trained feature extraction network and the standard feature extraction network meets the preset difference condition. Since the standard feature extraction network is trained by taking the third audio sample generated by the sample display device in a soundproof environment as input data, and taking the second difference feature between the third audio sample and the fourth audio sample generated by the sample display device in a soundproof environment after uninstalling the WiFi driver as a label, the feature extraction capability of the standard feature extraction network is good. Since the feature extraction capability of the standard feature extraction network in the soundproof room environment is good, it is not suitable for the conventional indoor environment, so the standard feature extraction network is used as a reference for the feature extraction network training process applied to the conventional indoor environment. When the difference between the trained feature extraction network and the standard feature extraction network meets the preset difference condition, the training of the feature extraction network is terminated, ensuring that the feature extraction capability of the feature extraction network is equivalent to that of the standard feature extraction network. Accordingly, the audio interference feature of the audio generated by the WiFi module of the display device in a network search condition is extracted based on the trained feature extraction network, improving the accuracy of the extracted audio interference feature.

[0056] In a seventh aspect, some embodiments further provide a computer readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the steps of the method provided in some embodiments of the third aspect or the fourth aspect.

[0057] In an eighth aspect, some embodiments further provide a computer program product comprising a computer program, the computer program being executable by a processor to implement the steps of the method provided in some embodiments of the third aspect or the fourth aspect. BRIEF DESCRIPTION OF DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0059] Figure 1 A schematic diagram of an operation scenario between a display device and a control device according to some embodiments of the present application;

[0060] Figure 2 A schematic diagram of a hardware configuration of a display device according to some embodiments of the present application;

[0061] Figure 3 A schematic diagram of a hardware configuration of a control device according to some embodiments of the present application;

[0062] Figure 4 A schematic diagram of a software configuration of a display device according to some embodiments of the present application;

[0063] Figure 5A A schematic diagram of an evolution process of a circuit board according to some embodiments of the present application;

[0064] Figure 5B A schematic diagram of a hardware structure of a display device according to some embodiments of the present application;

[0065] Figure 5C An interference spectrum diagram of a WiFi module on a microphone according to some embodiments of the present application;

[0066] Figure 6 A flowchart of an audio noise reduction method according to some embodiments of the present application;

[0067] Figure 7 A flowchart of a step of obtaining a target interference feature according to some embodiments of the present application;

[0068] Figure 8 A schematic diagram of a clipping result of an audio sample according to some embodiments of the present application;

[0069] Figure 9 A flowchart of a model training method provided for some embodiments of the present application;

[0070] Figure 10 A flowchart of an audio noise reduction method provided for some embodiments of the present application;

[0071] Figure 11 A structural block diagram of an audio noise reduction device provided for some embodiments of the present application;

[0072] Figure 12 A structural block diagram of a model training device provided for some embodiments of the present application;

[0073] Figure 13 An internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0074] The embodiments will be described in detail with reference to the drawings, wherein the same reference numerals can indicate the same or similar elements throughout the several views. The following detailed description is not representative of all embodiments consistent with the present application. It is merely exemplary of a system and method consistent with some aspects of the present application as recited by the claims.

[0075] It should be noted that the brief description of terms in the present application is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.

[0076] The terms "first", "second", "third", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar or like objects or entities, and do not necessarily mean a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances.

[0077] The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not necessarily limit to all components clearly listed, but can include other components not clearly listed or inherent to these products or devices.

[0078] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing a function associated with that element.

[0079] In the embodiments of the present application, the display device 200 generally refers to a device with picture display and data processing capabilities. For example, the display device 200 includes, but is not limited to, a smart television, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, and the like.

[0080] Figure 1 The present application provides a schematic diagram of the operation scenario between the display device and the control device. As shown in Figure 1 The user can operate the display device 200 through a touch operation, a mobile terminal 300, and a control device 100. For example, the control device 100 can be a remote controller, a stylus, a handle, and the like.

[0081] The mobile terminal 300 can be used as a control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device to establish a communication connection with the display device 200 and perform data interaction. In some embodiments, the mobile terminal 300 can install a software application on the display device 200, implement connection communication through a network communication protocol, and achieve the purpose of one-to-one control operation and data communication. The mobile terminal 300 can also transmit audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve a synchronous display function.

[0082] As shown in Figure 1 It is also shown in that the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 can be allowed to communicate through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0083] The display device 200 can provide a broadcast receiving television function, and can additionally provide a smart network television function with computer support, including but not limited to a network television, a smart television, an Internet protocol television (IPTV), and the like.

[0084] Figure 2 The present application provides a schematic diagram of the operation scenario between the display device and the control device. As shown in Figure 1 A hardware configuration block diagram of the display device 200 is shown in

[0085] In some embodiments, the display device 200 can include at least one of a tuning demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface 280.

[0086] In some embodiments, the detector 230 is configured to collect signals of the external environment or the external interaction. For example, the detector 230 includes a light receiver configured to collect ambient light intensity; or the detector 230 includes an image collector, such as a camera, configured to collect an external environment scene, a user attribute, or a user interaction gesture; or the detector 230 includes a sound collector, such as a microphone, configured to receive external sound, for example, to collect user voice data.

[0087] In some embodiments, the display 260 includes a display component configured to present a picture, and a driving component configured to drive the image display. The display 260 is configured to receive an image signal output from the controller 250 for display. For example, the display 260 can be configured to display video content, image content, and components of a menu control interface, a user control UI interface, and the like.

[0088] In some embodiments, the communication device 220 is a component configured to communicate with an external device or a server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication manners. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0089] The communication device 220 can be configured to connect the display device 200 to the external device or the server 400 in a wireless or wired manner. The wired connection can be achieved by connecting the display device 200 to the external device through a data line, an interface, or the like. The wireless connection can be achieved by connecting the display device 200 to the external device through a wireless signal or a wireless network. The display device 200 can be directly connected to the external device, or can be indirectly connected to the external device through a gateway, a router, a connection device, or the like.

[0090] In some embodiments, the controller 250 can include at least one of a central processor, a video processor, an audio processor, a graphics processor, a power processor, a first interface to an n-th interface for input / output, and the like. The controller 250 controls the operation of the display device and responds to user operations by controlling various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0091] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, i.e. the tuner demodulator 210 can also be located in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.

[0092] In some embodiments, the user can input user commands through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI).

[0093] In some embodiments, the audio output device 270 can be a native loudspeaker of the display device 200, or can be an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 can also be provided with an external audio output terminal, and the audio output device can be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0094] Figure 3 The hardware configuration block diagram of the control device provided in some embodiments of the present application is shown in Figure 1 FIG. 1. As shown in Figure 3 the control device 100 can include a controller 110, a communication interface 130, a user input / output interface 140, a power supply 180, and a memory 190.

[0095] The control device 100 is configured to control the display device 200, and can receive the input operation instructions of the user, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, thereby playing the role of an intermediary between the user and the display device 200.

[0096] In some embodiments, the control device 100 can be a smart device. For example, the control device 100 can install various applications for controlling the display device 200 according to the user's needs.

[0097] In some embodiments, as shown in Figure 1 the mobile terminal 300 or other smart electronic devices can play a similar function of the control device 100 after installing the application for controlling the display device 200.

[0098] The controller 110 includes a processor 112 and a RAM (Random Access Memory) 113 and a ROM (Read-Only Memory) 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, and the communication and cooperation between the internal components, as well as the data processing functions of the external and internal.

[0099] The communication interface 130, under the control of the controller 110, implements communication of control signals and data signals with the display device 200. The communication interface 130 can include at least one of a WiFi chip 131, a Bluetooth module 132, a NFC (Near Field Communication) module 133, and other near field communication modules.

[0100] The user input / output interface 140, wherein the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a key 144, and other input interfaces.

[0101] In some embodiments, the control device 100 includes at least one of the communication interface 130 and the input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC, etc. module, which can encode user input instructions through a WiFi protocol, or a Bluetooth protocol, or a NFC protocol and send them to the display device 200.

[0102] The storage 190 is used to store various running programs, data and applications for driving and controlling the control device 100 under the control of the controller. The storage 190 can store various control signal instructions input by the user.

[0103] The power supply 180 is used to provide operating power support for the elements of the control device 100 under the control of the controller.

[0104] In order to perform user interaction, in some embodiments, the display device 200 can run an operating system. The operating system is a computer program used to manage and control hardware resources and software resources in the display device 200. The operating system can provide a user interface to allow the user to interact with the display device 200 and support the running of various application programs.

[0105] It should be noted that the operating system can be a native operating system based on a specific operating platform, or a third-party operating system based on a specific operating platform, or an independent operating system specially developed for the display device.

[0106] The operating system can be divided into different modules or levels according to the functions implemented, for example, as shown in FIG. 1, in some embodiments, the system is divided into four layers, from top to bottom, the Applications layer (referred to as "application layer"), the Application Framework layer (referred to as "framework layer"), the system library layer and the kernel layer. Figure 4

[0107] ​In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0108] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0109] like Figure 4 As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0110] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0111] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions implemented by the framework layer.

[0112] In some embodiments, the kernel layer is a functional layer between the hardware and the software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, memory management, etc. For example, as shown in Figure 4 The kernel layer can be configured with hardware drivers, and the drivers contained in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB (Universal Serial Bus) driver, HDMI (High-Definition Multimedia Interface) driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power supply driver, etc.

[0113] It should be noted that the above examples are only a simple division of the functions of the operating system, and do not constitute a limitation on the specific operating system form of the display device 200 in the embodiments of the present application. According to the function of the display device, the type of the operating system, and other factors, the number and specific type of the layers contained in the operating system can be in other forms.

[0114] Referring to Figure 5A With the trend of miniaturization of product circuits, the display device 200 gradually evolves from the traditional microphone module (i.e. the aforementioned voice module) and the WiFi module being separately arranged on different circuit boards to the microphone module and the WiFi module being arranged on the same small circuit board to obtain a multi-in-one module.

[0115] In combination with the hardware structure diagram of the display device as shown in Figure 5B It can be known from the hardware structure diagram of the display device that the controller 110 sends a first control signal to the WiFi module to control the WiFi module to send a probe request signal at a certain frequency for network search, and the controller 110 sends a second control signal to the microphone to control the microphone to collect an audio signal and send the audio signal to the controller 110 for subsequent processing. Since the probe request signal is reflected after being reflected, if it is reflected into the microphone circuit, it will be mixed with the microphone signal and interfere with the journey, etc. Referring to Figure 5CAs shown in the interference spectrum diagram of the WiFi module to the microphone, since the transmission time of the probe request signal is usually relatively fixed (often set by software), the occurrence time of such interference corresponds to the time when the WiFi module sends the probe request signal. It is worth noting that since the scanning frequency of the WiFi module is different in different working modes, and the signal strength of the transmitted probe request signal is also different, under the same environment, due to the difference in the working mode of the WiFi module, the audio interference characteristics may also have certain differences.

[0116] In order to reduce the audio interference of the WiFi module, the controller 110 can pre-acquire the target interference characteristics of the WiFi module under the current working mode, and deduct it from the audio data collected by the microphone, thereby improving the clarity of the audio data, and further improving the voice performance and reducing the influence on the voice function of the display device 200.

[0117] In some optional embodiments, referring to Figure 6 , an audio noise reduction method is provided, characterized in that it is applied to the controller 110 in the display device 200, comprising:

[0118] S610, acquiring first audio data collected by the microphone; wherein the first audio data at least includes interference audio of the WiFi module under the current working mode.

[0119] The controller 110 can control the microphone to collect the first audio data in real time, and acquire the first audio data transmitted by the microphone. Since the WiFi module works, it will periodically transmit probe request signals according to the set frequency for network search, therefore, the first audio data collected by the microphone carries interference noise introduced by the probe request signal, that is, interference audio.

[0120] Optionally, the controller 110 can control the microphone to collect the first audio data in real time. Or optionally, the controller 110 can also control the microphone to collect the first audio data only when the voice function is turned on. The advantage of this is that when the voice function is not turned on, it reduces the hardware loss and unnecessary occupation of computing resources caused by irrelevant data collection.

[0121] S620, acquiring the audio interference characteristics of the WiFi module of the display device under the current working mode for network search as the target interference characteristics.

[0122] Wherein, the audio interference characteristics are used to represent the interference noise introduced by the above-mentioned WiFi module, and carry the key characteristics in the interference noise. Wherein, the feature frequency or feature strength of the audio interference characteristics under different working modes is different.

[0123] In some embodiments, the audio interference features corresponding to the network search in different working modes of the self-WiFi module can be stored in the memory 190 of the display device 200, and when audio noise reduction is needed, the audio interference feature in the current working mode of the WiFi module is acquired to obtain the target interference feature.

[0124] For example, since the WiFi module can work in different working modes, the scanning properties (such as scanning frequency and strength of probe request signal) of the WiFi module in different working modes can be different, for example, the transmission frequency of the transmitted probe request signal, that is, the scanning frequency, or the signal strength of the transmitted probe request signal, that is, the signal strength, and so on. Therefore, the audio interference features corresponding to the WiFi module in different working modes are also different, so the audio interference features matched with the corresponding working modes can be stored for different working modes to avoid the influence of mismatched working modes. Accordingly, the target interference feature can be acquired based on the current working mode of the WiFi module.

[0125] For example, the working mode can include a self-scanning mode and an application scanning mode. The self-scanning mode can be understood as the working mode entered when the display device 200 enters a wireless network coverage area and is not connected to a network. The application scanning mode can be understood as the working mode entered when a network search is performed based on a client that provides services in a network. Of course, the working mode of the WiFi module can also be configured with other modes according to the WiFi model or settings, and the present embodiment does not make any limitation thereto.

[0126] For example, the controller 110 pre-sets the working mode of the WiFi module when controlling the WiFi module to perform a network search, so the current working mode of the WiFi module can be directly used as the current working mode.

[0127] Optionally, the working mode identifier of the WiFi module can be acquired from the memory or local cache, and the working mode corresponding to the working mode identifier is used as the current working mode.

[0128] For example, the audio interference features in different working modes are pre-stored in the memory 190 of the display device 200. Accordingly, the controller 110 searches the audio interference feature corresponding to the current working mode from the memory 190 as the target interference feature.

[0129] It should be noted that the audio interference features in different working modes can also be stored in the server; accordingly, in the case of needing to acquire the audio interference features, the controller 110 can send an access request based on the current working mode to the server through the self-communication device 220, and receive the audio interference features corresponding to the current working mode fed back by the server as the target interference features for subsequent use.

[0130] Notably, by storing the audio interference features corresponding to different working modes in the memory 190, the network side does not need to access an external server to acquire the target interference features, improving the convenience and efficiency of target interference feature acquisition, and reducing the dependence on network functions, so that the target interference features can still be acquired when the display device 200 is not connected to the network, further improving the universality and convenience of target interference feature acquisition.S630, according to the feature frequency and feature intensity of the target interference feature, the first audio data is processed to obtain the second audio data.

[0131] Exemplarily, the target interference feature can be deducted from the first audio data, so as to realize the noise reduction processing of the first audio data and obtain the second audio data. It can be understood that since the noise interference of the WiFi module has been removed from the second audio data, the target audio is clearer than the first audio data.

[0132] In an optional embodiment, the target interference feature can represent the noise feature of a single working cycle of the WiFi module, and accordingly, the first audio data can be cut according to the emission time of the probe request signal of the WiFi module during the first audio data acquisition process to obtain at least one first audio segment; for each first audio segment, the target interference feature in the first audio segment is deducted to obtain a second audio segment corresponding to the first audio segment; the corresponding second audio segments are spliced in the order of the first audio segment division to obtain the second audio data.

[0133] In the optional embodiment, the target interference feature of the WiFi module in the current working mode is obtained, and the first audio data collected by the microphone is processed to obtain the second audio data, so that the second audio data can deduct the radiation interference generated by the WiFi module, and the intelligibility of the second audio data is improved. Since the above noise reduction process adopts a pure software mode, it does not need to set other hardware structures or materials such as shielding cloth, reduces the hardware cost, and does not need to adjust the relative position relationship between the modules arranged in the circuit board, so it does not need to perform complicated radiation path analysis, reduces the labor and time cost, and has better operability and universality. In addition, since the scanning frequency and / or scanning intensity of the WiFi module in different working modes are different, the feature frequency or feature intensity of the audio interference feature in different working modes also has certain differences, so the audio interference feature of the WiFi module in the current working mode is taken as the target interference feature, instead of using a unified audio interference feature, which can improve the accuracy of the selected target interference feature, thereby avoiding the situation that the key information is lost or the noise is not completely deducted, resulting in poor accuracy of the second audio data.

[0134] On the basis of the technical solutions of the above embodiments, in some embodiments, the target interference feature acquisition step is refined.

[0135] Referring to Figure 7 The target interference feature acquisition step shown in the figure includes:

[0136] S710, in the case that the indoor noise intensity is lower than the first intensity threshold, controlling the microphone to collect the interference audio of the WiFi module in the current working mode.

[0137] The first intensity threshold can be set by the technician according to the need or experience, or determined repeatedly through a large number of experiments.

[0138] The interference audio is used to represent the environmental noise of the current indoor environment of the display device 200, which includes the interference noise generated by the WiFi module in the current working mode. The current indoor environment can be understood as the environment in which the space where the display device 200 is installed is currently located.

[0139] For example, in the case that the indoor environment meets the preset quiet condition, the controller 110 controls the microphone to collect the audio data to obtain the interference audio.

[0140] Optionally, the audio acquisition instruction information is output through the user output interface 140 or the display interface of the display 260; the controller 110 controls the microphone to collect audio data in response to the audio acquisition operation based on the audio acquisition instruction information; the collected audio data is analyzed to identify whether the indoor environment meets the preset quiet condition, that is, whether the indoor noise intensity is lower than the first intensity threshold; and in the case where the preset quiet condition is met, the collected audio data is taken as the interference audio.

[0141] Optionally, whether the indoor environment meets the preset quiet condition can be determined based on the indoor noise intensity of the collected indoor environment. Specifically, in the case where the indoor noise intensity is lower than the first intensity threshold, it is determined that the indoor environment meets the preset quiet condition.

[0142] In some embodiments, the controller 110 can control the microphone to collect the interference audio in the case where the display device 200 is currently started for the first time.

[0143] Since the WiFi module of the same display device has a relatively small change range of the collected interference audio in a relatively unchanged indoor environment (for example, the set position change is less than a preset distance threshold) in a long period of time (for example, within one year), in another optional embodiment, the controller 110 can also control the microphone to collect the interference audio of the WiFi module in the current working mode in the case where the indoor environment meets the preset quiet condition, that is, the indoor noise intensity is lower than the first intensity threshold, according to a certain collection period, without frequent collection of the interference audio.

[0144] Since different display devices have different degrees of device wear and tear over time, the change range of the interference audio is also different. Therefore, in an optional embodiment, the controller 110 can also obtain the device type of the display device from the memory 190; and control the microphone to collect the interference audio of the WiFi module in the current working mode in the case where the indoor environment meets the preset quiet condition according to the collection period corresponding to the device type. For example, the collection period for an air conditioner device can be 2 years, the collection period for a television device can be 1 year, and the like.

[0145] S720, performing feature extraction on the interference audio based on the trained feature extraction network to obtain target interference features.

[0146] The feature extraction network can be understood as a network model specially used for performing audio interference feature extraction, and can be implemented based on a traditional machine learning model or a deep learning model. The network structure of the feature extraction network is not limited in the embodiment.

[0147] In some embodiments, a call request for the pre-trained feature extraction network can be sent to the server to enable the server to perform feature extraction on the interference audio based on the trained feature extraction network; the feature extraction result fed back by the server is acquired; and the target interference feature is included in the feature extraction result.

[0148] For example, the controller 110 can generate a call request for the trained feature extraction network based on the interference audio; send the call request to the server through the communication device 220; the server performs feature extraction on the interference audio based on the trained feature extraction network in response to the call request, and feeds back the feature extraction result to the display device 200; the communication device 220 of the display device 200 receives the feature extraction result fed back by the server; and the controller 110 parses the feature extraction result to obtain the target interference feature carried thereby.

[0149] It can be understood that by setting the trained feature extraction network on the server outside the display device 200, and by calling the server to replace the display device 200 to extract the target interference feature through an external interface, the requirement for the computing power of the display device 200 is reduced.

[0150] In other embodiments, the trained feature extraction network can also be stored in the memory 190 of the display device 200; accordingly, the controller 110 runs the trained feature extraction network stored in the memory 190 to extract the target interference feature from the interference audio.

[0151] It can be understood that by running the feature extraction network by the display device 200 to extract the target interference feature from the interference audio, the display device 200 does not need to rely on external devices, and the use of communication bandwidth is reduced. In addition, since the target interference feature of the same display device under the same working mode changes less in a long period of time (for example, within one year), the extraction of the target interference feature can be performed periodically according to the above period of time, the occupation of the computing power resources of the display device 200 is reduced, and the convenience of the extraction of the target interference feature is improved.

[0152] It is worth noting that no matter which way is used to extract the target interference feature, the extracted target interference feature can be stored in the memory 190 of the display device 200 for searching and acquiring the target interference feature in the audio noise reduction process before the next extraction of the target interference feature.

[0153] The optional embodiment described above can extract the target interference feature in real time or at a fixed time by introducing a trained feature extraction network, thereby effectively dealing with the case where the relative position between the WiFi module and the microphone circuit changes due to aging of the WiFi module or maintenance or replacement of the WiFi module, resulting in a large difference in the target interference feature, avoiding the influence of a fixed target interference feature on the audio noise reduction effect, and improving the flexibility of audio noise reduction.

[0154] In some embodiments, the feature extraction network described above can be trained on the display device 200 or on another device, such as a server, other than the display device 200.

[0155] In some embodiments, the initial audio sample generated during the operation of the sample display device and the reference audio sample generated during the operation of the sample display device after uninstalling the WiFi driver can be obtained under the condition that the indoor environment meets the preset quiet condition, i.e., the indoor noise intensity is lower than the first intensity threshold. The difference data between the initial audio sample and the reference audio sample is determined. The initial audio sample is used as the input data, and the difference data is used as the label to train the feature extraction network to be trained until the training termination condition is met. The training termination condition can include that the number of training samples is greater than a preset number threshold, the number of iterations meets a preset number threshold, or the feature extraction network converges, etc. The preset number threshold and the preset number threshold, etc. can be set or adjusted by the technician according to the need or experience, or repeatedly determined through a large number of experiments, and the present embodiment does not make any limitation thereto.

[0156] The first intensity threshold can be set by the technician according to the need or experience, or repeatedly determined through a large number of experiments. It is worth noting that the indoor environment corresponding to the indoor environment described herein can be the same as or have the same spatial structure as the indoor environment corresponding to the aforementioned indoor environment, and the present embodiment does not make any limitation thereto.

[0157] The sample display device can be a display device dedicated to model training, and the device type of the sample display device can be the same as or different from the device type of the aforementioned display device, and the present embodiment does not make any limitation thereto. The number of sample display devices can be at least one, and usually to ensure the performance of the trained feature extraction network, the number of sample display devices is usually multiple.

[0158] The initial audio sample can be understood as audio data collected by the microphone in the sample display device under the condition that the indoor environment meets the preset quiet condition and the WiFi module of the sample display device works normally. The audio data can carry noise generated by the WiFi module during operation and environment audio of the indoor environment. After the sample display device uninstalls the WiFi driver, the WiFi module will no longer work, that is, it will no longer periodically send detection request signals. In this case, the reference audio sample collected by the microphone of the sample display device during operation will no longer carry noise generated by the WiFi module during operation, but only include environment audio of the indoor environment.

[0159] The difference data is used to represent the difference between the initial audio sample and the reference audio sample. The difference between the two can be quantified by subtraction. Since the initial audio sample carries noise generated by the WiFi module during operation and environment audio, and the reference audio sample only carries environment audio, the difference data only carries noise generated by the WiFi module during operation.

[0160] Correspondingly, the initial audio sample can be used as input data, and the difference data can be used as a label to train the feature extraction network to be trained, so as to adjust the network parameters of the feature extraction network, so that the feature extraction network gradually has the ability to extract noise generated by the WiFi module during operation from the input data.

[0161] It is worth noting that in the case where the collection period of the initial audio sample corresponds to multiple working periods of the WiFi module, the transmission time of the probe request signal of the WiFi module can be used for rough positioning to determine the position of the interference in the audio. Then, the interference-free position with a value tending to 0 is removed according to the difference data, so as to cut the difference data into a single-period feature sequence corresponding to the working period of the WiFi module. Correspondingly, referring to the clipping result diagram of Figure 8 The initial audio sample can also be clipped according to the transmission time of the probe request signal of the WiFi module to obtain a single-period initial audio sample. Correspondingly, the single-period initial audio sample is used as input data, and the single-period difference data is used as a label to train the feature extraction network to be trained.

[0162] In order to improve the generality of the trained feature extraction network, the initial audio sample can be audio data collected when the WiFi module works in different working modes, thereby improving the sample diversity of the initial audio sample, and further helping to improve the generality and robustness of the trained feature extraction network.

[0163] It is worth noting that the training of the feature extraction network suitable for different types of display devices, such as different categories or models, can be carried out respectively, but this training method requires a large amount of human and time costs and poor universality. In an optional embodiment, a universal feature extraction network can also be used to train the network structure suitable for extracting audio interference features of different types of display devices.

[0164] In still other optional embodiments, referring to Figure 9 , a model training method is provided, applied to a server, comprising:

[0165] S910, acquiring a first audio sample generated by the sample display device during operation when the indoor noise intensity is lower than a first intensity threshold, and a second audio sample generated by the sample display device during operation after uninstalling the WiFi driver.

[0166] The first audio sample can be understood as the audio data collected by the microphone in the sample display device under the condition that the WiFi module in the sample display device works normally when the indoor noise intensity is lower than the first intensity threshold. The audio data can carry the noise generated by the WiFi module during operation, and also includes the environmental audio of the indoor environment.

[0167] The second audio sample can be understood as the audio data collected by the microphone in the sample display device under the condition that the WiFi module in the sample display device no longer works when the indoor noise intensity is lower than the first intensity threshold. The audio data no longer carries the noise generated by the WiFi module during operation, and only includes other audio such as the environmental audio of the indoor environment.

[0168] It is worth noting that in order to improve the accuracy of the acquired audio sample, multiple first audio samples under the same condition can be collected, and the average value is taken to reduce the interference of random noise, and the determined average value is used as the final first audio sample. Similarly, multiple second audio samples under the same condition can be collected, and the average value is taken to reduce the interference of random noise, and the determined average value is used as the final second audio sample.

[0169] S920, determining a first difference feature between the first audio sample and the second audio sample.

[0170] The first difference feature is used to measure the difference between the audio data of the two situations that the WiFi module in the sample display device works normally and the WiFi module no longer works under the condition that the indoor environment meets the preset quiet condition.

[0171] Exemplarily, the first difference feature between the first audio sample and the second audio sample can be determined in a way of subtraction.

[0172] S930, training the feature extraction network to be trained with the first audio sample as input data and the first difference feature as a label until a difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition.

[0173] The feature extraction network is configured to extract a target interference feature of audio generated by the WiFi module of the display device in a network search condition.

[0174] Exemplarily, the feature extraction network to be trained can be trained with the first audio sample as input data and the first difference feature as a label to adjust network parameters of the feature extraction network, so that the feature extraction network gradually has an extraction capability for noise generated by the WiFi module in operation carried in the input data.

[0175] Notably, in a case where a collection period of the first audio sample corresponds to multiple working periods of the WiFi module, the transmission time of the probe request signal of the WiFi module can be used for rough positioning to determine the position of the interference in the audio; then, the interference-free position with a value tending to 0 is removed according to the first difference feature, so as to cut the first difference feature into a single-period feature sequence corresponding to the working period of the WiFi module. Correspondingly, the first audio sample can also be cut according to the transmission time of the probe request signal of the WiFi module to obtain a single-period first audio sample. Correspondingly, the feature extraction network to be trained is trained with the single-period first audio sample as input data and the single-period first difference feature as a label.

[0176] In order to improve the generality of the trained feature extraction network, the first audio sample can be audio data collected when the WiFi module works in different working modes, so as to improve the sample diversity of the first audio sample, and further help to improve the generality and robustness of the trained feature extraction network.

[0177] In some embodiments, the standard feature extraction network is trained in the following manner: obtaining a third audio sample generated during running of a sample display device in a soundproof environment, and a fourth audio sample generated during running of the sample display device after uninstalling a WiFi driver; determining a second difference feature between the third audio sample and the fourth audio sample; training the standard feature extraction network to be trained with the third audio sample as input data and the second difference feature as a label until a difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition.

[0178] The preset difference condition can be that the similarity of the two output results is less than a preset similarity threshold. The similarity can be determined in at least one of the conventional manners, which is not limited in the embodiment. The preset similarity threshold can be set by the technician according to the need or experience, or determined repeatedly through a large number of tests, which is not limited in the embodiment.

[0179] The soundproof environment is an indoor environment in which the environmental sound intensity does not exceed the second intensity threshold. The second intensity threshold is less than the first intensity threshold, and can be set by the technician according to the need or experience, or determined repeatedly through a large number of tests.

[0180] The third audio sample can be understood as audio data collected by the microphone in the sample display device under the condition that the WiFi module in the sample display device works normally in the soundproof environment. The audio data can carry noise generated by the WiFi module in the working process, and also includes environmental audio of the indoor environment.

[0181] The fourth audio sample can be understood as audio data collected by the microphone in the sample display device under the condition that the WiFi module in the sample display device does not work in the soundproof environment. The audio data does not carry noise generated by the WiFi module in the working process, and only includes other audio such as environmental audio of the indoor environment.

[0182] It is worth noting that in order to improve the accuracy of the obtained audio sample, multiple third audio samples under the same condition can be collected, the average value is taken to reduce the interference of random noise, and the determined average value is used as the final third audio sample. Similarly, multiple fourth audio samples under the same condition can be collected, the average value is taken to reduce the interference of random noise, and the determined average value is used as the final fourth audio sample.

[0183] The second difference feature is used to measure the difference between the audio data under the condition that the WiFi module in the sample display device works normally and the condition that the WiFi module does not work in the soundproof environment.

[0184] Optionally, the second difference feature between the third audio sample and the fourth audio sample can be determined by difference.

[0185] Exemplarily, the third audio sample can be taken as the input data, the second difference feature can be taken as the label, and the standard feature extraction network to be trained can be trained to adjust the network parameters of the standard feature extraction network until a training stop condition is met, so that the standard feature extraction network gradually has the extraction capability of the noise generated by the WiFi module working carried in the input data.

[0186] The training stop condition can include that the number of training samples is greater than a preset number threshold, the number of iterations meets a preset number threshold, or the feature extraction network tends to converge. The preset number threshold and the preset number threshold can be set or adjusted by the technician according to the need or experience, or repeatedly determined through a large number of experiments, and the present embodiment does not make any limitation.

[0187] It is worth noting that the standard feature extraction network here can be understood as a network model specially used for audio interference feature extraction, which can be realized based on a traditional machine learning model or a deep learning model, and the present embodiment does not make any limitation on the network structure of the standard feature extraction network. Optionally, the network structure of the standard feature extraction network is the same as or at least partially different from that of the foregoing feature extraction network.

[0188] It can be understood that in the case that the collection period of the third audio sample corresponds to multiple working periods of the WiFi module, the transmission time of the probe request signal of the WiFi module can be used for rough positioning to determine the position of the interference in the audio; then, the interference-free position with the value tending to 0 is removed according to the second difference feature, so as to cut the second difference feature into a single-period feature sequence corresponding to the working period of the WiFi module. Correspondingly, the third audio sample can also be cut according to the transmission time of the probe request signal of the WiFi module to obtain a single-period third audio sample. Correspondingly, the single-period third audio sample is taken as the input data, the single-period second difference feature is taken as the label, and the standard feature extraction network to be trained is trained.

[0189] It should be noted that the execution device used in the training process of the standard feature extraction network can be the same as or different from the execution device used in the foregoing feature extraction network, and the present embodiment does not make any limitation.

[0190] Since the audio samples collected by the standard feature extraction network correspond to a noise-free environment, and the audio samples collected by the aforementioned feature extraction network correspond to a relatively quiet indoor environment, the feature extraction capability of the standard feature extraction network is better, but the working environment requirement of the display device is more stringent, and it is not suitable for the conventional device use scene. Therefore, directly using the trained standard feature extraction network to extract the audio interference feature will cause the extracted audio interference feature to be unmatched with the working environment of the display device, and affect the accuracy and clarity of the audio noise reduction result. In order to eliminate the interference of different indoor environments, in the above embodiment, the standard feature extraction network with better feature extraction capability is taken as the benchmark, and the feature extraction network suitable for interference feature extraction in a relatively quiet indoor environment is trained until the feature extraction capability of the trained feature extraction network is comparable to that of the standard feature extraction network, thereby helping to improve the indoor universality and feature extraction capability of the feature extraction network.

[0191] In order to improve the universality of the trained standard feature extraction network, the third audio sample can be audio data collected when the WiFi module works in different working modes, thereby improving the sample diversity of the third audio sample, and further helping to improve the universality and robustness of the trained standard feature extraction network.

[0192] In some embodiments, the aforementioned trained feature extraction network can be migrated to a server providing interference noise extraction services. The server can be the same as or different from the server performing the feature extraction network training process, and the present embodiment does not make any limitation thereon.

[0193] Optionally, if they are the same server, the server can also obtain a calling request of the display device for the trained feature extraction network; in response to the calling request, the trained feature extraction network is used to extract features of the interference audio; target interference features are obtained; the target interference features are sent to the display device, so that the display device performs noise reduction processing on the first audio data according to the feature frequency and feature intensity of the target interference features, and obtains second audio data; wherein the first audio data at least includes interference audio of the WiFi module in the current working mode.

[0194] Exemplarily, the display device 200 obtains, through the communication device in the server, a calling request for the trained feature extraction network; the calling request includes the interference audio of the WiFi module in the display device in the current working mode; in response to the calling request, the interference audio is subjected to feature extraction based on the trained feature extraction network to obtain target interference features; the server feeds back the target interference features to the display device 200 through the communication device in the server, so that the display device performs noise reduction processing on the first audio data according to the feature frequency and the feature intensity of the target interference features to obtain second audio data; wherein the first audio data at least includes the interference audio of the WiFi module in the current working mode.

[0195] In some other embodiments, the aforementioned trained feature extraction network can be directly migrated to the memory 190 of the display device 200 for local calling by the controller 110 of the display device 200, thereby improving the extraction efficiency and convenience of the target interference features.

[0196] On the basis of the technical solutions of the above-mentioned embodiments, in some optional embodiments, the audio noise reduction process is described taking the interaction process between the display device and the server as an example.

[0197] Referring to the audio noise reduction method shown in Figure 10 The audio noise reduction method comprises the following steps.

[0198] S1001, the display device obtains the device type of the display device;

[0199] S1002, the display device controls the microphone to collect interference audio in different working modes under the condition that the indoor environment meets the preset quiet condition according to the collection period corresponding to the device type;

[0200] S1003, the display device sends a calling request for the trained feature extraction network to the server;

[0201] S1004, the server subjects the interference audio in different working modes to feature extraction based on the trained feature extraction network to obtain audio interference features in different working modes;

[0202] S1005, the server sends the audio interference features in different working modes to the display device;

[0203] S1006, the display device stores the audio interference features in different working modes in the local memory in association;

[0204] S1007, in the case that the voice function is enabled, the display device controls the microphone to collect first audio data;

[0205] S1008, the display device looks up the audio interference feature in the current working mode of the WiFi module from the memory as the target interference feature;

[0206] S1009, the display device cuts the first audio data according to the detection request signal transmission time of the WiFi module during the first audio data collection, to obtain first audio segments;

[0207] S1010, the display device deducts the matched target interference feature from each first audio segment to obtain second audio segments;

[0208] S1011, the display device combines each second audio segment according to the cutting order of the corresponding first audio segment to obtain second audio data.

[0209] Based on the same inventive concept, the embodiments of the present application also provide an audio noise reduction device for implementing the above-mentioned audio noise reduction method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more audio noise reduction device embodiments provided below can refer to the limitations of the audio noise reduction method in the foregoing, which will not be repeated here.

[0210] In one exemplary embodiment, as shown in Figure 11 An audio noise reduction device is provided, comprising a first acquisition module 1110 and a noise reduction module 1120. Wherein,

[0211] The first acquisition module 1110 is configured to acquire first audio data collected by the microphone of the display device; wherein the first audio data at least includes interference audio of the WiFi module in the current working mode;

[0212] The first acquisition module 1110 is further configured to acquire the audio interference feature of the WiFi module of the display device in the current working mode for network search as the target interference feature; wherein the feature frequency or feature intensity of the audio interference feature in different working modes is different;

[0213] The noise reduction module 1120 is configured to perform noise reduction processing on the first audio data according to the feature frequency and feature intensity of the target interference feature, to obtain second audio data.

[0214] In some embodiments, the first acquisition module 1110 is further configured to control the microphone to collect interference audio of the WiFi module in the current working mode when the indoor noise intensity is lower than the first intensity threshold; the device further comprises a first extraction module configured to perform feature extraction on the interference audio based on the trained feature extraction network to obtain the target interference feature.

[0215] In some embodiments, the first extraction module comprises: a first sending unit configured to send a calling request for the trained feature extraction network to the server, so that the server performs feature extraction on the interference audio based on the trained feature extraction network; and a first obtaining unit configured to obtain a feature extraction result fed back by the server, wherein the feature extraction result comprises the target interference feature.

[0216] In some embodiments, the first extraction module comprises: a running-in-process unit configured to run the trained feature extraction network stored in the in-process memory to extract the target interference feature in the interference audio.

[0217] In some embodiments, the first obtaining module comprises: a second obtaining unit configured to obtain the device type of the display device from the memory; and a first control unit configured to control the microphone to collect the interference audio under the condition that the indoor environment meets the preset quiet condition according to the collection period corresponding to the device type.

[0218] In some embodiments, the first obtaining module 1110 comprises: a second control unit configured to control the microphone to collect the first audio data under the condition that the voice function is turned on.

[0219] Based on the same inventive concept, the embodiments of the present application also provide a model training device for implementing the above-mentioned model training method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more model training device embodiments provided below can be referred to the limitations of the model training method in the above text, which will not be repeated here.

[0220] In one exemplary embodiment, as shown in Figure 12 a model training device is provided, comprising: a second obtaining module 1210, a determination module 1220 and a training module 1230. Wherein,

[0221] The second obtaining module 1210 is configured to obtain a first audio sample generated by a sample display device during the running process under the condition that the indoor noise intensity is lower than a first intensity threshold, and a second audio sample generated by the sample display device during the running process after uninstalling the WiFi driver;

[0222] The determination module 1220 is configured to determine a first difference feature between the first audio sample and the second audio sample.

[0223] The training module 1230 is configured to take the first audio sample as input data, take the first difference feature as a label, train the feature extraction network to be trained, and stop until the difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition.

[0224] The feature extraction network is configured to extract a target interference feature of audio generated by a WiFi module in the display device under a network search condition.

[0225] The standard feature extraction network is trained in the following manner.

[0226] A third audio sample generated by the sample display device in a mute environment is obtained, and a fourth audio sample generated by the sample display device after uninstalling the WiFi driver is obtained.

[0227] A second difference feature between the third audio sample and the fourth audio sample is determined.

[0228] The third audio sample is taken as input data, and the second difference feature is taken as a label, and the standard feature extraction network to be trained is trained.

[0229] In some embodiments, the apparatus further includes a third acquisition module configured to acquire a calling request for the trained feature extraction network sent by the display device, wherein the calling request includes interference audio of the WiFi module in the display device under a current working mode; a request response module configured to, in response to the calling request, perform feature extraction on the interference audio based on the trained feature extraction network to obtain a target interference feature; and a feedback module configured to send the target interference feature to the display device, so that the display device performs noise reduction processing on first audio data according to a feature frequency and a feature intensity of the target interference feature to obtain second audio data, wherein the first audio data at least includes the interference audio of the WiFi module under the current working mode.

[0230] The various modules in the above audio noise reduction apparatus / model training apparatus can be all or partially implemented by software, hardware, and combinations thereof. The various modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the various modules.

[0231] In an exemplary embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 13As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium during the running process. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the external terminal in a wired or wireless manner. The wireless manner can be realized through WIFI, mobile cellular network, near field communication (NFC) or other technologies. The computer program is executed by the processor to realize an audio noise reduction method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0232] Those skilled in the art can understand that, Figure 13 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0233] In an optional embodiment, Figure 13 The computer device shown in the figure can be the aforementioned projection device, for example, can be a laser television.

[0234] In an exemplary embodiment, a computer device is provided, including a memory and a processor, the memory stores a computer program, and the processor executes the computer program to realize the steps in the above method embodiments. In an embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by the processor to realize the steps in the above method embodiments.

[0235] In an embodiment, a computer program product is provided, including a computer program, and the computer program is executed by the processor to realize the steps in the above method embodiments.

[0236] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0237] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. In the embodiments provided in the present application, any reference to memory, database or other medium can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (Resistive Random Access Memory, ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. The volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (Artificial Intelligence, AI) processor, etc., without being limited thereto.

[0238] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, however, any combination of these technical features is deemed to be within the scope of the present application.

[0239] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A display device, characterized by comprising: The display device comprises: a wireless fidelity (WiFi) module configured to perform network search in different working modes; a microphone disposed on a same circuit board as the WiFi module and configured to collect audio data; at least one controller connected to the microphone and the WiFi module and configured to: control the microphone to collect first audio data, wherein the first audio data at least comprises interference audio of the WiFi module in a current working mode; take an audio interference feature of the WiFi module in the current working mode as a target interference feature, wherein audio interference features in different working modes are different in feature frequency or feature intensity; perform noise reduction processing on the first audio data according to the feature frequency and the feature intensity of the target interference feature, to obtain second audio data.

2. The display device of claim 1, wherein, The controller is further configured to: in a case where indoor noise intensity is lower than a first intensity threshold, control the microphone to collect interference audio of the WiFi module in the current working mode; perform feature extraction on the interference audio based on a trained feature extraction network, to obtain the target interference feature.

3. The display device of claim 2, wherein, The display device further comprises: a communication device in communication connection with a server; the controller, when performing the feature extraction on the interference audio based on the trained feature extraction network to obtain the target interference feature, is configured to: send a calling request for the trained feature extraction network to the server, so that the server performs feature extraction on the interference audio based on the trained feature extraction network; and obtain a feature extraction result fed back by the server, wherein the target interference feature is included in the feature extraction result.

4. The display device of claim 2, wherein, The display device further comprises: a memory configured to store the trained feature extraction network; the controller, when performing the feature extraction on the interference audio based on the trained feature extraction network to obtain the target interference feature, is configured to: run the trained feature extraction network stored in the memory, to extract the target interference feature from the interference audio.

5. The display device of claim 2, wherein, The display device further comprises: a memory configured to store a device type of the display device; the controller, when controlling the microphone to collect interference audio in a case where indoor noise intensity is lower than a first intensity threshold, is configured to: obtain the device type of the display device from the memory; and control the microphone to collect interference audio in the case where indoor noise intensity is lower than the first intensity threshold, according to a collection period corresponding to the device type.

6. The display device according to any of claims 1-5, characterized in that, The controller, when controlling the microphone to collect first audio data, is configured to: in a case where a voice function is turned on, control the microphone to collect the first audio data.

7. The display device according to any of claims 2-5, characterized in that, The feature extraction network is trained in the following manner: obtain first audio samples generated in a running process of a sample display device in a case where indoor noise intensity is lower than a first intensity threshold, and second audio samples generated in a running process of the sample display device after uninstalling a WiFi driver; determine first difference features between the first audio samples and the second audio samples; and train the feature extraction network based on the first difference features. The first audio sample is used as input data, and the first difference feature is used as a label to train the feature extraction network to be trained until the difference between the trained feature extraction network and the standard feature extraction network meets a preset difference condition. The feature extraction network is used to extract an audio interference feature of audio generated by a WiFi module in a display device under network search.

8. The display device of claim 7, wherein, The standard feature extraction network is trained in the following manner: A third audio sample generated by the sample display device during operation in a soundproof environment and a fourth audio sample generated by the sample display device during operation after uninstalling the WiFi driver are obtained; A second difference feature between the third audio sample and the fourth audio sample is determined; The third audio sample is used as input data, and the second difference feature is used as a label to train the standard feature extraction network to be trained.

9. An audio noise reduction method, comprising: The application is applied to a display device, comprising: obtaining first audio data collected by a microphone of a display device; wherein the first audio data at least includes interference audio of a WiFi module in the display device under a current working mode; and obtaining an audio interference feature of the WiFi module of the display device under a current working mode for network search as a target interference feature; wherein the feature frequency or the feature intensity of the audio interference feature under different working modes is different; performing noise reduction processing on the first audio data according to the feature frequency and the feature intensity of the target interference feature to obtain second audio data; wherein the microphone and the WiFi module are arranged on the same circuit board.

10. The method of claim 9, wherein, The method further comprises: controlling the microphone to collect interference audio of the WiFi module under the current working mode when the indoor noise intensity is lower than a first intensity threshold; extracting a feature of the interference audio based on the trained feature extraction network to obtain a target interference feature.

Citation Information

Patent Citations

  • Sound-controlled interaction device, sound-controlled interaction method and television

    CN107396158A

  • Audio processing method and device, electronic equipment and computer readable storage medium

    CN115442710A