Control box for matching display with AI digital human

By integrating visual and acoustic positioning technologies into the control box, the problem of lack of unified planning of information between AI digital human devices is solved, realizing efficient and user-friendly multimodal interaction, improving the interactive experience between AI digital humans and users, and enhancing the energy efficiency of the devices.

CN223977692UActive Publication Date: 2026-03-06TPV ELECTRONICS (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

The current market lacks dedicated control boxes for AI digital humans, resulting in a lack of unified planning for information between data collection and interaction devices, incomplete voice interaction functions, limited interaction methods, and an inability to meet the needs of special groups. Furthermore, AI digital humans lack the ability to perceive emotions and cannot achieve friendly interaction.

Method used

Design a control box for a display that integrates visual and acoustic positioning technologies, including a camera, microphone, sound unit, and main control board. The main control board enables data fusion and device collaboration, supports directional sound pickup and playback, and combines proximity sensor and turntable technology to achieve precise positioning and interactive control.

Benefits of technology

It improves the efficiency and experience of AI digital human interaction with users, ensures consistency between voice and movement, enhances immersion, supports multimodal interaction, meets the needs of special groups, and achieves energy saving and cost reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN223977692U_ABST
    Figure CN223977692U_ABST
Patent Text Reader

Abstract

The utility model discloses a control box for matching a display with an AI digital human, which comprises a digital human accessory and at least one control box, and the digital human accessory is electrically connected with a matched main control board; the digital human accessory comprises a sound production unit, a proximity sensor, a camera and a microphone; the proximity sensor and the camera are arranged in the control box, at least one of the sound production unit and the microphone is arranged in the control box, and the control box is electrically connected with the display through a matched interface arranged on the display; the proximity sensor is used for proximity detection of an object; the camera is used for detecting an image picture of a target object; the microphone is used for sound pickup; the sounding unit is used for audio playing; and the main control board respectively controls the work of the sound production unit, the proximity sensor, the camera and the microphone. The utility model provides a digital human accessory for collecting visual and acoustic information data, and realizes more efficient and friendly human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This utility model relates to the field of display device technology, and in particular to a control box for matching AI digital humans with displays. Background Technology

[0002] Digital humans refer to 3D digital human images created through 3D graphic character modeling and using information science methods to virtually simulate the human body at different levels of form and function. With the maturity of technologies such as 3D, artificial intelligence, and virtual-real interaction, digital humans are beginning to transform from virtual idols into interactive service providers. Based on the two dimensions of anthropomorphism and automation, digital humans are divided into five levels, L1-L5. Among them, L4 and L5 level digital humans are collectively referred to as AI digital humans. They possess a high degree of anthropomorphism, approaching the level of real people in appearance, movement, and intelligence. AI digital humans with multimodal interaction capabilities can not only present multimedia information that traditional voice dialogue cannot convey, but also, by combining visual AI technology, complete multiple interactive tasks such as identity recognition, gesture recognition, and emotion recognition, making the interaction process richer and more efficient. AI digital humans are currently in their early stages and have a long way to go before market maturity, requiring significant technological accumulation. The current industry-wide technological breakthrough goal is "how to make digital humans more like people, and think like humans."

[0003] As the presentation terminal for digital humans, displays and related products not only need to vividly present the posture and micro-expressions of digital virtual humans, but also need to serve as interactive entry points for data collection and audio-visual signal transmission. Currently, the market lacks dedicated control boxes specifically for digital humans, which is reflected in the following aspects: 1. In the data collection and interaction section, information between different system devices lacks unified planning and management: acoustic and visual information are independent and not fused (including the mutual access and use of ranging data). 2. Voice interaction devices are not fully functional, such as lacking directional sound pickup and playback capabilities, which can easily cause disturbance to neighbors or affect other unrelated users in public places. 3. There are few interaction methods, failing to cover all users, especially the usage habits of special groups such as the deaf, mute, and elderly. 4. Due to incomplete data collection by terminal devices, AI digital humans lack the ability to perceive user emotions and fail to control their posture well, failing to provide users with the emotional rendering of friendly interaction such as the AI ​​digital human's attention and micro-expressions. Utility Model Content

[0004] The purpose of this invention is to provide a control box for matching AI digital humans with displays. Based on visual and acoustic positioning and AI technology, it integrates visual and acoustic information data to achieve more efficient and user-friendly human-computer interaction.

[0005] The technical solution adopted in this utility model is:

[0006] A control box for display-matched AI digital human includes a digital human accessory and at least one control box. The digital human accessory is electrically connected to a matching main control board. The digital human accessory includes a voice-emitting unit, a proximity sensor, a camera, and a microphone. The proximity sensor and camera are located within the control box, and at least one of the voice-emitting unit and microphone is located within the control box. The control box is electrically connected to the display via a matching interface on the display. The proximity sensor is used for object proximity detection. The camera is used to detect the image of the target object. The microphone is used for sound pickup. The voice-emitting unit is used for audio playback. The main control board controls the operation of the voice-emitting unit, proximity sensor, camera, and microphone respectively.

[0007] Furthermore, the main control board is equipped with a vision processing unit, an auditory processing unit, an interface unit, a power management unit, and a clock unit. The vision processing unit provides functional interfaces for image acquisition, preprocessing, and feature extraction, while the auditory processing unit provides functional interfaces for sound source localization, echo cancellation, and noise suppression. The interface unit includes a USB interface, an I2S interface, an I2C interface, an SPI interface, and a storage controller. The power management unit is used to power the device, and the clock unit provides the clock frequency required for the main control board to operate.

[0008] Furthermore, the vision processing unit and the hearing processing unit are integrated on a main control chip. The main control chip interacts with the interface unit, power management unit and clock unit through a bus to realize data access, storage and sharing; the bus includes USB, I2S and SPI buses.

[0009] Furthermore, the main control board is the main control board of the display, and the digital human accessory of the control box is detachably electrically connected to the main control board inside the display through a matching interface on the display; or, the main control board is the main control board set inside the control box; or the main control board is an independent main control board set outside the control box.

[0010] Furthermore, a single-lens camera can be used, which is less expensive, but the ranging accuracy will be affected; or a dual-lens camera can be used, which is a combination of an RGB camera and an IR camera, such as one RGB camera paired with one IR camera. This provides a more balanced overall performance in terms of environmental adaptability (such as lighting), ranging accuracy, power consumption, resolution, and frame rate; or a multi-lens camera can be used.

[0011] Furthermore, one of the sound-emitting unit and the microphone is located inside the control box, while the other of the sound-emitting unit and the microphone is independently installed outside the control box and electrically connected to the main control board; or the sound-emitting unit is a directional sound-emitting unit, and the sound-emitting unit and the microphone are both located inside the control box; or the sound-emitting unit is a non-directional sound-emitting unit, and the sound-emitting unit and the microphone are both located inside the control box, and the chambers where the sound-emitting unit and the microphone are located are physically isolated from each other.

[0012] Specifically, because directional sound generation and directional microphone may resonate and echo from the speaker, they cannot be housed in the same enclosure. The directional sound generation needs to be separated and placed in another location, including integration into the display. To mitigate the impact of resonance and echo from the sound generation unit on the microphone, the sound generation unit and microphone can be placed in separate enclosures. This can be achieved by adding vibration-damping brackets, using sound insulation and absorption materials, sealing the enclosure, or creating independent chambers. These enclosures can be independent, connected as a whole, or a single enclosure can be divided into different units to house different components. Alternatively, the sound generation unit can be integrated into the display and controlled in conjunction with the main control box.

[0013] Furthermore, the sound-emitting unit is a directional sound-emitting unit with two or more sound-emitting areas. The directional sound-emitting unit has its own mechanical steering device, and the main control board controls the movement of the mechanical steering device of the directional sound-emitting unit.

[0014] Alternatively, the sound-emitting unit may consist of two or more directional sound-emitting units. Multiple directional sound-emitting units with fixed angles or zone-direction functions can be combined to achieve area coverage, i.e., multiple directional sound-emitting units can be installed crosswise to achieve area coverage.

[0015] Alternatively, the sound-emitting unit may be a combination of a directional sound-emitting unit and an omnidirectional sound-emitting unit, wherein the directional sound-emitting unit adopts a fixed-angle sound-emitting unit or a zoned control sound-emitting unit with zoned steering function; the main control board controls the operation of the directional sound-emitting unit and the omnidirectional sound-emitting unit.

[0016] Specifically, as a feasible implementation of this utility model, a single directional sound unit can be combined with a camera, sensor, microphone, or other ranging device. The main control board performs target capture, direction finding, and positioning through the camera, sensor, microphone, or other ranging device combination. Then, by calling the software interface of the directional sound unit, the directional sound channel is adjusted to achieve directional following. The directional sound unit has its own mechanical device to achieve zone-based turning. The zone size depends on the hardware and requirements. Specifically, as an optional implementation, a 30-degree angle zone is used.

[0017] Alternatively, there may be two or more directional sound units. Multiple directional sound units with fixed angles or zoned steering functions can be combined for zoned control, that is, multiple directional sound units can be installed in a cross pattern to achieve area coverage. The main control board uses a combination of cameras, sensors, microphones or other ranging devices to locate the target by direction finding and ranging. The main control board controls the directional sound units in the corresponding area to emit sound or not to emit sound.

[0018] Furthermore, the microphone uses a directional pickup unit, and the directional pickup unit has two or more pickup areas. The directional pickup unit has a built-in mechanical steering device, and the main control board controls the action of the mechanical steering device of the directional pickup unit.

[0019] Alternatively, the microphone may consist of two or more directional pickup units. Multiple directional pickup units with fixed angles or zone-direction functions can be combined to achieve area coverage. In other words, multiple directional pickup units can be installed in a cross-sectional manner to achieve area coverage. The main control board controls the directional pickup units in the corresponding areas to work or not work.

[0020] Alternatively, the microphone may be a combination of a directional pickup unit and an omnidirectional pickup unit, wherein the directional pickup unit adopts a fixed-angle pickup unit or a zone control pickup unit with zone steering function; the main control board controls the directional pickup or omnidirectional pickup to work or not as needed.

[0021] Specifically, when the microphone uses a directional pickup unit, the main control board captures and locates the target by combining a camera, sensor, microphone, or other ranging device. Then, it calls the software interface of the directional pickup unit to control the mechanical steering device to adjust the direction and achieve directional following.

[0022] When the microphone uses two or more directional pickup units, the main control board uses a combination of camera, sensor, microphone, or other ranging devices to locate the target by direction finding and ranging. The main control board then controls the directional pickup units in the corresponding areas to operate or not operate. For example, power supply can be controlled via GPIO pins, or the system interface can be used to control whether the pickup units are working.

[0023] Specifically, the microphone employs a combination of directional and omnidirectional pickup units. The directional pickup unit can be a fixed-angle pickup unit or a zone-controlled pickup unit with zone-direction functionality. The main control board controls the operation of either directional or omnidirectional pickup as needed. To avoid noise sampling distortion, the system can simultaneously activate both directional and omnidirectional pickup units. The omnidirectional microphone is used to pick up ambient noise and ocean currents (requiring the directional and omnidirectional microphones to be positioned close together), facilitating the subsequent signal processing unit to obtain accurate noise signals.

[0024] Specifically, the microphone uses an omnidirectional pickup unit to achieve directional tracking through a zone enhancement and weakening strategy. That is, by recording voiceprints or training models such as wake word recognition, and combining them with cameras, sensors, microphones or other ranging devices, the target is captured, located, and then noise is reduced, non-target zones are weakened (such as eliminating shielding interference sounds), and target zones are enhanced.

[0025] Furthermore, the control box is fixedly installed on the target position on the monitor using locking accessories and is electrically connected to the monitor's main control board, thereby enabling interactive communication with the monitor's main control board.

[0026] Furthermore, it also includes a turntable, which comprises a platform, a main control unit, and a power unit. The control box is detachably mounted on the platform, and the power unit provides rotational energy to the turntable. The main control unit drives the turntable to rotate. The turntable has position servo control, speed servo control, and limit functions. The main control board controls the turntable's servo control system. Taking the control process of a motor servo drive, a single-axis turntable, and a control box assembly as an example, the main control board uses the camera or other sensors in the control box to achieve direction finding, ranging, and positioning, and then uses the turntable's servo control system to achieve precise steering.

[0027] This invention employs the above technical solutions, utilizing a proximity sensor to detect the proximity of objects; a camera to detect the image of the target object, enabling visual ranging based on actual needs and existing algorithms; a microphone for sound pickup; and a sound-emitting unit for audio playback, which can also achieve acoustic ranging based on actual needs and existing algorithms. Furthermore, it can be combined with other sensors, such as radar, IR, ultrasonic, and laser ranging. The display, through a digital human accessory, obtains the depth distance between the screen and the user. Using existing ranging and positioning technologies and virtual reality technologies, a virtual sound field can be generated, allowing the user to experience sounds from different directions, enhancing immersion and surround sound. The display, through the digital human accessory, detects and obtains the direction and intensity of sound based on the user's position and environment, controlling the sound-emitting and pickup units to achieve directional sound playback and pickup, reducing noise interference and impact on the external environment.

[0028] This invention enables displays to achieve precise synchronization between sound and the virtual human's image, ensuring consistency between sound and AI digital human movements, controlling the AI ​​digital human's posture, micro-expressions, and sensing user emotions, thus enhancing the interactive experience. Through this device, displays can better serve customers by integrating existing AI applications such as identity recognition, gesture recognition, and emotion recognition; combined with proximity sensors, sleep and wake-up functions can be more effectively implemented, achieving energy savings and cost reduction. Attached Figure Description

[0029] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments;

[0030] Figure 1 This is a schematic diagram of a display structure for matching AI digital humans in this utility model;

[0031] Figure 2 This is a schematic diagram of one embodiment of the present invention using a monocular camera;

[0032] Figure 3 This is a schematic diagram of another embodiment of the present invention using a binocular camera;

[0033] Figure 4 A schematic diagram of an embodiment of the present invention when a single directional sound-emitting unit is paired with a camera;

[0034] Figure 5 This is a schematic diagram of an embodiment of the present invention when multiple directional sound-emitting units are combined with a camera;

[0035] Figure 6 This is a schematic diagram of the turntable system structure of this utility model;

[0036] Figure 7 A schematic diagram of the core architecture of the main control board;

[0037] Figure 8 This is a schematic diagram of the system principle of a control box for matching an AI digital human with a display, according to this utility model. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0039] like Figures 1 to 8 As shown in the figure, this utility model discloses a control box for matching an AI digital human with a display, which includes a control box 10 and a digital human accessory disposed within the control box 10; the control box 10 is detachably electrically connected to the main control board inside the display 1 via a matching interface 2 on the display 1; the digital human accessory includes a directional sound unit 5, a proximity sensor 6, a camera 7, and a microphone 8; the proximity sensor 6 is used for object proximity detection so as to wake the device from sleep when a human body approaches; the camera 7 is used for target object detection and distance estimation; the microphone 8 is used for ultrasonic ranging and sound pickup; the directional sound unit 5 is used for audio playback; the main control board controls the device to sleep or wake up based on the signal from the proximity sensor 6; the main control board is used to fuse the data collected by the camera 7 and the microphone 8 to obtain accurate user location information, and realize gesture recognition, posture control, and directional sound pickup based on the user location information, while controlling the directional sound unit 5 to play sound in a specific direction.

[0040] The system includes a digital human accessory and at least one control box 10. The digital human accessory is electrically connected to a matching main control board. The digital human accessory includes a sound-emitting unit 5, a proximity sensor 6, a camera 7, and a microphone 8. The proximity sensor 6 and the camera 7 are located inside the control box. At least one of the sound-emitting unit 5 and the microphone 8 is located inside the control box 10. The control box 10 is electrically connected to the display 1 via a matching interface on the display 1. The proximity sensor 6 is used for object proximity detection. The camera 7 is used to detect the image of the target object. The microphone 8 is used for sound pickup. The sound-emitting unit 5 is used for audio playback. The main control board controls the operation of the sound-emitting unit, the proximity sensor, the camera, and the microphone.

[0041] Furthermore, the main control board is equipped with a vision processing unit, an auditory processing unit, an interface unit, a power management unit, and a clock unit. The vision processing unit provides functional interfaces for image acquisition, preprocessing, and feature extraction, while the auditory processing unit provides functional interfaces for sound source localization, echo cancellation, and noise suppression. The interface unit includes a USB interface, an I2S interface, an I2C interface, an SPI interface, and a storage controller. The power management unit is used for power supply management of the device, and the clock unit provides the clock frequency required for the main control board to operate.

[0042] Furthermore, the vision processing unit and the hearing processing unit are integrated on a main control chip. The main control chip interacts with the interface unit, power management unit and clock unit through a bus to realize data access, storage and sharing; the bus includes USB, I2S and SPI buses.

[0043] Furthermore, the main control board is the main control board of the display, and the digital human accessory of the control box is detachably electrically connected to the main control board inside the display through a matching interface on the display; or, the main control board is the main control board set inside the control box; or the main control board is an independent main control board set outside the control box.

[0044] Furthermore, camera 7 is a monocular camera, which is less expensive, but its ranging accuracy will be affected. Further, camera 7 is a binocular camera, consisting of two cameras 7. More specifically, binocular camera 7 is a combination of an RGB camera and an IR camera, such as one RGB camera paired with one IR camera, offering a more balanced overall performance in terms of environmental adaptability (e.g., lighting), ranging accuracy, power consumption, resolution, and frame rate. Finally, camera 7 is a multi-camera system.

[0045] Furthermore, one of the sound-emitting unit and the microphone is located inside the control box, while the other of the sound-emitting unit and the microphone is independently installed outside the control box and electrically connected to the main control board; or the sound-emitting unit is a directional sound-emitting unit, and the sound-emitting unit and the microphone are both located inside the control box; or the sound-emitting unit is a non-directional sound-emitting unit, and the sound-emitting unit and the microphone are both located inside the control box, and the chambers where the sound-emitting unit and the microphone are located are physically isolated from each other.

[0046] Specifically, considering the potential resonance and echo from the sound-generating unit, which could affect the microphone, the sound-generating unit and the microphone can be placed in separate enclosures. The impact can be reduced by adding vibration-damping brackets, using sound insulation and sound-absorbing materials, sealing the enclosure, or creating independent chambers. The enclosures can be independent, connected as a whole, or a single enclosure can be divided into different units to hold different accessories. Alternatively, the sound-generating unit can be integrated into the display and controlled in conjunction with the main control box.

[0047] Furthermore, there is one directional sound unit 5, and the sound area corresponding to the directional sound unit 5 is divided into two or more zones; the directional sound unit 5 is equipped with a mechanical steering device; the main control board controls the mechanical steering device to adjust the directional sound channel to achieve directional following.

[0048] Specifically, as one implementation method, a single directional sound unit 5 can be used in conjunction with a camera 7 or a proximity sensor 6. The main control board captures direction and positions the target object through the camera or other sensors, and then adjusts the directional sound channel by calling the software interface of the directional sound unit 5 to achieve directional following. The directional sound unit 5 has its own mechanical device to achieve zone-based turning. The zone size depends on the hardware and requirements. Specifically, as an optional implementation method, a 30-degree angle zone is used.

[0049] Furthermore, there are two or more directional sound units. The directional sound units with fixed angles or zoned steering functions are combined to achieve zoned control, that is, multiple directional sound units are installed crosswise to achieve area coverage.

[0050] The main control board uses a combination of cameras, sensors, microphones, or other ranging devices to locate and measure the direction of the target. The main control board then controls the directional sound units in the corresponding area to emit sound, remain silent, or change the volume.

[0051] Furthermore, the combination of directional sound units and omnidirectional sound units, wherein the directional sound units adopt fixed-angle sound units or zoned control sound units with zoned steering function; the main control board controls directional sound or omnidirectional sound as needed.

[0052] Furthermore, the microphone employs a directional pickup unit with two or more pickup areas; the directional pickup unit has a built-in mechanical steering device; the main control board uses a combination of camera, sensor, microphone, or other ranging devices to capture, locate, and position the target, and then calls the software interface of the directional pickup unit to control the mechanical steering device to adjust the direction and achieve directional following.

[0053] Furthermore, the microphone employs two or more directional pickup units, using a combination of multiple fixed-angle or zoned-directional pickup units for zoned control; that is, multiple directional pickup units are cross-installed to achieve area coverage. The main control board uses a combination of cameras, sensors, microphones, or other ranging devices to locate the target by direction finding and ranging. The main control board then controls the directional pickup units in the corresponding areas to operate or not operate. For example, power supply can be controlled via GPIO pins, or the system interface can control whether the pickup units are operational.

[0054] Furthermore, the microphone employs a combination of directional and omnidirectional pickup units. The directional pickup unit can be a fixed-angle pickup unit or a zone-controlled pickup unit with zone-direction functionality. The main control board controls the operation of directional or omnidirectional pickup as needed. To avoid noise sampling distortion, the system can simultaneously activate both directional and omnidirectional pickup units. The omnidirectional microphone is used to pick up ambient noise and ocean currents (requiring the directional and omnidirectional microphones to be positioned close to each other), facilitating the subsequent signal processing unit to obtain accurate noise signals.

[0055] Furthermore, the microphone employs an omnidirectional pickup unit to achieve directional tracking through a zone enhancement and weakening strategy. This involves recording voiceprints or training models such as wake word recognition, combining them with cameras, sensors, microphones, or other ranging devices to capture, locate, and position the target. Then, noise reduction and weakening of non-target zones, such as eliminating interference noise, and strengthening of target zones are employed.

[0056] Furthermore, it also includes a turntable 20, which includes a platform 11, a main control device 12, and an energy device. The control box is detachably mounted on the platform 11, and the energy device provides rotational energy to the turntable 20. The main control device 12 drives the turntable 20 to rotate. The turntable 20 system has two working modes: local control and remote control. The turntable 20 has position servo control, speed servo control, and limit functions. Taking the control process of a motor servo drive, a single-axis turntable, and a control box assembly as an example, the main control board achieves direction finding, distance finding, and positioning through the camera or other sensors in the control box, and then achieves precise steering through the servo control system of the turntable 20.

[0057] Furthermore, the control box 10 is fixedly installed on the target position on the display 1 by locking accessories and electrically connected to the main control board of the display, thereby connecting with the main control board of the display 1 to achieve interactive communication.

[0058] The specific principles of this utility model will be explained in detail below:

[0059] like Figure 1As shown, this utility model is a device for using a display 1 in conjunction with an AI digital human. Taking the corresponding control box design as an example, through an integrated design, the camera 7, array microphone 8, directional sound unit 5, sensor components, etc., are integrated as a whole into the control box 10, and connected via an internal USB interface 3 or an external USB interface 4 and I2S interface, etc. The advantage of being a whole component is that it is easy to connect and quick to assemble, but the disadvantage is that it requires higher design requirements for the mechanism. It includes the camera 7, array microphone 8 (MIC), directional sound unit 5, and may include various sensors (such as proximity sensor 6), etc. It can be connected to the display 1 via USB, I2S, etc., and fixed to an appropriate position on the display 1, such as the top, middle, bottom, or side of the body, depending on the application and body size. For larger models or situations where a single directional component cannot cover, directional following can be achieved by combining turntable technology or expanding combined zone control.

[0060] like Figure 2 and Figure 3 The diagram shows the structural design of the integrated digital human accessory control box 10. Taking the control box 10 with a monocular camera 7 and a binocular camera 7 as examples, the monocular camera 7 control box 10 contains only one camera 7, which is lower in cost, but the ranging accuracy will be affected. The binocular camera 7 control box 10 contains two cameras 7, which can be a combination of an RGB camera 7 and an IR camera 7, such as one RGB camera 7 paired with one IR camera 7. It has a more balanced overall performance in terms of environmental adaptability (such as lighting), ranging accuracy, power consumption, resolution, and frame rate. In addition to using a microphone 8 with directional pickup function, the microphone 8 can also use a combination of a directional microphone 8 and an omnidirectional microphone 8. The omnidirectional microphone 8 is used to pick up ambient noise and ocean tide sounds (requiring that the directional microphone 8 and the omnidirectional microphone 8 are located close to each other), which helps the signal processing unit to obtain the correct noise signal and avoids noise sampling distortion. To mitigate the potential impact of resonance and echo from the sound-generating unit on the microphone, the sound-generating unit and microphone can be placed in separate enclosures. This impact can be reduced by adding vibration-damping brackets, using sound-insulating and sound-absorbing materials, sealing the enclosures, or creating independent chambers. These enclosures can be independent, connected as a whole, or a single enclosure can be divided into different units to house different components. Alternatively, the sound-generating unit can be integrated into the display and controlled in conjunction with the main control box.

[0061] For larger models or situations where a single orientation component cannot cover the area, turntable technology or extended combination zone control can be used to achieve orientation following (i.e., simultaneously achieving precise positioning and area coverage).

[0062] The extended combination zone control method mainly achieves linkage control through interaction between the upper main control board and the directional sound unit 5, and can include the following schemes:

[0063] like Figure 4 As shown, in one preferred embodiment, a single component works with a camera 7 or a sensor. The main control board uses the camera 7 or other sensors to capture and locate the target, and then adjusts the directional sound channel by calling the software interface of the directional sound unit 5 to achieve directional tracking. The structural diagram is shown below. Figure 4 As shown, the directional sound unit 5 has a built-in mechanical device to achieve zone-based turning. The size of the zone depends on the hardware and requirements, and a 30-degree angle for each zone is recommended.

[0064] like Figure 5 As shown, in Scheme 2, a preferred implementation, multiple fixed-angle sound-emitting units are combined for zoned control. Each sound-emitting unit has a fixed directional angle, and area coverage is achieved through the cross-installation of multiple components. Combined with camera 7 or sensor-based direction finding and ranging, the main control board controls the corresponding area components to emit or not emit sound. A schematic diagram is shown below. Figure 5 As shown.

[0065] Furthermore, as a preferred implementation, Scheme Three combines Schemes One and Two, using multiple sound-emitting units with zoned steering functions for zoned control. This means that area coverage is achieved through the cross-installation of multiple sound-emitting components with zoned steering functions. Combined with camera 7 or sensor-based direction finding and ranging for positioning, the main control board controls the corresponding area components to emit or not emit sound. Compared to Scheme One, Scheme Three can achieve a more multi-dimensional and three-dimensional solution or reduce the performance and specification requirements of the zoned steering mechanism; compared to Scheme Two, Scheme Three can achieve more comprehensive area coverage, reduce blind spots, and may reduce the number of sound-emitting units, making installation and configuration more flexible.

[0066] Furthermore, as a preferred implementation method, Scheme Four employs a combination of directional sound-emitting units and omnidirectional sound-emitting units (such as loudspeakers). The directional sound-emitting unit can be a fixed-angle sound-emitting unit or a zone-controlled sound-emitting unit with zone-direction functionality. The system controls directional or omnidirectional sound emission as needed, for example, by controlling the power supply to the loudspeaker via GPIO pins or by controlling whether the sound-emitting unit emits sound or changes its volume via the system interface.

[0067] Furthermore, the sound-generating components of the microphone directional tracking function scheme are similar. Further, in one preferred embodiment, the microphone 8 employs a directional pickup unit with two or more pickup areas; the directional pickup unit has a built-in mechanical steering device; the main control board uses a combination of a camera, sensor, microphone, or other ranging device to capture, locate, and position the target, and then calls the software interface of the directional pickup unit to control the mechanical steering device to adjust the direction and achieve directional tracking.

[0068] Furthermore, in a preferred second implementation, microphone 8 employs two or more directional pickup units. Multiple directional pickup units with fixed angles or zoned steering functions are combined for zoned control; that is, multiple directional pickup units are cross-installed to achieve area coverage. The main control board uses a combination of a camera, sensor, microphone, or other ranging device to locate the target by direction finding and ranging. The main control board then controls the directional pickup units in the corresponding areas to operate or not operate. For example, power supply can be controlled via GPIO pins, or the system interface can control whether the pickup units operate.

[0069] Furthermore, in a preferred embodiment, Scheme 3, microphone 8 employs a combination of a directional pickup unit and an omnidirectional pickup unit. The directional pickup unit can be a fixed-angle pickup unit or a zone-controlled pickup unit with zone-direction functionality. The main control board controls the operation of either directional or omnidirectional pickup as needed. To avoid noise sampling distortion, the system can simultaneously activate both directional and omnidirectional pickup units. The omnidirectional microphone is used to pick up ambient noise and ocean tide sounds (requiring the directional and omnidirectional microphones to be positioned close together), facilitating the subsequent signal processing unit to obtain accurate noise signals.

[0070] Furthermore, the microphone employs an omnidirectional pickup unit to achieve directional tracking through a zone enhancement and weakening strategy. This involves recording voiceprints or training models such as wake word recognition, combining them with cameras, sensors, microphones, or other ranging devices to capture, locate, and position the target. Then, noise reduction and weakening of non-target zones, such as eliminating interference noise, and strengthening of target zones are employed.

[0071] like Figure 6 As shown, this utility model further provides a control box 10 structure utilizing turntable technology, mainly including components of the control box 10 and a turntable section. A schematic diagram of the turntable section is shown below. Figure 6As shown, the system mainly includes a platform, a main control unit 12, and an energy unit. The turntable system can operate in both local and remote control modes, and features functions such as position servo control, speed servo control, and limit switches. Taking the control flow of a single-axis turntable driven by a motor servo and equipped with a control box 10 as an example, the main control board uses the camera 7 or other sensors in the control box 10 to achieve direction finding, ranging, and positioning, and then uses the turntable's servo control system to achieve precise steering.

[0072] like Figure 7 The diagram shown is a schematic of the core architecture of the main control board, including descriptions of the main visual and acoustic functional modules on the terminal side, as well as descriptions of the main interfaces. The visual processing unit mainly includes functional interfaces related to image acquisition, preprocessing, and feature extraction, while the acoustic processing unit mainly includes functional interfaces for sound source localization, echo cancellation, and noise suppression. Both interact with other unit modules on the main control board via buses such as USB, I2S, and SPI to achieve data access, storage, and sharing.

[0073] like Figure 8 As shown in the system model diagram of this utility model, it includes the main technical principles of each component and the process and application of data fusion. Taking the fusion of ranging data from camera 7 and microphone 8 as an example, the system design includes planning data access interfaces, including a unified coordinate system, to enable mutual access and use of data. To further improve the accuracy of ranging, traditional sensing devices such as laser ranging, radar ranging, and infrared ranging can be added, and corresponding data access interfaces can be planned. The validity and timeliness of data are ensured by improving data acquisition timing, synchronization mechanisms, and fusion algorithms. Based on data fusion, the system can obtain more accurate user location information, thereby realizing functions such as gesture recognition, posture control, directional sound playback, and directional sound pickup. The start and stop of camera 7, microphone 8, and directional sound unit 5 can be triggered by proximity sensor 6, system configuration, and historical data, achieving energy saving and cost reduction.

[0074] This invention achieves effective and rapid information fusion through a pre-planned unified coordinate system, a public data structure and access interface, and an AI data fusion mechanism. Based on this, it enables more precise directional sound playback and pickup, digital human posture control, and synchronization of sound with the virtual human's facial features and movements. Furthermore, by incorporating sensors, such as laser, radar, and infrared ranging devices, it achieves data fusion from more sources, further improving ranging accuracy. The proximity sensor 6, system configuration, and historical data trigger the activation and deactivation of the camera 7, microphone 8, and speaker, achieving energy savings and cost reduction.

[0075] This invention employs the above technical solutions, utilizing a proximity sensor to detect the proximity of objects; a camera to detect the image of the target object, enabling visual ranging based on actual needs and existing algorithms; a microphone for sound pickup; and a sound-emitting unit for audio playback, which can also achieve acoustic ranging based on actual needs and existing algorithms. Furthermore, it can be combined with other sensors, such as radar, IR, ultrasonic, and laser ranging. The display, through a digital human accessory, obtains the depth distance between the screen and the user. Using existing ranging and positioning technologies and virtual reality technologies, a virtual sound field can be generated, allowing the user to experience sounds from different directions, enhancing immersion and surround sound. The display, through the digital human accessory, detects and obtains the direction and intensity of sound based on the user's position and environment, controlling the sound-emitting and pickup units to achieve directional sound playback and pickup, reducing noise interference and impact on the external environment.

[0076] This invention enables displays to achieve precise synchronization between sound and the virtual human's image, ensuring consistency between sound and AI digital human movements, controlling the AI ​​digital human's posture, micro-expressions, and sensing user emotions, thus enhancing the interactive experience. Through this device, displays can better serve customers by integrating existing AI applications such as identity recognition, gesture recognition, and emotion recognition; combined with proximity sensors, sleep and wake-up functions can be more effectively implemented, achieving energy savings and cost reduction.

[0077] Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Without conflict, the embodiments and features in the embodiments of this application can be combined with each other. The components of the embodiments of this application described and illustrated herein can generally be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of this application is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

Claims

1. A control box for a display matching AI digital person, characterized in that: It includes a digital human accessory and at least one control box, the digital human accessory is electrically connected with the matched main control board; the digital human accessory includes a sound emitting unit, a proximity sensor, a camera and a microphone; the proximity sensor and the camera are arranged in the control box, at least one of the sound emitting unit and the microphone is arranged in the control box, the control box is electrically connected with a display through a matched interface arranged on the display; the proximity sensor is used for proximity detection of an object; the camera is used for detecting an image of a target object; the microphone is used for sound pickup; and the sound emitting unit is used for audio playing. The main control board controls the working of the sound emitting unit, the proximity sensor, the camera and the microphone.

2. The control box for display-matched AI digital person of claim 1, wherein: The main control board is provided with a visual processing unit, an auditory processing unit, an interface unit, a power management unit and a clock unit; the visual processing unit provides a functional interface for image acquisition, preprocessing and feature extraction; the auditory processing unit provides a functional interface for sound source positioning, echo cancellation and noise suppression; the interface unit includes a USB interface, an I2S interface, an I2C interface, an SPI interface and a storage controller; the power management unit is used for power supply of the device; and the clock unit provides a clock frequency required by the working of the main control board.

3. The control box for display-matched AI digital person of claim 2, wherein: The visual processing unit and the auditory processing unit are integrated on a main control chip, the main control chip is interactively connected with the interface unit, the power management unit and the clock unit through a bus, so as to realize data access storage and sharing; the bus includes a USB, an I2S and an SPI bus.

4. The control box for display-matched AI digital person of claim 1, wherein: The main control board is a main control board of a display, the digital human accessory of the control box is detachably electrically connected with the main control board in the display through a matched interface on the display; or the main control board is a main control board arranged in the control box; or the main control board is an independent main control board arranged outside the control box.

5. The control box for display-matched AI digital person of claim 1, wherein: The camera is a monocular camera; or the camera is a binocular camera, the binocular camera is a combination of an RGB camera and an IR camera; or the camera is a multi-lens camera.

6. The control box for display-matched AI digital person of claim 1, wherein: One of the sound emitting unit and the microphone is arranged in the control box, and the other of the sound emitting unit and the microphone is independently installed outside the control box and electrically connected with the main control board; or the sound emitting unit is a directional sound emitting unit, and the sound emitting unit and the microphone are arranged in the control box; or the sound emitting unit is a non-directional sound emitting unit, and the sound emitting unit and the microphone are arranged in the control box and are physically isolated from each other.

7. The control box for display-matched AI digital people of claim 1, wherein: The sound emitting unit is a directional sound emitting unit, the directional sound emitting unit has two or more sound emitting areas, the directional sound emitting unit is provided with a mechanical steering device, and the main control board controls the mechanical steering device of the directional sound emitting unit to act; Or the sound emitting unit is two or more directional sound emitting units, and multiple directional sound emitting units with fixed angles or with partition steering function are matched and combined to realize area coverage, that is, multiple directional sound emitting units are cross-installed to realize area coverage; Or the sound emitting unit is a combination of a directional sound emitting unit and an omnidirectional sound emitting unit, wherein the directional sound emitting unit adopts a fixed-angle sound emitting unit or a partition control sound emitting unit with a partition steering function; and the main control board controls the working of the directional sound emitting unit and the omnidirectional sound emitting unit.

8. The control box for display-matched AI digital person of claim 1, wherein: The microphone adopts a directional sound pickup unit, and the directional sound pickup unit has two or more sound pickup areas. The directional sound pickup unit is provided with a mechanical steering device, and the main control board controls the mechanical steering device of the directional sound pickup unit to act. Or the microphone is two or more directional sound pickup units, and multiple fixed-angle or directional sound pickup units with partition steering function are combined to realize area coverage, that is, multiple directional sound pickup units are cross-installed to realize area coverage, and the main control board controls the directional sound pickup units in the corresponding area to work or not to work. Or the microphone adopts an omnidirectional sound pickup unit. Or the microphone is a combination of directional sound pickup units and omnidirectional sound pickup units, wherein the directional sound pickup unit adopts a fixed-angle sound pickup unit or a partition-controlled sound pickup unit with a partition steering function; and the main control board controls the directional sound pickup or omnidirectional sound pickup to work or not to work as needed.

9. The control box for display-matched AI digital person of claim 1, wherein: The control box is fixedly installed on the target position of the display through a locking accessory and is electrically connected with the main control board of the display, thereby realizing interactive communication with the main control board of the display.

10. The control box for display-matched AI digital person of claim 1, wherein: It also includes a turntable, the turntable includes a table body, a main control device and an energy device; the control box is detachably installed on the table body, the energy device provides the turntable with rotating energy; the main control device drives the turntable to rotate; the turntable has position servo control, speed servo control and limit function; and the main control board controls the servo control system of the turntable to work.