Display matched with AI digital human
By integrating components such as directional sound units, proximity sensors, and cameras, the display addresses the issues of insufficient data acquisition and interaction in AI digital human displays, achieving more efficient visual and auditory fusion and enhancing the interactive capabilities and user experience of AI digital humans.
Patent Information
- Application Number
- CN202520342789.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2035-02-28
AI Technical Summary
The current market lacks dedicated displays for AI digital humans, resulting in incomplete data collection and interaction, an inability to meet the information fusion needs between different system devices, insufficient voice interaction functions, limited interaction methods, inability to cover special groups, and a lack of emotional perception capabilities in AI digital humans.
A display designed to match AI digital humans is presented, integrating a directional sound unit, proximity sensor, camera, and microphone. Through visual and acoustic positioning technologies, combined with a main control board, visual processing unit, and auditory processing unit, data fusion and precise directional interaction are achieved.
It achieves more efficient and user-friendly human-computer interaction, enhances immersion and surround sound, ensures consistency between voice and AI digital human movements, and improves the interactive experience, especially the ability to interact with special groups.
Smart Images

Figure CN223911399U_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The utility model relates to display equipment technical field especially, it relates to a display that matches AI digital person. BACKGROUND
[0002] Digital person refers to 3D digital person image that through 3D figure person modeling, utilize the method of information science to the virtual simulation of human body in different level form and function. With the maturity of 3D, intelligent, virtual and real interaction technology, digital person starts from virtual idol IPization and changes into interactive service type. According to the personification and automation two dimensions, digital person is divided into L1-L5 five grades. Among them, L4 and L5 grade digital person are collectively referred to as " AI digital person ". They have high personification, and in the image, action and intelligence level are closer to the real person level. AI digital person with multi-modal interaction ability, not only can present the multimedia information that traditional voice dialogue cannot show, through the combination of visual AI technology, can also complete identity recognition, gesture recognition, emotion recognition and other interactive tasks, make the interactive process more rich and efficient. AI digital person is currently in the initial stage, there is still a long way to go to the mature market, and a lot of technical accumulation is needed. The current industry is breaking through the technical target of " how to make digital person more like a person, think like a person ".
[0003] Display as the presentation terminal of digital person, not only need lively present digital virtual person's posture form, micro-expression, also need play data acquisition and audio and video signal transmission and other interactive entry. The current market lacks the matching special display for digital person, which is embodied in the following aspects: 1. Data acquisition and interaction part, the information between different system devices lacks unified planning management: acoustic information and visual information are independent of each other, and data fusion (including mutual visit of ranging data) is not carried out. 2. The voice interaction device function is not perfect, such as lacking directional pickup and directional sound output function, which is easy to cause disturbance or affect other irrelevant users in public places. 3. The interactive mode is less, cannot cover all users, especially the use habits of special groups such as deaf-mutes and the elderly. 4. Limited by the terminal device acquisition data, the AI digital person lacks the emotion perception ability of the user, and also cannot well control the posture, so that the user cannot feel the friendly interactive emotion rendering of the AI digital person such as attention and micro-expression. UTILITY MODEL CONTENT
[0004] The utility model aims at providing a display that matches AI digital person, based on vision, acoustic positioning and AI technology, fusion visual and acoustic information data, realize more efficient friendly man-machine interaction.
[0005] The technical scheme adopted by the utility model is:
[0006] A display matched with an AI digital person, the display has a display screen, a shell and a main control board arranged inside the shell, and further comprises a digital person accessory electrically connected with the main control board; the display screen is used to display images, texts, video information, so as to provide a man-machine interactive interface; the digital person accessory comprises a directional sound emitting unit, a proximity sensor, a camera and a microphone arranged on the display shell, the proximity sensor is used to detect a nearby object; the camera is used to detect an image picture of a target object; the microphone is used to pick up sound; the sound emitting unit is used to play audio; the main control board controls the working of the sound emitting unit, the proximity sensor, the camera and the microphone respectively.
[0007] Further, the main control board is provided with a visual processing unit, an auditory processing unit, an interface unit, a power management unit and a clock unit; the visual processing unit provides a functional interface for image acquisition, preprocessing and feature extraction, the auditory processing unit provides a functional interface for sound source positioning, echo cancellation and noise suppression; the interface unit comprises a USB interface, an I2S interface, an I2C interface, an SPI interface and a storage controller; the power management unit is used for power supply management of the device, and the clock unit provides a clock frequency required by the working of the main control board.
[0008] Further, the visual processing unit and the auditory processing unit are integrated on a main control chip, the main control chip is interactively connected with the interface unit, the power management unit and the clock unit through a bus, so as to realize data access storage and sharing; the bus comprises a USB, an I2S, an SPI and the like.
[0009] Further, as shown in Figure 2 or 3, the display screen comprises a single-screen structure, a double-screen structure or a multi-screen structure; taking the double-screen structure as an example: it can be subdivided into double equal screens (the screen sizes are the same), a primary and secondary screen (one screen is large and the other is small); one of the screens can be used as a digital person interactive screen, and the other screen can be used for product introduction broadcast and the like, wherein the secondary screen such as the digital person interactive screen can adopt a sound collecting screen, i.e. a screen sound emitting technology; according to the integration condition, it can be divided into independent double screens (i.e. two screens, such as one is a main screen and the other is a secondary screen like a tablet) or integrated double screens (i.e. two screens are integrated into the same mechanism). The single-screen structure can utilize PIP or PBP technology to realize the segmentation of the display screen, so as to achieve the purpose of content split screen display. The multi-screen structure is mainly used for double-sided display or multi-sided display, such as adopting two pairs of letter screens in a back-to-back structure to realize double-sided display and control.
[0010] Further, the camera is a monocular camera, which has a lower cost but the ranging accuracy is affected.
[0011] Further, the camera is a binocular camera, which includes two cameras. Further, the binocular camera is a combination of an RGB camera and an IR camera, such as one RGB camera and one IR camera, which have a balanced performance in terms of environmental adaptability (such as light), ranging accuracy, power consumption, resolution, frame rate, and other parameters.
[0012] Further, the camera is a multi-view camera.
[0013] Further, the directional sound emitting unit is one, and the sound emitting area corresponding to the directional sound emitting unit is divided into two or more partitions. The directional sound emitting unit is provided with a mechanical steering device. The main control board controls the mechanical steering device to act.
[0014] Specifically, as an embodiment, a single directional sound emitting unit can be used in combination with a camera or a sensor. The main control board captures, locates, and positions the target through the combination of the camera or the sensor or a sound pickup or other ranging devices, and then adjusts the directional sound channel through the software interface of the directional sound emitting unit to achieve directional following. The directional sound emitting unit is provided with a mechanical device to realize partition steering. The partition size is determined according to the hardware and requirements. Specifically, as an optional embodiment, a 30-degree partition is used.
[0015] Further, there are two or more directional sound emitting units, which are combined to realize area coverage through a plurality of fixed-angle or partition steering directional sound emitting units. The directional sound emitting units are cross-installed to realize area coverage. The main control board controls the directional sound emitting units in the corresponding area to emit sound or not to emit sound or to change the volume.
[0016] The main control board captures, locates, and positions the target through the combination of the camera or the sensor or a sound pickup or other ranging devices. The main control board controls the directional sound emitting units in the corresponding area to emit sound or not to emit sound or to change the volume.
[0017] Further, the combination of the directional sound emitting unit and the omnidirectional sound emitting unit is used, in which the directional sound emitting unit is a fixed-angle sound emitting unit or a partition control sound emitting unit with a partition steering function. The main control board controls the directional sound emitting or omnidirectional sound emitting as needed.
[0018] Further, the microphone uses a directional sound pickup unit, and the sound emitting area corresponding to the directional sound pickup unit is divided into two or more partitions. The directional sound pickup unit is provided with a mechanical steering device. The main control board controls the mechanical steering device to act.
[0019] Specifically, as an embodiment, the main control board captures, locates, and positions the target through the combination of the camera or the sensor or a sound pickup or other ranging devices, and then controls the mechanical steering device through the software interface of the directional sound pickup unit to adjust the direction and achieve directional following.
[0020] Further, the microphone adopts two or more directional pickup units, and the regional coverage is realized by the combination of multiple directional pickup units with fixed angles or with partitioned steering function, i.e., the regional coverage is realized by the cross installation of multiple directional pickup units; the main control board controls the working or non-working of the directional pickup units.
[0021] Specifically, as an implementation, the main control board performs direction finding and positioning of the target by combining a camera or a sensor or a pickup or other ranging devices, and controls the directional pickup units in the corresponding region to work or not to work. For example, the working or non-working of the pickup units is controlled by the GPIO pin or the system interface.
[0022] Further, the microphone adopts an omnidirectional pickup unit. As an optional implementation, the omnidirectional pickup unit realizes directional following through a partitioned strengthening and weakening strategy.
[0023] Further, the microphone adopts a combination of directional pickup units and omnidirectional pickup units, wherein the directional pickup units adopt fixed-angle pickup units or partitioned control pickup units with partitioned steering function; the main control board controls the working or non-working of the directional pickup or omnidirectional pickup as needed.
[0024] Specifically, as an implementation, to avoid the situation of noise sampling distortion, the system can simultaneously start the directional pickup and omnidirectional pickup units, wherein the omnidirectional microphone is used to pick up environmental noise and sea tide sound (the directional microphone and the omnidirectional microphone are required to be close to each other), which is beneficial to obtaining correct noise signals by the later signal processing unit.
[0025] Further, the microphone adopts an omnidirectional pickup unit.
[0026] Specifically, as an implementation, the omnidirectional pickup unit realizes directional following through a partitioned strengthening and weakening strategy, i.e., through recording voiceprints or direction training models such as wake-up word recognition, combining a camera or a sensor or a pickup or other ranging devices to capture and locate the target, and then through noise reduction, weakening of non-target partitions such as eliminating shielding interference sound, and strengthening of target partitions.
[0027] Further, the shell of the display is provided with a number of digital human accessory interfaces corresponding to the number of components of the digital human accessory, and the directional sound unit, proximity sensor, camera and microphone of the digital human accessory are detachably installed on the display shell through the digital human accessory interfaces.
[0028] The utility model discloses above technical scheme, through visual ranging, sound perception ranging, and can combine other sensor, such as radar, IR, ultrasonic wave, laser ranging mode and through data fusion, obtain the depth distance of screen and user. Combining ranging positioning technology and virtual reality technology, can generate virtual sound field, make the user feel the sound from different directions, enhance the immersion and surround feeling, and according to user position and environment, the direction and intensity of sound are intelligently adjusted, realize directional playback and pickup, reduce noise interference and influence to external environment. Through data fusion, it can help to realize the accurate synchronization of sound and virtual person's portrait, ensure the consistency of sound and AI digital person action, and control AI digital person's posture, microexpression, perceive user's mood, improve interactive experience. Better service customer through the combination of identity recognition, gesture recognition, emotion recognition and other AI applications, and can more effectively realize the hibernation and wake -up function, realize energy saving and cost reduction. BRIEF DESCRIPTION OF DRAWINGS
[0029] The utility model makes further detailed description in combination with the drawings and specific embodiment to the utility model;
[0030] Figure 1 It is a display structure schematic drawing of matching AI digital person in the utility model;
[0031] Figure 2 It is the structure schematic drawing of the screen sounding technology sub - mother screen that the utility model adopts;
[0032] Figure 3 It is the single screen display structure schematic drawing of the utility model adopting double sound track piezoelectric directional sounding module;
[0033] Figure 4 It is the embodiment structure schematic drawing when the single directional sounding unit of the utility model is matched with camera;
[0034] Figure 5 It is the embodiment structure schematic drawing when the multiple directional sounding unit of the utility model is matched with camera;
[0035] Figure 6 It is the main control board core architecture schematic drawing;
[0036] Figure 7 It is the system principle schematic drawing of the display of matching AI digital person of the utility model. DETAILED DESCRIPTION
[0037] To make the purpose, technical scheme and advantage of the embodiment of the present application more clear, the technical scheme in the embodiment of the present application will be clearly and completely described below in combination with the drawings in the embodiment of the present application.
[0038] As Figures 1 to 7The utility model discloses a kind of display matched AI digital person, display has display screen, shell and main control board being located in the shell inside, it further include and the digital person accessory of main control board electric connection;Display screen is used to show image, text, video information, to provide man-machine interface;Digital person accessory includes setting in the directional sound unit of display shell, proximity sensor, camera and microphone, proximity sensor is used to detect adjacent object;Camera is used to detect the image picture of target object;Microphone is used for sound pickup;Sound unit is used for audio playback;Main control board controls the work of sound unit, proximity sensor, camera and microphone respectively.
[0039] Further, main control board is equipped with visual processing unit, auditory processing unit, interface unit, power management unit and clock unit;Visual processing unit provides the functional interface for image acquisition, pre-processing, feature extraction, auditory processing unit provides the functional interface for sound source positioning, echo cancellation, noise suppression;Interface unit includes USB interface, I2S interface, I2C interface, SPI interface and storage controller;Power management unit is used for the power supply management of equipment, clock unit provides the clock frequency required for main control board work.
[0040] Further, visual processing unit, auditory processing unit are integrated on a main control chip, and the main control chip is interactively connected with interface unit, power management unit and clock unit through bus, to realize data access storage and sharing;Bus includes USB, I2S, SPI bus.
[0041] Further, display screen includes single screen structure, double screen structure or multi-screen structure;Taking double screen structure as an example: it can be subdivided into double equal screen (screen size is same), child and mother screen (screen size screen one big one small);One of the screens can be used as digital person interactive screen 9, another screen 10 can be used for product introduction broadcast, etc., wherein the child screen, such as digital person interactive screen 9, can adopt sound gathering screen, that is, adopt screen sound emitting technology; According to integration condition, it can be divided into independent double screen (that is, independent as two screens, such as one as main screen, one as vice screen similar to tablet) or integrated double screen (that is, two screens are integrated into the same mechanism). Single screen structure can utilize PIP or PBP technology to realize the segmentation of display screen, to achieve the purpose of content split-screen display. Multi-screen structure is mainly used for double-sided display or multi-sided display, such as using two pairs of letter screen back-to-back structure to realize double-sided display and control.
[0042] Further, camera 7 is monocular camera, and the accuracy of distance measurement is affected.
[0043] Further, the camera 7 is a binocular camera, which includes two cameras 7. Further, the binocular camera 7 is a combination of an RGB camera and an IR camera, such as one RGB camera and one IR camera. In terms of environmental adaptability (such as light), ranging accuracy, power consumption, resolution, frame rate and other parameters, the comprehensive performance is relatively balanced.
[0044] Further, the camera 7 is a multi-view camera.
[0045] Further, the directional sound emitting unit 5 is one, and the sound emitting area corresponding to the directional sound emitting unit 5 is divided into two or more partitions. The directional sound emitting unit 5 is provided with a mechanical steering device. The main control board captures and locates the target through the camera 7, and then calls the software interface of the directional sound emitting unit 5 to control the mechanical steering device to adjust the directional sound channel to realize directional following.
[0046] Specifically, as an implementation, a single directional sound emitting unit 5 can be used in cooperation with a camera 7 or a proximity sensor 6. The main control board captures and locates the target through the camera or other sensors, and then calls the software interface of the directional sound emitting unit 5 to adjust the directional sound channel to realize directional following. The directional sound emitting unit 5 is provided with a mechanical device to realize partition steering. The partition size is determined according to hardware and requirements. Specifically, as an optional implementation, a 30-degree partition is used.
[0047] Further, the directional sound emitting unit 5 is two or more, and the partition control is realized through a combination of multiple directional sound emitting units with fixed angles or partition steering functions, that is, multiple directional sound emitting units are cross-installed to realize area coverage. The main control board captures and locates the target through the camera or the proximity sensor, and controls the directional sound emitting unit in the corresponding area to emit sound or not to emit sound.
[0048] Further, the directional sound emitting unit is combined with an omnidirectional sound emitting unit, wherein the directional sound emitting unit is a fixed-angle sound emitting unit or a partition control sound emitting unit with a partition steering function. The main control board controls the directional sound emission or omnidirectional sound emission as needed.
[0049] Further, the microphone uses a directional sound pickup unit, and the sound emitting area corresponding to the directional sound pickup unit is divided into two or more partitions. The directional sound pickup unit is provided with a mechanical steering device. The main control board captures and locates the target through the camera or the proximity sensor, and then calls the software interface of the directional sound pickup unit to control the mechanical steering device to adjust the direction to realize directional following.
[0050] Further, the microphone adopts two or more directional sound pickup units, and the directional sound pickup units are combined to realize partition control through multiple fixed angles or partition control with steering function, that is, multiple directional sound pickup units are cross-installed to realize area coverage; the main control board controls the directional sound pickup units in the corresponding area to work or not work through the target direction finding, distance measuring and positioning by the camera or proximity sensor.
[0051] Further, the microphone adopts a combination of directional sound pickup units and omnidirectional sound pickup units, wherein the directional sound pickup units adopt sound pickup units with fixed angles or partition control sound pickup units with steering function; the main control board controls the directional sound pickup or omnidirectional sound pickup to work or not work as needed. In order to avoid the situation of noise sampling distortion, the system can simultaneously start the directional sound pickup and omnidirectional sound pickup units, wherein the omnidirectional microphone is used to pick up environmental noise and sea tide sound (the directional microphone and the omnidirectional microphone are required to be close to each other), which is beneficial to the correct noise signal obtained by the post-processing unit.
[0052] Further, the microphone adopts an omnidirectional sound pickup unit to realize directional following through partition strengthening and weakening strategy, that is, through recording voice prints or direction training models such as wake-up word recognition, the target is captured, oriented and positioned by combining the camera or sensor or sound pickup or other distance measuring devices, and then the non-target partition is weakened through noise reduction, such as eliminating shielding interference sound, and the target partition is strengthened.
[0053] Further, the shell of the display is provided with a digital human accessory interface 2 corresponding to the number of components of the digital human accessory, and the directional sound unit, proximity sensor, camera and microphone of the digital human accessory are respectively detachably installed on the display shell through the digital human accessory interface 2.
[0054] The specific principles of the utility model will be described in detail as follows:
[0055] For example, Figure 1As shown, this utility model is a device for using a display 1 in conjunction with an AI digital human. The display is equipped with a camera 7, an array microphone 8 (MIC), a directional sound unit 5, and may include various sensors (such as a proximity sensor 6). These components can be integrated into the display 1 as a single unit or as separate components, and can be connected to the display 1 via interfaces such as USB or I2S, and fixed in appropriate positions on the display 1, such as the top, middle, bottom, or side of the device, depending on the application and device size. The advantage of using a single unit is that it is easy to connect and quick to assemble; the disadvantage is that it requires a higher level of design expertise, especially in addressing the impact of resonance or echo from the sound unit on the pickup unit (this can be achieved by adding a housing for isolation and then integrating, for example, placing the sound unit and the pickup unit in different housings, and reducing the impact by adding shock-absorbing brackets, using sound-insulating materials, sound-absorbing materials, sealing the housing, or using independent chambers). The advantage of using separate components directly integrated into the display 1 is flexible assembly, allowing for adjustment of component positions as needed (e.g., ...). Figure 4 As shown, the camera 7 and microphone 8 are placed on the top, and the sound unit 5 is placed on the side of the device; the disadvantage is that there are many wires and the assembly is more complicated.
[0056] Taking camera 7 as an example, which can be either a monocular or binocular camera, a monocular camera 7 contains only one camera 7, resulting in lower cost, but affecting ranging accuracy. A binocular camera contains two cameras, which can be a combination of an RGB camera and an IR camera, such as one RGB camera 7 paired with one IR camera 7. This provides a more balanced performance in terms of environmental adaptability (e.g., lighting), ranging accuracy, power consumption, resolution, and frame rate. In addition to using a directional microphone, microphone 8 can also employ a combination of a directional and an omnidirectional microphone. The omnidirectional microphone 8 is used to pick up ambient noise and ocean currents (requiring the directional and omnidirectional microphones to be positioned close to each other), which helps the signal processing unit obtain the correct noise signal and avoids noise sampling distortion.
[0057] For larger models or applications where a single directional component cannot cover the area, extended combination zone control and other methods can be used to achieve directional following (i.e., simultaneously achieving precise positioning and area coverage). Extended combination zone control primarily involves interactive linkage control between the host control board and the directional sound unit 5, and can include the following solutions:
[0058] like Figure 2 As shown, in one preferred embodiment, a single component works with a camera 7 or a sensor. The main control board uses the camera 7 or other sensors to capture and locate the target, and then adjusts the directional sound channel by calling the software interface of the directional sound unit 5 to achieve directional tracking. The structural diagram is shown below. Figure 4As shown, the directional sound unit 5 has a built-in mechanical device to achieve zone-based turning. The size of the zone depends on the hardware and requirements, and a 30-degree angle for each zone is recommended.
[0059] like Figure 3 As shown, in Scheme 2, a preferred implementation, multiple fixed-angle sound-emitting units are combined for zoned control. Each sound-emitting unit has a fixed directional angle, and area coverage is achieved through the cross-installation of multiple components. Combined with camera 7 or sensor-based direction finding and ranging, the main control board controls the corresponding area components to emit or not emit sound. A schematic diagram is shown below. Figure 5 As shown.
[0060] Furthermore, as a preferred implementation, Scheme Three combines Schemes One and Two, using multiple sound-emitting units with zone-direction functionality for zone control. This means that area coverage is achieved through the cross-installation of multiple sound-emitting components with zone-direction functionality. Combined with camera 7 or sensor-based direction finding and ranging for positioning, the main control board controls the corresponding area components to emit sound, remain silent, or adjust the volume. Compared to Scheme One, Scheme Three can achieve a more multi-dimensional and three-dimensional solution or reduce the performance and specification requirements of the zone-direction mechanical device; compared to Scheme Two, Scheme Three can achieve more comprehensive area coverage, reduce blind spots, and may reduce the number of sound-emitting units, making installation and configuration more flexible.
[0061] Furthermore, as a preferred implementation method, Scheme Four employs a combination of directional sound-emitting units and omnidirectional sound-emitting units (such as loudspeakers). The directional sound-emitting unit can be a fixed-angle sound-emitting unit or a zone-controlled sound-emitting unit with zone-direction functionality. The system controls directional or omnidirectional sound emission as needed, for example, by controlling the power supply to the loudspeaker via GPIO pins or by controlling whether the sound-emitting unit emits sound or changes its volume via the system interface.
[0062] Furthermore, the sound-generating components of the microphone directional tracking function scheme are similar. Further, in one preferred implementation scheme, the microphone 8 employs a directional pickup unit, and the sound-generating area corresponding to the directional pickup unit is divided into two or more zones; the directional pickup unit has a built-in mechanical steering device; the main control board uses a camera or proximity sensor to capture and locate the target, and then calls the software interface of the directional pickup unit to control the mechanical steering device to adjust the direction to achieve directional tracking.
[0063] Further, as a preferred embodiment of scheme two, the microphone 8 adopts two or more directional pickup units, and the directional pickup units are combined to realize partition control through multiple fixed angles or partition control with steering function. The directional pickup units are cross-installed to realize regional coverage. The main control board controls the directional pickup units in the corresponding region to work or not work through the camera or proximity sensor for target direction finding, distance measuring and positioning. For example, the GPIO pin controls the power supply or the system interface controls whether the pickup unit works or not.
[0064] Further, as a preferred embodiment of scheme three, the microphone 8 adopts a combination of directional pickup units and omnidirectional pickup units. The directional pickup units adopt fixed-angle pickup units or partition control pickup units with steering function. The main control board controls the directional pickup or omnidirectional pickup to work or not work as needed. In order to avoid the situation of noise sampling distortion, the system can start the directional pickup and omnidirectional pickup units at the same time. The omnidirectional microphone is used to pick up environmental noise and sea tide sound (the directional microphone and the omnidirectional microphone are required to be close to each other), which is beneficial to the correct noise signal obtained by the later signal processing unit.
[0065] Further, the microphone adopts omnidirectional pickup units to realize directional following through partition strengthening and weakening strategy, that is, through recording voice prints or direction training models such as wake-up word recognition, combining the camera or sensor or pickup or other distance measuring devices to capture and locate the target, and then through noise reduction, weakening non-target partition such as eliminating shielding interference sound, and strengthening target partition and other ways.
[0066] As shown in Figure 6 , it is a schematic diagram of the core architecture of the main control board, which includes the description of the main function modules of the terminal side vision and hearing, and the description of the main interface. The vision processing unit mainly includes image acquisition, preprocessing, feature extraction and other related function interfaces. The hearing processing courtyard mainly includes sound source positioning, echo cancellation, noise suppression and other function interfaces. The two are interacted with other unit modules of the main control board through USB, I2S, SPI and other buses to realize data access storage and sharing.
[0067] As shown in Figure 7As shown, the system model of the utility model contains the main technical principle of each component and the process and application of data fusion. Taking the fusion of the ranging data of the camera 7 and the microphone 8 as an example, during system design, the data access interface is planned, including a unified coordinate system, so as to realize mutual access and use of data. In order to further improve the ranging accuracy, traditional sensing devices including laser ranging, radar sensing, infrared ranging, etc. can be considered to be added, and the corresponding data access interface is planned, and the data acquisition timing, synchronization mechanism and fusion algorithm are improved, so that the data validity and timeliness are ensured. On the basis of data fusion, the system can obtain more accurate user position information, and then realize gesture recognition, posture control, directional sound playing, directional sound picking and other functions. The start and stop of the camera 7 and the microphone 8, the directional sound playing unit 5 can be triggered by the proximity sensor 6, system configuration and historical data, so that the effect of energy saving and cost reduction is realized.
[0068] The utility model discloses a mechanism of public data structure and access interface and AI data fusion in advance planning unified coordinate system, realizes information effective, fast fusion. On this basis, more accurate directional sound playing, sound picking, digital posture control and realization sound and virtual person's portrait action synchronization. Another can through the sensor, such as adding laser, radar, infrared ranging, etc. Sensing device, realize more source data fusion, further improve the ranging accuracy. The start and stop of the camera 7 and the microphone 8, the speaker are triggered by the proximity sensor 6, system configuration and historical data, so that the effect of energy saving and cost reduction is realized.
[0069] The utility model discloses the above technical scheme, through visual ranging, sound perception ranging, and can combine other sensors, such as radar, IR, ultrasonic, laser ranging, etc. Mode, and through data fusion, obtain the depth distance of screen and user. Combining ranging positioning technology and virtual reality technology, a virtual sound field can be generated, so that the user can feel the sound from different directions, enhance the sense of immersion and surround, and intelligently adjust the direction and intensity of the sound according to the user position and environment, realize directional sound playing and sound picking, reduce noise interference and influence on external environment. Through data fusion, it can help to realize the accurate synchronization of sound and virtual person's portrait, ensure the consistency of sound and AI digital person's action, and control the posture, micro-expression and emotion perception of AI digital person, improve the interactive experience. By combining identity recognition, gesture recognition, emotion recognition and other AI applications, the customer can be better served. Combined with the application of proximity sensor 6, the sleep and wake-up functions can be more effectively realized, and energy saving and cost reduction are realized.
[0070] It is apparent that the described embodiments are only some — but not all — of the embodiments of the present application. The embodiments described in this application and features in the embodiments can be combined with each other in cases without conflict. The components of the embodiments of the present application, which are generally described and shown in the accompanying drawings, can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work, shall fall within the scope of protection of the present application.
Claims
1. A display matching an AI digital person, the display having a display screen, a housing, and a main control board arranged inside the housing, characterized in that: It also includes a digital human accessory electrically connected with the main control board; The display screen is used to display images, texts, video information, so as to provide a man-machine interactive interface; the digital human accessory includes a directional sound unit, a proximity sensor, a camera and a microphone arranged on the display housing, the proximity sensor is used to detect the nearby object; The camera is used to detect the image picture of the target object; The microphone is used for sound pickup; The sound unit is used for audio playing; The main control board controls the working of the sound unit, the proximity sensor, the camera and the microphone respectively.
2. The display matching an AI digital person of claim 1, wherein: The display screen includes a single screen structure, a double screen structure or a multi-screen structure; one screen in the double screen structure or the multi-screen structure is used as a digital human interactive screen.
3. The display matching an AI digital person of claim 1, wherein: The main control board is provided with a visual processing unit, an auditory processing unit, an interface unit, a power management unit and a clock unit; the visual processing unit provides a functional interface for image acquisition, preprocessing and feature extraction, the auditory processing unit provides a functional interface for sound source positioning, echo cancellation and noise suppression; the interface unit includes a USB interface, an I2S interface, an I2C interface, an SPI interface and a storage controller; the power management unit is used for power supply management of the device, and the clock unit provides a clock frequency required by the working of the main control board.
4. The display matching an AI digital person of claim 3, wherein: The visual processing unit and the auditory processing unit are integrated on a main control chip, the main control chip is interactively connected with the interface unit, the power management unit and the clock unit through a bus, so as to realize data access storage and sharing; the bus includes a USB, an I2S and an SPI bus.
5. The display matching an AI digital person of claim 1, wherein: The camera is a monocular camera; or the camera is a binocular camera, which is a combination of an RGB camera and an IR camera; Or the camera is a multi-lens camera.
6. The display matching an AI digital person of claim 1, wherein: There is one directional sound unit, the sound area corresponding to the directional sound unit is divided into two or more subareas, the directional sound unit is provided with a mechanical steering device, and the main control board controls the mechanical steering device to adjust the directional sound channel and realize directional following; Or there are two or more directional sound units, a plurality of directional sound units with fixed angles or with partition steering function are combined to realize area coverage, that is, a plurality of directional sound units are cross-installed to realize area coverage, and the main control board controls the directional sound units in the corresponding area to sound or not to sound or to change the volume.
7. The display matching an AI digital person of claim 1, wherein: The combination of the directional sound unit and the omnidirectional sound unit, wherein the directional sound unit adopts a fixed-angle sound unit or a partition-controlled sound unit with a partition steering function; the main control board controls the directional sound or the omnidirectional sound as required.
8. The display matching an AI digital person of claim 1, wherein: The microphone adopts one directional pickup unit, and the directional pickup unit has two or more pickup areas, the directional pickup unit is provided with a mechanical steering device, and the main control board controls the mechanical steering device of the directional pickup unit; Or the microphone adopts two or more directional pickup units, a plurality of directional pickup units with fixed angles or with partition steering function are combined to realize area coverage, that is, a plurality of directional pickup units are cross-installed to realize area coverage, and the main control board controls the directional pickup units in the corresponding area to work or not to work; Or the microphone adopts an omnidirectional pickup unit.
9. The display matching an AI digital person of claim 1, wherein: The microphone is a combination of a directional sound pickup unit and an omnidirectional sound pickup unit, wherein the directional sound pickup unit adopts a sound pickup unit with a fixed angle or a partition control sound pickup unit with a partition steering function; and the main control board controls the directional sound pickup or the omnidirectional sound pickup to work or not to work as required.
10. The display matching an AI digital person of claim 1, wherein: The shell of the display is provided with a digital human accessory interface corresponding to the number of components of the digital human accessory, and the directional sound unit, the proximity sensor, the camera and the microphone of the digital human accessory are respectively detachably installed on the display shell through the digital human accessory interface.