Multimodal acquisition platform

CN224317068UActive Publication Date: 2026-06-02SHENZHEN YUEJIANG TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
SHENZHEN YUEJIANG TECH CO LTD
Filing Date
2025-07-31
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, single sound source sensors are easily affected by external interference, leading to inaccurate user location judgment and affecting the smoothness and accuracy of human-computer interaction.

Method used

Employing a multimodal acquisition platform that integrates visual sensors, human body sensors, and sound source sensors, the visual sensor captures environmental images, the human body sensor detects the presence of the user, and the sound source sensor collects bio-audio information, working together to accurately locate the user's position.

Benefits of technology

It improves the accuracy of user location determination, simplifies wiring and installation maintenance, and enhances the smoothness and accuracy of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN224317068U_ABST
    Figure CN224317068U_ABST
Patent Text Reader

Abstract

This application discloses a multimodal acquisition platform, including: a frame, a visual sensor, a human body sensor, and a sound source sensor. The visual sensor is disposed on one side of the frame; the human body sensor is disposed on the frame, with the human body sensor positioned closer to the side where the visual sensor is located, and is used to detect the presence of a user in the environment; the sound source sensor is disposed on the frame, and is used to determine the location of the sound source by collecting bio-audio information from the surrounding environment. This application's multimodal acquisition platform not only achieves a high degree of integration, thereby reducing the overall size of the multimodal acquisition platform, but also improves the accuracy of user location determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart terminal technology, specifically to a multimodal acquisition platform. Background Technology

[0002] With the development of technology, human-computer interaction devices are widely used in production and daily life. To improve user experience, environmental data is collected to determine the user's location, allowing the human-computer interaction device to turn towards the user. In related technologies, a single sound source sensor is used to collect the user's audio data. However, sound source sensors are susceptible to external interference, and a single sound source sensor cannot accurately determine the user's location, affecting the smoothness and accuracy of the interactive experience. Utility Model Content

[0003] This application provides a multimodal acquisition platform that not only achieves a high degree of integration, thereby reducing the overall size of the multimodal acquisition platform, but also improves the accuracy of user location determination.

[0004] This application provides a multimodal data acquisition platform, including:

[0005] frame;

[0006] A vision sensor is disposed on one side of the frame;

[0007] A human body sensor is disposed on the frame, with the human body sensor positioned closer to the side where the visual sensor is located. The human body sensor is used to detect the presence of a user in the environment.

[0008] A sound source sensor is disposed on the frame, and the sound source sensor is used to determine the location of the sound source by collecting biological audio information of the surrounding environment.

[0009] In some embodiments, the frame has opposing top and bottom surfaces, and a side surface connecting the top and bottom surfaces;

[0010] The visual sensor is disposed on the side of the frame, the human body sensor is disposed on the top surface of the frame and close to the side where the visual sensor is located, and the sound source sensor is disposed near the middle of the top surface of the frame.

[0011] In some embodiments, the number of sound source sensors is one or more, and when the number of sound source sensors is multiple, the multiple sound source sensors are arranged in an array.

[0012] In some embodiments, the multimodal acquisition platform further includes an environmental sensing component, which includes one or more of a temperature sensor, a humidity sensor, and an olfactory sensor, and the environmental sensing component is disposed inside the frame;

[0013] The temperature sensor is used to detect the ambient temperature, the humidity sensor is used to detect the ambient humidity, and the olfactory sensor is used to detect the gases in the environment.

[0014] In some embodiments, the multimodal acquisition platform further includes a tactile sensor disposed on the top surface of the frame, and the number of the tactile sensors is one or more.

[0015] When there are multiple tactile sensors, the multiple tactile sensors are distributed in an array with regular or irregular geometric shapes on the top surface.

[0016] In some embodiments, the number of tactile sensors is four, and the four tactile sensors are respectively located near the four corners of the top surface of the frame.

[0017] In some embodiments, the multimodal acquisition platform further includes a speaker disposed around the vision sensor for emitting sound.

[0018] In some embodiments, the multimodal acquisition platform further includes an auxiliary light source disposed around the vision sensor, the auxiliary light source being used to supplement the vision sensor with light.

[0019] In some embodiments, the multimodal acquisition platform further includes a control module, which is communicatively connected to the visual sensor, the human body sensor, and the sound source sensor. The control module is used to process the environmental data acquired by the visual sensor, the human body sensor, and the sound source sensor and generate action commands.

[0020] In some embodiments, a drive device is further included, which is tractively connected to the frame and is capable of driving the frame to move.

[0021] Beneficial effects: Compared with the prior art, the multimodal acquisition platform provided in this application embodiment not only achieves a high degree of integration by setting multiple sensors on the frame, thereby reducing the overall size of the multimodal acquisition platform, simplifying wiring, and reducing the difficulty of installation and maintenance, but also features a human body sensor positioned close to the side where the vision sensor is located to detect the presence of a user in the environment. The human body sensor and the vision sensor can move synchronously, and the sound source sensor collects audio data from the user to accurately locate the user's position. The collaborative work of multiple sensors improves the accuracy of user position determination, thereby enhancing the user experience. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a three-dimensional structural diagram of the multimodal acquisition platform of this application;

[0024] Figure 2 This is a schematic diagram of the planar structure of the multimodal acquisition platform of this application from one perspective;

[0025] Figure 3 This is a schematic diagram of the planar structure of the multimodal acquisition platform of this application from another perspective;

[0026] Figure 4 It is along Figure 3 A cross-sectional view of line AA in the middle.

[0027] Explanation of reference numerals in the attached figures:

[0028] 100. Multimodal acquisition platform; 1. Frame; 11. Top surface; 12. Bottom surface; 13. Side surface; 131. First side surface; 132. Second side surface; 133. Third side surface; 134. Fourth side surface; 2. Sensing components; 21. Visual sensor; 22. Human body sensor; 23. Sound source sensor; 24. Tactile sensor; 3. Environmental sensing components; 31. Temperature sensor; 32. Humidity sensor; 33. Olfactory sensor; 4. Output components; 41. Speaker; 42. Auxiliary light source; 421. Bracket; 5. Control module; 6. Drive device. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. In addition, it should be understood that the specific embodiments described herein are only for illustration and explanation of this application and are not intended to limit this application. In this application, unless otherwise stated, directional terms such as "up," "down," "left," and "right" generally refer to up, down, left, and right in the actual use or working state of the device, specifically the drawing directions in the accompanying drawings.

[0030] In this application, unless otherwise expressly specified and limited, the terms "connected," "linked," "stacked," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two elements or the interaction between two elements. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0031] This application provides a multimodal data acquisition platform, which will be described in detail below. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments of this application. Furthermore, the descriptions of each embodiment have their own emphasis; parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments.

[0032] Reference Figure 1 and Figure 2 One embodiment of this application provides a multimodal acquisition platform 100, including a frame 1 and a sensing component 2, the sensing component 2 being disposed on the frame 1. (Refer to...) Figure 3 and Figure 4 The multimodal data acquisition platform 100 may further include an environmental sensing component 3, an output component 4, and a control module 5. These components may also be mounted on the frame 1. The control module 5 can communicate with the environmental sensing component 3, the sensing component 2, and the output component 4. The multimodal data acquisition platform 100 may also include a drive device 6, which is connected to the frame 1 and drives the frame 1 to move. The drive device 6 can also communicate with the control module 5. The sensing component 2 is used to collect environmental data, and the control module 5 is used to process the environmental data collected by the sensing component 2, generate action commands, and control the output component 4 and / or the drive device 6 to achieve intelligent operation of the multimodal data acquisition platform 100.

[0033] Reference Figure 1 and Figure 2 The frame 1 can be shaped like a cuboid, cylinder, or other similar shapes. In this application, a cuboid is used as an example for illustration. The frame 1 has a top surface 11 and a bottom surface 12, and side surfaces 13 connecting the top surface 11 and the bottom surface 12. The side surfaces 13 may include a first side surface 131 and a second side surface 132, and a third side surface 133 and a fourth side surface 134, which are arranged opposite to each other.

[0034] Reference Figure 1The sensing component 2 may include, but is not limited to, the following sensors as needed: a visual sensor 21, a human body sensor 22, and a sound source sensor 23. Specifically, the visual sensor 21 is located on one side of the frame 1. The visual sensor 21, for example, is a camera, and is used to capture images of the surrounding environment. For example, the visual sensor 21 can take photos and videos of the environment and the user. When taking photos and videos of the user, the visual sensor 21 can obtain the user's commands, thereby enabling human-computer interaction. The human body sensor 22 is located on the frame 1, closer to the side where the visual sensor 21 is located. The human body sensor 22 is used to detect the presence of a user in the environment. The human body sensor 22 may employ technologies such as infrared sensors, detecting infrared signals emitted by the human body to determine if a person is within the monitoring range. The sound source sensor 23 is located on the frame 1. The sound source sensor 23 is used to determine the location of the sound source by collecting bio-audio information from the surrounding environment. Bio-audio information may include, for example, the user's audio information. Once the location of the sound source is determined, the driving device 6 can drive the visual sensor 21 to face the location of the sound source, that is, the driving device 6 can drive the visual sensor 21 to face the user. The sound source sensor 23 can also obtain user commands through the user's audio, thereby realizing human-computer interaction.

[0035] The control module 5 is communicatively connected to the vision sensor 21, the human body sensor 22, and the sound source sensor 23. The control module 5 is, for example, a microprocessor. The control module 5 processes the environmental data collected by the vision sensor 21, the human body sensor 22, and the sound source sensor 23 and generates action commands. (See reference...) Figure 4 The frame 1 is used to mount the drive device 6. For example, the drive device 6 is connected to the bottom surface 12 of the frame 1. The drive device 6 may include a motor, etc. The drive device 6 can drive the multimodal acquisition platform 100 to move based on the detection data of the vision sensor 21, the human body sensor 22, and the sound source sensor 23. That is, the drive device 6 receives instructions generated by the control module 5 based on the environmental data collected by the sensing components 2, thereby driving the multimodal acquisition platform 100 to move. As an example, when the sound source sensor 23 determines the location of the user through the user's audio data, the control module 5 controls the drive device 6 to drive the multimodal acquisition platform 100 to move so that the vision sensor 21 faces the user. In some embodiments, the drive device 6 can also control the movement of the multimodal acquisition platform 100 according to the instructions given by the user.

[0036] The human body sensor 22 can serve as a wake-up device for the sound source sensor 23. In an unoccupied state, only the human body sensor 22 is active, while the sound source sensor 23 can remain in sleep mode, thus reducing the energy consumption of the multimodal acquisition platform 100. When the human body sensor 22 detects the presence of a user in the environment, the sound source sensor 23 switches from sleep mode to active mode, ensuring timely and accurate determination of the user's location, thereby directing the visual sensor 21 towards the user. Simultaneously, the human body sensor 22 can not only detect the presence of a user in the environment but also determine the user's location, preventing misjudgments by the sound source sensor 23 due to non-user audio. The human body sensor 22 also enhances the anti-interference capability of the sound source sensor 23, preventing non-user audio from affecting the accuracy of its judgments.

[0037] In other embodiments of this application, the human body sensor 22 can determine the user's location, thereby causing the visual sensor 21 to face the user. For example, when the user is in a silent state, the human body sensor 22 can locate the user, thus causing the visual sensor 21 to face the user. Furthermore, both the human body sensor 22 and the sound source sensor 23 can locate the user, which not only improves the accuracy of the location but also ensures that the multimodal acquisition platform 100 can still locate the user even if one of the human body sensor 22 or the sound source sensor 23 fails.

[0038] In this application, by setting multiple sensors on the frame 1, not only is the multimodal acquisition platform 100 highly integrated, thereby reducing the overall size of the multimodal acquisition platform 100, simplifying wiring, and reducing the difficulty of installation and maintenance, but also the human body sensor 22 is set close to the side where the vision sensor 21 is located and detects whether there is a user in the environment. The human body sensor 22 and the vision sensor 21 can move synchronously. The sound source sensor 23 collects audio data from the user and accurately locates the user's position. The collaborative work of multiple sensors improves the accuracy of user position judgment, thereby improving the user experience.

[0039] In some implementations, refer to Figure 1The visual sensor 21 is disposed on the side 13 of the frame 1, more specifically, on the first side 131 of the frame 1. This allows the visual sensor 21 to face the user, providing a wider field of view and facilitating the acquisition of visual information about the surrounding environment, especially the user's visual information. The human body sensor 22 is disposed on the top surface 11 of the frame 1, close to the side 13 where the visual sensor 21 is located; that is, the human body sensor 22 is disposed on the top surface 11 of the frame 1, close to the first side 131 where the visual sensor 21 is located. This allows the human body sensor 22 to more accurately detect users approaching the monitoring range of the visual sensor 21. The sound source sensor 23 is disposed near the center of the top surface 11 of the frame 1. This allows the sound source sensor 23 to acquire audio data from all directions, more accurately determining the location of the sound source, thereby improving the accuracy of user location judgment.

[0040] In some implementations, refer to Figure 1 and Figure 3 The number of sound source sensors 23 can be one or more. When there are multiple sound source sensors 23, they are arranged in an array. The sound source sensors 23 can be, for example, microphones. In this embodiment, the sound source sensors 23 are microphones, and multiple sound source sensors 23 form a microphone array. By analyzing the sound signals received by microphones at different locations, the location of the sound source is determined using principles such as the time difference between the arrival times of the sound at different microphones. The array distribution can be, for example, a circular array or a rectangular array. In this embodiment, the multiple sound source sensors 23 are arranged in a circular array, which allows for more accurate determination of the user's location. As an example, there are eight sound source sensors 23, which are evenly distributed circumferentially.

[0041] Reference Figure 4 The multimodal acquisition platform 100 also includes an environmental sensing component 3, which includes one or more of a temperature sensor 31, a humidity sensor 32, and an olfactory sensor 33. The environmental sensing component 3 is located inside the frame 1 and can communicate with the control module 5. The temperature sensor 31, for example, is a thermistor, used to detect the ambient temperature. The humidity sensor 32, for example, is a capacitive humidity sensor, used to detect the ambient humidity. The olfactory sensor 33, for example, is a metal oxide semiconductor gas sensor, used to detect gases in the environment, such as smoke, harmful gases, and odors. By incorporating the temperature sensor 31, humidity sensor 32, and olfactory sensor 33, the multimodal acquisition platform 100 can perceive environmental information from more dimensions, further improving its integration level.

[0042] Reference Figure 1 The multimodal acquisition platform 100 also includes a tactile sensor 24, meaning the sensing component 2 may also include a tactile sensor 24. The tactile sensor 24 is disposed on the top surface 11 of the frame 1 and can communicate with the control module 5. There can be one or more tactile sensors 24. When there are multiple tactile sensors 24, they can be arranged in a regular or irregular geometric array on the top surface 11 of the frame 1.

[0043] Different functions can be switched by pressing the tactile sensor 24. The tactile sensor 24, such as a pressure sensor, generates a pressure signal when pressed by the user, which can be recognized by the multimodal acquisition platform 100, thereby triggering the corresponding function switching operation, increasing the interactivity and ease of operation of the multimodal acquisition platform 100. In this application, users can achieve human-computer interaction through voice, body movements, or the tactile sensor 24.

[0044] As an example, refer to Figure 1 There are four tactile sensors 24, which are respectively positioned near the four corners of the top surface 11 of the frame 1. This allows users to easily operate the tactile sensors 24 from different positions, improving the convenience and comfort of operation, thereby enhancing the user experience.

[0045] Reference Figures 1 to 3 The output component 4 may include a speaker 41 and an auxiliary light source 42, meaning the multimodal acquisition platform 100 may also include a speaker 41 and an auxiliary light source 42. The speaker 41 is positioned around the vision sensor 21, and may be located on the first side 131 of the frame 1, communicating with the control module 5. The speaker 41 is used to emit sound, i.e., for audio output. The speaker 41 can be used to play prompts, alarms, music, etc., enhancing user interaction. There may be one or more speakers 41; in this embodiment, there are two speakers 41. The two speakers 41 are located on either side of the vision sensor 21 along the direction from the third side 133 to the fourth side 134 of the frame 1, and preferably are symmetrically distributed. The two speakers 41 can achieve a stereo effect, improving the audio output performance and interactive experience of the multimodal acquisition platform 100.

[0046] Reference Figures 1 to 3The auxiliary light source 42, such as an LED light or LED array, is positioned around the vision sensor 21 and is communicatively connected to the control module 5. The auxiliary light source 42 provides supplemental lighting for the vision sensor 21; that is, it is used for light output. In low-light environments, the auxiliary light source 42 can provide sufficient light to the vision sensor 21, ensuring clear acquisition of visual images and improving the adaptability of the multimodal acquisition platform 100 under different lighting conditions. In some embodiments, refer to... Figure 4 The auxiliary light source 42 can also emit light of different colors and brightness according to the action commands issued by the control module 5, providing users with better lighting effects.

[0047] The auxiliary light source 42 can be one or more; in this embodiment, there are two auxiliary light sources 42. The auxiliary light source 42 is preferably an LED array, i.e., there are two LED arrays. Along the direction from the third side 133 to the fourth side 134 of the frame 1, the two auxiliary light sources 42 are respectively located on both sides of the vision sensor 21, and the two auxiliary light sources 42 are preferably symmetrically distributed. The two auxiliary light sources 42 provide better illumination for the vision sensor 21. The auxiliary light sources 42 can be connected to the frame 1 via brackets 421. The brackets 421 can be L-shaped, with one end of each bracket connected to the third side 133 and the fourth side 134 of the frame 1, and the other end connected to each of the two auxiliary light sources 42.

[0048] In one embodiment, to enable the auxiliary light source 42 to automatically supplement the light for the vision sensor 21, the sensing component 2 may further include a light sensor (not shown). The light sensor is disposed around the vision sensor 21 and is communicatively connected to the control module 5. The light sensor is used to detect the light intensity in the environment. When the light intensity is lower than a preset value, the auxiliary light source 42 can be automatically turned on, which not only facilitates user operation but also provides timely supplementary light for the vision sensor 21.

[0049] The above provides a detailed description of the multimodal acquisition platform provided by this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A multimodal acquisition platform, characterized in that, include: frame; A vision sensor is disposed on one side of the frame; A human body sensor is disposed on the frame, and the human body sensor is disposed on the side close to the visual sensor. The human body sensor is used to detect whether there is a user in the environment. as well as A sound source sensor is disposed on the frame, and the sound source sensor is used to determine the location of the sound source by collecting biological audio information of the surrounding environment.

2. The multimodal acquisition platform according to claim 1, characterized in that, The frame has a top surface and a bottom surface opposite each other, and a side surface connecting the top surface and the bottom surface; The visual sensor is disposed on the side of the frame, the human body sensor is disposed on the top surface of the frame and close to the side where the visual sensor is located, and the sound source sensor is disposed near the middle of the top surface of the frame.

3. The multimodal acquisition platform according to claim 1, characterized in that, The number of sound source sensors can be one or more, and when the number of sound source sensors is multiple, the multiple sound source sensors are arranged in an array.

4. The multimodal acquisition platform according to claim 1, characterized in that, The multimodal acquisition platform also includes an environmental sensing component, which includes one or more of a temperature sensor, a humidity sensor, and an olfactory sensor, and is disposed inside the frame. The temperature sensor is used to detect the ambient temperature, the humidity sensor is used to detect the ambient humidity, and the olfactory sensor is used to detect the gases in the environment.

5. The multimodal acquisition platform according to claim 1, characterized in that, The multimodal acquisition platform also includes a tactile sensor, which is disposed on the top surface of the frame, and the number of the tactile sensor is one or more. When there are multiple tactile sensors, the multiple tactile sensors are distributed in an array with regular or irregular geometric shapes on the top surface.

6. The multimodal acquisition platform according to claim 5, characterized in that, The number of tactile sensors is four, and the four tactile sensors are respectively set near the four corners of the top surface of the frame.

7. The multimodal acquisition platform according to claim 1, characterized in that, The multimodal acquisition platform also includes a speaker, which is positioned around the vision sensor and is used to emit sound.

8. The multimodal acquisition platform according to claim 1, characterized in that, The multimodal acquisition platform also includes an auxiliary light source, which is positioned around the vision sensor and is used to supplement the vision sensor with light.

9. The multimodal acquisition platform according to claim 1, characterized in that, The multimodal acquisition platform also includes a control module, which is communicatively connected to the visual sensor, the human body sensor, and the sound source sensor. The control module is used to process the environmental data acquired by the visual sensor, the human body sensor, and the sound source sensor and generate action commands.

10. The multimodal acquisition platform according to any one of claims 1-9, characterized in that, It also includes a drive device, which is tractively connected to the frame and is capable of driving the frame to move.