Computer AI intelligent interaction device and use method thereof

By using a concave main body design and multimodal sensing components, combined with rotation drive and cleaning adjustment functions, the problems of inconvenient rotation and difficult cleaning of existing devices are solved, thereby improving immersion and interactive usability.

CN121300583APending Publication Date: 2026-01-09JIANGXI INST OF FASHION TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511471613.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing computer AI intelligent interactive devices suffer from a fixed base that makes rotation inconvenient, resulting in poor immersion and interactive experience, and the cleaning components are difficult to replace.

Method used

It features a U-shaped main body design, equipped with omnidirectional wheels and a rotary drive component. It combines a high-definition camera, microphone array, and millimeter-wave radar for multimodal perception, integrates cleaning and adjustment components, and the control module adopts a four-layer architecture for data processing and privacy protection.

Benefits of technology

It enables flexible rotation and following of the device, enhancing immersion and interactive experience. The cleaning components are easy to replace, privacy protection is refined, and the overall interactive usability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121300583A_ABST
    Figure CN121300583A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence interaction, and particularly relates to a computer AI intelligent interaction device and a use method thereof.The device comprises a body, a base, an interaction assembly, a multi-mode sensing assembly, a cleaning assembly, an adjusting assembly and a control module. The main body is of a concave arc-shaped structure, and the rotation stability is improved through gear-gear ring transmission and auxiliary shaft and ball supporting; the multi-mode sensing assembly integrates a camera, a microphone array and a millimeter wave radar, and realizes face recognition, voice acquisition and position tracking. The control module adopts a four-layer architecture, the edge layer processes simple instructions, the cloud layer processes complex requirements, and the response efficiency and the intelligent precision are balanced. According to the invention, the structural reliability, the sensing accuracy and the privacy security are improved, and the method is suitable for intelligent interaction scenes such as families and offices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence interaction technology, specifically relating to a computer AI intelligent interaction device and its usage method. Background Technology

[0002] With the development of artificial intelligence and the Internet of Things (IoT) technologies, computer AI intelligent interactive devices have been widely used in scenarios such as home services, office assistance, and public services, becoming an important carrier of human-computer interaction. Existing interactive devices typically integrate components such as display screens, cameras, and microphones, and realize the reception and response of user commands through voice, touch, and image recognition, thus initially meeting basic interaction needs.

[0003] Chinese patent application CN120368176A discloses an AI interactive device and system based on cloud services, including an adjustment component. The adjustment component comprises a slide groove, a first lead screw, a slot, and a limiting post. The slide groove is formed within the interactive device, and the first lead screw is rotatably connected within the slide groove. This invention, through the use of a personalized component, can automatically provide personalized interactive content for different users and effectively protect user privacy. By using the adjustment component, the height of the auxiliary device can be automatically adjusted according to the user's height, and the angle of the auxiliary device can be easily adjusted and fixed, improving the user experience. By using the auxiliary component, the interactive device can automatically rotate with the user, providing good support during rotation. By using a protective component, the screen area of ​​the interactive device can be well protected when not in use, and the screen can be automatically cleaned, improving practicality.

[0004] However, in actual use, the aforementioned patents have drawbacks. Due to the fixed base, it is inconvenient to rotate and follow the device, resulting in a poor sense of immersion and interactive experience. Furthermore, the cleaning components are inconvenient to replace. Summary of the Invention

[0005] In view of the above-mentioned shortcomings in the prior art, the present invention provides a computer AI intelligent interaction device and its usage method to solve the problems in the background art.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A computer AI intelligent interactive device includes a main body with a concave cross-section, wherein one concave side is arc-shaped, and mounting grooves are provided on both sides of the main body; a base connected to the bottom of the main body, including casters at the bottom and a rotation drive assembly connected to the main body at the top; an interactive component mounted on the arc-shaped side of the main body, including a high-definition curved and touch-sensitive screen supporting text input and touch operation; a multimodal perception component mounted on the upper part of the interactive component, including: a high-definition wide-angle camera for face recognition, gesture recognition, and user location tracking; a microphone array for collecting voice commands; a millimeter-wave radar for assisting user location positioning and movement trajectory detection; and a cleaning component including a first motor, a first lead screw connected to the output end of the first motor, and... The system includes a first slider mounted on a first lead screw, a connecting block fixedly connected to it, and a cleaning rod inserted into the connecting block; an adjustment assembly including a second motor, a second lead screw connected to the output end of the second motor, a second slider mounted on the second lead screw, and an auxiliary machine fixedly connected to it; and a control module employing a four-layer architecture: a perception layer, an edge layer, a cloud layer, and an application layer. The perception layer is connected to a multimodal perception component for quickly receiving and initially processing raw perception data. The edge layer deploys a lightweight large language model for quickly parsing user commands and generating basic semantic results locally. The cloud layer communicates with the edge layer via a network and uses a large language model to process complex commands. The application layer is connected to an interaction component for accurately outputting interaction results and service interfaces.

[0007] Furthermore, the control module also includes: a data collection unit for collecting user interaction data, including facial features, voice features, and operation preferences; a user matching unit for matching existing users or registering new users based on the information from the data collection unit, and establishing a related user profile database; and a privacy protection unit that uses the AES-256 encryption algorithm to locally encrypt and store user facial and voiceprint data, with encryption triggered when the user matching unit identifies a new user and when data is updated during user interaction; and supports automatic clearing of temporary interaction data when the user leaves, provided that the high-definition wide-angle camera cannot recognize the face and the millimeter-wave radar has no signal for ≥3 minutes. The temporary interaction data includes the input command record of this interaction, unencrypted intermediate semantic results, and temporary screen display content.

[0008] Furthermore, the rotary drive assembly includes a support block, a drive shaft extending through the top of the support block, a rotary motor embedded within the support block, and a gear sleeved on the outside of the drive shaft; the bottom of the main body is provided with a gear ring that meshes with the gear; one end of the drive shaft is fixedly connected to the output end of the rotary motor, and the other end is rotatably connected to the bottom of the main body; the gear drives the main body to rotate by meshing with the gear ring, and the rotation angle is calculated using the following formula: Where θ is the rotation angle. and Based on the same two-dimensional coordinate system, with the center of the device base as the origin, the positive x-axis is horizontally to the right and the positive y-axis is vertically forward; The real-time coordinates of the user are collected jointly by a high-definition wide-angle camera and millimeter-wave radar. These are the coordinates of the device.

[0009] Furthermore, the base is provided with a sliding groove, and the bottom of the main body is designed with four auxiliary shafts; ball bearings are installed at the bottom of the auxiliary shafts.

[0010] Furthermore, the lightweight large language model at the edge layer collaborates with the large language model at the cloud layer: when the semantic matching degree S ≥ 0.8, the response is generated independently by the edge layer, where S is calculated using the following formula: ;in, For user command word vectors, The preset action template word vectors are used; when S < 0.8, the edge layer uploads the instructions and user history interaction data to the cloud layer, and the large language model generates personalized strategies and returns them to the application layer.

[0011] Furthermore, the control module also includes a sleep mode control unit; when the sleep mode control unit detects that there has been no user interaction for ≥30 minutes, it automatically turns off the backlight of the high-definition wide-angle camera and the screen, and only retains the microphone array for low-power listening; it resumes normal operation within 1.5 seconds of receiving a wake-up command.

[0012] A method of using a computer AI intelligent interactive device includes: Initialization steps: The high-definition wide-angle camera of the multimodal perception component collects the user's facial features, and the microphone array collects the voiceprint features; the data collection unit transmits the feature data to the user matching unit. If an existing user is matched, their profile is loaded; if no match is found, the new user is guided to complete the registration. Based on the user's initial position determined by the high-definition wide-angle camera, the control module calculates the rotation angle θ and drives the rotation drive component to initially face the screen towards the user. At the same time, based on the height data in the user profile, the control module drives the second motor of the adjustment component to work, and the second lead screw rotates to move the second slider and auxiliary machine to the appropriate height.

[0013] Multimodal interaction steps: Receiving user input: Receiving voice commands through a microphone array, receiving text or touch input through the screen, and recognizing gesture commands through a high-definition wide-angle camera; Data processing steps: The lightweight large language model at the edge layer parses the input instructions and calculates the semantic matching degree S; Response generation steps: If S≥0.8, the edge layer directly generates basic semantic results and outputs them through the application layer; if S<0.8, the results are uploaded to the cloud layer and combined with the user's historical data to generate a personalized strategy before being returned.

[0014] Furthermore, it also includes intelligent following steps: When the high-definition wide-angle camera and millimeter-wave radar are working normally and detect that the user is within the effective range (i.e., less than 2 meters away from the device), intelligent tracking continues; the high-definition wide-angle camera and millimeter-wave radar collect the user's position coordinates in real time. The control module calculates θ using a rotation angle formula and drives the rotation drive component to adjust the main body's orientation, ensuring the screen always faces the user. Based on a trajectory prediction formula, it anticipates the user's movement trend and initiates rotation 0.2-0.5 seconds in advance. The formula is: ;in To predict the location, This represents the historical position of the past three moments. As weight and .

[0015] Furthermore, it also includes cleaning steps: The cleaning step can be initiated by the user by starting the first motor, driving the first lead screw to rotate, and moving the first slider, connecting block and cleaning rod to clean the screen.

[0016] Furthermore, it also includes privacy protection measures; After the user matching unit verifies the user's identity, the privacy protection unit decrypts the user's historical data. If the high-definition wide-angle camera cannot recognize the face and the millimeter-wave radar has no signal for ≥3 minutes, it is determined that the user has left, and the screen display content and the temporary cache of this interaction are automatically cleared. In sleep mode, the feature collection of the data collection unit is stopped, and only the locally encrypted user profile database is retained.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. The main body adopts a concave and arc-shaped recessed design to provide structural support for the installation of curved screens and enhance the immersive experience of using the device. The mounting slots on both sides provide rigid support for the cleaning and adjustment components. The bottom auxiliary shaft and the base slide groove are connected by ball bearings to assist the rotation drive component and provide support force. Rolling friction replaces sliding friction, which reduces rotational resistance and enhances the structural reliability for long-term use. In terms of the transmission structure, the precise meshing of gears and gear rings and the rigid connection between sliders and lead screws clearly define the force transmission path and avoid the hidden dangers of transmission gaps or shaking.

[0018] 2. A three-dimensional perception network is formed by high-definition cameras, microphone arrays, and millimeter-wave radar, with clear division of labor and collaborative verification of user information, solving the problems of functional overlap or lack of coordination. In the intelligent following stage, the effective range and advance prediction time are quantified, and the weight setting of the trajectory prediction algorithm is combined to make the following response more timely; multimodal interaction is based on semantic matching degree to achieve layered processing, with simple commands responding quickly locally and complex commands being deeply optimized in the cloud, balancing response speed and interaction accuracy.

[0019] 3. The cleaning components adopt a detachable plug-in structure, facilitating component replacement and reducing maintenance costs; the adjustment function focuses on height adaptation for the user, covering the needs of users of different heights; the privacy protection mechanism refines the encryption trigger conditions and the scope of temporary data erasure, ensuring the security of sensitive data throughout its entire lifecycle through quantified rules. The coordinate reference definition in the initialization phase, the multi-scenario cleaning triggering methods, and the low-power wake-up design in sleep mode make the overall interaction more in line with user habits, enhancing the practical value of the device. Attached Figure Description

[0020] Figure 1 This is a three-dimensional structural diagram of a computer AI intelligent interaction device according to the present invention (viewpoint 1). Figure 2 This is a three-dimensional structural diagram of a computer AI intelligent interaction device according to the present invention (viewpoint 2). Figure 3 This is a partial structural schematic diagram of the base of a computer AI intelligent interaction device according to the present invention; Figure 4 This is a partial three-dimensional structural diagram of the rotation drive component of a computer AI intelligent interaction device according to the present invention; Figure 5 This is a partial cross-sectional view of the base of a computer AI intelligent interaction device according to the present invention. Figure 6 for Figure 4 A magnified view of a section at point A in the middle; Figure 7 This is a schematic diagram showing the structure between the various parts of the present invention; Figure 8 This is a schematic diagram of the structure of the multimodal sensing component of the present invention; Figure 9 This is a schematic diagram of the control module of the present invention; The reference numerals in the accompanying drawings include: 1. Main body; 11. Mounting slot; 12. Auxiliary shaft; 121. Ball bearing; 13. Gear ring; 2. Base; 21. Caster wheel; 22. Rotary drive assembly; 221. Support block; 222. Drive shaft; 223. Gear; 23. Slide groove; 3. Interaction assembly; 31. Screen; 4. Multimodal sensing assembly; 41. Connecting rod; 42. High-definition wide-angle camera; 43. Microphone array; 44. Millimeter-wave radar; 5. Cleaning assembly; 51. Cleaning rod; 511. Mounting slot; 52. Connecting block; 521. Protrusion; 53. First slider; 54. First lead screw; 55. First motor; 6. Adjustment assembly; 61. Second slider; 62. Second lead screw; 63. Second motor; 64. Auxiliary machine; 7. Control module. Detailed Implementation

[0021] To enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0022] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this patent. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0023] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present patent. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0024] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0025] Example 1: like Figure 1-9As shown, this invention provides a computer AI intelligent interactive device, whose overall structure includes a main body 1, a base 2, an interactive component 3, a multimodal sensing component 4, a cleaning component 5, an adjustment component 6, and a control module 7. These components work together to achieve intelligent interactive functions.

[0026] In this embodiment, the main body 1 serves as the core support structure of the device. Its horizontal cross-section is U-shaped, with one concave side being arc-shaped. This arc-shaped structure adapts to the installation requirements of the interactive component 3. Mounting grooves 11 are provided on both sides of the main body 1. One side connects to the cleaning component 5, providing stable support for the screw drive structure of the cleaning component. The other side connects to the adjustment component 6, ensuring the structural stability of the adjustment component during the screen lifting and lowering process. A gear ring 13 is provided at the bottom of the main body 1. This gear ring 13 cooperates with the rotation drive component 22 on the base 2 to achieve rotational adjustment of the main body 1. Simultaneously, several auxiliary shafts 12 (preferably four for better support) are designed at the bottom of the main body 1 corresponding to the sliding groove 23 of the base 2. Ball bearings 121 are installed at the bottom of the auxiliary shafts 12. When the main body 1 rotates, the ball bearings 121 roll along the sliding groove 23, providing stable support for the main body 1 and reducing rotational friction.

[0027] In this embodiment, the base 2 is connected to the bottom of the main body 1, and a caster wheel 21 is installed on its bottom to facilitate the overall movement of the device; a rotary drive assembly 22 is provided on the top to drive the main body 1 to rotate. The rotary drive assembly 22 includes a support block 221, a drive shaft 222, a rotary motor, and a gear 223. The support block 221 is fixed on the base 2, the rotary motor is embedded in the support block 221, the drive shaft 222 passes through the top of the support block 221, one end of which is fixedly connected to the output end of the rotary motor, and the other end is rotatably connected to the bottom of the main body 1; the gear 223 is sleeved on the outside of the drive shaft 222 and meshes with the gear ring 13 at the bottom of the main body 1. When the rotary motor is working, it drives the gear 223 to rotate through the drive shaft 222, and then drives the main body 1 to rotate through the meshing of the gear 223 and the gear ring 13.

[0028] In this embodiment, the interactive component 3 is installed on the curved side of the main body 1. Its core component is a high-definition curved and touch-sensitive screen 31. The screen 31 supports text input and touch operation and is the main interface for visual interaction between the device and the user. The curved design of the screen 31 enhances the immersive experience of the interaction.

[0029] In this embodiment, the multimodal perception component 4 is installed on the upper part of the interaction component 3, integrating a high-definition wide-angle camera 42, a microphone array 43, and a millimeter-wave radar 44. The high-definition wide-angle camera 42 is used for face recognition, gesture recognition, and user location tracking; the microphone array 43 is used to accurately collect the user's voice commands; and the millimeter-wave radar 44 is used to assist in user location positioning and movement trajectory detection. The three components work together to achieve the perception of multi-dimensional information about the user.

[0030] In this embodiment, the cleaning component 5 is used to clean the screen 31. Its structure includes a first motor 55, a first lead screw 54, a first slider 53, a connecting block 52, and a cleaning rod 51. The output end of the first motor 55 is connected to the first lead screw 54. The first slider 53 is mounted on the first lead screw 54. The connecting block 52 is fixedly connected to the first slider 53. The cleaning rod 51 is inserted into the connecting block 52. Figure 6 As shown, a mounting groove 511 is provided at the connection point between the end of the cleaning rod 51 and the connecting block 52, and a corresponding protrusion 521 is provided on the connecting block 52. The cooperation between the mounting groove 511 and the protrusion 521 allows the cleaning rod 51 to be replaced and positioned. When the first motor 55 is started, it drives the first lead screw 54 to rotate, which in turn drives the first slider 53 to move along the lead screw direction, thereby driving the cleaning rod 51 to move on the surface of the screen 31 through the connecting block 52 to complete the cleaning.

[0031] In this embodiment, the adjustment component 6 is used to adjust the height of the interactive component 3 to accommodate the usage needs of users of different heights. It includes a second motor 63, a second lead screw 62, a second slider 61, and an auxiliary machine 64. The output end of the second motor 63 is connected to the second lead screw 62, the second slider 61 is mounted on the second lead screw 62, and the auxiliary machine 64 is fixedly connected to the second slider 61 and rigidly connected to the interactive component 3. When the second motor 63 is working, it drives the second lead screw 62 to rotate, which in turn drives the second slider 61 and the auxiliary machine 64 to move in the vertical direction, thereby driving the interactive component 3 to rise and fall.

[0032] In this embodiment, the control module 7 adopts a four-layer architecture consisting of a perception layer, an edge layer, a cloud layer, and an application layer, serving as the core of the device's control. The perception layer connects to the multimodal perception component 4, responsible for quickly receiving and initially processing raw perception data. The edge layer deploys a lightweight large language model for rapidly parsing user commands and generating basic semantic results locally. The cloud layer communicates with the edge layer via the network, utilizing the large language model to process complex commands. The application layer connects to the interaction component 3, used for accurately outputting interaction results and service interfaces. Furthermore, the control module 7 also includes a data collection unit, a user matching unit, a privacy protection unit, and a sleep mode control unit. The data collection unit collects user facial features, voice features, and operational preferences, among other interaction data. The user matching unit matches existing users or registers new users based on the collected data, establishing a related user profile database. The privacy protection unit uses the AES-256 encryption algorithm to locally encrypt and store sensitive user data, and automatically clears temporary interaction data after the user leaves. The sleep mode control unit automatically enters a low-power state when it detects no user interaction for ≥30 minutes.

[0033] How to use: The working process of this computer AI intelligent interactive device mainly includes initialization steps, multimodal interaction steps, data processing steps, response generation steps, etc. Each step is coordinated by the control module 7 to achieve intelligent operation.

[0034] During the initialization phase, after the device is started, the high-definition wide-angle camera 42 of the multimodal perception component 4 automatically collects the user's facial features, and the microphone array 43 collects the user's voiceprint features. The data collection unit transmits these feature data to the user matching unit. The user matching unit analyzes the feature data. If a match is found with an existing user in the database, the user's profile data is directly loaded; otherwise, the new user is guided to complete registration and a new user profile is created. Simultaneously, the control module 7 uses the user's initial location coordinates determined by the high-definition wide-angle camera 42. and the device's own coordinates Through the rotation angle formula The rotation angle θ is calculated, where the coordinates are based on a two-dimensional coordinate system with the center of the device base 2 as the origin, the positive X-axis pointing horizontally to the right, and the positive Y-axis pointing vertically forward. After the calculation is completed, the rotation drive component 22 is driven to work, so that the screen 31 initially faces the user. In addition, the control module 7 controls the second motor 63 of the adjustment component 6 to start according to the height data in the user profile. The second lead screw 62 rotates, driving the second slider 61 and the auxiliary machine 64 to move, adjusting the screen 31 to the appropriate height.

[0035] During the multimodal interaction phase, the device acquires user input through various means, including receiving user voice commands via microphone array 43, receiving text or touch input via screen 31, and recognizing gesture commands via high-definition wide-angle camera 42.

[0036] In the data processing stage, the lightweight large language model at the edge layer parses the input commands and converts the user commands into word vectors. and with preset action template word vectors Using semantic matching degree formula Calculate the semantic matching degree S.

[0037] In the response generation step, S is judged and compared. When S≥0.8, the edge layer independently generates basic semantic results and outputs the interaction results on screen 31 through the application layer. When S<0.8, the edge layer uploads the instructions and user historical interaction data to the cloud layer. The large language model in the cloud layer processes complex instructions and generates personalized strategies, which are then returned to the application layer for output.

[0038] During the intelligent tracking phase, as long as the high-definition wide-angle camera 42 and the millimeter-wave radar 44 are working normally and detect that the user is within effective range (less than 2 meters from the device), the device continues to perform intelligent tracking. The high-definition wide-angle camera 42 and the millimeter-wave radar 44 collect the user's position coordinates in real time. The control module 7 calculates the rotation angle θ in real time using a rotation angle formula, driving the rotation drive component 22 to adjust the orientation of the main body 1, ensuring that the screen 31 always faces the user. Simultaneously, the device is based on a trajectory prediction formula... Predicting user movement trends, among which To predict the location, This represents the historical position of the past three moments. As weight and It starts rotating 0.2-0.5 seconds in advance, improving the timeliness of following.

[0039] During the cleaning process, the device can activate the cleaning function based on user voice commands, touch commands, or preset timer commands. At this time, the first motor 55 starts, driving the first lead screw 54 to rotate, and the first slider 53 moves along the lead screw. Through the connecting block 52, it drives the cleaning rod 51 to move back and forth on the surface of the screen 31 to complete the screen cleaning.

[0040] During the privacy protection phase, after the user matching unit verifies the user's identity, the privacy protection unit decrypts the user's historical data for use in interactions. The privacy protection unit uses the AES-256 encryption algorithm to locally encrypt and store user facial and voiceprint features. Encryption is triggered when a new user is identified and when data is updated during interaction (data update refers to a change in user feature similarity exceeding 10% or the addition of 3 or more operation preferences). When the high-definition wide-angle camera 42 fails to recognize a face and the millimeter-wave radar 44 has no signal for ≥3 minutes, it is determined that the user has left, and the system automatically clears the input command records, unencrypted intermediate semantic results, temporary screen display content, and other temporary interaction data. When the sleep mode control unit detects that there has been no user interaction for ≥30 minutes, it automatically turns off the backlight of the high-definition wide-angle camera 42 and the screen 31, retaining only the microphone array 43 for low-power listening. Upon receiving a wake-up command, it resumes normal operation within 1.5 seconds.

[0041] The above are merely embodiments of the present invention. The circuits, electronic components, and modules involved are all prior art, fully achievable by those skilled in the art, and require no further explanation. The scope of protection in this application does not involve improvements to the software and methods. Commonly known structures and characteristics in the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all prior art in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.

Claims

1. A computer AI intelligent interactive device, characterized in that: include The main body (1) has a concave cross section, with one concave side being arc-shaped, and mounting grooves (11) are provided on both sides of the main body (1). The base (2) is connected to the bottom of the main body (1) and includes a caster wheel (21) at the bottom and a rotary drive assembly (22) connected to the main body (1) at the top. Interactive component (3), installed on the curved side of the main body (1), includes a high-definition curved and touch-sensitive screen (31) that supports text input and touch operation; A multimodal perception component (4) is installed on the upper part of the interaction component (3) and includes: a high-definition wide-angle camera (42) for face recognition, gesture recognition and user location tracking; a microphone array (43) for collecting voice commands; and a millimeter-wave radar (44) for assisting user location positioning and movement trajectory detection. The cleaning assembly (5) includes a first motor (55), a first lead screw (54) connected to the output end of the first motor (55), a first slider (53) mounted on the first lead screw (54), a connecting block (52) fixedly connected thereto, and a cleaning rod (51) inserted into the connecting block (52). The adjustment assembly (6) includes a second motor (63), a second lead screw (62) connected to the output end of the second motor (63), a second slider (61) mounted on the second lead screw (62), and an auxiliary machine (64) fixedly connected thereto. The control module (7) adopts a four-layer architecture consisting of a perception layer, an edge layer, a cloud layer, and an application layer. The perception layer is connected to the multimodal perception component (4) and is used to quickly receive and initially process the raw perception data. The edge layer is equipped with a lightweight large language model, which is used to quickly parse user commands and generate basic semantic results locally. The cloud layer communicates with the edge layer through the network and uses a large language model to process complex commands. The application layer is connected to the interaction component (3) and is used to accurately output interaction results and service interfaces.

2. The computer AI intelligent interaction device as described in claim 1, characterized in that: The control module (7) further includes: The data collection unit is used to collect user interaction data, including facial features, voice features, and operation preferences; The user matching unit matches existing users or newly registered users based on information from the data collection unit, and establishes a related user profile database. The privacy protection unit uses the AES-256 encryption algorithm to locally encrypt and store user face, voiceprint features and other data. The encryption trigger condition is when the user matching unit identifies a new user and when the data is updated during user interaction. It also supports the automatic clearing of temporary interaction data when the user leaves, if the high-definition wide-angle camera (42) cannot identify the face and the millimeter-wave radar (44) has no signal for ≥3 minutes. The temporary interaction data includes the input command record of this interaction, the unencrypted intermediate semantic result and the temporary display content on the screen.

3. The computer AI intelligent interaction device as described in claim 1, characterized in that: The rotary drive assembly (22) includes a support block (221), a drive shaft (222) passing through the top of the support block (221), a rotary motor embedded in the support block (221), and a gear (223) sleeved on the outside of the drive shaft (222); the bottom of the main body (1) is provided with a gear ring (13) meshing with the gear (223); one end of the drive shaft (222) is fixedly connected to the output end of the rotary motor, and the other end is rotatably connected to the bottom of the main body (1); the gear (223) drives the main body (1) to rotate by meshing with the gear ring (13), and the rotation angle is calculated by the following formula: ,in, θ For rotation angle, and Based on the same two-dimensional coordinate system, with the center of the device base (2) as the origin, the horizontal direction to the right is the positive x-axis, and the vertical direction forward is the positive y-axis; The real-time coordinates of the user are collected jointly by a high-definition wide-angle camera (42) and a millimeter-wave radar (44). These are the coordinates of the device.

4. The computer AI intelligent interaction device as described in claim 3, characterized in that: The base (2) is provided with a sliding groove (23), and the bottom of the main body (1) is designed with four auxiliary shafts (12); the bottom of the auxiliary shafts (12) is equipped with ball bearings (121).

5. The computer AI intelligent interaction device as described in claim 1, characterized in that: The lightweight large language model in the edge layer works collaboratively with the large language model in the cloud layer: when the semantic matching degree S ≥ 0.8, the edge layer generates the response independently, where S is calculated using the following formula: ;in, For user command word vectors, The preset action template word vectors are used; when S < 0.8, the edge layer uploads the instructions and user history interaction data to the cloud layer, and the large language model generates personalized strategies and returns them to the application layer.

6. The computer AI intelligent interaction device as described in claim 2, characterized in that: The control module (7) also includes a sleep mode control unit; when the sleep mode control unit detects that there is no user interaction for ≥30 minutes, it automatically turns off the backlight of the high-definition wide-angle camera (42) and the screen (31), and only retains the microphone array (43) for low-power listening; and resumes normal operation within 1.5 seconds of receiving a wake-up command.

7. A method of using a computer AI intelligent interactive device, applied to the device according to any one of claims 1-6, characterized in that, Includes the following steps: Initialization steps: The high-definition wide-angle camera (42) of the multimodal perception component (4) collects the user's facial features, and the microphone array (43) collects the voiceprint features; The data collection unit transmits the feature data to the user matching unit. If an old user is matched, their profile is loaded. If no match is found, the new user is guided to complete the registration. The control module (7) calculates the rotation angle θ based on the user position initially located by the high-definition wide-angle camera (42) and drives the rotation drive component (22) to make the screen (31) initially face the user. At the same time, the control module (7) drives the second motor (63) of the adjustment component (6) to work according to the height data in the user profile. The second lead screw (62) rotates and drives the second slider (61) and the auxiliary machine (64) to move to the appropriate height. Multimodal interaction steps: Receiving user input: Receiving voice commands through the microphone array (43), receiving text or touch input through the screen (31), and recognizing gesture commands through the high-definition wide-angle camera (42); Data processing steps: The lightweight large language model at the edge layer parses the input instructions and calculates the semantic matching degree S; Response generation steps: If S≥0.8, the edge layer directly generates basic semantic results and outputs them through the application layer; if S<0.8, the results are uploaded to the cloud layer and combined with the user's historical data to generate a personalized strategy before being returned.

8. The method of using a computer AI intelligent interaction device as described in claim 7, characterized in that: It also includes intelligent following steps: When the high-definition wide-angle camera (42) and millimeter-wave radar (44) are working normally and detect that the user is within the effective range, i.e. less than 2 meters away from the device, intelligent tracking is continuously performed; the high-definition wide-angle camera (42) and millimeter-wave radar (44) collect the user's position coordinates in real time. The control module (7) calculates θ using the rotation angle formula and drives the rotation drive component (22) to adjust the orientation of the main body (1) to ensure that the screen (31) always faces the user; based on the trajectory prediction formula, it predicts the user's movement trend and starts the rotation 0.2-0.5 seconds in advance. The formula is: ;in To predict the location, This represents the historical position of the past three moments. As weight and .

9. The method of using a computer AI intelligent interactive device as described in claim 7, characterized in that: It also includes cleaning steps: The cleaning step can be performed by starting the first motor (55) according to the user's instructions, driving the first lead screw (54) to rotate, and moving the first slider (53), connecting block (52) and cleaning rod (51) to clean the screen (31).

10. The method of using a computer AI intelligent interaction device as described in claim 7, characterized in that: It also includes privacy protection steps: After the user matching unit verifies the user's identity, the privacy protection unit decrypts the user's historical data. When the high-definition wide-angle camera (42) cannot recognize the face and the millimeter-wave radar (44) has no signal for ≥3 minutes, it is determined that the user has left and the screen (31) display content and temporary cache of this interaction are automatically cleared. In sleep mode, the feature collection of the data collection unit is stopped, and only the locally encrypted user profile database is retained.

Citation Information

Patent Citations

  • Artificial intelligence interaction device and system based on cloud service

    CN120368176A