Visual data learning system for human-computer interaction
By designing a visual data learning system for human-computer interaction, using dynamic visual data of user's eyebrows and mouth, the intelligent analysis of user's real semantics is achieved, and the problems of limited and inaccurate communication channels in the existing technology are solved, and the reliability and effectiveness of communication channels are improved.
Patent Information
- Application Number
- CN202510132993.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing human-computer interaction system has limited and single communication channels, making it difficult to accurately analyze the patient's real semantic information.
Design a visual data learning system, and through targeted artificial intelligence models, use dynamic visual data of current user's eyebrows and mouth to achieve intelligent analysis of user's real semantics. The system includes a content grabbing mechanism, a segmented construction mechanism, an object analysis mechanism, a data capture device and a detection execution device. Through multiple training of radial-based neural networks, combined with preset frame rate and resolution, visual data is synchronized to the AI detection model and output user's string data.
By intelligently analyzing the dynamic visual data of users' eyebrows and mouth, we create more and more reliable communication channels, which improves the effectiveness and reliability of the communication channels.
Smart Images

Figure CN119987560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction, and more specifically, to a visual data learning system for human-computer interaction. Background Art
[0002] Human-computer interaction involves multiple disciplines and aims to create an efficient interactive environment that conforms to human cognition. Its core elements include user-centered design, usability, interactive systems, interface design, user experience, etc. Future development directions include brain-computer interface, augmented reality, multimodal interaction, etc. Human-computer interaction is not only a technical tool, but also an intelligent partner, which requires interdisciplinary cooperation and ethical considerations. Computer human-computer interaction is a field that intersects multiple disciplines such as computer science, design, psychology and sociology. It focuses on how to make computer technology better serve human needs. This field not only focuses on technical optimization, but also attaches importance to the harmonious coexistence of technology and human activities. The goal is to create an interactive environment that is both efficient and in line with human cognitive habits.
[0003] Although the human-computer interaction system in the existing technology provides a variety of communication channels for patients with limited mobility and poor communication, such as voice communication channels, expression communication channels, etc., these communication channels are limited after all, and they are all single-channel communication channels, resulting in the communication information obtained may not be the semantic information truly expressed by the patient. Therefore, a richer and more accurate communication channel is needed to achieve reliable and accurate analysis of the patient's true semantics. Summary of the invention
[0004] In order to solve the technical problems in the prior art, the present invention provides a visual data learning system for human-computer interaction, which can adopt a targeted artificial intelligence model to realize intelligent analysis of the real semantics of the current user based on the dynamic visual data of the eyebrows and the dynamic visual data of the mouth of the current user in a set time segment, thereby creating more communication channels for human-computer interaction and improving the reliability and effectiveness of the communication channels.
[0005] According to the present invention, a visual data learning system for human-computer interaction is provided, the system comprising: The content capture mechanism is arranged opposite to the face of the current user and captures the real-time captured images corresponding to each consecutive moment on the time axis at a preset frame rate within a set time segment, wherein the each moment is evenly distributed on the time axis; A stepwise construction mechanism is used to perform multiple training actions on the radial basis neural network to obtain a radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism; An object analysis mechanism connected to the content capture mechanism, used to identify an eyebrow sub-image in a real-time captured image based on a grayscale value distribution interval corresponding to the eyebrow part of a human body, and also to identify a mouth sub-image in a real-time captured image based on an appearance imaging feature corresponding to the mouth part of a human body; A data capture device, connected to the step-by-step construction mechanism and the object analysis mechanism, respectively, for outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment; A detection execution device, connected to the sub-construction mechanism and the data capture device, respectively, for synchronously inputting the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI detection model, so as to execute the AI detection model and obtain the character string data expressed by the current user in the set time segment output by the AI detection model; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the human eyebrow part, and the mouth sub-image in the real-time captured image is identified based on the appearance imaging feature corresponding to the human mouth, including: the grayscale value distribution interval corresponding to the human eyebrow part is a value interval limited by the grayscale upper limit value corresponding to the human eyebrow part and the grayscale lower limit value corresponding to the human eyebrow part; Among them, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the eyebrow part of the human body, and the mouth sub-image in the real-time captured image is identified based on the shape imaging features corresponding to the human mouth, which also includes: the shape imaging features corresponding to the human mouth are the respective reference mouth patterns corresponding to the human mouth at various opening and closing degrees.
[0006] Therefore, the present invention has at least the following three beneficial technical effects: Firstly, multiple training actions are performed on the radial basis neural network to obtain the radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism, thereby completing the targeted design of the AI detection model for human-computer interaction; Secondly: adopting a content capture mechanism arranged opposite to the face of the current user, capturing real-time captured images corresponding to each continuous moment on the time axis at a preset frame rate within a set time segment, identifying the eyebrow sub-image in the real-time captured image based on the grayscale value distribution interval corresponding to the eyebrow part of the human body, identifying the mouth sub-image in the real-time captured image based on the appearance imaging features corresponding to the mouth of the human body, and outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment, thereby completing the customized screening of the visualization data at each moment for subsequent intelligent detection; Again: the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism and the various visual data corresponding to each moment in the set time segment are synchronously input into the AI detection model to execute the AI detection model, and obtain the character string data expressed by the current user in the set time segment output by the AI detection model, thereby realizing intelligent semantic analysis based on the dynamic visual data of the eyebrows and the dynamic visual data of the mouth of the current user in the set time segment, and creating more communication channels for human-computer interaction.
[0007] The visual data learning system for human-computer interaction of the present invention has a compact structure and stable operation. Since a targeted artificial intelligence model can be used, the real semantics of the current user can be intelligently analyzed based on the dynamic visual data of the eyebrows and the dynamic visual data of the mouth of the current user in a set time segment, thereby creating more communication channels for human-computer interaction and improving the reliability and effectiveness of the communication channels. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Those skilled in the art may better understand the numerous advantages of the present invention by referring to the accompanying drawings, in which: Figure 1 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the first embodiment of the present invention.
[0009] Figure 2 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the second embodiment of the present invention.
[0010] Figure 3 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the third embodiment of the present invention. DETAILED DESCRIPTION
[0011] The implementation scheme of the visual data learning system for human-computer interaction of the present invention will be described in detail below with reference to the accompanying drawings.
[0012] Figure 1is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to a first embodiment of the present invention, the system comprising: The content capture mechanism is arranged opposite to the face of the current user and captures the real-time captured images corresponding to each consecutive moment on the time axis at a preset frame rate within a set time segment, wherein the each moment is evenly distributed on the time axis; For example, the content capture mechanism is arranged opposite to the face of the current user, and captures the real-time captured images corresponding to the continuous moments on the time axis at a preset frame rate within a set time segment, and the moments are evenly distributed on the time axis. The content capture mechanism includes: a built-in frame rate setting unit, a capture execution unit and a timing unit; A stepwise construction mechanism is used to perform multiple training actions on the radial basis neural network to obtain a radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism; An object analysis mechanism connected to the content capture mechanism, used to identify an eyebrow sub-image in a real-time captured image based on a grayscale value distribution interval corresponding to the eyebrow part of a human body, and also to identify a mouth sub-image in a real-time captured image based on an appearance imaging feature corresponding to the mouth part of a human body; A data capture device, connected to the step-by-step construction mechanism and the object analysis mechanism, respectively, for outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment; A detection execution device, connected to the sub-construction mechanism and the data capture device, respectively, for synchronously inputting the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI detection model, so as to execute the AI detection model and obtain the character string data expressed by the current user in the set time segment output by the AI detection model; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the human eyebrow part, and the mouth sub-image in the real-time captured image is identified based on the appearance imaging feature corresponding to the human mouth, including: the grayscale value distribution interval corresponding to the human eyebrow part is a value interval limited by the grayscale upper limit value corresponding to the human eyebrow part and the grayscale lower limit value corresponding to the human eyebrow part; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the gray value distribution interval corresponding to the eyebrow part of the human body, and the mouth sub-image in the real-time captured image is identified based on the shape imaging feature corresponding to the human mouth, and further includes: the shape imaging feature corresponding to the human mouth is each reference mouth pattern corresponding to the human mouth at various opening and closing degrees; Among them, the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism and the various visual data corresponding to each moment in the set time segment are synchronously input into the AI detection model to execute the AI detection model, and the character string data expressed by the current user in the set time segment output by the AI detection model is obtained, including: using programmable logic devices to realize simulation and emulation of the AI detection model.
[0013] Figure 2 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the second embodiment of the present invention.
[0014] and Figure 1 different, Figure 2 The visual data learning system for human-computer interaction in may also include the following components: A field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the field control interface is used to respectively realize the synchronous driving control of each of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device.
[0015] Figure 3 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the third embodiment of the present invention.
[0016] and Figure 1 different, Figure 3 The visual data learning system for human-computer interaction in may also include the following components: A parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the parameter service mechanism is used to provide the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device with their respective required operating current values.
[0017] Next, the specific structure of the visual data learning system for human-computer interaction of the present invention will be further described.
[0018] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: An FPGA chip is used to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
[0019] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing maximum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
[0020] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing minimum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
[0021] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing median filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
[0022] And in a visual data learning system for human-computer interaction according to any embodiment of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing edge sharpening processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
[0023] In addition, in the visual data learning system for human-computer interaction, the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment are synchronously input into the AI detection model to execute the AI detection model, and obtaining the character string data expressed by the current user in the set time segment output by the AI detection model includes: using a synchronous driving mechanism to synchronously input the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI detection model; Many details of the present invention may be changed without departing from its spirit and scope. In addition, the description of the preferred embodiments of the present invention is provided for illustrative purposes only, not for limiting the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A visual data learning system for human-computer interaction, characterized in that: The system comprises: The content capture mechanism is arranged opposite to the face of the current user and captures the real-time captured images corresponding to each consecutive moment on the time axis at a preset frame rate within a set time segment, wherein the each moment is evenly distributed on the time axis; A stepwise construction mechanism is used to perform multiple training actions on the radial basis neural network to obtain a radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism; An object analysis mechanism connected to the content capture mechanism, used to identify an eyebrow sub-image in a real-time captured image based on a grayscale value distribution interval corresponding to the eyebrow part of a human body, and also to identify a mouth sub-image in a real-time captured image based on an appearance imaging feature corresponding to the mouth part of a human body; A data capture device, connected to the step-by-step construction mechanism and the object analysis mechanism, respectively, for outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment; A detection execution device, connected to the sub-construction mechanism and the data capture device, respectively, for synchronously inputting the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI detection model, so as to execute the AI detection model and obtain the character string data expressed by the current user in the set time segment output by the AI detection model; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the human eyebrow part, and the mouth sub-image in the real-time captured image is identified based on the appearance imaging feature corresponding to the human mouth, including: the grayscale value distribution interval corresponding to the human eyebrow part is a value interval limited by the grayscale upper limit value corresponding to the human eyebrow part and the grayscale lower limit value corresponding to the human eyebrow part; Among them, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the eyebrow part of the human body, and the mouth sub-image in the real-time captured image is identified based on the shape imaging features corresponding to the human mouth, which also includes: the shape imaging features corresponding to the human mouth are the respective reference mouth patterns corresponding to the human mouth at various opening and closing degrees.
2. The visual data learning system for human-computer interaction according to claim 1, characterized in that: The preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism and the respective visual data corresponding to each moment in the set time segment are synchronously input into the AI detection model to execute the AI detection model, and the character string data expressed by the current user in the set time segment output by the AI detection model is obtained, including: using a programmable logic device to realize simulation and emulation of the AI detection model.
3. The visual data learning system for human-computer interaction as claimed in claim 2, characterized in that: The system further comprises: A field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the field control interface is used to respectively realize the synchronous driving control of each of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device.
4. The visual data learning system for human-computer interaction according to claim 2, characterized in that: The system further comprises: A parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the parameter service mechanism is used to provide the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device with their respective required operating current values.
5. The visual data learning system for human-computer interaction according to any one of claims 2 to 4, characterized in that: An FPGA chip is used to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
6. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing maximum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
7. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing minimum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
8. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing median filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
9. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing edge sharpening processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.
Citation Information
Patent Citations
Machine interaction starting triggering method and system
CN109582139A
Anchor face delicate degree detection system
CN118823849A
Intelligent vehicle filtering processing parting system
CN118941875A
Zebra crossing passing scene shooting picture recognition system
CN118942055A
Irrigation field data detection system
CN119027877A