Visual data learning system for human-computer interaction

By designing a visual data learning system for human-computer interaction, using dynamic visual data of user's eyebrows and mouth, the intelligent analysis of user's real semantics is achieved, and the problems of limited and inaccurate communication channels in the existing technology are solved, and the reliability and effectiveness of communication channels are improved.

CN119987560AInactive Publication Date: 2025-05-13NANJING YUSHUIDOU TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510132993.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing human-computer interaction system has limited and single communication channels, making it difficult to accurately analyze the patient's real semantic information.

Method used

Design a visual data learning system, and through targeted artificial intelligence models, use dynamic visual data of current user's eyebrows and mouth to achieve intelligent analysis of user's real semantics. The system includes a content grabbing mechanism, a segmented construction mechanism, an object analysis mechanism, a data capture device and a detection execution device. Through multiple training of radial-based neural networks, combined with preset frame rate and resolution, visual data is synchronized to the AI ​​detection model and output user's string data.

Benefits of technology

By intelligently analyzing the dynamic visual data of users' eyebrows and mouth, we create more and more reliable communication channels, which improves the effectiveness and reliability of the communication channels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987560A_ABST
    Figure CN119987560A_ABST
Patent Text Reader

Abstract

The invention relates to a visual data learning system for human-computer interaction, which comprises a data capturing device and a visual data learning device, the visual data output module is used for outputting the field depth value and the coordinate value of each pixel point of the eyebrow sub-image and the field depth value and the coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as visual data corresponding to the moment; and the detection execution device is used for intelligently analyzing the character string data expressed by the current user in the set time segment by adopting an AI detection model. The visual data learning system for human-computer interaction is compact in structure and stable in operation. Due to the fact that the artificial intelligence model designed in a targeted mode can be adopted, intelligent analysis of real semantics of the current user is achieved based on the dynamic visual data of the eyebrow portion and the dynamic visual data of the mouth portion of the current user in the set time segment, and therefore more communication channels are created for man-machine interaction. And the reliability and effectiveness of the alternating current channel are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction, and more specifically, to a visual data learning system for human-computer interaction. Background Art

[0002] Human-computer interaction involves multiple disciplines and aims to create an efficient interactive environment that conforms to human cognition. Its core elements include user-centered design, usability, interactive systems, interface design, user experience, etc. Future development directions include brain-computer interface, augmented reality, multimodal interaction, etc. Human-computer interaction is not only a technical tool, but also an intelligent partner, which requires interdisciplinary cooperation and ethical considerations. Computer human-computer interaction is a field that intersects multiple disciplines such as computer science, design, psychology and sociology. It focuses on how to make computer technology better serve human needs. This field not only focuses on technical optimization, but also attaches importance to the harmonious coexistence of technology and human activities. The goal is to create an interactive environment that is both efficient and in line with human cognitive habits.

[0003] Although the human-computer interaction system in the existing technology provides a variety of communication channels for patients with limited mobility and poor communication, such as voice communication channels, expression communication channels, etc., these communication channels are limited after all, and they are all single-channel communication channels, resulting in the communication information obtained may not be the semantic information truly expressed by the patient. Therefore, a richer and more accurate communication channel is needed to achieve reliable and accurate analysis of the patient's true semantics. Summary of the invention

[0004] In order to solve the technical problems in the prior art, the present invention provides a visual data learning system for human-computer interaction, which can adopt a targeted artificial intelligence model to realize intelligent analysis of the real semantics of the current user based on the dynamic visual data of the eyebrows and the dynamic visual data of the mouth of the current user in a set time segment, thereby creating more communication channels for human-computer interaction and improving the reliability and effectiveness of the communication channels.

[0005] According to the present invention, a visual data learning system for human-computer interaction is provided, the system comprising: The content capture mechanism is arranged opposite to the face of the current user and captures the real-time captured images corresponding to each consecutive moment on the time axis at a preset frame rate within a set time segment, wherein the each moment is evenly distributed on the time axis; A stepwise construction mechanism is used to perform multiple training actions on the radial basis neural network to obtain a radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism; An object analysis mechanism connected to the content capture mechanism, used to identify an eyebrow sub-image in a real-time captured image based on a grayscale value distribution interval corresponding to the eyebrow part of a human body, and also to identify a mouth sub-image in a real-time captured image based on an appearance imaging feature corresponding to the mouth part of a human body; A data capture device, connected to the step-by-step construction mechanism and the object analysis mechanism, respectively, for outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment; A detection execution device, connected to the sub-construction mechanism and the data capture device, respectively, for synchronously inputting the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI ​​detection model, so as to execute the AI ​​detection model and obtain the character string data expressed by the current user in the set time segment output by the AI ​​detection model; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the human eyebrow part, and the mouth sub-image in the real-time captured image is identified based on the appearance imaging feature corresponding to the human mouth, including: the grayscale value distribution interval corresponding to the human eyebrow part is a value interval limited by the grayscale upper limit value corresponding to the human eyebrow part and the grayscale lower limit value corresponding to the human eyebrow part; Among them, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the eyebrow part of the human body, and the mouth sub-image in the real-time captured image is identified based on the shape imaging features corresponding to the human mouth, which also includes: the shape imaging features corresponding to the human mouth are the respective reference mouth patterns corresponding to the human mouth at various opening and closing degrees.

[0006] Therefore, the present invention has at least the following three beneficial technical effects: Firstly, multiple training actions are performed on the radial basis neural network to obtain the radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism, thereby completing the targeted design of the AI ​​detection model for human-computer interaction; Secondly: adopting a content capture mechanism arranged opposite to the face of the current user, capturing real-time captured images corresponding to each continuous moment on the time axis at a preset frame rate within a set time segment, identifying the eyebrow sub-image in the real-time captured image based on the grayscale value distribution interval corresponding to the eyebrow part of the human body, identifying the mouth sub-image in the real-time captured image based on the appearance imaging features corresponding to the mouth of the human body, and outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment, thereby completing the customized screening of the visualization data at each moment for subsequent intelligent detection; Again: the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism and the various visual data corresponding to each moment in the set time segment are synchronously input into the AI ​​detection model to execute the AI ​​detection model, and obtain the character string data expressed by the current user in the set time segment output by the AI ​​detection model, thereby realizing intelligent semantic analysis based on the dynamic visual data of the eyebrows and the dynamic visual data of the mouth of the current user in the set time segment, and creating more communication channels for human-computer interaction.

[0007] The visual data learning system for human-computer interaction of the present invention has a compact structure and stable operation. Since a targeted artificial intelligence model can be used, the real semantics of the current user can be intelligently analyzed based on the dynamic visual data of the eyebrows and the dynamic visual data of the mouth of the current user in a set time segment, thereby creating more communication channels for human-computer interaction and improving the reliability and effectiveness of the communication channels. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Those skilled in the art may better understand the numerous advantages of the present invention by referring to the accompanying drawings, in which: Figure 1 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the first embodiment of the present invention.

[0009] Figure 2 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the second embodiment of the present invention.

[0010] Figure 3 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the third embodiment of the present invention. DETAILED DESCRIPTION

[0011] The implementation scheme of the visual data learning system for human-computer interaction of the present invention will be described in detail below with reference to the accompanying drawings.

[0012] Figure 1is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to a first embodiment of the present invention, the system comprising: The content capture mechanism is arranged opposite to the face of the current user and captures the real-time captured images corresponding to each consecutive moment on the time axis at a preset frame rate within a set time segment, wherein the each moment is evenly distributed on the time axis; For example, the content capture mechanism is arranged opposite to the face of the current user, and captures the real-time captured images corresponding to the continuous moments on the time axis at a preset frame rate within a set time segment, and the moments are evenly distributed on the time axis. The content capture mechanism includes: a built-in frame rate setting unit, a capture execution unit and a timing unit; A stepwise construction mechanism is used to perform multiple training actions on the radial basis neural network to obtain a radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism; An object analysis mechanism connected to the content capture mechanism, used to identify an eyebrow sub-image in a real-time captured image based on a grayscale value distribution interval corresponding to the eyebrow part of a human body, and also to identify a mouth sub-image in a real-time captured image based on an appearance imaging feature corresponding to the mouth part of a human body; A data capture device, connected to the step-by-step construction mechanism and the object analysis mechanism, respectively, for outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment; A detection execution device, connected to the sub-construction mechanism and the data capture device, respectively, for synchronously inputting the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI ​​detection model, so as to execute the AI ​​detection model and obtain the character string data expressed by the current user in the set time segment output by the AI ​​detection model; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the human eyebrow part, and the mouth sub-image in the real-time captured image is identified based on the appearance imaging feature corresponding to the human mouth, including: the grayscale value distribution interval corresponding to the human eyebrow part is a value interval limited by the grayscale upper limit value corresponding to the human eyebrow part and the grayscale lower limit value corresponding to the human eyebrow part; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the gray value distribution interval corresponding to the eyebrow part of the human body, and the mouth sub-image in the real-time captured image is identified based on the shape imaging feature corresponding to the human mouth, and further includes: the shape imaging feature corresponding to the human mouth is each reference mouth pattern corresponding to the human mouth at various opening and closing degrees; Among them, the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism and the various visual data corresponding to each moment in the set time segment are synchronously input into the AI ​​detection model to execute the AI ​​detection model, and the character string data expressed by the current user in the set time segment output by the AI ​​detection model is obtained, including: using programmable logic devices to realize simulation and emulation of the AI ​​detection model.

[0013] Figure 2 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the second embodiment of the present invention.

[0014] and Figure 1 different, Figure 2 The visual data learning system for human-computer interaction in may also include the following components: A field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the field control interface is used to respectively realize the synchronous driving control of each of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device.

[0015] Figure 3 It is a schematic diagram of the internal structure of a visual data learning system for human-computer interaction according to the third embodiment of the present invention.

[0016] and Figure 1 different, Figure 3 The visual data learning system for human-computer interaction in may also include the following components: A parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the parameter service mechanism is used to provide the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device with their respective required operating current values.

[0017] Next, the specific structure of the visual data learning system for human-computer interaction of the present invention will be further described.

[0018] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: An FPGA chip is used to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

[0019] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing maximum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

[0020] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing minimum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

[0021] In a visual data learning system for human-computer interaction according to any one of the embodiments of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing median filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

[0022] And in a visual data learning system for human-computer interaction according to any embodiment of the present invention: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing edge sharpening processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

[0023] In addition, in the visual data learning system for human-computer interaction, the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment are synchronously input into the AI ​​detection model to execute the AI ​​detection model, and obtaining the character string data expressed by the current user in the set time segment output by the AI ​​detection model includes: using a synchronous driving mechanism to synchronously input the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI ​​detection model; Many details of the present invention may be changed without departing from its spirit and scope. In addition, the description of the preferred embodiments of the present invention is provided for illustrative purposes only, not for limiting the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A visual data learning system for human-computer interaction, characterized in that: The system comprises: The content capture mechanism is arranged opposite to the face of the current user and captures the real-time captured images corresponding to each consecutive moment on the time axis at a preset frame rate within a set time segment, wherein the each moment is evenly distributed on the time axis; A stepwise construction mechanism is used to perform multiple training actions on the radial basis neural network to obtain a radial basis neural network after multiple training actions and output it as an AI detection model, wherein the number of training actions performed by the radial basis neural network is monotonically positively correlated with the value of the preset frame rate corresponding to the content capture mechanism; An object analysis mechanism connected to the content capture mechanism, used to identify an eyebrow sub-image in a real-time captured image based on a grayscale value distribution interval corresponding to the eyebrow part of a human body, and also to identify a mouth sub-image in a real-time captured image based on an appearance imaging feature corresponding to the mouth part of a human body; A data capture device, connected to the step-by-step construction mechanism and the object analysis mechanism, respectively, for outputting the depth of field value and coordinate value of each pixel point of the eyebrow sub-image and the depth of field value and coordinate value of each pixel point of the mouth sub-image in the real-time captured image corresponding to each moment as the visualization data corresponding to the moment; A detection execution device, connected to the sub-construction mechanism and the data capture device, respectively, for synchronously inputting the preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism, and the respective visual data corresponding to each moment in the set time segment into the AI ​​detection model, so as to execute the AI ​​detection model and obtain the character string data expressed by the current user in the set time segment output by the AI ​​detection model; Wherein, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the human eyebrow part, and the mouth sub-image in the real-time captured image is identified based on the appearance imaging feature corresponding to the human mouth, including: the grayscale value distribution interval corresponding to the human eyebrow part is a value interval limited by the grayscale upper limit value corresponding to the human eyebrow part and the grayscale lower limit value corresponding to the human eyebrow part; Among them, the eyebrow sub-image in the real-time captured image is identified based on the grayscale value distribution interval corresponding to the eyebrow part of the human body, and the mouth sub-image in the real-time captured image is identified based on the shape imaging features corresponding to the human mouth, which also includes: the shape imaging features corresponding to the human mouth are the respective reference mouth patterns corresponding to the human mouth at various opening and closing degrees.

2. The visual data learning system for human-computer interaction according to claim 1, characterized in that: The preset frame rate corresponding to the content capture mechanism, the resolution of the content capture mechanism and the respective visual data corresponding to each moment in the set time segment are synchronously input into the AI ​​detection model to execute the AI ​​detection model, and the character string data expressed by the current user in the set time segment output by the AI ​​detection model is obtained, including: using a programmable logic device to realize simulation and emulation of the AI ​​detection model.

3. The visual data learning system for human-computer interaction as claimed in claim 2, characterized in that: The system further comprises: A field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the field control interface is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the field control interface is used to respectively realize the synchronous driving control of each of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device.

4. The visual data learning system for human-computer interaction according to claim 2, characterized in that: The system further comprises: A parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively; Among them, the parameter service mechanism is arranged near the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device and is respectively connected to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device, including: the parameter service mechanism is used to provide the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device with their respective required operating current values.

5. The visual data learning system for human-computer interaction according to any one of claims 2 to 4, characterized in that: An FPGA chip is used to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

6. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing maximum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

7. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing minimum value filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

8. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing median filtering processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

9. The visual data learning system for human-computer interaction according to claim 5, characterized in that: Using an FPGA chip to perform image data processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively includes: performing edge sharpening processing on the output data of the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device to obtain the output processing data corresponding to the object analysis mechanism, the detection execution device, the step-by-step construction mechanism and the data capture device respectively.

Citation Information

Patent Citations

  • Machine interaction starting triggering method and system

    CN109582139A

  • Anchor face delicate degree detection system

    CN118823849A

  • Intelligent vehicle filtering processing parting system

    CN118941875A

  • Zebra crossing passing scene shooting picture recognition system

    CN118942055A

  • Irrigation field data detection system

    CN119027877A