Human-machine interaction method, apparatus and device, and storage medium

By converting sensor signals into clues, objects and scenarios are constructed in the construction space, and the underlying common language is formed, the problem of lack of physical representation of human-computer interaction methods in the existing technology is solved, and the subjective interpretation and transparent interaction capabilities of the machine are realized.

WO2025176156A1PCT designated stage Publication Date: 2025-08-28CHENGDU YUANJI TONGZHI TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/078131
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-23
Filing Date
2025-02-19
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

The human-computer interaction mode in the prior art lacks underlying physical representation, resulting in poor adaptability of the interactive system environment, unable to perform complex two-way communication and fine motion control, and there are "illusion" and "black box" problems.

Method used

By converting sensor signals into clues, using the concept library to evoke relevant concepts, construct objects and scenarios in the construction space, form the underlying common language, and generate hardware control programs to achieve interaction.

Benefits of technology

It realizes the machine's subjective interpretation ability of input data, has the ability to observe and interpret, breaks the barriers of physical space and conceptual space, and provides transparent interactive understanding and highly flexible communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025078131_28082025_PF_FP_ABST
    Figure CN2025078131_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a human-machine interaction method, which comprises: by means of low-level intuitive processing, converting signals and interaction signals received by a sensor into clues; using the clues to evoke related concepts in a concept library, and storing the related concepts in a constructed space; by using the evoked concepts as templates and the related clues as materials, constructing objects or scenes with new perspectives and poses in the constructed space, and perceiving the presence and states of the objects; searching for new clues for filling and updating; determining the similarities and differences between the objects and the related concepts, and forming new concepts on the basis of the existing concepts so as to expand the concept library; executing multiple rounds of the steps, so as to observe and analyze the scenes and the objects more meticulously; generating objects and scenes in the constructed space, forming a low-level universal language, and interacting with a user by means of an interaction interface; and for a machine having a manipulation interface, a mobility interface or other types of hardware interfaces, after a universal language statement instruction of the user is received, generating an action sequence that can meet requirements, and generating a hardware control program for execution.
Need to check novelty before this filing date? Find Prior Art

Description

Human-computer interaction method, device, equipment and storage medium Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a human-computer interaction method, device, equipment and storage medium. Background Art

[0002] After evolving through the traditional command-line interface (CLI) and graphical interface (GUI) eras, human-computer interaction has entered the natural user interface (NUI) era, enabling people to interact with devices in more direct ways, such as through voice, gestures, body language, movement, and eye contact. New large-scale model prompt interaction methods allow users to directly tell the model what to do, and the large model then guides the system to perform specific tasks. While these new interaction methods have greatly alleviated the limitations of human-computer interaction, they still suffer from "weak interaction"—interaction systems require specialized design and training based on specific functions, resulting in poor environmental adaptability. Furthermore, they are limited to simple command interactions and lack the ability for complex two-way communication and fine motor control. The widespread "illusions" and "black box" problems of large models have also created significant challenges in practical applications. If human-computer interaction is considered an interaction language, then a language lacking underlying physical representation is like a sentence without context; no interpretation is correct. Therefore, intelligent human-computer interaction systems that use an underlying interaction language and can accurately convey mutual intentions are crucial. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a human-computer interaction method, device, equipment and storage medium, which solve the technical problems existing in the prior art.

[0004] In order to solve the above technical problems, the present invention is implemented in the following ways:

[0005] A human-computer interaction method, comprising:

[0006] 1) The signals received by the sensor and the interaction signals are converted into clues through intuitive processing at the bottom layer;

[0007] 2) Use clues to recall relevant concepts in the concept library and store them in the construction space;

[0008] 3) Using the evoked concepts as templates and relevant clues as materials, construct objects or scenes from new perspectives and postures in the constructed space, and perceive the existence and status of the objects;

[0009] 4) Find new clues to fill, update, support and enrich the constructed objects;

[0010] 5) Determine the similarities and differences between the object and related concepts in the concept library, form new concepts based on existing concepts, and expand the concept library;

[0011] 6) Repeat steps 1) to 5) multiple times to observe and analyze the scene and objects in more detail;

[0012] 7) Generate objects and scenes within the construction space, form the underlying common language and interact with users through the interactive interface;

[0013] 8) For machines with grasping, moving or other hardware interfaces, after receiving the user's general language statement instructions, it generates an action sequence that meets the requirements and generates a hardware control program for execution.

[0014] As a preferred embodiment of the above method, step 1) specifically includes:

[0015] Step 11) Acquire, sample, and preprocess signals. These signals include visual, auditory, tactile, and olfactory signals received by sensors, as well as user interaction signals. Preprocessing includes various common signal preprocessing methods, such as image resolution change, signal enhancement, denoising, and filtering.

[0016] Step 12) Acquiring preset attention parameters, including attention points, attention feeling type and value range, minimum attention granularity and other calculation parameters related to attention in the intuitive process;

[0017] Step 13) Based on the preset attention parameters, a region segmentation algorithm is used to select the plane or spatial region of interest in the signal digital space;

[0018] Step 14) The selected area is structured and converted into a set of meta-points composed of meta-points as basic elements, and stored in the construction space as a clue.

[0019] Furthermore, in step 14), the meta-points represent objects or object components perceived from the data in an abstract form, and the meta-point set is represented as follows:

[0020]

[0021] Among them, P i Represents the location attribute, R i Indicates the range attribute, S i Indicates sensory attributes, E i Indicates extended attributes, A i Indicates the target area pointed by the origin, O i Indicates other attributes;

[0022] The said element points are simplified by using skeleton line extraction or detection and extraction methods of various key points and key lines in image processing such as corner points, centroids, edges, etc., or by using various methods such as point cloud estimation trained by neural network and Gaussian distribution estimation, and the extraction results are converted into structured data;

[0023] The set of element points is converted into a corresponding clue pattern and stored in the construction space as a visual subsidiary image of the corresponding clue.

[0024] As a preferred embodiment of the above method, step 3) specifically includes:

[0025] After the concept is evoked, the meta-points in the concept are fully or partially filled with the corresponding clue meta-points according to the matching relationship between the concept and the clues, and the attributes of each meta-point in the concept are updated. In this way, an object or scene is constructed in a new perspective and posture in the space-time coordinate system and stored in the construction space.

[0026] As a preferred embodiment of the above method, the method for obtaining new clues in step 4) is:

[0027] Among the clues found in step 1), attention is paid to clues that were not used in the arousal process or to the unused elements in the clues, or new clues are found again by changing the attention parameters in the original input signal.

[0028] As a preferred embodiment of the above method, step 5) specifically includes:

[0029] Step 51) In the case of the object obtained in step 4), determine the similarities and differences between it and the relevant concepts in the concept library, and update the object and the corresponding clues to form a new concept based on the existing concept in a way that simplifies and highlights the differences, and stores it in the concept library as a high-level concept category;

[0030] Step 52) Users can directly proofread and modify concepts in the concept library, or view newly generated concepts in the construction space and modify them through human-computer interaction in subsequent steps; users can also add symbolic names, typical feature descriptions, and derivative relationships with existing concepts to the machine-generated concept data structure.

[0031] As a preferred embodiment of the above method, step 6) specifically includes:

[0032] Based on the original sensor input signals and the user's interactive commands, or the clues, concepts, and objects obtained in the constructed space, steps 1) to 5) are repeated multiple times as the attention parameters change, continuously updating objects and scenes, carefully observing the basis of existing objects and scenes, and generating new objects and new scenes with the newly evoked concepts after evoking more similar concepts.

[0033] As a preferred embodiment of the above method, step 7) specifically includes:

[0034] According to the graphical form of the objects and scene contents generated in the construction space in the previous steps, all or selected parts of the content are directly output as sentences in the common language and output to the user through the interactive interface; and the user also sends his or her own needs to the machine in the form of sentences in the common language, and the machine understands them after generating the corresponding objects and scenes in the construction space through the previous steps.

[0035] As a preferred embodiment of the above method, step 8) specifically includes:

[0036] Step 81) For machines with grasping, moving, or other hardware interfaces, after receiving instructions sent by the user in a universal language, they pre-generate action sequences that meet the requirements within the constructed scene and generate hardware control programs to operate;

[0037] Step 82) As the operation of the hardware device changes the state of the target object and the surrounding environment, the machine needs to re-perceive or in real time and compare the environmental changes with the predicted scenario results; if the task is not completed, the machine either autonomously generates a new action sequence and performs it, or repeats it again in the new environment under user supervision in a common language form.

[0038] In a second aspect, the present invention provides a human-computer interaction device, comprising:

[0039] The input module is used to obtain various sensor signals and the interaction information distributed by the interaction module;

[0040] The universal language model module is used to convert signals and interaction information from the physical or data space to the conceptual space, generating abstract concept interaction language sentences embedded with the underlying physical representation. It is the core module for achieving harmonious interaction between humans and machines at the underlying level.

[0041] An interaction module, used to receive and distribute bidirectional interaction information from users and machines;

[0042] The control module is used to convert the action sequence information generated by user instructions and target requirements into a hardware control program to operate and control the mechanical hardware of the machine.

[0043] In a third aspect, the present invention provides a human-computer interaction device, comprising: at least one processor and a memory,

[0044] The memory is used to store computer-executable instruction programs;

[0045] The processor is used to execute the computer-executable instruction program stored in the memory, so that the processor performs each step of a human-computer interaction method.

[0046] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer-executable instruction program. When a processor executes the computer-executable instruction program, the steps of the human-computer interaction method are implemented.

[0047] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the human-computer interaction method as described in the first aspect and various possible designs of the first aspect.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] a) The human-computer interaction method of the present invention is based on intuitive and conceptual understanding of objects. The objects and scene content generated in the construction space are simulated in the machine with intuitive and abstract capabilities, opening up a channel from physical or data space to conceptual space. The machine can not only perform passive statistical calculations on the input data, but also has the subjective motivation to observe and interpret, providing the feasibility of machine construction of world models and human-computer conceptual interaction.

[0050] b) Within a machine agent, its thought processes and results are completely transparent, and it possesses the ability to interact, understand, and communicate with humans as a low-level universal language. This universal language, embedded in the abstract conceptual interaction language of the underlying physical representation, exists because of data homogeneity formalized using primitives and graphs, and data identity to a certain degree of abstraction. The universal language model breaks down the barriers between physical and conceptual space, offering maximum expressive flexibility, precise and contextualized intent expression, transparency, security, and complete interpretability.

[0051] c) The system is small in scale, low in cost, and has strong learning and environmental adaptability, and can be applied to various artificial intelligence scenarios and various machine terminal systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0053] FIG1 is a schematic diagram of the system composition of a human-computer interaction method of the present invention;

[0054] FIG2 is a schematic flow chart of a human-computer interaction method according to the present invention;

[0055] FIG3 is a schematic diagram of an input image according to Embodiment 1 of the present invention;

[0056] FIG4 is a schematic diagram of various elements extracted from an input signal according to embodiment 1 of the present invention;

[0057] FIG5 is a schematic diagram of a clue pattern of a one-element point set conversion according to an embodiment of the present invention;

[0058] FIG6 is a schematic diagram of a triangle point set and a conceptual diagram in a concept library according to embodiment 1 of the present invention;

[0059] FIG7 is a schematic diagram of a circular element point set and a conceptual diagram in a concept library according to embodiment 1 of the present invention;

[0060] FIG8 is a schematic diagram of multiple preliminary objects constructed in Example 1 of the present invention;

[0061] FIG9 is a schematic diagram of details of finding new clues and enriching objects according to Example 1 of the present invention;

[0062] FIG10 is a schematic diagram of an input picture according to Embodiment 2 of the present invention;

[0063] FIG11 is a schematic diagram of a set of element points and clue patterns in Example 2 of the present invention;

[0064] FIG12 is a schematic diagram of a test picture of Example 2 of the present invention;

[0065] FIG13 is a schematic diagram of a set of element points and clue patterns of a test image according to Example 2 of the present invention;

[0066] FIG14 is a schematic diagram of a process of perceiving an object in a test picture by concepts and clues according to Embodiment 2 of the present invention;

[0067] FIG15 is a schematic diagram of a stick figure drawing according to embodiment 2 of the present invention;

[0068] FIG16 is a schematic diagram of a set of primitive points and a pattern corresponding to a stick figure according to Example 2 of the present invention;

[0069] FIG17 is a schematic diagram of an image captured by a camera in accordance with Embodiment 3 of the present invention;

[0070] FIG18 is a schematic diagram of a three-dimensional scene generated in the construction space in Example 3 of the present invention;

[0071] FIG19 is a schematic diagram of a scene converted from a three-dimensional scene to a two-dimensional scene in accordance with Embodiment 3 of the present invention;

[0072] FIG20 is a schematic diagram of a command sent by a user in a common language according to Example 3 of the present invention;

[0073] FIG21 is a schematic diagram showing the prediction of the execution result of a user command by a robot according to Example 3 of the present invention;

[0074] FIG22 is a schematic structural diagram of a human-computer interaction device provided by the present invention;

[0075] FIG23 is a schematic structural diagram of a human-computer interaction device provided by the present invention.

[0076] The above drawings illustrate specific embodiments of the present invention, which will be described in more detail below. These drawings and the accompanying description are not intended to limit the scope of the present invention in any way, but rather to illustrate the concept of the present invention to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0077] Here, exemplary embodiments will be described in detail, and the embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are only examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0078] The terms "first," "second," "third," etc. (if any), or "step 1," "step 2," "step 3," etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are not necessarily used to describe a specific order or sequential sequence. It should be understood that the terms used in this way are interchangeable under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0079] Example 1

[0080] The present invention is described in detail below in terms of object perception, scene understanding, and human-computer interaction with reference to the accompanying drawings and specific embodiment 1.

[0081] As shown in Figures 1 and 2, a human-computer interaction method includes the following specific steps:

[0082] 1) The signals received by the sensor and the interaction signals are converted into clues through intuitive processing at the bottom layer;

[0083] 2) Use clues to recall relevant concepts in the concept library and store them in the construction space;

[0084] 3) Using the evoked concepts as templates and relevant clues as materials, construct objects or scenes from new perspectives and postures in the constructed space, and perceive the existence and status of the objects;

[0085] 4) Find new clues to fill, update, support and enrich the constructed objects;

[0086] 5) Determine the similarities and differences between the object and related concepts in the concept library, form new concepts based on existing concepts, and expand the concept library;

[0087] 6) Repeat steps 1) to 5) multiple times to observe and analyze the scene and objects in more detail;

[0088] 7) Generate objects and scenes within the construction space, form the underlying common language and interact with users through the interactive interface;

[0089] 8) For machines with grasping, moving or other hardware interfaces, after receiving the user's general language statement instructions, it generates an action sequence that meets the requirements and generates a hardware control program for execution.

[0090] In the present invention, “constructing” objects and scenes means constructing or generating abstract objects and scenes in a conceptual space, and using these abstract objects and scenes to understand real things in the external world.

[0091] The step 1) specifically includes:

[0092] Step 11) Acquire signals and perform sampling and preprocessing. The signals include various modal signals such as visual, auditory, tactile, and olfactory signals received by the sensor, as well as received user interaction signals. Sampling includes various common digital signal sampling methods to convert the signals from the physical space to a suitable digital space. Preprocessing includes various common signal preprocessing methods such as changing image resolution, signal enhancement, denoising, and filtering. In this embodiment, the input image is shown in Figure 3, with a pixel size of 279*260.

[0093] Step 12) Obtain preset attention parameters and call parameters that are adjusted in real time as attention changes. The preset attention parameters include attention points, attention sensation type and value range, minimum attention granularity, and other calculation parameters related to attention in the intuitive process. In this embodiment, the attention point covers the entire image, the sensation parameter is the grayscale value, and the minimum attention granularity is a circle with a radius of 4.

[0094] Step 13) Based on the preset attention parameters, a region segmentation algorithm is used to select the plane or spatial region of interest in the signal digital space. This region segmentation algorithm includes various traditional image segmentation methods and image segmentation methods using deep learning. The segmented regions are represented by marking the regions with binary masks or extracting edges in the form of closed curves or surfaces. The plane or spatial region includes the region range and corresponding sensory parameters. The sensory parameters are the sensory type and value range based on which the region is divided. In this embodiment, binary quantization is used to mark each black region with a binary mask as the plane region of interest.

[0095] Step 14) The selected area is structured and converted into a set of meta-points composed of meta-points as basic elements. This is stored as a clue in the construction space. Meta-points represent objects or object parts perceived from the data in an abstract form. The meta-point set is represented as follows:

[0096]

[0097] Among them, P i Represents the position attribute of the element point in the space-time coordinate system, R i Indicates the extended range attribute of the element point in the space-time coordinate system, S i Indicates the sensory type and value range of the element point, sensory attribute, E i It represents extended attributes including connection attributes and dynamic attributes. Connection attributes are static attributes that indicate the connection relationship between element points, and obtain the abstract perception of the topological relationship between the various components of the object; dynamic attributes describe the extension trend of element points in space in the form of straight lines and curves and their change trend in time, and obtain the abstract characteristic perception of the straight lines and curves of the object in the space-time coordinate system; A i Indicates the target area pointed by the origin, O i Indicates other attributes;

[0098] The said element points are simplified by using skeleton line extraction or detection and extraction methods of various key points and key lines in image processing such as corner points, centroids, edges, etc., or by using various methods such as point cloud estimation and Gaussian distribution estimation trained by neural network, and the extraction results are converted into structured data; the said element point set is converted into the corresponding clue graph, which is stored in the construction space as a visual auxiliary image of the corresponding clue. The graph is a graph converted from the element point set (which can be expressed as ), which can explicitly depict various extended attributes within a set of nodes. For example, topological relationships represented by connection attributes are represented as nodes and edges, and spatial and temporal extensions and changing trends represented by dynamic attributes are represented as lines and curves. Graphics are similar to commonly used abstract representations such as stick figures and schematic diagrams, making them highly understandable and user-friendly.

[0099] In this embodiment, the skeleton line algorithm for obtaining the maximum inscribed circle and the corner point detection method are combined to extract the meta-points. According to the preset attention parameters, the minimum inscribed circle radius is set to 4. The obtained meta-points are shown in Figure 4, which include two types of meta-points: static meta-points (meta-points whose dynamic attributes are empty) are marked with "×", and dynamic meta-points (meta-points whose dynamic attributes are not empty) are marked with solid circles.

[0100] As an example of a point, the following table shows the data structure of the three points in Figure 4, listing the position, range, connection, and dynamic attribute values. The position attribute value is the coordinates of the point, the range attribute value is the radius of the point, and the connection attribute value is the other connected points. The dynamic attribute value has various forms. Using the method of simulating the movement of the point, the extension trend of the point is expressed as , where the deviation has three values: -1 means deflecting to the left along the direction of travel, 0 means going straight, and 1 means deflecting to the right.

[0101]

[0102] The clue graph G(V) of the transformation of the element point set is shown in Figure 5, which shows the machine's initial perception of the input signal after the underlying intuitive processing, and perceives various straight lines, arcs, concave objects, etc.

[0103] Said step 2) specifically includes, after obtaining the clues, evoking the relevant concepts in the concept library and storing them in the construction space. In this embodiment, the concept library stores some basic geometric figure concepts obtained by presetting or learning. According to the various clues (various element point sets) found in step 1), evoking is to find and call out the relevant concepts from the concept library, including the straight line, angle, triangle, arc, circle and other concepts in the concept library. As shown in Figure 4, the element point D has two straight line extension directions, and the element point D' and the element point D" also have two straight line extension directions respectively, and the extension directions of these three elements are exactly opposite to each other, which is consistent with the triangle concept in the concept library. According to the principle of feature matching, the concept of "triangle" as shown in Figure 6 is evoked. As shown in Figure 4, the connected element points A and C have the same curvature radius in their dynamic properties, which can evoke the concept of "circle" as shown in Figure 7.

[0104] Each plane geometry concept in the concept library, such as a triangle or circle, can be subdivided into two categories: solid or hollow. These concepts are then selected based on the clues used during subsequent object construction. While there are many similarities between the corresponding clues and the properties of the concept elements, there are also many inconsistencies. Therefore, common sense concepts such as missing components and occlusion are also evoked in the concept library and used in subsequent object and scene construction.

[0105] Step 3) specifically includes: after evoking a concept, based on the matching relationship between the concept and the clue, fully or partially filling the concept's meta-points with the corresponding clue meta-points, and updating the location, scope, feeling, extension, and other attributes of each meta-point in the concept. In this way, an object or scene is constructed from a new perspective and attitude in the space-time coordinate system and stored in the construction space. Constructing a specific object through clues and concepts allows the machine to perceive the object's existence and state, and to recognize the object, its components, and its name.

[0106] Filling refers to establishing a matching relationship between the concept and the clue's meta-points in key points, lines, and diagrams during the evocation process, allowing the attributes of the meta-points in the clue to be used in place of the meta-point attributes in the concept. The construction of foreground and background or multiple objects can form a scene, which can be a static scene in a plane or space, or a dynamic continuous or sequentially sliced ​​scene in space and time. Construction works in a predictive manner, requiring the machine to construct and interpret the surrounding scene environment at the time, and to construct and predict the subsequent state of objects and scenes through relevant concepts. If subsequent developments deviate from the previously predicted results when interpreting and understanding scenes in a predictive manner, it indicates that the scene construction is not fully consistent with reality, requiring a re-examination of the clues and evoked concepts, and an update of the object and scene construction.

[0107] In this embodiment, during the evocation process, a correspondence is established between the concept's primitives and the clue's primitives. This allows the primitive attributes in the clue to be used to replace the primitive attributes in the concept. The primitive attributes in the triangle and circle concepts evoked in step 2) are then filled with the corresponding primitive attributes in the clue in step 1). This results in the construction of multiple new objects within the construction space, with new perspectives and postures, such as the two triangles and three circles shown in Figure 8 (the primitives are omitted in the figure, only the object outlines are drawn). These constructed objects can be stored as materials in the construction space. Because the triangle and circle concepts shown in Figures 6 and 7 also have textual names for some of their components in the concept library, the machine assigns these and their components textual names to the objects constructed in Figure 8.

[0108] Step 4) specifically includes: After the object is constructed, if there's a mismatch between clues and concept points, resulting in missing parts, or if the object requires more detail, the machine will search for new clues to fill in and update the object, thereby supporting and enriching the constructed object. In this embodiment, for the solid upright triangle object in Figure 8, the machine uses the method from step 1) to search for clues based on the triangle concept's classification of hollow and solid triangles. The machine calculates and finds a high probability of solid white areas. As shown in Figure 9, the large white area (marked by the striped area) is selected as the clue area, and the previously missed clue point set is calculated.

[0109] The step 6) specifically includes:

[0110] Based on raw sensor input signals and user interaction commands, or clues, concepts, and objects acquired within the constructed space, steps 1) through 5) are repeated multiple times as attention parameters change, continuously updating objects and scenes. The machine carefully observes the existing objects and scenes, and after evoking more similar concepts, uses these newly evoked concepts to generate new objects and scenes. Changes in attention parameters include shifts in attention points, focusing or generalization, sensory type, and attention granularity. These parameters can be changed autonomously by the machine during object and scene construction, or by user commands during interaction.

[0111] In this embodiment, objects and scenes are updated predictively within the constructed space based on the constructed objects and evoked common-sense concepts such as missing components and occlusions. For example, as shown in Figure 3, the real-world scene or original input signal may not exist within the white upright triangle. Instead, it is constructed within the conceptual space and the machine's mind. Although it does not actually exist, this phenomenon often occurs in human visual perception when viewed, and is generally referred to as an "optical illusion." In practice, based on attention, the concept library, and the material within the constructed space, the machine's predictions or interpretations of scenes and objects can be flexible and diverse. As long as the machine's understanding is consistent with its own cognition and it can interpret the input signal in its own way, any interpreted scene can be used as an alternative to achieve the user's specific purpose through interaction with the user in step 7).

[0112] The step 7) specifically includes:

[0113] Based on the previous steps, graphical representations of objects and scene content have been generated in the build space. These abstract graphics, similar to commonly used stick figures, diagrams, and flowcharts, are easily understood by users. Therefore, all or selected portions of these contents are directly output as statements in a universal language. Users can also send their needs to the machine using stick figures, diagrams, flowcharts, and the like. The machine, after generating the corresponding objects and scenes in the build space through the previous steps, can also understand them. The objects and scene content generated in the build space clearly demonstrate the machine's thought processes and results, essentially forming the machine's underlying language, serving as the universal language for machine thinking, program control, and user interaction. In this embodiment, the objects and scenes generated by the machine in the build space are transmitted to the interactive interface in the form of a multi-layered scene, such as Figures 8 or 9, to communicate the user's understanding of the scene of the input signal. If the user needs to continue communicating with the machine or issue instructions, they can send these abstract universal language forms, such as stick figures, sketches, or diagrams, to the machine's clue module through the interactive interface, allowing the machine to understand their intentions and respond and take action.

[0114] Example 2

[0115] The following describes the present invention in detail in terms of high-level concept learning, object recognition, and human-computer interaction, with reference to the accompanying drawings and Example 2. Example 2 is divided into two parts. First, the machine learns the concept of "human," and then the machine uses this learned concept of "human" to interact with the user.

[0116] Part 1: Machine Learning of the Concept of “Human”:

[0117] The step 1) specifically includes:

[0118] Step 11) Acquire the signal, perform sampling and preprocessing, input the image shown in Figure 10, and convert the image into a grayscale image.

[0119] Step 12) Obtain the preset attention parameters, with the minimum attention fine-grained radius being a circle of 7.

[0120] Step 13) Based on the preset attention parameters, a region segmentation algorithm is used to select the plane or spatial region of interest in the signal digital space. A certain grayscale threshold is used to mark the dark area with a binary mask as the plane area that needs attention.

[0121] Step 14) The selected area is structured and converted into a set of meta-points composed of meta-points as basic elements, and stored as clues in the construction space. The maximum inscribed circle skeleton line algorithm is used to extract the meta-points. The obtained meta-points and the converted clue patterns are shown in Figure 11. The meta-point set and the corresponding patterns are stored as clues in the construction space.

[0122] The step 2) specifically includes: after obtaining the clue, evoking the relevant concepts in the concept library and storing them in the construction space. At this stage, the machine concept library only stores low-level concepts and some basic geometric concepts, and has not yet stored relevant high-level concepts. Therefore, it is impossible to evoke valid concepts. Here, steps 3) and 4) are skipped, and the clue is directly treated as an object that the machine has never seen before.

[0123] The step 5) specifically includes:

[0124] Step 51) Once the object has been obtained through the preceding steps, the similarities and differences between the object and related concepts are determined. The object and its corresponding clues are simplified and their differences are highlighted. This is then updated or a new concept is formed based on the existing concept. This concept is then stored in the concept library as a high-level concept. In this embodiment, the object obtained in the preceding steps is stored in the concept library as a new high-level concept.

[0125] In step 52, users can directly review and modify concepts in the concept library or view newly generated concepts in the construction space. These can then be modified through human-computer interaction in the subsequent step 7). Users can also add symbolic names, typical feature descriptions, and derivative relationships with existing concepts to the machine-generated concept data structure. This additional information facilitates concept search and enhances auxiliary recall methods such as knowledge graphs. In this embodiment, users directly view newly generated concepts in the concept library as meta-points and graphs, and add textual identifiers to the concept attributes, such as the concept name "person," meta-point identifiers "head," "chest," "abdomen," "ear," "hand," and "foot," and connection attribute identifiers "arm" and "leg."

[0126] The step 6) specifically includes:

[0127] Based on a variety of materials, as attention parameters are changed, steps 1) through 5) are repeated multiple times, continuously updating the objects and scenes, thereby deepening the understanding and interpretation of the input signal or the surrounding real-world scene. In this embodiment, the attention parameters are changed, such as selecting smaller local areas and a finer granularity of attention to observe the original image, thereby continuously enriching the details of the object and the concept of "person"; shifting attention to the head area, thereby forming sub-concepts such as "hair," "eyes," "nose," and "mouth" within the concept.

[0128] In the second part, the machine uses the learned concept of "human" to interact with the user:

[0129] To test whether the machine has learned the concept of "person," we input a test image as shown in Figure 12. The various element points obtained in step 1) and the transformed clue patterns are shown in Figure 13.

[0130] In step 2), the concept of “person” as shown in FIG11 is evoked in the concept library based on the structural similarity between the clue and the meta-point set in the concept.

[0131] As shown in Figure 14, in step 3), based on the matching relationship between the concept and the clue's meta-points, the concept's meta-points are filled with the corresponding meta-points in the clue, thereby updating the position and other attributes of the corresponding meta-points in the concept. In this way, a "person" object is constructed in the space-time coordinate system with a new perspective and posture. The machine thus recognizes the object in Figure 12 as a "person," as well as its various parts, such as the head, chest, abdomen, arms, and legs.

[0132] In step 4), the two “ears” in the concept of “person” shown in Figure 11 do not have matching points in the generated object. Therefore, a finer focus granularity is used to specifically search for “ear” clues in the corresponding area of ​​the test image in Figure 12 and update the object.

[0133] In step 5), in the clues and constructed objects, there are two meta-points (shoulder positions) that have no corresponding points in the concept. Therefore, it is possible to update the existing "person" concept based on the concept library, update or expand the new "person" concept, and the user can add the part name "shoulder".

[0134] Step 7) specifically includes: the user can communicate with the machine using simple drawings, schematic diagrams, and other methods commonly used in daily life. For example, the stick figure shown in Figure 15 is used to communicate some information about "people" to the machine; as shown in Figure 16, the machine can easily understand the user's intention through clue extraction and concept comparison; the abstract method allows both humans and machines to clearly understand the referenced object and the other party's intention without relying on text language, clearly indicating the machine's thinking process and results, and serving as a common language for machine thinking, program control, and interaction with humans.

[0135] Example 3

[0136] The present invention is described in detail below in terms of three-dimensional scene understanding, motion control and human-computer interaction in conjunction with the accompanying drawings and specific embodiment 3.

[0137] In this embodiment, the robot is equipped with binocular cameras, a robotic arm, and wheels, enabling it to move around a room, observe the environment, and place objects. The binocular cameras capture signals from steps 1) to 6), as shown in Figure 17, capturing an image from one of the cameras. After clue search and concept recall, a three-dimensional scene graph (Figure 18) is generated in the three-dimensional constructed space. In this embodiment, the robot has only learned simple object concepts, generating room structures such as walls, ceilings, and floors, as well as objects within the room such as tables and spheres. Other objects, such as outlet panels, closets, fire safety signs on closets, and hand sanitizer, are not perceived by the robot. Figure 19 shows the scene graph converted from a three-dimensional scene to a two-dimensional form, as seen in Figure 17. The clue search process in step 1) utilizes a three-dimensional point extraction method, first extracting a point on a plane from two monocular images, and then using a binocular visual spatial position estimation algorithm to obtain the spatial coordinates of the point.

[0138] In step 7, the user instantly views the robot's scene construction results on the interactive interface and, based on this, sends instructions and commands to the robot in a common language. For example, if the user wants the robot to move a sphere from the center of a table to a corner, as shown in Figure 20, an arrow is added to Figure 19 and sent to the robot. The arrow instruction here is a concept already in the concept library, representing the movement or operation of an object in a space-time coordinate system, that is, the spatial and temporal process of the object moving from the arrow's starting point to the arrow's ending point. After receiving the command, the robot predictively generates the object's movement process and results in the construction space and provides feedback to the user on the interactive interface in the form of a dynamic diagram or a static diagram, as shown in Figure 21. The user can check whether the robot has accurately understood the command and provide additional commands if necessary.

[0139] In step 8, the robot has clearly understood the user's intent and therefore converts the action sequence generated within the build space into a hardware control program. It then uses its wheels to maneuver to the table and uses its robotic arm to move the sphere to the desired location. During this operation, the robot perceives changes in the object and environment in real time and compares them with the predicted scenario outcome. If the task fails, it re-formulates the action sequence and executes the operation, either unsupervised or with user supervision.

[0140] This embodiment demonstrates that, by using an abstract concept interaction language embedded in underlying physical representations, the machine can accurately understand user intent in new tasks in unknown environments and complete user instructions in a safe manner.

[0141] The embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referenced to each other.

[0142] As shown in Figure 22, the present invention also provides a human-computer interaction device, including: an input module 2201, a universal language model module 2202, an interaction module 2203, and a control module 2204. It should be noted that the division of these modules is only a logical division of functions, and physically these modules can be integrated or independent.

[0143] The input module is used to obtain various sensor signals and the interaction information distributed by the interaction module;

[0144] The universal language model module is used to convert signals from physical or data space to conceptual space and generate an abstract conceptual interaction language embedded in the underlying physical representation. It is the core module for achieving harmonious interaction between humans and machines.

[0145] An interaction module, for receiving and distributing interaction information;

[0146] The control module is used to convert the action sequence information generated by user instructions and target requirements into a hardware control program to operate and control the mechanical hardware of the machine.

[0147] More specifically, the input module 2201 is used to receive various sensor modal signals and received user-sent interaction signals; the general language model module 2202 functionally includes various possible processing of the input signal in steps 1) to 6) described in the above-mentioned method embodiments, wherein the general language model module is divided into three sub-modules: a clue sub-module, a concept sub-module, and an object / scene sub-module, and the general language model module integrates multiple steps; the interaction module 2203 functionally includes various possible processing in step 7) described in the above-mentioned method embodiments; the control module 2204 functionally includes various possible processing in step 8) described in the above-mentioned method embodiments.

[0148] As shown in Figure 23 , the present invention also provides a human-computer interaction device comprising a processor 2301 and a memory 2302. These components are interconnected via a bus and can be mounted on a common motherboard or in other ways as needed. Processor 2301 processes instructions executed within the human-computer interaction device, including instructions for storing graphical information in or on the memory for display on external input / output devices. In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple storage devices if desired. Figure 23 illustrates a single processor 2301 as an example.

[0149] Memory 2302, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs, and modules, such as the program instructions / modules corresponding to the methods of the human-computer interaction device in various embodiments of the present invention (for example, the universal language model module 2202 shown in FIG. 22 ). Processor 2301 executes the non-transitory software programs, instructions, and modules stored in memory 2302 to execute various functional applications and human-computer interaction methods, thereby implementing the methods of the human-computer interaction device in the aforementioned method embodiments.

[0150] The human-computer interaction device may also include: an input device 2303, an output device 2304 and a control device 2305; the processor 2301, the memory 2302, the input device 2303, the output device 2304 and the control device 2305 may be connected via a bus or other means, with bus connection being used as an example in FIG23 .

[0151] The input device 2303 can receive input digital or character information, audio, image, or video information, as well as generate signal input related to user settings and function control of the human-computer interaction device. Examples include a touch screen, keypad, mouse or multiple mouse buttons, trackball, joystick, camera, microphone, and other input devices. The output device 2304 can be an output device such as a display device of the human-computer interaction device. Such display devices may include, but are not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device may be a touch screen. The control device 2305 can use program instructions to implement motion control, object grasping, and other operations for various connected machine hardware.

[0152] The human-computer interaction device of the embodiment of the present invention can be used to execute the technical solutions in the above-mentioned method embodiments of the present invention. Its implementation principles and technical effects are similar and will not be repeated here.

[0153] The present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement any of the above-mentioned human-computer interaction methods.

[0154] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms. In addition, the functional units in the various embodiments of the present invention can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0155] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that the present disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and that various modifications and variations may be made without departing from the scope thereof.

Claims

1. A human-computer interaction method, characterized in that: include: 1) The signals received by the sensor and the interaction signals are converted into clues through intuitive processing at the bottom layer; 2) Use clues to recall relevant concepts in the concept library and store them in the construction space; 3) Using the evoked concepts as templates and relevant clues as materials, construct objects or scenes from new perspectives and postures in the constructed space, and perceive the existence and status of the objects; 4) Find new clues to fill, update, support and enrich the constructed objects; 5) Determine the similarities and differences between the object and related concepts in the concept library, form new concepts based on existing concepts, and expand the concept library; 6) Repeat steps 1) to 5) multiple times to observe and analyze the scene and objects in more detail; 7) Generate objects and scenes within the construction space, form the underlying common language and interact with users through the interactive interface; 8) For machines with grasping, moving or other hardware interfaces, after receiving the user's general language statement instructions, it generates an action sequence that meets the requirements and generates a hardware control program for execution.

2. A human-computer interaction method according to claim 1, characterized in that: The step 1) specifically includes: Step 11) Acquire, sample, and preprocess signals. These signals include visual, auditory, tactile, and olfactory modal signals received by sensors, as well as user interaction signals. Preprocessing includes changing image resolution, signal enhancement, denoising, filtering, and other common signal preprocessing methods. Step 12) Acquiring preset attention parameters, including attention point, attention feeling type and value range, and attention-related calculation parameters in the minimum attention fine-grained intuitive process; Step 13) Based on the preset attention parameters, a region segmentation algorithm is used to select the plane or spatial region of interest in the signal digital space; Step 14) The selected area is structured and converted into a set of meta-points composed of meta-points as basic elements, and stored in the construction space as a clue.

3. A human-computer interaction method according to claim 2, characterized in that: In step 14), the meta-points represent the objects or object components perceived from the data in an abstract form. The meta-point set is represented as follows: ; Among them, P i Represents the location attribute, R i Indicates the range attribute, S i Indicates sensory attributes, E i Indicates extended attributes, A i Indicates the target area pointed by the origin, O i Indicates other attributes; The said element points are simplified by using skeleton line extraction or detection and extraction methods of various key points and key lines in corner point, centroid and edge image processing, or by using various methods of point cloud estimation and Gaussian distribution estimation trained by neural network, and the extraction results are converted into structured data; The set of element points is converted into a corresponding clue pattern and stored in the construction space as a visual subsidiary image of the corresponding clue.

4. The human-computer interaction method according to claim 1, wherein: The step 3) specifically includes: After the concept is evoked, based on the matching relationship between the concept and the clue, the meta-points in the concept are fully or partially filled with the corresponding clue meta-points, and the attributes of each meta-point in the concept are updated. In this way, an object or scene is constructed in a new perspective and posture in the space-time coordinate system and stored in the construction space.

5. The human-computer interaction method according to claim 1, characterized in that: The step 5) specifically includes: Step 51) Based on the object obtained above, determine its similarities and differences with related concepts in the concept library, and simplify and highlight the differences between the object and its corresponding clues, thereby updating the existing concepts to form new concepts and storing them in the concept library as high-level concept categories; Step 52) Users can directly proofread and modify concepts in the concept library, or view newly generated concepts in the construction space and modify them through human-computer interaction in subsequent steps; users can also add symbolic names, typical feature descriptions, and derivative relationships with existing concepts to the machine-generated concept data structure.

6. The human-computer interaction method according to claim 1, characterized in that: The step 6) specifically includes: Based on the original sensor input signals and the user's interactive commands, or the clues, concepts, and objects obtained in the constructed space, steps 1) to 5) are repeated multiple times as the attention parameters change, continuously updating objects and scenes, carefully observing the basis of existing objects and scenes, and generating new objects and new scenes with the newly evoked concepts after evoking more similar concepts.

7. The human-computer interaction method according to claim 1, characterized in that: The step 8) specifically includes: Step 81) For machines with grasping, moving, or other hardware interfaces, after receiving instructions sent by the user in a universal language, they pre-generate action sequences that meet the requirements within the constructed scene and generate hardware control programs to operate; Step 82) As the operation of the hardware device changes the state of the target object and the surrounding environment, the machine needs to re-perceive or in real time and compare the environmental changes with the predicted scenario results; if the task is not completed, the machine either autonomously generates a new action sequence and performs it, or repeats it again in the new environment under user supervision in a common language form.

8. A human-computer interaction device, characterized in that: include: The input module is used to obtain various sensor signals and the interaction information distributed by the interaction module; The universal language model module is used to convert signals and interaction information from the physical or data space to the conceptual space, generating abstract concept interaction language sentences embedded with the underlying physical representation. It is the core module for achieving harmonious interaction between humans and machines at the underlying level. An interaction module, used to receive and distribute bidirectional interaction information from users and machines; The control module is used to convert the action sequence information generated by user instructions and target requirements into a hardware control program to operate and control the mechanical hardware of the machine.

9. A human-computer interaction device, characterized in that: include: at least one processor and memory, The memory is used to store computer-executable instruction programs; The processor is configured to execute the computer-executable instruction program stored in the memory, so that the processor performs the method steps of claims 1-7.

10. A computer-readable storage medium characterized by: The computer-readable storage medium stores a computer-executable instruction program. When a processor executes the computer-executable instruction program, the method steps of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Concept learning-based thorough perception and dynamic understanding method

    CN110287941A

  • Dynamic target recognition and scene memory cognition method and system based on visual perception

    CN115082717A

  • Human-model interactive interpretation guiding method based on visual concept graph representation, electronic equipment and storage medium

    CN115797498A

  • Instruction grabbing method and system based on concept learning and priori knowledge

    CN116935025A

  • Human-computer interaction method, device and equipment and storage medium

    CN118226999A