Digital human interaction control method, system and equipment based on screen recognition
Through the digital human interaction control method based on screen recognition, the knowledge graph and function search tree of the application system are built, which solves the problem that digital humans in the existing technology are difficult to perform complex system operations, and realizes efficient interaction and control of the smart park management platform.
Patent Information
- Application Number
- CN202510550543.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The digital human technology in the existing smart park management system is difficult to perform complex system operations, and the smart assistant has poor adaptability to different application systems, which limits its application in smart park scenarios.
Using a digital human interaction control method based on screen recognition, a knowledge graph is constructed by analyzing the information of the application system, user input data is obtained for intent classification, screen interface images are captured for semantic segmentation, key-value dictionary is generated, and function search trees are constructed, and operation instruction sequences are generated for anthropomorphic operations.
It improves the flexibility and applicability of digital human interaction control, enables the system to perform complex system operations and adapts to different application systems, and realizes efficient interaction and control of the smart park management platform.
Smart Images

Figure CN120066282A_ABST
Abstract
Description
Background Art
[0002] In the prior art, the intelligent park management system usually adopts digital human technology to realize voice interaction between users and anthropomorphic images. However, the existing digital human technology is mainly based on speech recognition and natural language processing, enabling users to obtain park data through voice commands, such as the number of visitors, vehicle access records, etc. However, this voice interaction-based method is limited to data query and cannot directly execute system operations, such as modifying the status of park equipment, adjusting access control permissions, or managing the parking system. Since such tasks usually involve complex interface operations and multi-step processes, the traditional voice-driven method is difficult to meet higher-level control requirements.
[0003] In addition, existing intelligent assistants usually rely on interfaces pre-opened by software. Due to the large differences in interface rules among various application systems, the adaptability of intelligent assistants to different applications is poor, and it is difficult to achieve unified automated operations. The above problems limit the wider promotion and application of digital human technology in the intelligent park scenario. Therefore, there is an urgent need to provide a more flexible and general digital human interaction control method to achieve more efficient interaction and control of the intelligent park management platform.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information related to technologies that are not known to those of ordinary skill in the art. Summary of the Invention
[0005] An object of the embodiments of the present disclosure is to provide a digital human interaction control method based on screen recognition, a digital human interaction control system based on screen recognition, an electronic device, and a computer-readable storage medium, thereby at least to some extent improving the flexibility and applicability of the digital human interaction control method.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0007] According to the first aspect of the embodiments of the present disclosure, there is provided a digital human interaction control method based on screen recognition, the method including: based on the information parsing of the application system, constructing an application system knowledge graph; obtaining user input data, and classifying the intent of the user input data based on the application system knowledge graph; in response to the result of the intent classification being a system operation type intent, capturing the interface image of the current screen, and performing semantic segmentation on the interface image to generate a key-value dictionary including the function description and coordinate position of the interface component, and constructing an application system function search tree based on the key-value dictionary; based on the application system knowledge graph and the application system function search tree, generating an operation instruction sequence from the current page to the target function page, and performing anthropomorphic operations according to the operation instruction sequence.
[0008] In some embodiments of the present disclosure, based on the foregoing solution, the digital human interaction control method based on screen recognition further includes: in response to the result of the intent classification being an inquiry and answer type intent, extracting keywords from the user input data, and generating a database query instruction based on the keywords; obtaining query result data based on the query instruction, and using a multimodal large model to parse and generate natural language for the query result data to generate corresponding reply content.
[0009] In some embodiments of the present disclosure, based on the foregoing solution, constructing an application system knowledge graph based on the information parsing of the application system includes: extracting buttons, form fields and their hierarchical relationships from the function description documents, user manuals, and interface elements of the application system, and performing semantic annotation on the UI icons, texts, and coordinate information of the buttons and form fields; using a pre-trained model to perform entity recognition on the semantically annotated text data, and extracting the verb-object structure relationship of the text data through dependency syntactic analysis; based on the result of the verb-object structure relationship extraction, constructing the application system knowledge graph including entity types, relationship types, and attribute constraints.
[0010] In some embodiments of the present disclosure, based on the foregoing solution, obtaining user input data and performing intent classification on the user input data based on the application system knowledge graph includes: collecting text data, image data, and voice data input by the user to generate multimodal input data; extracting keywords from the multimodal input data, and retrieving function, operation, and entity information associated with the keywords based on the application system knowledge graph; inputting the retrieval result and the multimodal input data into a pre-trained classification model to generate an intent classification result.
[0011] In some embodiments of the present disclosure, based on the foregoing solution, performing semantic segmentation on the interface image to generate a key-value dictionary including the function description and coordinate position of the interface components, and constructing an application system function search tree based on the key-value dictionary includes: performing semantic segmentation on the interface image to determine the target regions corresponding to each interface component; extracting the text features and icon features of the target regions to generate a key-value dictionary of the function description and coordinate position of each interface component; constructing nodes of the application system function search tree based on the key-value dictionary, and gradually expanding the nodes according to the interaction operations to generate the application system function search tree.
[0012] In some embodiments of the present disclosure, based on the foregoing solution, extracting the text features and icon features of the target area to generate a key-value dictionary of the function description and coordinate position of each interface component includes: performing optical character recognition on the target area to obtain a text feature vector; using a convolutional neural network to extract image features of the icons in the target area to generate an icon feature vector; fusing the text feature vector and the icon feature vector to generate a fused feature vector; calculating the function category of the interface component based on the fused feature vector: , where represents the classifier weight matrix, represents the classifier bias term, represents the normalization function, represents the th function category of the interface component, represents the fused feature vector; combining the function category and the coordinate position of the target area to generate a key-value dictionary of the function description and coordinate position of the interface component.
[0013] In some embodiments of the present disclosure, based on the foregoing solution, constructing nodes of the application system function search tree based on the key-value dictionary and gradually expanding the nodes according to interaction operations to generate the application system function search tree includes: using the function category in the key-value dictionary as the function search tree node identifier and the coordinate position as the node attribute to construct function search tree nodes; after performing corresponding interaction operations on the function search tree nodes, updating the interface image and generating a new key-value dictionary, and expanding new function search tree nodes; connecting the newly generated function search tree nodes to the existing function search tree nodes and gradually expanding to construct a complete application system function search tree.
[0014] In some embodiments of the present disclosure, based on the foregoing solution, generating an operation instruction sequence from the current page to the target function page based on the application system knowledge graph and the application system function search tree includes: obtaining the keywords corresponding to the target function page according to the system operation class intention; retrieving based on the keywords in the application system knowledge graph to obtain the entities, relationships, and attributes associated with the keywords; inputting the keywords, the entities, relationships, and attributes, and the application system function search tree into a multi-modal large model to generate an operation path from the current page to the target function page; generating a corresponding operation instruction sequence according to the operation path.
[0015] According to the second aspect of the embodiments of the present disclosure, a digital human interaction control system based on screen recognition is provided. The system includes: a knowledge graph module for constructing an application system knowledge graph based on the information parsing of the application system; an intent classification module for obtaining user input data and classifying the intent of the user input data based on the application system knowledge graph; a screen capture module for capturing an interface image of the current screen and performing semantic segmentation on the interface image to generate a key-value dictionary containing the function description and coordinate position of the interface components in response to the result of the intent classification being a system operation type intent, and constructing an application system function search tree based on the key-value dictionary; and an interaction execution module for generating an operation instruction sequence from the current page to the target function page based on the application system knowledge graph and the application system function search tree, and performing anthropomorphic operations according to the operation instruction sequence.
[0016] According to the third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the above-mentioned digital human interaction control method based on screen recognition is implemented.
[0017] According to the fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned digital human interaction control method based on screen recognition is implemented.
[0018] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: In the digital human interaction control method based on screen recognition in the exemplary embodiments of the present disclosure, on the one hand, based on the information parsing of the application system, an application system knowledge graph is constructed, enabling the system to extract and organize the function information, operation logic, and interface structure of the application system. By parsing the information of the application system, a structured knowledge graph is formed, enabling the system to generalize the function levels in different application environments and providing a basis for subsequent interactions, thereby supporting the semantic understanding and task classification of user input data. On the other hand, user input data is obtained and the intent of the user input data is classified based on the application system knowledge graph, enabling the system to distinguish user interaction requirements and match corresponding processing methods to ensure the pertinence and accuracy of subsequent operations.
[0019] On the other hand, in response to the result of intent classification being a system operation type intent, capture the interface image of the current screen, perform semantic segmentation on the interface image, and generate a key-value dictionary containing the function descriptions and coordinate positions of interface components, enabling the system to parse interface information and identify operable components. Through interface image analysis and component segmentation, the system can directly extract interface elements from the screen content without additional data input, and establish an operation mapping relationship to provide visual recognition support for interaction execution. Further, based on the application system knowledge graph and the application system function search tree, generate an operation instruction sequence from the current page to the target function page, enabling the system to autonomously plan an interaction path and form a coherent operation process. Combining the knowledge graph and the search tree structure of interface components enables the system to establish functional associations between interfaces and determine the execution path, enabling interface operations to proceed hierarchically, thus ensuring the smooth execution of tasks.
[0020] Finally, perform anthropomorphic operations according to the operation instruction sequence, enabling the system to perform interface operations based on the identified interface information in accordance with the user's interaction method. By identifying interface components and generating instructions in combination with the operation path, the system can simulate the user's interaction methods such as clicking and inputting to complete automated control based on screen recognition.
[0021] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0023] Figure 1 Schematically shows a flowchart of a digital human interaction control method based on screen recognition according to some embodiments of the present disclosure.
[0024] Figure 2 Schematically shows a flowchart of another digital human interaction control method based on screen recognition according to some embodiments of the present disclosure.
[0025] Figure 3 Schematically shows a block diagram of a digital human interaction control system based on screen recognition according to some embodiments of the present disclosure.
[0026] Figure 4 Schematically shows a structural diagram of a computer system of an electronic device according to some embodiments of the present disclosure.
[0027] Figure 5 A schematic diagram schematically shows a computer-readable storage medium according to some embodiments of the present disclosure.
[0028] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners
[0029] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0030] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0031] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0032] Now, the exemplary embodiments will be described more fully with reference to the drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0033] In addition, the drawings are only schematic diagrams and are not necessarily drawn to scale. The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0034] In related technologies, the intelligent park management system usually relies on digital human technology to achieve information communication between users and the system through voice interaction, mainly for querying and broadcasting park situation data. However, this type of technology only supports data query based on preset instructions and cannot deeply control and operate the management system. For example, users can obtain park access records through voice commands, but cannot directly adjust the device status, such as setting the vehicle barrier to be normally closed, resulting in great limitations when digital humans perform actual management tasks.
[0035] At the same time, the related technologies mainly rely on voice and text interactions, and the interaction method with application software is usually based on API calls. However, since the API rules of different application systems are different, and some systems do not open standardized interfaces, it is difficult for AI assistants to adapt to all software, resulting in limited control capabilities, insufficient generalization capabilities and flexibility. In addition, AI assistants lack the ability to directly understand the content of device screens and cannot recognize and operate UI components when dealing with tasks involving interface interactions, so users still need to complete them manually when searching for information or performing operations in specific applications, which limits the scope of application of automated control. In addition, the related digital human technology only supports data query instructions, depends on background operation when executing command line operations, and lacks visual feedback, so users cannot supervise the execution process. At the same time, this method requires a high technical ability of users, limits the use of non-professional users, and further reduces the operability of digital human technology in complex management scenarios.
[0036] To solve all or part of the above technical problems in related technologies, Figure 1 A schematic flowchart of a digital human interaction control method based on screen recognition according to some embodiments of the present disclosure is schematically shown. Refer to Figure 1 As shown, the digital human interaction control method based on screen recognition may include the following steps: Step S110, constructing an application system knowledge graph based on information parsing of the application system; Step S120, obtaining user input data and classifying the intent of the user input data based on the application system knowledge graph; Step S130, in response to the result of intent classification being a system operation type intent, capturing the interface image of the current screen, performing semantic segmentation on the interface image to generate a key-value dictionary containing the function description and coordinate position of the interface component, and constructing an application system function search tree based on the key-value dictionary; Step S140, generating an operation instruction sequence from the current page to the target function page based on the application system knowledge graph and the application system function search tree, and performing anthropomorphic operations according to the operation instruction sequence.
[0037] In the actual operation process, the digital human interaction control method based on screen recognition in this embodiment first performs information parsing on the application system, extracts the system functions, interface structures, and interaction logics, and organizes them into structured data to construct the knowledge graph of the application system. During this process, the system parses the hierarchical relationships, functional attributes, and associated operations of the interface elements to provide a basis for subsequent interactions. After the user initiates an interaction request, the user input data is obtained, and the application system knowledge graph is used to analyze the input data to identify the user's intention. Through semantic parsing and knowledge matching of the input data, the user's intention is classified into different task types to determine the subsequent processing method. When the user's intention belongs to the system operation type task, the screen image is captured according to the current interface state, and semantic segmentation is performed on it to extract the functional information and position coordinates of the interface components, generating a key-value dictionary of component function descriptions and coordinates. Based on this key-value dictionary, a functional search tree of the application system is established, enabling the interface components to form a retrievable organization method in the hierarchical structure to support subsequent function positioning and operation path planning.
[0038] Based on the generated functional search tree of the application system, combined with the application system knowledge graph, the target function page corresponding to the user's intention is inferred, and based on the association relationship between the current interface state and the target function, an operation instruction sequence from the current page to the target page is generated. By sequentially executing the generated operation instruction sequence, the anthropomorphic operation is performed by simulating the user's interaction method to complete the task execution. Through this process, the system can achieve automated interaction control of the application system based on the screen content, enabling the user to manipulate the system functions through natural input without relying on preset APIs or manual interface operations, thereby improving the intelligence level of the interaction.
[0039] Next, the digital human interaction control method based on screen recognition in the above exemplary embodiment will be further described.
[0040] In step S110, based on the information parsing of the application system, the knowledge graph of the application system is constructed.
[0041] Among them, the application system can represent a software platform that includes specific functions, operation logics, and interface elements. Information parsing can represent the data extraction and analysis of the functional description documents, user manuals, and interface elements of the application system to generate structured information. The knowledge graph of the application system can represent a data structure established based on the information parsing of the application system, which includes the functional relationships, component hierarchies, and interaction rules of the application system and can be used to support the functional understanding and operation reasoning of the application system.
[0042] In some embodiments, based on the information parsing of the application system, an application system knowledge graph is constructed, which specifically includes the following technical steps: extracting buttons, form fields and their hierarchical relationships from the function description documents, user manuals, and interface elements of the application system, and performing semantic annotation on the UI icons, texts, and coordinate information of the buttons and form fields; using a pre-trained model to perform entity recognition on the semantically annotated text data, and extracting verb-object structure relationships from the text data through dependency syntactic analysis; based on the results of the verb-object structure relationship extraction, constructing an application system knowledge graph including entity types, relationship types, and attribute constraints.
[0043] Among them, the function description document can represent the document data used to record the function structure, operation process, and related descriptions of the application system. The user manual can represent the document resources providing the usage methods, function descriptions, and interaction guidelines of the application system. The interface elements can represent the visual components in the application system interface, including buttons, form fields, UI icons, and text information, etc. Buttons can represent the interactive components used to trigger specific functions or execute specific instructions. Form fields can represent the interface input areas used to input, display, or modify data. The hierarchical relationship can represent the parent-child nesting relationship, superior-subordinate relationship, or the relationship of the operation process sequence, etc. among different user interface components in the structural organization. Entity recognition can represent detecting and classifying the vocabulary or phrases with specific semantic attributes from the text data. Dependency syntactic analysis can represent parsing the syntactic relationships between words in the text data based on the syntactic structure to identify semantic dependency relationships such as subject-predicate and verb-object. The verb-object structure relationship can represent the grammatical connection between the verb representing the action and its affected object in the sentence. Entity types can represent different categories of concepts in the knowledge graph, such as functions, operations, or interface components. Relationship types can represent the association methods between different entities in the knowledge graph, such as function calls, interface hierarchies, or data input-output relationships. Attribute constraints can represent the conditions or constraint rules used to limit the entity features in the application system knowledge graph. It shows the structured organization method between interface elements, reflecting the nesting relationship and interaction priority of components.
[0044] Specifically, when constructing the application system knowledge graph, it is first necessary to extract the core information from the function description documents, user manuals, and interface elements of the application system to form a structured knowledge model for describing functions, operations, and data. In the information extraction stage, the interface elements are parsed to extract buttons, form fields, and their hierarchical relationships. The UI icons, texts, and coordinate data in the interface are associated, and semantic annotation is combined with the function modules to clarify the functions and interaction logics of different interface elements. This process includes data cleaning, converting the original text and interface information into a machine-processable format, removing redundant information, and segmenting and clause-splitting through natural language processing methods to ensure the accuracy of subsequent analysis.
[0045] After the structured data is generated, a pre-trained model is used to perform entity recognition on the semantically annotated text data, extract key concepts, and match them to predefined entity categories. For interface text and instruction documents, pattern matching techniques are applied to identify specific terms, and statistical methods are combined to improve the recognition accuracy. For the verb and noun structures in the function description, dependency syntactic analysis is adopted to extract the verb-object structure, so as to obtain the relationship between the function operation and the affected object, and construct a preliminary function association network. After the verb-object structure relationship is extracted, based on the parsed entity and relationship information, the semantic framework of the knowledge base is defined, and a knowledge representation including entity types, relationship types, and attribute constraints is established. According to the extracted function modules, interface elements, and operation logic, an application system knowledge graph is constructed to support reasoning and query, providing a basis for subsequent user intention recognition and interaction control.
[0046] Exemplarily, when performing entity recognition on the semantically annotated text data, other suitable pre-trained models such as the Bidirectional Encoder Representations from Transformers (BERT) model and the Bidirectional Long Short-Term Memory with Conditional Random Fields (BiLSTM-CRF) model can be used.
[0047] In step S120, the user input data is obtained, and the user input data is classified according to intention based on the application system knowledge graph.
[0048] Among them, the user input data can represent the interaction information provided by the user in the forms of text, voice, or image, and this information is used to express query requirements or operation instructions. Intention classification can represent the analysis of the user input data to determine the target category it expresses, including various types of interaction requirements such as consultation and answering, system operation, or data query.
[0049] In some embodiments, obtaining the user input data and classifying the user input data according to intention based on the application system knowledge graph specifically includes the following technical steps: collecting the text data, image data, and voice data input by the user to generate multimodal input data; extracting the keywords in the multimodal input data, and retrieving the function, operation, and entity information associated with the keywords based on the application system knowledge graph; inputting the retrieval result and the multimodal input data into a pre-trained classification model to generate an intention classification result.
[0050] First, collect the text data, image data, and voice data input by the user to generate multi-modal input data. For the text data, perform text preprocessing, including operations such as word segmentation, stop word removal, and word vectorization; for the image data, use optical character recognition (OCR) to extract text information and use a convolutional neural network to extract visual features; for the voice data, use speech recognition to convert speech into text and extract audio features such as Mel-Spectrogram at the same time.
[0051] Then, extract keywords from the multi-modal input data and retrieve the function, operation, and entity information associated with the keywords based on the application system knowledge graph. Specifically, use named entity recognition and dependency syntactic analysis to extract key entities with semantic information from the text data, match the corresponding nodes in the knowledge graph, and obtain relevant function and operation information. At the same time, for the keyword extraction of image and voice data, combine the text transcription content to ensure the accuracy of knowledge retrieval.
[0052] Next, fuse the retrieval results with the multi-modal input data and input them into a pre-trained classification model to generate an intention classification result. During the fusion process, use a multi-modal fusion method based on the attention mechanism to perform weighted combination of the feature vectors of the text data, image data, and voice data to enhance the complementarity and information relevance of different modal information.
[0053] Specifically, define a multi-modal feature vector set: Among them, represents the text feature vector, represents the image feature vector, represents the voice feature vector, represents the feature vector of the function, operation, and entity information retrieved from the application system knowledge graph.
[0054] To enhance the interaction of different modal features, use the self-attention mechanism to calculate the weighted representation of each modality: Among them, represents the attention score of each modal feature, represents the attention weight matrix, represents the bias term, is used to perform normalization processing on different modal features.
[0055] Then, calculate the weighted feature after fusion: Among them, Represents an element-wise weighted operation, represents the fused multi-modal feature vector.
[0056] Subsequently, the fused feature vector is input into a pre-trained classification model for intent prediction, using a classification method based on a fully connected layer: Among them, represents the final intent classification result, and represent the hidden layer weights and biases, and represent the classification layer weights and biases, is the activation function, used for multi-class classification. Through the above method, not only the fusion of multi-modal information is achieved, but also the importance distribution of different modal information is enhanced by using the attention mechanism, improving the accuracy and robustness of the intent classification result.
[0057] In some embodiments of the present disclosure, when extracting features from the text data and speech data input by the user, the following technical steps can be adopted: First, preprocess the text data to ensure data consistency and parsability. The preprocessing of text data includes removing special symbols, punctuation normalization, and case normalization to reduce the impact of noise on subsequent feature extraction. Subsequently, use sub-word level or character level tokenization techniques to tokenize the text to improve the model's ability to handle out-of-vocabulary words while maintaining semantic integrity. The tokenized text data is mapped to a high-dimensional vector space, and the text feature vector is defined as follows: Among them, represents the text feature vector, represents the word vector of the th word, represents the number of words in the text, represents the dimension of the word vector. To enhance text semantic information, pre-trained word vectors or context dynamic word vectors are used for text representation.
[0058] Furthermore, use a Transformer-based encoding method to perform deep semantic modeling on the text data to capture global dependencies. First, calculate the self-attention representation of the text: Among them, represents the self-attention weighted matrix of the text features, respectively represent the query matrix, key matrix, and value matrix of the text, Denote the scaling factor to balance the numerical range. Then, perform a feed-forward neural network transformation on the text features to enhance the non-linear expression ability: Where, Denote the encoded text feature vector, And Denote the weight matrix and bias term respectively, Denote the non-linear activation function. To preserve the order information of the text, further introduce positional encoding so that the Transformer can perceive the word order information and improve the model's temporal modeling ability.
[0059] When processing the speech data input by the user, first perform endpoint detection to determine the start and end times of the speech segment and remove the silent part to ensure the validity of the input speech data. Subsequently, use a convolutional neural network or self-attention model to perform noise reduction processing on the speech data to remove background noise and improve the clarity of the speech signal. Further, to ensure the consistency of the speech data format, perform normalization processing on the speech data with different sampling rates.
[0060] Next, convert the processed speech data into a time-frequency domain representation, and use the short-time Fourier transform to transform the speech signal to extract the time-frequency information of the speech signal. To further enhance the time-frequency resolution of the speech features, use a Mel filter for conversion to obtain Mel spectrogram features: , where, Denote the Mel spectrogram features, Denote the Mel filter matrix, Denote the speech signal after STFT transformation.
[0061] Then, use an audio pre-trained model for feature extraction and use a convolutional front-end to extract local features to generate a preliminary embedding: Where, Is the extracted speech feature vector, Is the convolutional kernel, Denote the convolutional operation, Is the bias term.
[0062] Finally, use the Transformer to perform global modeling on the extracted speech features and perform feature transformation through a feed-forward neural network: Where, Denote the finally extracted speech feature vector, Denote the weight matrix and bias term respectively, Denote the feature transformation activation function.
[0063] In step S130, in response to the result of intent classification being a system operation type intent, capture an interface image of the current screen, perform semantic segmentation on the interface image to generate a key-value dictionary containing the function descriptions and coordinate positions of interface components, and construct an application system function search tree based on the key-value dictionary.
[0064] Among them, the interface image can represent the visual content of the current display interface of the application system, including the overall presentation of interface components, text information, and other graphical elements. Semantic segmentation can represent pixel-level analysis of the interface image to distinguish different functional regions and identify information such as the category, text, and position of interface components therein. The function description of the interface component can represent the specific role of the interface component in the application system, including the interaction functions of components such as buttons, forms, and menus and their associated operations. The coordinate position can represent the spatial position information of the interface component in the interface image, including the center point coordinates, boundary range, or relative position relationship of the component. The key-value dictionary can represent a data structure composed of the function description of the interface component as the key value and the coordinate position as the corresponding value, used to store the function information of the interface component and its spatial coordinates. The application system function search tree can represent a tree-structured data model constructed based on the hierarchical relationship and interaction logic of interface components, where each node represents different functional components and their position relationships in the interface to support the retrieval of interface components and the inference of operation paths.
[0065] In some embodiments, performing semantic segmentation on the interface image to generate a key-value dictionary containing the function descriptions and coordinate positions of interface components and constructing an application system function search tree based on the key-value dictionary specifically includes the following technical steps: performing semantic segmentation on the interface image to determine the target regions corresponding to each interface component; extracting the text features and icon features of the target regions to generate a key-value dictionary of the function descriptions and coordinate positions of each interface component; constructing the nodes of the application system function search tree based on the key-value dictionary, and gradually expanding the nodes according to the interaction operations to generate the application system function search tree.
[0066] Specifically, when constructing the application system function search tree, first obtain the current screen image by calling the operating system interface, and input the obtained interface image into the visual recognition module. The visual recognition module performs semantic segmentation on the interface image to determine the target areas corresponding to each interface component, and segments the target areas through a region extraction algorithm so that each area contains only a single interface component or its related text description to ensure the accuracy of subsequent feature extraction. After the target areas are determined, text features and icon features are extracted from the target areas. For the target areas containing text, an optical character recognition method is used to recognize the text information and generate a text feature vector. For the target areas containing icons, a deep learning model is used to extract the icon feature vector to represent its function category and visual features. By combining the text features and icon features, a function description of the interface component is generated, and based on the position information of the target area, a key-value dictionary containing the function description and coordinate position is constructed to form a structured representation of the interface component.
[0067] Subsequently, based on the key-value dictionary, an initial node of the application system function search tree is constructed. Each node of the function search tree corresponds to a target area in the interface, and the node information includes the text recognition result, the icon feature vector, and the position coordinates. By modeling the hierarchical relationship of the interface components, the superior-subordinate relationship between each component is determined so that the function search tree can accurately reflect the interface structure. When performing an interaction operation, the system simulates click, input, or other interaction operations according to the coordinate information of the search tree node, and after the interface state is updated, semantic segmentation of the interface image and target area recognition are performed again to create new sub-nodes of the function search tree. During the interaction process, by gradually expanding the search tree nodes, the identified interface components and their hierarchical relationships in different interface states are continuously recorded so that the function search tree can cover the complete interface structure of the application system. After sufficient interaction operations and interface recognition, a complete representation of the application system function search tree is finally formed to provide support for subsequent interaction control.
[0068] In some embodiments, extracting the text features and icon features of the target area and generating a key-value dictionary of the function description and coordinate position of each interface component specifically includes the following technical steps: performing optical character recognition on the target area to obtain a text feature vector; using a convolutional neural network to extract image features of the icons in the target area to generate an icon feature vector; fusing the text feature vector and the icon feature vector to generate a fused feature vector; calculating the function category of the interface component based on the fused feature vector: , where represents the classifier weight matrix, represents the classifier bias term, represents the normalization function, represents the th function category of the interface component. Represents a fused feature vector; combines the functional category with the target area coordinate position to generate a key-value dictionary of the functional description and coordinate position of the interface component.
[0069] Among them, by using optical character recognition and convolutional neural network to extract the text and icon features of the interface component and perform feature fusion, the integrity and accuracy of component recognition can be improved. The Softmax normalization method is used to calculate the functional category to achieve automatic classification of the components. Combining the classification result with the coordinate information to generate a key-value dictionary associates the functional description of the interface component with the spatial position information, providing structured support for subsequent interface interaction and automated operations.
[0070] In some embodiments, nodes of the application system function search tree are constructed based on the key-value dictionary, and the nodes are gradually expanded according to the interaction operations to generate the application system function search tree. The specific technical steps are as follows: use the functional category in the key-value dictionary as the identifier of the function search tree node, and the coordinate position as the node attribute to construct the function search tree node; after performing the corresponding interaction operation on the function search tree node, update the interface image and generate a new key-value dictionary to expand new function search tree nodes; connect the newly generated function search tree nodes to the existing function search tree nodes and expand them step by step to construct a complete application system function search tree.
[0071] Specifically, when constructing the application system function search tree based on the key-value dictionary, first, use the functional category in the key-value dictionary as the identifier of the function search tree node, and the corresponding coordinate position as the attribute of the node to create an initial function search tree node. This node contains the functional description, coordinate information, and interaction type of the interface component, providing a basis for subsequent interface structure modeling. When performing an interaction operation, according to the identifier information of the current function search tree node, simulate user input, click, or other interaction methods to change the interface. After the interaction operation is completed, re-capture and parse the updated interface image to generate a new key-value dictionary. The new key-value dictionary contains the functional component information on the current interface and its coordinate position, providing a basis for the dynamic expansion of the function search tree. Subsequently, generate new function search tree nodes according to the new key-value dictionary and determine their association relationship with the existing nodes. If the newly generated node represents a component at the same functional level, it is associated as a child node of the current node; if it represents a change in the interface level, a cross-level link is established to ensure that the search tree can accurately reflect the interface structure and functional navigation path. By repeatedly performing interaction operations, interface updates, and search tree expansions, the newly generated function search tree nodes are gradually connected to the existing function search tree, gradually improving the entire search structure. After multiple rounds of iteration, a complete application system function search tree is finally formed.
[0072] In addition, in other embodiments of the present disclosure, when the result of intent classification is a Q&A intent, keywords are extracted from the user input data, and a database query instruction is generated based on the keywords; query result data is obtained based on the query instruction, and a multimodal large model is used to parse and generate natural language for the query result data to generate corresponding response content. Among them, the keyword can represent the core vocabulary or phrase extracted from the user input data and is used to match relevant content in the database. The database query instruction can represent a standardized query statement generated based on the extracted keywords, and this statement can be used to retrieve relevant information in the database to obtain query result data. The multimodal large model can represent a deep learning model trained based on multiple data types, such as text, images, speech, etc.
[0073] Specifically, when the result of intent classification is a Q&A intent, first, keywords are extracted from the user input data. Through natural language processing techniques, the user input data is tokenized, stop words are removed, and part-of-speech tagging is performed to extract core keywords with semantic information, and based on the API interface rules of the application system, standardized query parameters available for database query are determined. After obtaining the keywords, a database query instruction is generated according to the API rules of the application system. The query interface of the park system database is called to map the extracted keywords to the database query statement to match relevant information in the database. The database query instruction can adopt the structured query language (SQL) or the data request format based on the API to ensure that query result data that meets the user's needs can be obtained.
[0074] After the database query instruction is executed, query result data is obtained. The returned data is preprocessed, including format conversion, deduplication, sorting, and aggregation, to meet the user's query requirements. If the query result contains numerical values, statistical information, or time-series data, further data visualization processing is performed to generate charts or data summaries suitable for intuitive display. The processed query result data is input into the multimodal large model, and the multimodal large model parses and generates natural language for the query result data. The multimodal large model performs semantic analysis on the structured or unstructured query result data, extracts key information, and generates natural language text response content that meets the user's intent based on context reasoning. If the query result data contains images, charts, or other visual data, the model will generate a supporting visualization response in combination with the text information. Finally, the generated response content is fed back to the user in different forms. The full response content is presented through a screen window, including detailed text descriptions and visualization data, such as icons, charts, or tables. At the same time, key information in the response content is extracted to generate a simplified content summary, which is broadcast by a digital human in voice, enabling the user to quickly obtain the key answers.
[0075] In step S140, an operation instruction sequence from the current page to the target function page is generated based on the application system knowledge graph and the application system function search tree, and anthropomorphic operations are performed according to the operation instruction sequence. The operation instruction sequence can represent a sequence composed of multiple operation instructions arranged in order. This sequence is generated based on the application system knowledge graph and the application system function search tree and is used to guide the execution path of interface operations. Anthropomorphic operations can represent automated interactions performed based on the operation instruction sequence, including simulating user regular interaction methods such as mouse clicks, keyboard inputs, touch operations, or interface scrolling to achieve control of the application system.
[0076] In some embodiments, generating an operation instruction sequence from the current page to the target function page based on the application system knowledge graph and the application system function search tree specifically includes the following technical steps: obtaining the keyword corresponding to the target function page according to the system operation type intention; retrieving in the application system knowledge graph based on the keyword to obtain the entities, relationships, and attributes associated with the keyword; inputting the keyword, entities, relationships, and attributes into a multi-modal large model together with the application system function search tree to generate an operation path from the current page to the target function page; and generating a corresponding operation instruction sequence according to the operation path.
[0077] Specifically, when generating the operation instruction sequence, first, according to the system operation type intention, the user input data is parsed to obtain the keyword corresponding to the target function page. The keyword can be extracted from the instruction text input by the user or inferred based on a semantic parsing model to ensure the accuracy of the match. Subsequently, a retrieval is performed in the application system knowledge graph based on the keyword to obtain the entities, relationships, and attributes associated with the keyword. Entities can represent function modules, menu items, or interface components in the application system. Relationships can represent the association methods between different function components, such as hierarchical structures or interaction paths. Attributes can include information such as menu types, supported operations, and access conditions. Through this retrieval process, a knowledge representation of the target function page is constructed to support path reasoning.
[0078] After the retrieval is completed, the keywords, retrieved content, and the application system function search tree are input into the multi-modal large model. The multi-modal large model combines the entity relationship information provided by the knowledge graph and the hierarchical structure of the function search tree to infer the operation path from the current page to the target function page. This path may include multiple intermediate interaction steps to ensure a smooth navigation from the current interface to the target function page, and the page hierarchical relationship and interaction rules are combined during the inference process to optimize the rationality of the operation path. Finally, a corresponding operation instruction sequence is generated according to the operation path. The operation instruction sequence includes specific operations such as clicks, inputs, scrolls, etc. required for interface interaction, and is arranged in the order of page flow. The system performs anthropomorphic operations according to the instruction sequence to achieve automated interaction from the current page to the target function page, thereby accurately executing the system operations corresponding to the user's intention.
[0079] In addition, as shown in Figure 2 In other embodiments of the present disclosure, another digital human interaction control method based on screen recognition is also provided, which specifically includes the following steps: Step S201, user voice input. Specifically, the user issues an instruction to the digital human assistant through voice input, and the instruction may involve data query or system operation.
[0080] Step S202, the large model obtains the user's intention. Specifically, the multi-modal large model analyzes the user's voice input to extract key information.
[0081] Step S203, perform instruction classification. Specifically, the application system knowledge graph classifies the user input to determine whether the input belongs to an inquiry and answer type intention or a system operation type intention. If the instruction belongs to the inquiry and answer type, it enters the inquiry and answer process and jumps to step S204; if the instruction belongs to the system operation type, it enters the system operation process and jumps to step S208.
[0082] Step S204, call the database query interface and the system knowledge graph interface. Specifically, when the instruction classification result is an inquiry and answer type intention, keywords are extracted from the user input data, and a database query instruction is generated based on the keywords. Subsequently, the database query interface is called to obtain the query result data, and the data is input into the multi-modal large model for analysis and processing.
[0083] Step S205, the large model analyzes the query data. Specifically, based on the query results returned by the database, the multi-modal large model is used for data parsing, and the user's query intention is inferred in combination with the system knowledge graph to generate a natural language reply that meets the user's needs.
[0084] Step S206, generate response content. Specifically, the multimodal large model generates a natural language response that can be understood by the user based on the query data and the semantic analysis results, and combines visual data to enhance the information presentation effect.
[0085] Step S207, the digital human broadcasts / generates response content. Specifically, the digital human assistant displays the complete information on the screen according to the response content generated in Step S206, and at the same time extracts the key content for voice broadcast.
[0086] Step S208, call the system interface to take a screenshot. Specifically, when the instruction classification result is a system operation type intention, call the operating system interface to obtain the interface image of the current screen for interface content parsing.
[0087] Step S209, image recognition. Specifically, perform semantic segmentation on the intercepted interface image, extract the function description and coordinate information of the interface components, generate a key-value dictionary to construct a structured representation of the interface components.
[0088] Step S210, determine whether it is the first operation on this page. Specifically, determine whether the current interface is the first operation. If it is the first operation, jump to Step S213. If there is a historical interaction record, jump to Step S211.
[0089] Step S211, determine whether there is a change in the recognition result compared with the function search tree. Specifically, compare the current interface parsing result with the existing function search tree to determine whether the structure of the interface components has changed. If there is no change, jump to Step S214. If there is a change, jump to Step S212.
[0090] Step S212, update the function search tree. Specifically, if the interface components or the hierarchical structure change, update the function search tree to ensure accurate subsequent operation path planning.
[0091] Step S213, create a function search tree for the current page. Specifically, when the interface is the first operation page, construct a function search tree based on the interface parsing result, and record the hierarchical relationship and interaction path of each component.
[0092] Step S214, input the user intention, function search tree, and system knowledge graph into the training large model. Specifically, input the user intention classification result, function search tree, and system knowledge graph into the pre-trained large model, and combine historical data for learning and optimization to improve the future task inference ability.
[0093] Step S215: Generate a sequence of screen operation instructions. Specifically, based on the application system knowledge graph, the function search tree, and the predicted path of the target function page, infer the operation path from the current page to the target page, and generate the corresponding operation instruction sequence, including anthropomorphic operation instructions such as clicks and inputs.
[0094] Step S216: Execute the instructions. Specifically, perform interface operations in sequence according to the operation instruction sequence to complete the task requested by the user.
[0095] Step S217: Determine whether all instructions are completed. Specifically, determine whether the currently executed instruction sequence has completed all steps. If not, jump to step S208. If so, the task execution is completed. When all instructions are executed, the task process ends, returns to the standby state, and waits for new user input.
[0096] In the digital human interaction control method based on screen recognition in the exemplary embodiments of the present disclosure, on the one hand, based on the information parsing of the application system, an application system knowledge graph is constructed, enabling the user input data to establish associations with the functions, operations, and data of the application system. By parsing the function description documents, user manuals, and interface elements of the application system, buttons, form fields, and their hierarchical relationships are extracted, and combined with semantic annotation and entity recognition methods, they are stored in a structured manner. The pre-trained model is used to perform entity recognition on the text data, and the verb-object structure relationship is extracted through dependency syntactic analysis, making the relevance between different interface elements clear. In the above way, the application system knowledge graph can effectively represent the function hierarchy and component relationships of different application systems, providing a basis for subsequent user intention classification and interface operation planning.
[0097] On the other hand, an application system function search tree is constructed based on the key-value dictionary to establish the hierarchical relationship of interface operations. By using the function category of the interface component as the search tree node identifier and combining the coordinate position to construct the function search tree node, the rational organization of different interface components in the tree structure is ensured. In addition, based on the application system knowledge graph and the application system function search tree, an operation instruction sequence from the current page to the target function page is generated to realize the inference of the interaction path. After inputting keywords, entities, relationships, and attributes into the multi-modal large model, combined with the hierarchical relationship of the function search tree, an operation path is generated and further converted into a screen interaction instruction sequence. This method enables the system to simulate user operations, gradually execute interface interactions from the current interface, and realize the automated operation of the target function. Thus, the anthropomorphic control of the management system is realized, reducing the user learning cost and usage threshold. Through operation supervision via the user interface, the risks and potential hazards are reduced compared to background operations, ensuring the user's right to know. This method provides a new digital human assistant assistance mode, with higher generalization ability and transferability, without the need for complex API adaptation, and can learn and reason about the operations and controls of the new system based on screen graphics and text.
[0098] It should be noted that although the steps of the methods in the present disclosure are described in a specific order in the accompanying drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0099] Next, in the embodiments of the present disclosure, a digital human interaction control system based on screen recognition is also provided. As shown in Figure 3 the digital human interaction control system 300 based on screen recognition may be composed of a knowledge graph module 301, an intention classification module 302, a screen capture module 303, and an interaction execution module 304. Among them: the knowledge graph module can be used to construct an application system knowledge graph based on the information parsing of the application system; the intention classification module can be used to obtain user input data and classify the intention of the user input data based on the application system knowledge graph; the screen capture module can be used to capture the interface image of the current screen in response to the result of the intention classification being a system operation type intention, perform semantic segmentation on the interface image, generate a key-value dictionary containing the function description and coordinate position of the interface component, and construct an application system function search tree based on the key-value dictionary; the interaction execution module can be used to generate an operation instruction sequence from the current page to the target function page based on the application system knowledge graph and the application system function search tree, and perform anthropomorphic operations according to the operation instruction sequence.
[0100] In addition, in other embodiments of the present disclosure, the digital human interaction control system based on screen recognition further includes a consultation and Q&A module, which is used to, in response to the result of the intention classification being a consultation and Q&A type intention, extract keywords from the user input data, and generate a database query instruction based on the keywords; obtain query result data based on the query instruction, and use a multi-modal large model to parse and generate natural language for the query result data to generate corresponding reply content.
[0101] It should be noted that the specific details of each part in the above digital human interaction control system based on screen recognition have been described in detail in the implementation manners of the digital human interaction control method part. The details not disclosed can be referred to the implementation manners of the method part, and thus will not be elaborated here.
[0102] In addition, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above digital human interaction control method based on screen recognition is also provided.
[0103] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0104] Next, refer to Figure 4 to describe the electronic device 400 according to this embodiment of the present disclosure. Figure 4 The illustrated electronic device 400 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0105] As Figure 4 shown, the electronic device 400 is presented in the form of a general computing device. The components of the electronic device 400 may include but are not limited to: the above at least one processing unit 410, the above at least one storage unit 420, a bus 430 connecting different system components (including the storage unit 420 and the processing unit 410), and a display unit 440.
[0106] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 410, so that the processing unit 410 executes the steps according to various exemplary embodiments of the present disclosure described in the above "exemplary method" part of this specification.
[0107] The storage unit 420 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 421 and / or a cache storage unit 422, and may further include a read-only storage unit (ROM) 423.
[0108] The storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425. Such program modules 425 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0109] The bus 430 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0110] The electronic device 400 may also communicate with one or more external devices 470 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 400, and / or may communicate with any device that enables the electronic device 400 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 450. Also, the electronic device 400 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 460. As shown in the figure, the network adapter 460 communicates with other modules of the electronic device 400 through the bus 430. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0111] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0112] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above methods of this specification is stored. In some possible embodiments, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0113] Reference Figure 5 As shown, a program product 500 for implementing the above-described digital human interaction control method based on screen recognition according to an embodiment of the present disclosure is described. It may be in the form of a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0114] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0115] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0116] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, electromagnetic wave, etc., or any suitable combination of the above.
[0117] Program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0118] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0119] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0120] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A digital human interactive control method based on screen recognition, characterized in that: include: Based on the information analysis of the application system, build the knowledge graph of the application system; Obtaining user input data, and classifying the user input data by intent based on the application system knowledge graph; In response to the result of the intent classification being a system operation intent, capturing an interface image of the current screen, performing semantic segmentation on the interface image, generating a key-value dictionary including interface component function descriptions and coordinate positions, and constructing an application system function search tree based on the key-value dictionary; Based on the application system knowledge graph and the application system function search tree, an operation instruction sequence from the current page to the target function page is generated, and anthropomorphic operations are performed according to the operation instruction sequence.
2. The digital human interactive control method based on screen recognition according to claim 1, characterized in that: Also includes: In response to the result of the intent classification being a consultation question and answer type intent, extracting keywords from the user input data, and generating a database query instruction based on the keywords; Query result data is obtained based on the query instruction, and the query result data is parsed and natural language is generated using a multimodal large model to generate corresponding reply content.
3. The digital human interactive control method based on screen recognition according to claim 1, characterized in that: The information analysis based on the application system and the construction of the application system knowledge graph include: Extract buttons, form fields and their hierarchical relationships from the functional description documents, user manuals and interface elements of the application system, and semantically annotate the UI icons, texts and coordinate information of the buttons and form fields; Using a pre-trained model to perform entity recognition on the semantically annotated text data, and extracting the subject-object structure relationship of the text data through dependency syntactic analysis; Based on the results of subject-object structural relationship extraction, a knowledge graph of the application system including entity types, relationship types and attribute constraints is constructed.
4. The digital human interactive control method based on screen recognition according to claim 1, characterized in that: The obtaining of user input data and classifying the user input data by intent based on the application system knowledge graph includes: Collect text data, image data and voice data input by the user to generate multimodal input data; Extracting keywords from the multimodal input data, and retrieving functions, operations, and entity information associated with the keywords based on the application system knowledge graph; The retrieval results and the multimodal input data are input into a pre-trained classification model to generate an intent classification result.
5. The digital human interactive control method based on screen recognition according to claim 1, characterized in that: The semantic segmentation of the interface image is performed to generate a key-value dictionary including the function description and coordinate position of the interface components, and the application system function search tree is constructed based on the key-value dictionary, including: Performing semantic segmentation on the interface image to determine a target area corresponding to each interface component; Extracting text features and icon features of the target area, and generating a key-value dictionary of function description and coordinate position of each interface component; The nodes of the application system function search tree are constructed based on the key-value dictionary, and the nodes are expanded step by step according to the interactive operation to generate the application system function search tree.
6. The digital human interactive control method based on screen recognition according to claim 5, characterized in that: The step of extracting text features and icon features of the target area and generating a key-value dictionary of function description and coordinate position of each interface component includes: Performing optical character recognition on the target area to obtain a text feature vector; A convolutional neural network is used to extract image features of icons in the target area and generate icon feature vectors; Fusing the text feature vector and the icon feature vector to generate a fused feature vector; Based on the fused feature vector, the functional category of the interface component is calculated: ,in, represents the classifier weight matrix, represents the classifier bias term, represents the normalization function, Indicates Functional categories of interface components, represents the fused feature vector; The function category and the target area coordinate position are combined to generate a key-value dictionary of interface component function description and coordinate position.
7. The digital human interactive control method based on screen recognition according to claim 6 is characterized in that: The step of constructing nodes of the application system function search tree based on the key-value dictionary and expanding the nodes step by step according to the interactive operation to generate the application system function search tree includes: Use the function category in the key-value dictionary as the function search tree node identifier and the coordinate position as the node attribute to construct a function search tree node; After the function search tree node performs a corresponding interactive operation, the interface image is updated and a new key-value dictionary is generated, and a new function search tree node is expanded; Connect the newly generated function search tree nodes to the existing function search tree nodes, expand them level by level, and build a complete application system function search tree.
8. The digital human interactive control method based on screen recognition according to claim 1, characterized in that: The generating of an operation instruction sequence from a current page to a target function page based on the application system knowledge graph and the application system function search tree includes: According to the system operation intent, obtain keywords corresponding to the target function page; Searching the knowledge graph of the application system based on the keyword to obtain entities, relationships and attributes associated with the keyword; Input the keywords, entities, relationships and attributes and the application system function search tree into the multimodal macro model to generate an operation path from the current page to the target function page; A corresponding operation instruction sequence is generated according to the operation path.
9. A digital human interactive control system based on screen recognition, characterized in that: include: The knowledge graph module is used to parse the information of the application system and build the knowledge graph of the application system; An intent classification module, used to obtain user input data and perform intent classification on the user input data based on the application system knowledge graph; A screen capture module, for capturing an interface image of the current screen in response to the result of the intent classification being a system operation intent, and performing semantic segmentation on the interface image to generate a key-value dictionary containing interface component function descriptions and coordinate positions, and constructing an application system function search tree based on the key-value dictionary; The interactive execution module is used to generate an operation instruction sequence from the current page to the target function page based on the application system knowledge graph and the application system function search tree, and perform anthropomorphic operations according to the operation instruction sequence.
10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the digital human interaction control method based on screen recognition as described in any one of claims 1-8 by executing the executable instructions.
Citation Information
Patent Citations
Intelligent question answering system for automobile field
CN113505209A
Intelligent question answering system and method based on knowledge graph
CN115525751A
Corpus labeling method, device and system, knowledge extraction method, device and system and graph construction method, device and system
CN118152577A
Internet mass data accurate search method and system based on AI technology
CN119557500A
Analyzing graphical user interfaces to facilitate automatic interaction
US20220050661A1
Cited By
Power industry dynamic knowledge base construction method and system based on large language model
CN120705130A
Closed-loop reasoning method, device and equipment based on multi-modal large model
CN120996209A
Conversational interaction method and device suitable for tobacco machinery and medium thereof
CN121256001A
A dialog-based interaction method, device and medium suitable for tobacco machinery
CN121256001B
Menu navigation method and device, equipment and medium
CN122152180A