A human-computer interaction method, device and medium based on large model

By combining large-scale model technology with visual information on industrial production lines and building a multi-level knowledge base, we have solved the problems of existing systems in understanding complex sentences and slow knowledge base updates, achieved fast and professional question-and-answer support, and improved user experience and system efficiency.

CN117992587BActive Publication Date: 2025-09-23GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410084390.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-09-23
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

Existing human-computer question-answering systems for industrial production lines perform poorly when processing complex sentences, informal expressions, and written errors. They are unable to correctly understand and analyze information, and their knowledge base is updated slowly, making it impossible to provide accurate and professional real-time technical support.

Method used

Large model technology is used to combine visual information and question type classification to build local and external knowledge bases. Visual information collected by cameras and large models are used to analyze user questions and provide accurate and professional answers.

Benefits of technology

It improves the efficiency of the system and user experience, enhances the relevance and accuracy of information, can quickly provide professional answers, and supports the continuous updating and expansion of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117992587B_ABST
    Figure CN117992587B_ABST
Patent Text Reader

Abstract

The present invention discloses a human-computer interaction method, device and medium based on a large model, which is applied to a human-computer question-answering system on an industrial production line. The method comprises: obtaining a question input by a user, and in a scenario question-answering mode, using a large model to perform structured analysis on the text of the question to obtain the question type of the question; wherein the question type includes step type, object attribute type, object definition type, verification type and other types; according to the question type and the scene selected by the user, obtaining visual information collected by the camera and a pre-built knowledge base, and outputting the answer corresponding to the question based on the large model; wherein, when the question type is other types, first using the large model to answer the question, and then outputting guidance information and recommendation information related to the knowledge base. The present invention can provide more accurate and professional answers to different categories of questions, realize comprehensive interactive question-answering, improve the efficiency of the system, and enhance the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a large model-based human-computer interaction method, device and medium. Background Art

[0002] With the rapid development of intelligent technology, human-machine question-answering systems on industrial production lines are becoming increasingly intelligent, enhancing operators' decision-making capabilities and improving efficiency while also providing real-time technical support. Existing question-answering systems typically build knowledge bases by combining local knowledge bases with web search technology. When local knowledge bases cannot answer questions, they search knowledge bases on the Internet. In the context of industrial production line applications, the current construction of industrial-related knowledge bases faces problems such as highly specialized knowledge, rapid technological updates, multilingualism, and cross-culturalism. In addition, existing question-answering systems have poor processing effects on complex sentences, informal expressions, and written errors. They are unable to correctly understand, analyze, and synthesize information during human-computer interaction, and lack accuracy and professionalism when providing real-time teaching and technical support. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a human-computer interaction method, device and medium based on a large model. By introducing large model technology, combining visual information and finely classifying question types, it can provide more accurate and professional answers to questions of different categories, realize comprehensive interactive question and answer, improve the efficiency of the system, and enhance the user experience.

[0004] An embodiment of the present invention provides a human-computer interaction method based on a large model, comprising:

[0005] Obtain the question input by the user and determine whether there is matching information for the question in the system knowledge base; if so, output the corresponding answer based on the matching information; if not, enter the scenario question and answer mode;

[0006] In the scenario question-answering mode, a large model is used to perform structured analysis on the text of the question to obtain the question type of the question; wherein the question type includes step type, object attribute type, object definition type, verification type and other types;

[0007] According to the question type and the scenario selected by the user, the visual information collected by the camera and the pre-built knowledge base are obtained, and based on the large model, the answer corresponding to the question is output; wherein, when the question type is other types, the large model is first used to answer the question, and then the guidance information and recommendation information related to the knowledge base are output.

[0008] As an improvement of the above solution, the knowledge base includes a local knowledge base and an external knowledge base; wherein, the local knowledge base includes a scene-specific knowledge base, a large model prompt base and the system knowledge base; the scene-specific knowledge base includes a visual knowledge base and a behavioral knowledge base.

[0009] As an improvement to the above solution, obtaining a question input by a user and determining whether there is matching information for the question in the system knowledge base specifically includes:

[0010] Obtaining a question input by the user;

[0011] Preprocessing the text of the question to obtain first parsing information;

[0012] Performing entity recognition and keyword extraction on the first parsed information to obtain second parsed information;

[0013] Converting the second parsed information into a vector form to obtain third parsed information;

[0014] The third parsed information is matched with information in the system knowledge base, and it is determined whether there is matching information for the problem in the system knowledge base.

[0015] As an improvement to the above solution, after entering the scenario question-and-answer mode, the method further includes:

[0016] Determining whether the user has selected a scene based on historical question and answer information;

[0017] If the user has not selected a scene, outputting scene selection guidance information to enable the user to select a scene;

[0018] If the user has selected a scenario, the scenario selected by the user is used as the application scenario of the current question.

[0019] As an improvement to the above solution, the method obtains visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputs the answer corresponding to the question based on the large model, specifically including:

[0020] When the question type is the step type, obtaining information in the behavior knowledge base and prompt information related to the question in the large model prompt library according to the scenario selected by the user;

[0021] Acquiring first visual information collected by the camera;

[0022] Inputting the prompt information and the first visual information into the large model to obtain first text description information corresponding to the first visual information;

[0023] The first text description information is matched with the information in the behavior knowledge base to determine the user's current operation steps and output an answer corresponding to the question.

[0024] As an improvement to the above solution, the method obtains visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputs the answer corresponding to the question based on the large model, specifically including:

[0025] When the question type is the object attribute type or the object definition type, obtaining object information in the visual knowledge base according to the scene selected by the user;

[0026] If the question type is the object attribute type, the large model is used to extract object keywords from the text of the question; if the question type is the object definition type, second visual information captured by the camera is obtained, the large model is used to convert the second visual information into second text description information, and object keywords related to the question are extracted from the second text description information;

[0027] The object keywords are matched with the object information in the visual knowledge base, and the answer corresponding to the question is output according to the matching result.

[0028] As an improvement to the above solution, the method obtains visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputs the answer corresponding to the question based on the large model, specifically including:

[0029] When the question type is the verification type, obtaining information in the behavior knowledge base according to the scenario selected by the user;

[0030] Acquiring third visual information collected by the camera; wherein the third visual information is a video saved in segments according to the user's operation frequency;

[0031] Extracting key frames from the video, and inputting the obtained key frames into the large model to obtain third text description information corresponding to each key frame;

[0032] The third text description information is matched with the information in the behavior knowledge base to determine the correctness of the user operation and output an answer corresponding to the question.

[0033] As an improvement to the above solution, the method further includes:

[0034] Get industry-standard videos uploaded by professionals;

[0035] Extracting key frames of the industrial standard video, and performing behavior recognition on the key frames using the large model to obtain behavior descriptions corresponding to the key frames;

[0036] Outputting the behavior description so that the professional can judge whether the behavior description complies with the specification;

[0037] Obtain the judgment result of the professional. If the behavior description does not meet the specifications, return to the step of extracting the key frames of the industrial specification video; if the behavior description meets the specifications, update the behavior description to the scene-specific knowledge base.

[0038] The embodiment of the present invention further provides a large-scale model-based human-computer interaction device, which is applied to a human-computer question-answering system on an industrial production line, including:

[0039] The question acquisition module is used to obtain the question input by the user and determine whether there is matching information for the question in the system knowledge base; if so, output the corresponding answer based on the matching information; if not, enter the scenario question and answer mode;

[0040] A question classification module is used to perform structured analysis on the text of the question using a large model in the scenario question-answering mode to obtain the question type of the question; wherein the question type includes step type, object attribute type, object definition type, verification type and other types;

[0041] The answer output module is used to obtain visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and output the answer corresponding to the question based on the large model; wherein, when the question type is other types, the large model is first used to answer the question, and then the guidance information and recommendation information related to the knowledge base are output.

[0042] An embodiment of the present invention also provides a terminal device, including a processor and a memory, wherein a computer program is stored in the memory, and the computer program is configured to be executed by the processor, and when the processor executes the computer program, it implements any of the above-mentioned large model-based human-computer interaction methods.

[0043] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any of the above-mentioned large model-based human-computer interaction methods.

[0044] Compared with the existing technology, the beneficial effects of the large-scale model-based human-computer interaction method, device and medium provided by the embodiment of the present invention are: by adopting large-scale model technology, it can better integrate and refine information, and handle multilingual and multicultural problems, and also enable the knowledge base to have the ability to continuously expand and update, enhance the scalability of the system, and provide long-lasting and effective information support; by classifying the question types in detail, it can provide accurate and professional answers for different categories of questions, enhancing the technical effect of the system in specific fields; through the question guidance mechanism, the user's interactive experience is enhanced, and the relevance and accuracy of the information are ensured; through the selection of specific scenarios, the system can answer faster and more efficiently; by introducing visual information and combining video, graphics and text content, the question-answering system can more accurately assess the needs of learners and provide more personalized and intuitive learning guidance, thereby improving educational effects and enhancing user experience. The embodiment of the present invention can provide comprehensive, accurate and professional answers in human-computer interactive question-answering in industrial production lines, greatly enhancing the user's interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flowchart of a large-model-based human-computer interaction method provided by an embodiment of the present invention;

[0046] Figure 2 This is a flowchart of a problem classification method based on a large model provided by an embodiment of the present invention;

[0047] Figure 3 This is a learning flow chart of a large-model-based human-computer interaction method provided by an embodiment of the present invention;

[0048] Figure 4 This is a schematic structural diagram of a large-scale model-based human-computer interaction device provided by an embodiment of the present invention;

[0049] Figure 5 It is a structural diagram of a terminal device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0051] Large models are machine learning models with large parameters and complex computational structures. These models are typically built using deep neural networks and have billions or even hundreds of billions of parameters. Large models are capable of handling more complex tasks and data, including natural language processing, computer vision, and speech recognition.

[0052] See also Figure 1 , Figure 1 This is a flow chart of a large-scale model-based human-computer interaction method provided by an embodiment of the present invention. The large-scale model-based human-computer interaction method is applied to a human-computer question-answering system on an industrial production line, and includes:

[0053] S1: Obtain the question input by the user and determine whether there is matching information for the question in the system knowledge base; if so, output the corresponding answer based on the matching information; if not, enter the scenario question and answer mode;

[0054] S2: In the scenario question-answering mode, a large model is used to perform structured analysis on the text of the question to obtain the question type of the question; wherein the question type includes step type, object attribute type, object definition type, verification type, and other types;

[0055] S3: According to the question type and the scenario selected by the user, the visual information collected by the camera and the pre-built knowledge base are obtained, and based on the large model, the answer corresponding to the question is output; wherein, when the question type is other types, the large model is first used to answer the question, and then the guidance information and recommendation information related to the knowledge base are output.

[0056] Specifically, the human-computer interaction method provided by an embodiment of the present invention is applied to a human-computer interaction system on an industrial production line. The system includes a question-and-answer module and a learning module. The learning module can learn knowledge and build a knowledge base. The question-and-answer module can answer questions raised by learners based on the knowledge base to achieve the purpose of teaching.

[0057] When a user uses a human-computer question-answering system for human-computer interactive learning, the user first enters a question in the system's interactive interface. It should be noted that the user can enter the question in text or voice. When using voice input, the system automatically converts the user's voice input into text. The system receives the question entered by the user and determines whether the question is relevant to the system knowledge. If so, it outputs a relevant answer based on the system knowledge base. If not, it enters the scenario question-answering mode to answer questions in a specific scenario. Please refer to Figure 2 , Figure 2This is a flowchart of a question classification method based on a large model of human-computer interaction provided by an embodiment of the present invention. The embodiment of the present invention first uses a large model to perform structured analysis on the question text input by the user to obtain structured text. Specifically, the definitions and examples of five pre-set questions are provided to the large model. The large model analyzes and classifies the structured text based on the provided information to obtain the type of question. In the specific scenario selected by the user, the relevant knowledge base is retrieved according to the question type, and images or videos captured by the camera are obtained according to the question type and input into the large model to obtain text description information output by the large model. The text description information is extracted and queried in the relevant knowledge base, and the answer to the question is output based on the query result.

[0058] In the application scenario of industrial production lines, the embodiment of the present invention divides questions into step types, object attribute types, object definition types, verification types and other types. Among them, step-type questions involve specific step instructions for the operation process, such as: "How should the entire operation process be done?", "What should the action of a certain step be?", "What is the next operation?", etc.; object attribute-type questions focus on the specific attributes of related objects in the operation process, such as: "What does an object look like?", "What are the characteristics of an object?", etc.; object definition-type questions are object knowledge that users need to understand; verification-type questions are verification requests initiated by users to the system regarding the correctness of the operation, such as: "Does the operation I just performed comply with the operating specifications you taught me?", "Is the action I just performed correct?", etc.; other types of questions are problems that cannot be directly solved through the local knowledge base, such as "What are the safety hazards of this operation?", "Where is this operation applied?", etc. Other types of questions also include external problems that do not directly involve the system knowledge base, such as "What is the weather like today?", etc.

[0059] Among them, when the question type is other types, the system can perform intelligent guidance through the large model, and can naturally guide the user to topics related to the system. Specifically, for other types of questions, the system will first make a preliminary answer through the large model, and output guidance information and recommendation information after the answer. Among them, the guidance information can naturally guide the user's next question and make it closer to the content of the system knowledge base; the recommendation information is relevant information of the various functions of the system, such as learning, teaching, supervision and operating specifications. In an embodiment of the present invention, when the system deals with questions that are not related to the system, it can maintain user participation by giving brief answers based on the questions, while preventing users from deviating from the main functions and knowledge scope of the system. The system's intelligent guidance and recommendation can not only improve the relevance of the conversation and enhance the consistency of the user experience, but also increase the user's understanding and utilization of the system.

[0060] For example, when a user's question is unrelated to the system's functions, such as "How is the weather today?", the system will first give a brief answer, such as: "The weather is very good today!", and then guide the user to focus on the core functions of the system, such as: "If you have any questions about learning, teaching, supervision, and operating specifications, I will be here to help you at any time." etc.

[0061] As one of the optional embodiments, the knowledge base includes a local knowledge base and an external knowledge base; wherein, the local knowledge base includes a scene-specific knowledge base, a large model prompt base and the system knowledge base; the scene-specific knowledge base includes a visual knowledge base and a behavioral knowledge base.

[0062] Specifically, the knowledge base constructed by the system consists of two parts: the local knowledge base and the external knowledge base. The local knowledge base is divided into three sub-bases: the scene-specific knowledge base, the large model prompt base, and the system knowledge base.

[0063] The scenario-specific knowledge base stores professional and standardized knowledge related to industrial production lines. Its content is classified and stored for different industrial scenarios to facilitate rapid retrieval and application after the scenario is selected. Among them, the scenario-specific knowledge base includes a visual knowledge base and a behavioral knowledge base. The visual knowledge base mainly stores knowledge related to objects and object attributes, such as images of the appearance of objects, material and color information of objects, etc. The behavioral knowledge base mainly stores the actions and process specifications involved in each scenario, for example: the first step in a certain scenario requires the placement of an object, a certain operation needs to be performed when placing an object, and a sequence of images of changes between the hand and the operated object during a certain operation.

[0064] The large model prompt library stores large model prompt knowledge collected and maintained by the system, which includes pre-input prompt information customized according to different usage requirements. This prompt information can optimize the performance and output quality of large models.

[0065] The system knowledge base is the core part of the system. It provides relevant introductions and user manuals of the system and is a basic guide for users to understand and operate the system.

[0066] The external knowledge base primarily includes the large model knowledge base, which provides the system with extensive external information and data support, making the system's knowledge base more comprehensive. The large model knowledge base is the internal information and data storage of the large model. It is the knowledge set constructed internally during the model training process and is embedded in the large model in the form of model parameters.

[0067] As one of the optional embodiments, obtaining a question input by a user and determining whether there is matching information for the question in a system knowledge base specifically includes:

[0068] Obtaining a question input by the user;

[0069] Preprocess the text of the question to obtain the first parsing information;

[0070] Perform entity recognition and keyword extraction on the first parsing information to obtain the second parsing information;

[0071] Convert the second parsing information into a vector form to obtain the third parsing information;

[0072] Match the third parsing information with the information in the system knowledge base, and determine whether there is matching information for the question in the system knowledge base.

[0073] Specifically, in the knowledge Q&A between the user and the system, after the user inputs a question, the system first performs intelligent encoding and parsing on the question text to provide accurate and efficient information feedback. After obtaining the question input by the user, first preprocess the text of the question. The preprocessing includes: removing irrelevant characters such as punctuation marks and special symbols, text normalization (for example, converting all characters to lowercase), word segmentation (for languages such as Chinese, Japanese, etc.), and removing stop words (such as "de", "he", etc.). After text preprocessing, perform entity recognition and keyword extraction. Entity recognition (NER, Named Entity Recognition) is to identify the key entities in the question text, such as person names, place names, organization names, etc., which helps to understand the specific object of the user's query; keyword extraction is to extract keywords containing the core information of the user's query from the question text. Finally, perform encoding to convert the text into a format that the model can understand, that is, convert the processed text into a vector form, specifically through word embeddings, such as Word2Vec, GloVe or BERT embeddings. After converting the text into a vector form, it can be matched with the built-in system information library to query whether the user's question is related to the preset knowledge of the system. When there is matching information for the question in the system knowledge base, it means that the user's question is directly related to the system knowledge, and the system can quickly provide an accurate answer according to the information in the preset system knowledge base. When there is no matching information for the question in the system knowledge base, the system enters the scenario Q&A mode.

[0074] As an optional embodiment, after entering the scenario Q&A mode, the method further includes:

[0075] Judge whether the user has selected a scenario according to the historical Q&A information;

[0076] If the user has not selected a scenario, output guiding information for scenario selection to enable the user to select a scenario;

[0077] If the user has selected a scenario, use the scenario selected by the user as the application scenario of the current question.

[0078] Specifically, after entering the scenario question and answer mode, the user needs to first determine the scenario corresponding to the question. The system can determine whether the user has selected the scenario based on the historical question and answer information. If the user has not selected the scene, the system will output guidance information for scene selection. The guidance information will provide the scenarios currently supported by the system, thereby guiding the user to select a specific scene, and then conduct questions and answers for the relevant scenarios. If the user has already selected a scene in the previous question and answer information, the question and answer learning for the corresponding scene will start directly. It should be noted that the user can also change the selected scene by entering a question. By selecting the corresponding scene, the embodiment of the present invention enables the system to query the relevant answers in the knowledge base faster and more accurately when answering questions.

[0079] As one of the optional embodiments, the method of obtaining visual information collected by a camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputting an answer corresponding to the question based on the large model, specifically includes:

[0080] When the question type is the step type, obtaining information in the behavior knowledge base and prompt information related to the question in the large model prompt library according to the scenario selected by the user;

[0081] Acquiring first visual information collected by the camera;

[0082] Inputting the prompt information and the first visual information into the large model to obtain first text description information corresponding to the first visual information;

[0083] The first text description information is matched with the information in the behavior knowledge base to determine the user's current operation steps and output an answer corresponding to the question.

[0084] Specifically, for step-based questions, the system is required to observe the user's current learning progress, thus requiring the acquisition of primary visual information captured by the camera. For these types of questions, the system extracts answers from the behavioral knowledge base within the scenario-specific knowledge base, and extracts relevant prompt information from the large model prompt library. During specific processing, the system captures the user's real-time image through the camera and inputs the captured image into the large model. The large model, aided by the prompt information, generates a text description of the current operation scenario. The generated text description is then matched with the information in the behavioral knowledge base to determine the user's current operation steps and progress, and the corresponding answer to the question is provided based on the matching results.

[0085] For example, the prompt information for a step-type question is as follows: "'The general process of the heat sink installation specification is as follows: First, place the base plate: pick up the base plate by hand and place it on the operation panel; then, place the circuit board: pick up the circuit board by hand and place it in the groove of the base plate; then, place the cover plate: pick up the cover plate by hand and place it on the base plate; finally, install the heat sink: pick up the heat sink by hand, tear off the sticker on the heat sink, and place it in the groove of the cover plate.' Please refer to this operation specification to describe the pictures I provide you. The focus of the description is mainly on the objects involved and the operation status."

[0086] As one of the optional embodiments, the method of obtaining visual information collected by a camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputting an answer corresponding to the question based on the large model, specifically includes:

[0087] When the question type is the object attribute type or the object definition type, obtaining object information in the visual knowledge base according to the scene selected by the user;

[0088] If the question type is the object attribute type, the large model is used to extract object keywords from the text of the question; if the question type is the object definition type, second visual information captured by the camera is obtained, the large model is used to convert the second visual information into second text description information, and object keywords related to the question are extracted from the second text description information;

[0089] The object keywords are matched with the object information in the visual knowledge base, and the answer corresponding to the question is output according to the matching result.

[0090] Specifically, for questions about object attributes, the system answers them through the visual knowledge base of the scene-specific knowledge base. It parses the question text through the large model, extracts object keywords, and matches the object keywords with the object information in the knowledge base to provide relevant object information. For questions about object definitions, the system also answers them through the visual knowledge base of the scene-specific knowledge base. It also needs to obtain secondary visual information captured by the camera, input the captured image into the large model, obtain the corresponding secondary text description information, and extract object keywords related to the question from the second text description information. It then matches the object keywords with the object information in the knowledge base and outputs the answer to the question based on the matching results.

[0091] As one of the optional embodiments, the method of obtaining visual information collected by a camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputting an answer corresponding to the question based on the large model, specifically includes:

[0092] When the question type is the verification type, obtaining information in the behavior knowledge base according to the scenario selected by the user;

[0093] Acquiring third visual information collected by the camera; wherein the third visual information is a video saved in segments according to the user's operation frequency;

[0094] Extracting key frames from the video, and inputting the obtained key frames into the large model to obtain third text description information corresponding to each key frame;

[0095] The third text description information is matched with the information in the behavior knowledge base to determine the correctness of the user operation and output an answer corresponding to the question.

[0096] Specifically, for verification-type questions, the system needs to obtain a complete operation video to determine whether the user's operation complies with the specifications. First, the system obtains the third visual information collected by the camera. The system performs real-time image acquisition at startup and saves the operation video in segments according to the user's operation frequency. The operation frequency is determined by monitoring whether the user's hand is in the operation area of ​​the camera. The user's hand enters and leaves the operation area once, which is an operation. The system collects and saves each operation video through the camera. After the video is saved, key frames are extracted from each video segment, and the key frames are input into the large model for behavior recognition. The third text description information corresponding to each key frame is obtained, thereby forming a description corresponding to each video segment. It is then matched with the operation step knowledge in the behavior knowledge base to determine whether the user's operation complies with the specifications. Finally, the correctness of the user's operation is answered based on the matching results.

[0097] As an optional embodiment, the method further includes:

[0098] Get industry-standard videos uploaded by professionals;

[0099] Extracting key frames of the industrial standard video, and performing behavior recognition on the key frames using the large model to obtain behavior descriptions corresponding to the key frames;

[0100] Outputting the behavior description so that the professional can judge whether the behavior description complies with the specification;

[0101] Obtain the judgment result of the professional. If the behavior description does not meet the specifications, return to the step of extracting the key frames of the industrial specification video; if the behavior description meets the specifications, update the behavior description to the scene-specific knowledge base.

[0102] Specifically, in addition to the question-and-answer module, the system also includes a learning module. This module updates the scenario-specific knowledge base within the local knowledge base to ensure the accuracy and timeliness of professional knowledge. The system's learning function relies on the combined application of natural language processing, behavior recognition, and visual segmentation technologies to effectively absorb and integrate professional knowledge on industrial installation specifications. Experts and technicians can upload videos with detailed explanations of industrial installation specifications for the system to learn from. For example, videos detailing how to install a heat sink on a system motherboard or the assembly specifications for a computer case can be found.

[0103] See also Figure 3 , Figure 3 This is a learning flowchart of a human-computer interaction method based on a large model provided by an embodiment of the present invention. During the learning process, the system first obtains the industrial standard video uploaded by professionals, extracts key frames from the video to obtain picture sequence frames, and inputs the large model for behavior recognition, obtains and outputs the behavior description related to the picture, and then performs expert interpretation. Professionals verify and evaluate the learning results of the system at this stage to ensure the accuracy of the learned content, and can also provide learning suggestions. The system obtains the results of the expert interpretation. If the professional determines that the behavior description output by the system does not meet the requirements, the video key frame extraction and behavior recognition are re-performed according to the information fed back by the professional; if the professional determines that the behavior description meets the requirements, the behavior description is updated to the corresponding scene-specific knowledge base. The embodiment of the present invention enhances the learning results of the system through an interactive learning process, and also lays a solid foundation for building a comprehensive and accurate knowledge base.

[0104] The beneficial effects of a large-scale model-based human-computer interaction method provided by an embodiment of the present invention are: by adopting large-scale model technology, it can better integrate and refine information, and handle multilingual and multicultural problems, and also enable the knowledge base to have the ability to continuously expand and update, thereby enhancing the scalability of the system and providing lasting and effective information support; by classifying question types in detail, it can provide accurate and professional answers to questions of different categories, thereby enhancing the technical effect of the system in specific fields; through the question guidance mechanism, the user's interactive experience is enhanced, and the relevance and accuracy of the information are ensured; through the selection of specific scenarios, the system's answer speed can be made faster and more efficient; by introducing visual information and combining video, graphics and text content, the question-answering system can more accurately assess the needs of learners and provide more personalized and intuitive learning guidance, thereby improving educational effects and enhancing user experience. The embodiment of the present invention can provide comprehensive, accurate and professional answers in human-computer interactive question-answering in industrial production lines, greatly enhancing the user's interactive experience.

[0105] Correspondingly, the present invention also provides a large-model-based human-computer interaction device, which can implement all processes of the large-model-based human-computer interaction method in the above embodiment.

[0106] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of a large-scale model-based human-computer interaction device provided by an embodiment of the present invention. The large-scale model-based human-computer interaction device is applied to a human-computer question-answering system on an industrial production line, and includes:

[0107] The question acquisition module 401 is used to obtain the question input by the user and determine whether there is matching information for the question in the system knowledge base; if so, output the corresponding answer based on the matching information; if not, enter the scenario question and answer mode;

[0108] The question classification module 402 is used to perform structural analysis on the text of the question using a large model in the scenario question-answering mode to obtain the question type of the question; wherein the question type includes step type, object attribute type, object definition type, verification type, and other types;

[0109] The answer output module 403 is used to obtain the visual information collected by the camera and the pre-built knowledge base according to the question type and the scene selected by the user, and output the answer corresponding to the question based on the large model; wherein, when the question type is other types, the large model is first used to answer the question, and then the guidance information and recommendation information related to the knowledge base are output.

[0110] Preferably, the knowledge base includes a local knowledge base and an external knowledge base; wherein, the local knowledge base includes a scene-specific knowledge base, a large model prompt base and the system knowledge base; the scene-specific knowledge base includes a visual knowledge base and a behavioral knowledge base.

[0111] Preferably, in the question acquisition module 401, the step of acquiring the question input by the user and determining whether there is matching information for the question in the system knowledge base specifically includes:

[0112] Obtaining a question input by the user;

[0113] Preprocessing the text of the question to obtain first parsing information;

[0114] Performing entity recognition and keyword extraction on the first parsed information to obtain second parsed information;

[0115] Converting the second parsed information into a vector form to obtain third parsed information;

[0116] The third parsed information is matched with information in the system knowledge base, and it is determined whether there is matching information for the problem in the system knowledge base.

[0117] Preferably, the large model-based human-computer interaction device further includes:

[0118] The scenario selection module 404 is used to determine whether the user has selected a scenario based on historical question and answer information after entering the scenario question and answer mode; if the user has not selected a scenario, output guidance information for scenario selection to enable the user to select a scene; if the user has selected a scene, use the scenario selected by the user as the application scenario for the current question.

[0119] Preferably, in the answer output module 403, the visual information collected by the camera and the pre-built knowledge base are obtained according to the question type and the scene selected by the user, and the answer corresponding to the question is output based on the large model, specifically including:

[0120] When the question type is the step type, obtaining information in the behavior knowledge base and prompt information related to the question in the large model prompt library according to the scenario selected by the user;

[0121] Acquiring first visual information collected by the camera;

[0122] Inputting the prompt information and the first visual information into the large model to obtain first text description information corresponding to the first visual information;

[0123] The first text description information is matched with the information in the behavior knowledge base to determine the user's current operation steps and output an answer corresponding to the question.

[0124] Preferably, in the answer output module 403, the visual information collected by the camera and the pre-built knowledge base are obtained according to the question type and the scene selected by the user, and the answer corresponding to the question is output based on the large model, specifically including:

[0125] When the question type is the object attribute type or the object definition type, obtaining object information in the visual knowledge base according to the scene selected by the user;

[0126] If the question type is the object attribute type, the large model is used to extract object keywords from the text of the question; if the question type is the object definition type, second visual information captured by the camera is obtained, the large model is used to convert the second visual information into second text description information, and object keywords related to the question are extracted from the second text description information;

[0127] The object keywords are matched with the object information in the visual knowledge base, and the answer corresponding to the question is output according to the matching result.

[0128] Preferably, in the answer output module 403, the visual information collected by the camera and the pre-built knowledge base are obtained according to the question type and the scene selected by the user, and the answer corresponding to the question is output based on the large model, specifically including:

[0129] When the question type is the verification type, obtaining information in the behavior knowledge base according to the scenario selected by the user;

[0130] Acquiring third visual information collected by the camera; wherein the third visual information is a video saved in segments according to the user's operation frequency;

[0131] Extracting key frames from the video, and inputting the obtained key frames into the large model to obtain third text description information corresponding to each key frame;

[0132] The third text description information is matched with the information in the behavior knowledge base to determine the correctness of the user operation and output an answer corresponding to the question.

[0133] Preferably, the large model-based human-computer interaction device further includes:

[0134] The video acquisition module 405 is used to acquire the industrial standard video uploaded by professionals;

[0135] The behavior recognition module 406 is used to extract key frames of the industrial standard video and perform behavior recognition on the key frames using the large model to obtain behavior descriptions corresponding to the key frames;

[0136] An expert judgment module 407 is configured to output the behavior description so that the professional can judge whether the behavior description complies with the specification;

[0137] The information update module 408 is used to obtain the judgment result of the professional. If the behavior description does not meet the specifications, it returns to the step of extracting the key frames of the industrial standard video; if the behavior description meets the specifications, it updates the behavior description to the scene-specific knowledge base.

[0138] In the specific implementation, the working principle, control process and technical effect achieved by the large model-based human-computer interaction device provided in the embodiment of the present invention are the same as those of the large model-based human-computer interaction method in the above embodiment, and will not be repeated here.

[0139] See also Figure 5 , Figure 5 5 is a schematic diagram of the structure of a terminal device provided in an embodiment of the present invention. The terminal device includes a processor 501, a memory 502, and a computer program stored in the memory 502 and configured to be executed by the processor 501. When the processor 501 executes the computer program, it implements the large model-based human-computer interaction method described in any of the above embodiments.

[0140] Preferably, the computer program can be divided into one or more modules / units (e.g., computer program 1, computer program 2, ...), which are stored in the memory 502 and executed by the processor 501 to implement the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0141] The processor 501 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor, or the processor 501 can be any conventional processor. The processor 501 is the control center of the terminal device, and uses various interfaces and lines to connect the various parts of the terminal device.

[0142] The memory 502 mainly includes a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function, and the data storage area can store related data. In addition, the memory 502 can be a high-speed random access memory or a non-volatile memory, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, and a Flash Card. Alternatively, the memory 502 can be other volatile solid-state memory devices.

[0143] It should be noted that the above terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 5The structural diagram is only an example of the above-mentioned terminal device and does not constitute a limitation on the above-mentioned terminal device. It may include more or fewer components than shown in the figure, or combine certain components, or different components.

[0144] An embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the large model-based human-computer interaction method described in any of the above embodiments.

[0145] The embodiment of the present invention provides a human-computer interaction method, device and medium based on a large model. By adopting large model technology, it can better integrate and refine information, and handle multilingual and multicultural problems. It also enables the knowledge base to have the ability to continuously expand and update, enhances the scalability of the system, and can provide long-lasting and effective information support; by classifying question types in detail, it can provide accurate and professional answers for different categories of questions, enhancing the technical effect of the system in specific fields; through the question guidance mechanism, it enhances the user's interactive experience and ensures the relevance and accuracy of information; through the selection of specific scenarios, it can make the system's answer faster and more efficient; by introducing visual information and combining video, graphics and text content, the question-answering system can more accurately assess the needs of learners and provide more personalized and intuitive learning guidance, thereby improving educational effects and enhancing user experience. The embodiment of the present invention can provide comprehensive, accurate and professional answers in human-computer interactive question-answering in industrial production lines, greatly enhancing the user's interactive experience.

[0146] It should be noted that the system embodiment described above is merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the system embodiment provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive work.

[0147] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A large-scale model-based human-computer interaction method, applied to a human-computer question-answering system on an industrial production line, characterized in that: include: Obtain the question input by the user and determine whether there is matching information for the question in the system knowledge base; If so, output a corresponding answer based on the matching information; If not, enter the scenario question and answer mode; In the scenario question-answering mode, a large model is used to perform structured analysis on the text of the question to obtain the question type of the question; wherein the question type includes step type, object attribute type, object definition type, verification type and other types; Based on the question type and the scenario selected by the user, visual information collected by the camera and a pre-built knowledge base are obtained, and an answer corresponding to the question is output based on the large model. If the question type is other than the large model, the large model is first used to answer the question, and then guidance information and recommended information related to the knowledge base are output; The knowledge base includes a local knowledge base and an external knowledge base; the local knowledge base includes a scene-specific knowledge base, a large model prompt base, and the system knowledge base; the scene-specific knowledge base includes a visual knowledge base and a behavioral knowledge base; The method of obtaining visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputting the answer corresponding to the question based on the large model, specifically includes: When the question type is the object attribute type or the object definition type, obtaining object information in the visual knowledge base according to the scene selected by the user; If the question type is the object attribute type, the large model is used to extract object keywords from the text of the question; if the question type is the object definition type, second visual information captured by the camera is obtained, the large model is used to convert the second visual information into second text description information, and object keywords related to the question are extracted from the second text description information; Matching the object keyword with the object information in the visual knowledge base, and outputting an answer corresponding to the question based on the matching result; The method of obtaining visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputting the answer corresponding to the question based on the large model, specifically includes: When the question type is the verification type, obtaining information in the behavior knowledge base according to the scenario selected by the user; Acquiring third visual information collected by the camera; wherein the third visual information is a video saved in segments according to the user's operation frequency; Extracting key frames from the video, and inputting the obtained key frames into the large model to obtain third text description information corresponding to each key frame; The third text description information is matched with the information in the behavior knowledge base to determine the correctness of the user operation and output an answer corresponding to the question.

2. The human-computer interaction method based on a large model according to claim 1, characterized in that: The step of obtaining the question input by the user and determining whether there is matching information for the question in the system knowledge base specifically includes: Obtaining a question input by the user; Preprocessing the text of the question to obtain first parsing information; Performing entity recognition and keyword extraction on the first parsed information to obtain second parsed information; Converting the second parsed information into a vector form to obtain third parsed information; The third parsed information is matched with information in the system knowledge base, and it is determined whether there is matching information for the problem in the system knowledge base.

3. The human-computer interaction method based on a large model according to claim 1, characterized in that: After entering the scenario question-and-answer mode, the method further includes: Determining whether the user has selected a scene based on historical question and answer information; If the user has not selected a scene, outputting scene selection guidance information to enable the user to select a scene; If the user has selected a scenario, the scenario selected by the user is used as the application scenario of the current question.

4. The human-computer interaction method based on a large model according to claim 1, characterized in that: The method of obtaining visual information collected by the camera and a pre-built knowledge base based on the question type and the scenario selected by the user, and outputting an answer corresponding to the question based on the large model, specifically includes: When the question type is the step type, obtaining information in the behavior knowledge base and prompt information related to the question in the large model prompt library according to the scenario selected by the user; Acquiring first visual information collected by the camera; Inputting the prompt information and the first visual information into the large model to obtain first text description information corresponding to the first visual information; The first text description information is matched with the information in the behavior knowledge base to determine the user's current operation steps and output an answer corresponding to the question.

5. The human-computer interaction method based on a large model according to claim 1, characterized in that: The method further comprises: Get industry-standard videos uploaded by professionals; Extracting key frames of the industrial standard video, and performing behavior recognition on the key frames using the large model to obtain behavior descriptions corresponding to the key frames; Outputting the behavior description so that the professional can judge whether the behavior description complies with the specification; Obtain the judgment result of the professional. If the behavior description does not meet the specifications, return to the step of extracting the key frames of the industrial specification video; if the behavior description meets the specifications, update the behavior description to the scene-specific knowledge base.

6. A large-scale model-based human-computer interaction device, applied to a human-computer question-answering system on an industrial production line, characterized in that: include: The question acquisition module is used to obtain the question input by the user and determine whether there is matching information for the question in the system knowledge base; If yes, then output the corresponding answer according to the matching information; if no, then enter the scenario question and answer mode; A question classification module is used to perform structured analysis on the text of the question using a large model in the scenario question-answering mode to obtain the question type of the question; wherein the question type includes step type, object attribute type, object definition type, verification type and other types; an answer output module, configured to obtain visual information captured by the camera and a pre-built knowledge base based on the question type and the scenario selected by the user, and output an answer corresponding to the question based on the large model; wherein, when the question type is other than the large model, the large model is first used to answer the question, and then guidance information and recommended information related to the knowledge base are output; The knowledge base includes a local knowledge base and an external knowledge base; the local knowledge base includes a scene-specific knowledge base, a large model prompt base, and the system knowledge base; the scene-specific knowledge base includes a visual knowledge base and a behavioral knowledge base; The method of obtaining visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputting the answer corresponding to the question based on the large model, specifically includes: When the question type is the object attribute type or the object definition type, obtaining object information in the visual knowledge base according to the scene selected by the user; If the question type is the object attribute type, the large model is used to extract object keywords from the text of the question; if the question type is the object definition type, second visual information captured by the camera is obtained, the large model is used to convert the second visual information into second text description information, and object keywords related to the question are extracted from the second text description information; Matching the object keyword with the object information in the visual knowledge base, and outputting an answer corresponding to the question based on the matching result; The method of obtaining visual information collected by the camera and a pre-built knowledge base according to the question type and the scene selected by the user, and outputting the answer corresponding to the question based on the large model, specifically includes: When the question type is the verification type, obtaining information in the behavior knowledge base according to the scenario selected by the user; Acquiring third visual information collected by the camera; wherein the third visual information is a video saved in segments according to the user's operation frequency; Extracting key frames from the video, and inputting the obtained key frames into the large model to obtain third text description information corresponding to each key frame; The third text description information is matched with the information in the behavior knowledge base to determine the correctness of the user operation and output an answer corresponding to the question.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the device where the computer-readable storage medium is located executes the computer program, the large model-based human-computer interaction method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Target tracking method and device based on time sequence prediction

    CN110827320A

  • Method and system for optimizing man-machine conversation based on LLM model

    CN117271735A