AI interaction method and system, electronic equipment and storage medium
By using pre-built large language model and proprietary sub-models, combining context understanding and historical interaction information, and generating feedback content, the problem of insufficient intelligence of AI interaction is solved, and a more intelligent and immersive interactive experience is achieved.
Patent Information
- Application Number
- CN202411990933.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-09
Smart Images

Figure CN119962669A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to AI interaction methods, systems, electronic devices and storage media. Background Art
[0002] With the widespread application of large language models (LLMs) in areas such as dialogue generation and entertainment interaction, home AI devices have gradually become an important part of users' daily lives.
[0003] In related technologies, a single AI model cannot simultaneously meet users' needs in entertainment, knowledge questions and answers, personalized recommendations, etc., and when multiple AI systems switch AI models or services, they often cannot inherit previous interaction memories, resulting in a gap in experience and insufficiently intelligent interaction. In addition, traditional remote control interaction methods lack immersion and are difficult to meet users' entertainment needs.
[0004] Currently, no effective solution has been proposed to address the problem of insufficient intelligence in AI interactions in related technologies. Summary of the invention
[0005] The embodiments of the present application provide an AI interaction method, system, electronic device, and storage medium to at least solve the problem of insufficient intelligence of AI interaction in related technologies.
[0006] In a first aspect, an embodiment of the present application provides an AI interaction method, the method comprising:
[0007] Acquire user voice and convert the user voice into text information;
[0008] The pre-built large language model determines the interaction type corresponding to the text information through contextual understanding functions and historical interaction information;
[0009] The proprietary sub-model corresponding to the interaction type is called to generate feedback content of the text information, and the feedback content is output to the user, wherein the large language model includes at least two proprietary sub-modules.
[0010] In some embodiments, the pre-built large language model determines the interaction type corresponding to the text information through context understanding function and historical interaction information, including:
[0011] Acquiring the historical interaction information from a memory management module;
[0012] Performing intent recognition on the text information to obtain key information;
[0013] The large language model analyzes the key information through the context understanding function and the historical interaction information to determine the corresponding interaction type.
[0014] In some embodiments, calling the dedicated sub-model corresponding to the interaction type to generate the feedback result of the text information includes:
[0015] Calling a dedicated sub-model corresponding to the interaction type, wherein each interaction type is pre-associated with one or more dedicated sub-models;
[0016] Analyzing the key information based on the proprietary sub-model and the historical interaction information to obtain an analysis result;
[0017] Generate feedback content according to the analysis result.
[0018] In some embodiments, generating feedback content according to the analysis result includes:
[0019] In the case where the interaction type corresponds to a dedicated submodule, directly generating feedback content based on the analysis result;
[0020] In the case where the interaction type corresponds to at least two proprietary sub-modules, the analysis results output by each of the proprietary sub-modules are integrated, redundant information is removed, and feedback content is generated.
[0021] In some embodiments, the method further comprises:
[0022] The feedback content is sent to the memory management module for storage.
[0023] In some embodiments, outputting the feedback content to the user includes:
[0024] Display the feedback content through an interactive large screen; and / or
[0025] Based on the speech synthesis algorithm, the feedback content is converted into speech information and played.
[0026] In some embodiments, the method further comprises:
[0027] Receiving a user's touch command based on the interactive large screen to control the interaction process or input the interaction content; and / or
[0028] The user's gestures are recognized through the projector and the interactive large screen to control the interactive process or input the interactive content.
[0029] In a second aspect, an embodiment of the present application provides an AI interaction system, the system comprising:
[0030] An acquisition module, used for acquiring user voice and converting the user voice into text information;
[0031] A classification module, which is used for the pre-built large language model to determine the interaction type corresponding to the text information through context understanding function and historical interaction information;
[0032] A feedback module is used to call the proprietary sub-model corresponding to the interaction type to generate feedback content of the text information and output the feedback content to the user, wherein the large language model includes at least two proprietary sub-modules.
[0033] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the AI interaction method as described in the first aspect above is implemented.
[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the AI interaction method as described in the first aspect above.
[0035] Compared with the related art, the AI interaction method provided in the embodiment of the present application obtains user voice and converts the user voice into text information. The pre-built large language model uses the context understanding function and historical interaction information to determine the interaction type corresponding to the text information, and calls the proprietary sub-model corresponding to the interaction type to generate feedback content of the text information, and outputs the feedback content to the user. The large language model includes at least two proprietary sub-modules, which solves the problem of insufficient intelligence of AI interaction. The large language model analyzes the input content through the context understanding function and historical interaction information, combined with multiple different sub-models, improves the analysis efficiency and the accuracy of the analysis results, thereby improving the intelligence of AI interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0037] Figure 1 is a flow chart of an AI interaction method according to an embodiment of the present application;
[0038] Figure 2 is a flow chart of an AI interaction method according to an embodiment of the present application;
[0039] Figure 3 is a schematic diagram of information processing according to an embodiment of the present application;
[0040] Figure 4 is a structural block diagram of an AI interaction system according to an embodiment of the present application;
[0041] Figure 5 It is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0043] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.
[0044] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0045] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0046] This embodiment provides an AI interaction method. Figure 1 is a flow chart of an AI interaction method according to an embodiment of the present application, such as Figure 1 As shown, the process includes the following steps:
[0047] Step S101, obtaining user voice, and converting the user voice into text information.
[0048] The user's voice input supports both far-field and near-field modes. The input voice is converted into text and then processed by a pre-built large language model.
[0049] Step S102: The pre-built large language model determines the interaction type corresponding to the text information through context understanding functions and historical interaction information.
[0050] The large language model includes multiple proprietary models, which understand user input through context understanding functions and historical interaction information, ensuring that there is no "disconnection" when switching between multiple models. It can also dynamically adjust memory to provide a coherent conversation experience.
[0051] In some embodiments, step S102 specifically includes:
[0052] Step S1021, obtaining historical interaction information from the memory management module.
[0053] Step S1022, perform intent recognition on the text information to obtain key information.
[0054] Step S1023: The large language model analyzes key information through context understanding functions and historical interaction information to determine the corresponding interaction type.
[0055] After the big model receives the text information corresponding to the user's input content, it will first perform intent recognition on the text information through the master language model (Master LLM), convert it into structured data, and extract core information (such as question type, keywords, context dependencies). The interaction type is determined based on the extracted core information, and the dedicated model is called to process it based on the interaction type.
[0056] Step S103, calling a specific sub-model corresponding to the interaction type to generate feedback content of the text information, and output the feedback content to the user, wherein the large language model includes at least two specific sub-modules.
[0057] Classification is performed based on the intent recognition results, and specific questions are distributed to proprietary models for processing based on the classification results. Proprietary models include but are not limited to models responsible for knowledge question answering, models focusing on entertainment interaction, and models for implementing personalized recommendations.
[0058] In some embodiments, the dedicated sub-model corresponding to the interaction type is called in step S103 to generate a feedback result of text information, including:
[0059] Step S1031 , calling a dedicated sub-model corresponding to the interaction type, wherein each interaction type is pre-associated with one or more dedicated sub-models.
[0060] Step S1032: Analyze the key information based on the proprietary sub-model and the historical interaction information to obtain analysis results.
[0061] Step S1033: generating feedback content according to the analysis result.
[0062] Each interaction type is pre-associated to one or more dedicated models. The most suitable model is selected first according to the task complexity and model adaptability. At the same time, for multiple models with redundant functions, tasks can be dynamically allocated according to the current system load.
[0063] In some embodiments, step S1033 specifically includes:
[0064] Step S201 : when the interaction type corresponds to a dedicated submodule, feedback content is directly generated based on the analysis result.
[0065] Step S202: When the interaction type corresponds to at least two proprietary sub-modules, the analysis results output by each proprietary sub-module are integrated to remove redundant information and generate feedback content.
[0066] For tasks completed by multiple models in collaboration, the memory management module can be used to integrate the subtask results, remove redundant information, and ensure unified output. Figure 2 It is a flowchart of an AI interaction method according to an embodiment of the present application.
[0067] For example, the user input is: "Recommend me a good book and tell me the author's background". The input content is identified and analyzed, and the corresponding interaction type is determined to be recommendation and knowledge question answering. The personalized recommendation model and the knowledge question answering model are called for processing respectively. The personalized recommendation model is responsible for book recommendations, and the knowledge question answering model is responsible for providing the author's background. The results of both are displayed to the user in a unified manner, and optionally, combined with speech synthesis and projection functions for display.
[0068] Through the above steps, the user's voice is obtained and converted into text information. The pre-built large language model uses the context understanding function and historical interaction information to determine the interaction type corresponding to the text information, and calls the proprietary sub-model corresponding to the interaction type to generate feedback content of the text information and output the feedback content to the user. The large language model includes at least two proprietary sub-modules, which solves the problem of insufficient intelligence of AI interaction. The large language model analyzes the input content through the context understanding function and historical interaction information, combined with multiple different sub-models, improves the analysis efficiency and the accuracy of the analysis results, thereby improving the intelligence of AI interaction.
[0069] In some embodiments, the method further comprises:
[0070] Step S104, sending the feedback content to the memory management module for storage.
[0071] Figure 3 is a schematic diagram of information processing according to an embodiment of the present application, such as Figure 2 As shown in the figure, before the user has a conversation, the big model retrieves the historical interaction information from the memory management model, and the memory management model fills the historical interaction information into the big model to be called; when the user asks a question to the big model, the big model analyzes the historical interaction information, determines the required proprietary model from model A, model B and model C, and generates an answer based on the proprietary model. After the conversation ends, the conversation message continues to be saved in the memory management model.
[0072] With the memory scheduling function, users can switch between different tasks at any time without worrying about the context being interrupted. Whether it is used in home assistant, learning and education, or daily entertainment and smart home control, it can bring a smoother and smarter experience.
[0073] In some embodiments, outputting feedback content to the user in step S103 includes:
[0074] Step S1034, displaying the feedback content through the interactive large screen.
[0075] Step S1035: Based on the speech synthesis algorithm, the feedback content is converted into speech information and played.
[0076] The feedback results can be displayed on the interactive large screen in the form of text, or broadcast through voice output.
[0077] In some embodiments, the method further comprises:
[0078] Step S105, receiving the user's touch command based on the interactive large screen to control the interactive process or input interactive content.
[0079] Step S106, identifying user gestures through the projector and the interactive large screen to control the interaction process or input the interaction content.
[0080] The interactive large screen can not only be used to display results, but also supports gesture or touch operations to increase the entertainment and fun of the interaction. Through the image acquisition device of the projector or the interactive large screen, the user's gesture instructions are recognized, for example, the feedback content is slid up and down or turned, and the voice information of the feedback content is fast-forwarded or paused. In addition, the interactive content can be input by identifying the user's gesture content. For example, a user's gesture represents the interactive interface that wakes up AI; or, through the recognition of gestures, the content to be processed is input, for example, the position of the user's gesture in space is collected to realize typing in the virtual keyboard of the interactive interface to complete the input of the interactive content.
[0081] The above method combines the capabilities of large language models and proprietary models through memory management, multi-model collaboration and immersive interaction, thereby improving the intelligence and efficiency of AI interaction.
[0082] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0083] This embodiment also provides an AI interactive system, which is used to implement the above embodiments and preferred implementations, and will not be repeated here. As used below, the terms "module", "unit", "subunit", etc. can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0084] Figure 4 is a structural block diagram of an AI interaction system according to an embodiment of the present application, such as Figure 4 As shown, the system includes:
[0085] The acquisition module 41 is used to acquire the user's voice and convert the user's voice into text information.
[0086] The classification module 42 is used to use the pre-built large language model to understand the function and historical interaction information through context, and determine the interaction type corresponding to the text information.
[0087] The feedback module 43 is used to call the dedicated sub-model corresponding to the interaction type to generate feedback content of text information and output the feedback content to the user, wherein the large language model includes at least two dedicated sub-modules.
[0088] In some embodiments, the classification module 42 includes:
[0089] A history acquisition module is used to obtain historical interaction information from the memory management module;
[0090] The recognition module is used to identify the intent of text information and obtain key information;
[0091] The type determination module is used by the large language model to understand functions and historical interaction information through context, analyze key information, and determine the corresponding interaction type.
[0092] In some embodiments, the feedback module 43 includes:
[0093] The model calling module is used to call the dedicated sub-model corresponding to the interaction type, wherein each interaction type is pre-associated with one or more dedicated sub-models.
[0094] The analysis module is used to analyze key information based on the proprietary sub-model and historical interaction information to obtain analysis results.
[0095] The result generation module is used to generate feedback content according to the analysis results.
[0096] In some embodiments, the result generation module includes:
[0097] The first result generating module is used to directly generate feedback content based on the analysis result when the interaction type corresponds to a dedicated sub-module.
[0098] The second result generating module is used to integrate the analysis results output by each proprietary submodule, remove redundant information, and generate feedback content when the interaction type corresponds to at least two proprietary submodules.
[0099] In some embodiments, the system further includes: an interactive storage module, configured to send the feedback content to the memory management module for storage.
[0100] In some embodiments, the feedback module 43 includes:
[0101] The first result display module is used to display the feedback content through an interactive large screen; and / or
[0102] The second result display module is used to convert the feedback content into voice information based on the speech synthesis algorithm and play it.
[0103] In some embodiments, the system further comprises:
[0104] A touch control module, used to receive a user's touch control command based on the interactive large screen to control the interactive process or input interactive content; and / or
[0105] The gesture recognition module is used to recognize user gestures through the projector and interactive large screen to control the interaction process or input the interaction content.
[0106] Through the above system, the acquisition module 41 acquires the user's voice and converts the user's voice into text information. The large language model pre-built by the classification module 42 determines the interaction type corresponding to the text information through the context understanding function and historical interaction information. The feedback module 43 calls the proprietary sub-model corresponding to the interaction type to generate feedback content of the text information and output the feedback content to the user, wherein the large language model includes at least two proprietary sub-modules. , which solves the problem of insufficient intelligence of AI interaction. The large language model analyzes the input content through the context understanding function and historical interaction information, combined with multiple different sub-models, to improve the analysis efficiency and accuracy of the analysis results.
[0107] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0108] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0109] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0110] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0111] S1, obtaining user voice and converting the user voice into text information.
[0112] S2, the pre-built large language model understands the function and historical interaction information through context, and determines the interaction type corresponding to the text information.
[0113] S3, calling a proprietary sub-model corresponding to the interaction type to generate feedback content of text information, and output the feedback content to the user, wherein the large language model includes at least two proprietary sub-modules.
[0114] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0115] In one embodiment, Figure 5 is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application, such as Figure 5 As shown, an electronic device is provided, which may be a server, and its internal structure diagram may be as shown in Figure 5 As shown. The electronic device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store data. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an AI interaction method is implemented.
[0116] Those skilled in the art will understand that Figure 5 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.
[0117] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0118] Those skilled in the art should understand that the technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0119] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. An AI interaction method, characterized in that: The method comprises: Acquire user voice and convert the user voice into text information; The pre-built large language model determines the interaction type corresponding to the text information through contextual understanding functions and historical interaction information; The proprietary sub-model corresponding to the interaction type is called to generate feedback content of the text information, and the feedback content is output to the user, wherein the large language model includes at least two proprietary sub-modules.
2. The method according to claim 1, characterized in that The pre-built large language model determines the interaction type corresponding to the text information through context understanding function and historical interaction information, including: Acquiring the historical interaction information from a memory management module; Performing intent recognition on the text information to obtain key information; The large language model analyzes the key information through the context understanding function and the historical interaction information to determine the corresponding interaction type.
3. The method according to claim 2, characterized in that The calling of the proprietary sub-model corresponding to the interaction type to generate the feedback result of the text information includes: Calling a dedicated sub-model corresponding to the interaction type, wherein each interaction type is pre-associated with one or more dedicated sub-models; Analyzing the key information based on the proprietary sub-model and the historical interaction information to obtain an analysis result; Generate feedback content according to the analysis result.
4. The method according to claim 3, characterized in that Generating feedback content according to the analysis result includes: In the case where the interaction type corresponds to a dedicated submodule, directly generating feedback content based on the analysis result; In the case where the interaction type corresponds to at least two proprietary sub-modules, the analysis results output by each of the proprietary sub-modules are integrated, redundant information is removed, and feedback content is generated.
5. The method according to claim 2, characterized in that: The method further comprises: The feedback content is sent to the memory management module for storage.
6. The method according to claim 1, characterized in that The outputting the feedback content to the user comprises: Display the feedback content through an interactive large screen; and / or Based on the speech synthesis algorithm, the feedback content is converted into speech information and played.
7. The method according to claim 6, characterized in that The method further comprises: Receiving a user's touch command based on the interactive large screen to control the interaction process or input the interaction content; and / or The user's gestures are recognized through the projector and the interactive large screen to control the interactive process or input the interactive content.
8. An AI interactive system, characterized in that: The system comprises: An acquisition module, used for acquiring user voice and converting the user voice into text information; A classification module, which is used for the pre-built large language model to determine the interaction type corresponding to the text information through context understanding function and historical interaction information; A feedback module is used to call the proprietary sub-model corresponding to the interaction type to generate feedback content of the text information and output the feedback content to the user, wherein the large language model includes at least two proprietary sub-modules.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the AI interaction method according to any one of claims 1 to 7 is implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the AI interaction method according to any one of claims 1 to 7 is implemented.