A human-computer interaction method, device and equipment

By acquiring user questions and operation records, and combining them with artificial intelligence processing, the problem of accurately locating user doubts during the reading process was solved, resulting in more efficient auxiliary answers.

CN122432282APending Publication Date: 2026-07-21BEIJING KEYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING KEYI TECH CO LTD
Filing Date
2026-04-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing auxiliary tools lack the ability to accurately identify user doubts during the reading process, resulting in auxiliary answers that are out of touch with the actual reading context, affecting their relevance and accuracy.

Method used

By acquiring user questions and operation records within a preset time period, and processing them using an artificial intelligence server, combined with the user's emotional information and operational behavior, the system accurately identifies the user's doubts and provides answers in voice and/or text form.

Benefits of technology

It improves the relevance and accuracy of auxiliary answers, allowing users to quickly understand and locate the problem, thus improving the speed and quality of problem-solving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432282A_ABST
    Figure CN122432282A_ABST
Patent Text Reader

Abstract

The application discloses a human-computer interaction method, device and equipment, so that the auxiliary tool can more accurately understand the user's doubts and improve the pertinence and accuracy of the auxiliary answer. The method comprises the following steps: acquiring question information, wherein the question information comprises a question raised by a user and a time when the question is raised; when it is determined that relevant context information needs to be extracted based on the question, extracting operation records in a preset time period from pre-cached context operation records, wherein the operation records comprise operation events and time when the operation events occur, and the operation events comprise image data of operation behaviors and / or operation areas; uploading the question information and the operation records in the preset time period to an artificial intelligence server for processing, and receiving a processing result returned by the artificial intelligence server for the question; and displaying the processing result in the form of voice and / or text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a human-computer interaction method, apparatus, and device. Background Technology

[0002] In digital office and learning scenarios, desktop assistants and document understanding tools have become important means to improve users' information processing efficiency. These tools typically rely on technologies such as natural language processing and large language models to provide corresponding explanations, summaries, translations, or knowledge expansions based on the user's input.

[0003] However, in real-world scenarios where users seek help and receive comprehension assistance, their confusion often arises during the reading, reviewing, or previewing of documents. Existing assistance tools typically only acquire the user's input question text, which results in a lack of accurate judgment of the user's doubts. The generated answers are often detached from the user's actual reading context, making it difficult to accurately pinpoint the user's confusion and provide the most reasonable explanation and help, thus affecting the relevance and accuracy of the assistance responses. Summary of the Invention

[0004] The purpose of this application is to provide a human-computer interaction method, apparatus, and device to enable auxiliary tools to more accurately understand user questions and improve the relevance and accuracy of auxiliary answers.

[0005] In a first aspect, embodiments of this application provide a human-computer interaction method, the method comprising: Obtain question information, which includes the question raised by the user and the time the question was raised; When it is determined that relevant context information needs to be extracted based on the problem, operation records within a preset time period are extracted from the pre-cached context operation records. The operation records include operation events and the time when the operation events occur. The operation events include operation behaviors and / or image data of the operation area. The question information and the operation records within the preset time period are uploaded to the artificial intelligence server for processing, and the processing result for the question is returned by the artificial intelligence server. The processing results are displayed in the form of voice and / or text.

[0006] In the human-computer interaction method provided in this application embodiment, after obtaining the question information, when it is determined that relevant context information needs to be extracted based on the question, operation records within a preset time period are extracted from the pre-cached context operation records. The question information and the operation records within the preset time period are uploaded to an artificial intelligence server for processing, and the processing result returned by the artificial intelligence server is received. The processing result is then displayed in the form of voice and / or text. Compared with related technologies that only obtain the question text input by the user, by pre-caching the context operation records and uploading the question information and the operation records within the preset time period to the artificial intelligence server for processing, the artificial intelligence server can accurately locate the content that the user is currently processing when processing the user's question, combining the operation records within the preset time period, and more accurately understand the user's doubts, thereby improving the pertinence and accuracy of the assisted answer.

[0007] In addition, in the human-computer interaction method provided in this application embodiment, the processing results of the artificial intelligence server are displayed in the form of voice and / or text. Displaying the processing results in the form of voice makes it easier for users to quickly understand the processing results without interrupting the user's operation; while displaying the processing results in the form of text makes it easier for users to quickly locate and understand the original content, and also makes it easier for users to save, thereby significantly improving the speed and quality of users' problem processing.

[0008] In some possible embodiments, uploading the question information and the operation records within the preset time period to an artificial intelligence server for processing includes: Based on pre-set evaluation rules, calculate the correlation evaluation value between each operation event in the operation record and the problem within the preset time period; Select the target operation event whose relevance evaluation value meets the preset requirements, and upload the question information and the operation record corresponding to the target operation event to the artificial intelligence server for processing.

[0009] In the human-computer interaction method provided in this application embodiment, when uploading the question information and the operation records within a preset time period to the artificial intelligence server for processing, the correlation evaluation value between each operation event and the question in the operation records within the preset time period is first calculated based on the preset evaluation rules. Then, the target operation event whose correlation evaluation value meets the preset requirements is selected, and the question information and the operation records corresponding to the target operation event are uploaded to the artificial intelligence server for processing. This reduces the amount of data uploaded to the server, significantly reduces bandwidth and computing power costs, and reduces the risk of sensitive information leakage.

[0010] In some possible embodiments, calculating the correlation evaluation value between each operation event in the operation record within the preset time period and the problem based on pre-set evaluation rules includes: Centered on the time when the problem was raised, the operation events are sorted according to the occurrence time of each operation event in the operation record within the preset time period to obtain the sorting result; Based on the pre-set evaluation rules and the ranking results, the correlation evaluation value between each operation event in the ranking results and the problem is calculated.

[0011] In some possible embodiments, the method further includes: When the user's operation behavior while browsing content on the terminal device meets the preset conditions, the recording frequency of the operation event is increased and / or the recording range of the image data of the operation area is increased within the preset time window, and the preset time window is marked as a question window.

[0012] In some possible embodiments, uploading the question information and the operation records within the preset time period to an artificial intelligence server for processing includes: If a question window exists before the time the question is asked, the operation records contained in the question window are merged with the operation records within the preset time period to obtain an operation record set. The question information and the operation record set are then uploaded to the artificial intelligence server for processing.

[0013] The human-computer interaction method provided in this application embodiment indicates that the content being processed by the user may be content of doubt when the user's operation behavior while browsing content on a terminal device meets preset conditions. Therefore, the recording frequency of operation events and / or the recording range of image data in the operation area are increased within a preset time window, and the preset time window is marked as a doubt window.

[0014] Thus, when the question information and operation records within a preset time period are uploaded to the artificial intelligence server for processing, if it is determined that a doubt window existed before the question was asked, the operation records contained in the doubt window are merged with the operation records within the preset time period to obtain an operation record set. The question information and operation record set are then uploaded to the artificial intelligence server for processing. This allows the artificial intelligence server to more accurately understand and locate the user's doubt when processing the user's question, by combining the operation records in the doubt window, thereby improving the pertinence and accuracy of the assisted answers.

[0015] In some possible embodiments, the method further includes: Record the user's emotional information, which includes the user's emotional state, the time when the emotional state occurred, and the duration of the emotional state. The step of uploading the question information and the operation records within the preset time period to the artificial intelligence server for processing includes: Filter emotional information generated within the preset time period to generate an emotional information set; The question information, the operation records within the preset time period, and the set of emotion information are uploaded to the artificial intelligence server for processing.

[0016] In the human-computer interaction method provided in this application embodiment, when uploading the question information and the operation record within a preset time period to the artificial intelligence server for processing, the user's emotional state, the time of occurrence of the emotional state, and the duration of the emotional state are also uploaded to the artificial intelligence server for processing simultaneously. In this way, when processing the questions raised by the user, the user's emotional state can be taken into account, making it easier to hit the cause of the user's confusion and the scope of their concern, reducing repeated questioning and secondary positioning.

[0017] In some possible embodiments, uploading the question information, the operation records within the preset time period, and the set of emotion information to the artificial intelligence server for processing includes: The question information, the operation records within the preset time period, and the emotional information set are sorted in chronological order of occurrence to generate aligned contextual information. The aligned context information is uploaded to the artificial intelligence server for processing.

[0018] In some possible embodiments, recording the user's emotional information includes: Obtain video data containing the user's facial information; Based on the video data, the user's emotional state is determined using a pre-configured analysis model, and the time and duration of the emotional state are determined.

[0019] In some possible embodiments, after obtaining the query information and before retrieving the operation records within a preset time period from the pre-cached context operation records, the method further includes: The problem is semantically identified using a pre-configured large language model, and the identification results are generated. When determining that the question contains an indicator pronoun based on the recognition result, it is necessary to extract relevant contextual information to determine the question.

[0020] In some possible embodiments, after obtaining the query information and before retrieving the operation records within a preset time period from the pre-cached context operation records, the method further includes: When it is determined that no relevant contextual information needs to be extracted based on the question, the question information is uploaded to the artificial intelligence server, and the processing result for the question is returned by the artificial intelligence server, and the processing result is displayed in the form of voice and / or text.

[0021] In some possible embodiments, the pre-cached context operation record is triggered by the user performing a preset operation when browsing content on the terminal device.

[0022] In the human-computer interaction method provided in this application embodiment, the pre-cached context operation record is triggered by the user performing preset operation behavior when browsing content on the terminal device, and does not record all of the user's operation behavior, thereby avoiding the computing power and storage overhead caused by continuous collection.

[0023] In some possible embodiments, the question information and the context operation record are collected by the same terminal device or by different terminal devices respectively.

[0024] In the human-computer interaction method provided in this application embodiment, the question information and context operation records can be collected by the same terminal device or by different terminal devices, making it more convenient for users to ask questions and improving the flexibility of the interaction method.

[0025] Secondly, embodiments of this application provide a human-computer interaction device, the device comprising: The acquisition module is used to acquire question information, which includes the question raised by the user and the time when the question was raised; The processing module is used to extract operation records within a preset time period from pre-cached context operation records when it is determined that relevant context information needs to be extracted based on the problem. The operation records include operation events and the time when the operation events occur. The operation events include operation behaviors and / or image data of the operation area. The communication module is used to upload the question information and the operation records within the preset time period to the artificial intelligence server for processing, and to receive the processing result of the question returned by the artificial intelligence server; The display module is used to display the processing results in the form of voice and / or text.

[0026] Thirdly, another embodiment of this application also provides a human-computer interaction device, which includes a processor and a memory. The memory is used to store programs executable by the processor, and the processor is used to read the programs in the memory and execute any human-computer interaction method provided in the embodiments of this application.

[0027] Fourthly, another embodiment of this application also provides a computer storage medium storing a computer program for causing a computer to execute any of the human-computer interaction methods provided in the embodiments of this application.

[0028] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart illustrating a human-computer interaction method according to an embodiment of this application; Figure 2 This is an overall flowchart of a human-computer interaction method according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating the practical application of a human-computer interaction method according to an embodiment of this application; Figure 4 This is a structural diagram of a human-computer interaction device according to an embodiment of this application; Figure 5 This is a structural diagram of a human-computer interaction device according to an embodiment of this application. Detailed Implementation

[0031] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on conventional or non-inventive effort. For steps that do not logically have a necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application. In actual processing or when the control device executes the method, it may be executed sequentially or in parallel according to the method shown in the embodiments or drawings.

[0032] It is understood that the following specific embodiments of this application involve the collection of user operation behavior, the collection of operation area image data, the collection of image data and / or video data containing user facial information, the collection of user voice, etc. When the various embodiments of this application are applied to specific products or technologies, relevant licenses or consents need to be obtained, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0033] Given that related technologies typically only acquire the user's input question text, these tools lack the ability to accurately judge the user's doubts. The generated answers often deviate from the user's actual reading context, making it difficult to precisely pinpoint the user's confusion and provide the most reasonable explanation and assistance, thus affecting the relevance and accuracy of the auxiliary answers. This application provides a human-computer interaction method, apparatus, and device to enable auxiliary tools to more accurately understand user doubts and improve the relevance and accuracy of auxiliary answers.

[0034] It should be noted that the auxiliary tools mentioned in this application can be desktop (computer desktop, i.e., the display interface of a computer device) assistants, document understanding assistants, terminal virtual assistants, or terminal intelligent assistants, etc., which can provide corresponding auxiliary functions such as explanation, summary, translation, or knowledge expansion based on the user's input questions. Taking terminal devices including computers and handheld devices as an example, the auxiliary tools can be deployed in the computer or in the handheld device, and the auxiliary tools in the computer can communicate with the auxiliary devices in the handheld device.

[0035] All the moments mentioned in the embodiments of this application, such as the moment when a question is raised, the moment when an operation event occurs, the moment when image data is collected, and the moment when an emotional state occurs, are collected based on the absolute system time in the terminal device.

[0036] See Figure 1 The diagram shown is an implementation flowchart of the human-computer interaction method provided in this application embodiment, which includes the following steps: See Figure 1 The diagram shown is an implementation flowchart of the human-computer interaction method provided in this application embodiment, which includes the following steps: S101, Obtain question information, which includes the question raised by the user and the time the question was raised.

[0037] In practice, the auxiliary tool remains online on the terminal device to receive questions from users via voice or text and record the time the question is asked.

[0038] Of course, it should be noted that if a user asks a question in text form, the question text can be directly extracted and the input time of the question text can be used as the question's submission time; if a user asks a question in voice form, the user's voice can be converted into text and recorded as the question text. In this case, the start or end time of the voice can be used as the question's submission time.

[0039] In practical applications, after obtaining the question information, this embodiment of the application can use a pre-configured large language model to perform semantic recognition on the question, generate recognition results, and determine whether it is necessary to extract relevant contextual information based on the semantic recognition results. The large language model can be any large language model from related technologies, and this embodiment of the application does not limit its use.

[0040] Specifically, based on the semantic recognition results, it is determined whether relevant contextual information needs to be extracted. If the recognition results indicate that the question contains demonstrative pronouns, then it is determined that the question requires the extraction of relevant contextual information. Demonstrative pronouns are pronouns that indicate a concept, have a specific meaning, and can serve as an indicator or substitute for a mentioned noun. They can include, but are not limited to, "this," "these," "that," and "those." When a question contains demonstrative pronouns, it is necessary to determine their specific referential meaning by combining relevant contextual information. Therefore, such questions require the extraction of relevant contextual information.

[0041] For example, if a user asks a question like "What does this picture mean?" or "What does this sentence mean?", then such questions contain demonstrative pronouns. When processing such questions, it is necessary to determine what "this picture" and "this sentence" specifically refer to. Therefore, it is necessary to extract relevant contextual information for such questions.

[0042] Based on the semantic recognition results, the system determines whether relevant contextual information needs to be extracted. If the recognition results indicate that the question does not contain demonstrative pronouns or is a general question, then it is determined that no relevant contextual information needs to be extracted. General questions are typically those unrelated to the user's current processing content.

[0043] For example, if a user asks a question like "How's the weather today?" or "What is the definition of artificial intelligence?", then such questions do not contain demonstrative pronouns and are general questions. These types of questions do not require extracting relevant contextual information.

[0044] Of course, in other embodiments of this application, the large language model can directly output the decision result on whether or not context information needs to be extracted, and this application does not limit this.

[0045] In practice, if it is determined that no relevant contextual information needs to be extracted based on the question, the question information can be directly uploaded to the artificial intelligence server, and the processing result returned by the artificial intelligence server can be received and displayed in the form of voice and / or text; if it is determined that relevant contextual information needs to be extracted based on the question, the following steps S202-S204 are executed.

[0046] S102, when it is determined that relevant context information needs to be extracted based on the problem, the operation record within a preset time period is extracted from the pre-cached context operation record. The operation record includes the operation event and the time when the operation event occurs. The operation event includes the operation behavior and / or the image data of the operation area.

[0047] In this embodiment, the pre-cached context operation record is triggered by a user performing a preset operation while browsing content on a terminal device. The operation record includes the operation event and the time of occurrence of the operation event. The operation event includes the operation behavior and / or image data of the operation area. The preset operation behaviors include, but are not limited to: selection range change operation, copy operation, page scrolling operation, active window switching operation, file switching operation, preview position change, pause, etc. Of course, the operation record may also record the operation event type, user identifier, and necessary lightweight features (such as features extracted using a convolutional neural network).

[0048] It should be noted that the image data of the operation area can be collected when the user performs each operation, or it can be collected when the user performs certain key operations. This application embodiment does not limit this. Among them, key operations include, but are not limited to: operations of changing the selection range, copying, pausing, and repeated selection.

[0049] In practical applications, when caching context operation records, new data can overwrite old data to cache context operation records for a fixed period, thereby reducing the storage overhead of cached data. For example, if context operation records within 60 minutes are cached, the most recently collected operation record will overwrite the earliest collected operation record in the cache, provided that operation records within the last 60 minutes are already cached.

[0050] In practice, when it is determined that relevant context information needs to be extracted based on the question raised by the user, the operation records within a preset time period are extracted from the pre-cached context operation records. The preset time period includes the time when the question was raised. The specific preset time period can be set according to experience, such as setting a preset time period [T1, T2], where T1 is before the time when the question was raised and T2 is after the time when the question was raised.

[0051] Specifically, when extracting operation records within a preset time period from the pre-cached context operation records, you can filter the operation records whose operation events occur within the preset time period.

[0052] S103 uploads the question information and operation records within a preset time period to the artificial intelligence server for processing, and receives the processing results for the question returned by the artificial intelligence server.

[0053] In specific implementation, after retrieving the operation records within a preset time period from the pre-cached context operation records, the query information and the operation records within the preset time period can be uploaded to the artificial intelligence server for processing, and the processing result returned by the artificial intelligence server can be received. Specifically, the artificial intelligence server can be any artificial intelligence server in related technologies, and the processing procedure of the artificial intelligence server can adopt the processing methods in related technologies; this application embodiment does not limit these aspects.

[0054] In practical applications, when uploading the question information and the operation records within a preset time period to the artificial intelligence server for processing, the user's question and the operation records within the preset time period can be sorted on a unified timeline according to the time of occurrence, generating aligned context information, and then the aligned context information can be uploaded to the artificial intelligence server for processing.

[0055] S104, display the processing results in the form of voice and / or text.

[0056] In practical applications, after receiving the processing results from the AI ​​server, the results are displayed to the user in the form of voice and / or text. Displaying the results in voice format allows users to quickly understand them without interrupting their operation; displaying them in text format can be placed near the file the user is viewing or the selected paragraph, allowing for quick location and understanding, and facilitating saving, thus significantly improving the speed and quality of problem-solving.

[0057] In specific implementation, when uploading the query information and operation records within a preset time period to the artificial intelligence server for processing, the embodiments of this application may further increase or decrease the amount of data uploaded, including at least the following implementation schemes.

[0058] Implementation Plan 1

[0059] To reduce the amount of data uploaded to the server, significantly reduce bandwidth and computing costs, and reduce the risk of sensitive information leakage, this application embodiment, when uploading the question information and operation records within a preset time period to the artificial intelligence server for processing, can calculate the correlation evaluation value between each operation event and the question in the operation records within the preset time period based on pre-set evaluation rules, select the target operation event whose correlation evaluation value meets the preset requirements, and upload the question information and the operation record corresponding to the target operation event to the artificial intelligence server for processing.

[0060] The pre-defined evaluation rules may include, but are not limited to: the closer the time of the operation event is to the time of the question being asked, the higher the relevance of the operation event; selection, copying, and pausing operations are highly relevant to questions like "What does this mean?"; if the same content is selected repeatedly within a short period of time or the previously selected content is selected again after a rollback operation, then this part of the content is highly relevant to the question being asked.

[0061] When calculating the correlation evaluation value, one or more weight coefficients can be set based on the pre-set evaluation rules, and the correlation evaluation value of each operation event and the problem in the operation record within the preset time period can be calculated by weighted summation from multiple dimensions. Alternatively, a neural network model can be trained based on the pre-set evaluation rules, and then the neural network model can be used to calculate the correlation evaluation value of each operation event and the problem in the operation record within the preset time period. This application embodiment does not limit this.

[0062] In practical applications, since the correlation evaluation between operation events and questions is often closely related to the time when the operation events occur, in order to facilitate the calculation of the correlation evaluation value between each operation event and the question in the operation record within a preset time period, this embodiment of the application uses the time when the question is raised as the anchor point to uniformly sort the operation records within the preset time period to form an aligned time axis, and then calculates the correlation evaluation value between each operation event and the question based on the aligned time axis.

[0063] Specifically, this application embodiment takes the time when the problem is raised as the center, sorts each operation event according to the occurrence time of each operation event in the operation record within a preset time period, obtains the sorting result, and then calculates the correlation evaluation value between each operation event and the problem in the sorting result based on the preset evaluation rules and the sorting result.

[0064] After calculating the correlation evaluation value between each operation event and the question in the operation record within a preset time period, the target operation event whose correlation evaluation value meets the preset requirements can be selected, and the question information and the operation record corresponding to the target operation event can be uploaded to the artificial intelligence server. The preset requirements can be set according to actual needs, such as operation events whose correlation evaluation value is greater than the preset evaluation value threshold, or operation events whose correlation evaluation value is among the top N (N is a positive integer), etc.

[0065] Implementation Plan 2

[0066] In this embodiment of the application, when a user's operation behavior while browsing content on a terminal device meets preset conditions, the recording frequency of the operation event is increased and / or the recording range of the image data of the operation area is increased within a preset time window, and the preset time window is marked as a question window.

[0067] It should be noted that when a user's actions while browsing content on a terminal device meet preset conditions, such as the user repeatedly selecting the same or highly similar content; the user frequently performing scrolling or back-and-forth scrolling operations to view a certain part of the content; the user performing a selection, copy, search, or reselect operation after pausing for a preset time; the user switching windows or tabs multiple times and returning to the same content or the same preview position, it can be determined that the user is confused about this part of the content. In this case, the recording frequency of operation events and / or the recording range of image data in the operation area can be increased within the preset time window, and the preset time window can be marked as the confusion window.

[0068] In practical applications, when generating multiple question windows, if there are multiple question windows, this application embodiment can also extract the features of the content (such as the content selected by the user) in the operation area, and determine whether multiple question windows are for the same content based on the extracted features. If they are for the same content, only one question window can be recorded. If they are not for the same content, they can be recorded as multiple question windows one by one.

[0069] Considering the existence of the question window, in this embodiment of the application, when uploading the question information and the operation records within a preset time period to the artificial intelligence server for processing, if it is determined that a question window existed before the time the question was raised, the operation records contained in the question window are merged with the operation records within the preset time period to obtain an operation record set, and the question information and the operation record set are uploaded to the artificial intelligence server for processing.

[0070] In this way, the question information and operation records are uploaded to the artificial intelligence server for processing. When the artificial intelligence server processes the questions raised by the user, it can combine the operation records in the question window to more accurately understand and locate the user's questions, thereby improving the pertinence and accuracy of the assisted answers.

[0071] Implementation Method 3

[0072] Considering the user's emotional state, this application embodiment can also help determine the user's location of doubt. It can also record the user's emotional information, including the user's emotional state, the time when the emotional state occurs, and the duration of the emotional state. When uploading the question information and the operation records within the preset time period to the artificial intelligence server for processing, the emotional information whose occurrence time is within the preset time period is filtered to generate an emotional information set. The question information, the operation records within the preset time period, and the emotional information set are then uploaded to the artificial intelligence server for processing.

[0073] In this way, when dealing with user questions, we can take into account the user's emotional state, making it easier to pinpoint the cause of the user's confusion and the scope of their concerns, thus reducing the need for repeated questioning and re-identification.

[0074] Specifically, when uploading the question information, operation records within a preset time period, and emotional information set to the artificial intelligence server for processing, the question information, operation records within a preset time period, and emotional information set can be sorted according to the chronological order of their occurrence to generate aligned context information, and then the aligned context information can be uploaded to the artificial intelligence server for processing.

[0075] When specifically recording a user's emotional information, video data containing the user's facial information can be acquired through an image acquisition device in the terminal device in a periodic or on-demand manner. Then, based on the video data, a pre-configured analysis model is used to determine the user's emotional state, including the time of occurrence and duration of the emotional state. The analysis model can be any model from related technologies, and this application embodiment does not limit its application to this specific model.

[0076] In some implementations, the user's emotional information can be cached together with the context operation record, or it can be cached separately. This application embodiment does not limit this.

[0077] In practical applications, the above-mentioned implementation methods one, two, and three can be implemented individually or in combination, and the combined implementation methods are also within the protection scope of this application.

[0078] It should be noted that, in this embodiment, the question information, context operation records, and video data containing user facial information can be collected by the same terminal device or by different terminal devices, making it easier for users to ask questions and improving the flexibility of the interaction method. Of course, when the question information and context operation records are collected and recorded by different terminal devices, auxiliary tools need to be deployed on different terminal devices, and the different terminal devices need to be connected to each other.

[0079] In one example, taking terminal devices including computers and handheld devices as an example, the question information can be collected by the handheld device, the context operation record can be collected by the computer, and the video data can be collected by the handheld device.

[0080] The above has provided a detailed description of each implementation step of the human-computer interaction method provided in the embodiments of this application. The following will be combined with... Figure 2 The specific implementation process of the human-computer interaction method provided in the embodiments of this application will be described.

[0081] like Figure 2As shown, the specific implementation process of the human-computer interaction method provided in this application embodiment includes: S201: When a user browses content on a terminal device, they perform a preset operation, triggering the recording of the operation record. The context operation record is then cached for a fixed duration by overwriting the old data with new data.

[0082] The preset operation behaviors include, but are not limited to: changing the selected area, copying, scrolling the page, switching the active window, switching files, changing the preview position, and pausing.

[0083] The operation log includes operation events and the time of occurrence of the operation events. Operation events include operation behaviors and / or image data of the operation area. Image data of the operation area can be collected when the user performs each operation behavior, or when the user performs certain key operation behaviors; this application embodiment does not limit this. Key operation behaviors include, but are not limited to: changing the selection range, copying, pausing, and repeatedly selecting, etc.

[0084] S202: Acquire video data containing the user's facial information. Based on the video data, use a pre-configured analysis model to determine the user's emotional state, the time of occurrence of the emotional state, and the duration of the emotional state, and record the user's emotional information accordingly.

[0085] S203, when the user's operation behavior while browsing content on the terminal device meets the preset conditions, increase the recording frequency of the operation event and / or increase the recording range of the image data of the operation area within the preset time window, and mark the preset time window as a question window.

[0086] S204: Obtain the question raised by the user and record the time when the question was raised.

[0087] S205: Determine whether relevant context information needs to be extracted based on the problem. If yes, execute S206; otherwise, execute S213.

[0088] S206: Extract operation records within a preset time period from the pre-cached context operation records.

[0089] S207. Before the question is raised, there is a question window. The operation records contained in the question window are merged with the operation records within a preset time period to obtain an operation record set.

[0090] S208, based on the pre-set evaluation rules, calculate the correlation evaluation value between each operation event and the problem in the operation record set, and select the target operation event whose correlation evaluation value meets the preset requirements.

[0091] S209, filter emotional information generated within a preset time period and generate a set of emotional information.

[0092] S210 uploads the question information, the operation record corresponding to the target operation event, and the emotional information set to the artificial intelligence server for processing.

[0093] S211, Receive the processing result of the artificial intelligence server for the problem.

[0094] S212, display the processing results in the form of voice and / or text.

[0095] S213, upload the problem to the artificial intelligence server for processing, and execute S211.

[0096] The above combination Figure 2 The overall implementation process of the human-computer interaction method provided in the embodiments of this application has been described below. Figure 3 Taking a specific example, the actual application process of the human-computer interaction method provided in the embodiments of this application will be described.

[0097] like Figure 3 As shown, the terminal device 30 is equipped with a desktop assistant 31. The desktop assistant 31 stays online while the user browses content. The desktop assistant 31 can collect image data and / or video data containing the user's facial information through the image acquisition device 32 to record the user's emotional state. The desktop assistant 31 triggers recording and caches context operation records based on the user's operation behavior through the mouse 33.

[0098] If a user frowns while previewing a file on terminal device 30, selects a piece of text with mouse 33, and asks desktop assistant 31, "What does this mean?", desktop assistant 31 will analyze the question and determine that contextual information needs to be extracted. It will then retrieve operation records within a preset time period from pre-cached context operation records and obtain the user's emotional information within that time period. Finally, it will upload the user's question information (the question asked and the time it was asked), the operation records within the preset time period, and the user's emotional information to the artificial intelligence server.

[0099] The AI ​​server locates the user's question based on the information uploaded by the terminal device 30, processes the question, generates a processing result, and then sends the processing result back to the terminal device. After receiving the processing result, the terminal device 30 displays the processing result in the form of voice broadcast 34 and text preview 35.

[0100] As can be seen from the above examples, the human-computer interaction method provided in this application, after obtaining the question information, when it is determined that relevant context information needs to be extracted based on the question, extracts the operation records within a preset time period from the pre-cached context operation records, and extracts the user's emotional information. Then, it uploads the question information, the operation records within the preset time period, and the user's emotional information to the artificial intelligence server for processing, and receives the processing result for the question returned by the artificial intelligence server, and displays the processing result in the form of voice and / or text.

[0101] In this way, when processing user questions, the AI ​​server can combine operation records within a preset time period with the user's emotional information to accurately locate the content the user is currently processing, and more accurately understand the user's doubts, thereby improving the pertinence and accuracy of the assisted answers.

[0102] In addition, in the human-computer interaction method provided in this application embodiment, the processing results of the artificial intelligence server are displayed in the form of voice and / or text. Displaying the processing results in the form of voice makes it easier for users to quickly understand the processing results without interrupting the user's operation; while displaying the processing results in the form of text makes it easier for users to quickly locate and understand the original content, and also makes it easier for users to save, thereby significantly improving the speed and quality of users' problem processing.

[0103] Based on the same inventive concept, this application also provides a human-computer interaction device, such as... Figure 4 As shown, the device includes: The acquisition module 401 is used to acquire question information, which includes the question raised by the user and the time the question was raised. Processing module 402 is used to extract operation records within a preset time period from pre-cached context operation records when it is determined that relevant context information needs to be extracted based on the problem. The operation records include operation events and the time when the operation events occur. The operation events include operation behaviors and / or image data of the operation area. The communication module 403 is used to upload the question information and the operation records within a preset time period to the artificial intelligence server for processing, and to receive the processing results of the question returned by the artificial intelligence server. Display module 404 is used to display the processing results in the form of voice and / or text.

[0104] In one possible implementation, the communication module 403 is specifically used for: Based on pre-defined evaluation rules, calculate the correlation evaluation value between each operation event and the problem in the operation record within a preset time period; Select the target operation event whose relevance evaluation value meets the preset requirements, and upload the question information and the operation record corresponding to the target operation event to the artificial intelligence server for processing.

[0105] In one possible implementation, the communication module 403 is specifically used for: Centered on the moment the problem is raised, the operation events are sorted according to the occurrence time of each operation event in the operation record within a preset time period to obtain the sorting result; Based on the pre-defined evaluation rules and ranking results, the correlation evaluation value between each operation event and the problem in the ranking results is calculated.

[0106] In one possible implementation, the processing module 402 is further configured to: When a user performs an operation while browsing content on a terminal device that meets preset conditions, the recording frequency of the operation event is increased and / or the recording range of the image data in the operation area is increased within a preset time window, and the preset time window is marked as a question window.

[0107] In one possible implementation, the communication module 403 is specifically used for: If a question window exists before the question is raised, the operation records contained in the question window are merged with the operation records within a preset time period to obtain an operation record set. The question information and the operation record set are then uploaded to the artificial intelligence server for processing.

[0108] In one possible implementation, the processing module 402 is further configured to: Record the user's emotional information, including the user's emotional state, the time when the emotional state occurred, and the duration of the emotional state; Communication module 403 is specifically used for: Filter emotional information generated within a preset time period and generate a set of emotional information. The question information, operation records within a preset time period, and emotional information are uploaded to the artificial intelligence server for processing.

[0109] In one possible implementation, the communication module 403 is specifically used for: The question information, operation records within a preset time period, and emotional information are sorted in chronological order of occurrence to generate aligned contextual information. The aligned context information is then uploaded to the AI ​​server for processing.

[0110] In one possible implementation, the processing module 402 is specifically used for: Acquire video data containing user facial information; Based on video data, a pre-configured analysis model is used to determine the user's emotional state, as well as the time and duration of the emotional state.

[0111] In one possible implementation, the processing module 402 is further configured to: The question is semantically identified using a pre-configured large language model, and the identification results are generated. When determining that a question contains demonstrative pronouns based on the recognition results, it is necessary to extract relevant contextual information to determine the question.

[0112] In one possible implementation, the communication module 403 is further configured to: When it is determined that no relevant contextual information needs to be extracted based on the question, the question information is uploaded to the artificial intelligence server, and the processing result for the question is returned by the artificial intelligence server.

[0113] In one possible implementation, the pre-cached context operation record is triggered when a user performs a preset operation while browsing content on a terminal device.

[0114] In one possible implementation, the question information and context operation records are collected by the same terminal device or by different terminal devices respectively.

[0115] After introducing the human-computer interaction method and apparatus according to exemplary embodiments of this application, another human-computer interaction device according to an exemplary embodiment of this application will be introduced next.

[0116] Those skilled in the art will understand that various aspects of this application can be implemented as system, method, or program products.

[0117] In some possible implementations, such as Figure 5 As shown, the human-computer interaction device according to this application includes a processor 500 and a memory 501. The memory 501 is used to store programs executable by the processor 500, and the processor 500 is used to read the programs in the memory 501 and execute the steps in the human-computer interaction methods according to the various exemplary embodiments of this application described above.

[0118] In some possible implementations, various aspects of the human-computer interaction method provided in this application can also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to cause the computer device to perform the steps of a human-computer interaction method according to various exemplary embodiments of this application described above.

[0119] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0120] The program product for human-computer interaction according to the embodiments of this application can be a portable compact disc read-only memory (CD-ROM) and include program code, and can run on an electronic device. However, the program product of this application is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0121] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0122] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0123] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's electronic device, partially on the user's device, as a standalone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the user's electronic device via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external electronic device (e.g., via the Internet using an Internet service provider).

[0124] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0125] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0126] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This application is described with reference to flowchart illustrations and block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block and / or segment of the flowchart illustrations and block diagrams, as well as combinations of blocks and segments in the flowchart illustrations and block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0131] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A human-computer interaction method, characterized in that, The method includes: Obtain question information, which includes the question raised by the user and the time the question was raised; When it is determined that relevant context information needs to be extracted based on the problem, operation records within a preset time period are extracted from the pre-cached context operation records. The operation records include operation events and the time when the operation events occur. The operation events include operation behaviors and / or image data of the operation area. The question information and the operation records within the preset time period are uploaded to the artificial intelligence server for processing, and the processing result for the question is returned by the artificial intelligence server. The processing results are displayed in the form of voice and / or text.

2. The method according to claim 1, characterized in that, The step of uploading the question information and the operation records within the preset time period to the artificial intelligence server for processing includes: Based on pre-set evaluation rules, calculate the correlation evaluation value between each operation event in the operation record and the problem within the preset time period; Select the target operation event whose relevance evaluation value meets the preset requirements, and upload the question information and the operation record corresponding to the target operation event to the artificial intelligence server for processing.

3. The method according to claim 2, characterized in that, The process of calculating the correlation evaluation value between each operation event in the operation record within the preset time period and the problem, based on pre-set evaluation rules, includes: Centered on the time when the problem was raised, the operation events are sorted according to the occurrence time of each operation event in the operation record within the preset time period to obtain the sorting result; Based on the pre-set evaluation rules and the ranking results, the correlation evaluation value between each operation event in the ranking results and the problem is calculated.

4. The method according to claim 1, characterized in that, The method further includes: When the user's operation behavior while browsing content on the terminal device meets the preset conditions, the recording frequency of the operation event is increased and / or the recording range of the image data of the operation area is increased within the preset time window, and the preset time window is marked as a question window.

5. The method according to claim 4, characterized in that, The step of uploading the question information and the operation records within the preset time period to the artificial intelligence server for processing includes: If a question window exists before the time the question is asked, the operation records contained in the question window are merged with the operation records within the preset time period to obtain an operation record set. The question information and the operation record set are then uploaded to the artificial intelligence server for processing.

6. The method according to claim 1, characterized in that, The method further includes: Record the user's emotional information, which includes the user's emotional state, the time when the emotional state occurred, and the duration of the emotional state. The step of uploading the question information and the operation records within the preset time period to the artificial intelligence server for processing includes: Filter emotional information generated within the preset time period to generate an emotional information set; The question information, the operation records within the preset time period, and the set of emotion information are uploaded to the artificial intelligence server for processing.

7. The method according to claim 6, characterized in that, The step of uploading the question information, the operation records within the preset time period, and the set of emotion information to the artificial intelligence server for processing includes: The question information, the operation records within the preset time period, and the emotional information set are sorted in chronological order of occurrence to generate aligned contextual information. The aligned context information is uploaded to the artificial intelligence server for processing.

8. The method according to claim 6, characterized in that, The recording of the user's emotional information includes: Obtain video data containing the user's facial information; Based on the video data, the user's emotional state is determined using a pre-configured analysis model, and the time and duration of the emotional state are determined.

9. The method according to claim 1, characterized in that, After obtaining the query information, and before extracting the operation records within a preset time period from the pre-cached context operation records, the method further includes: The problem is semantically identified using a pre-configured large language model, and the identification results are generated. When determining that the question contains an indicator pronoun based on the recognition result, it is necessary to extract relevant contextual information to determine the question.

10. The method according to any one of claims 1-9, characterized in that, After obtaining the query information, and before extracting the operation records within a preset time period from the pre-cached context operation records, the method further includes: When it is determined that no relevant contextual information needs to be extracted based on the question, the question information is uploaded to the artificial intelligence server, and the processing result for the question is returned by the artificial intelligence server, and the processing result is displayed in the form of voice and / or text.

11. The method according to any one of claims 1-9, characterized in that, The pre-cached context operation record is triggered when the user performs a preset operation while browsing content on the terminal device.

12. The method according to any one of claims 1-9, characterized in that, The question information and the context operation record may be collected by the same terminal device or by different terminal devices.

13. A human-computer interaction device, characterized in that, The device includes: The acquisition module is used to acquire question information, which includes the question raised by the user and the time when the question was raised; The processing module is used to extract operation records within a preset time period from pre-cached context operation records when it is determined that relevant context information needs to be extracted based on the problem. The operation records include operation events and the time when the operation events occur. The operation events include operation behaviors and / or image data of the operation area. The communication module is used to upload the question information and the operation records within the preset time period to the artificial intelligence server for processing, and to receive the processing result of the question returned by the artificial intelligence server; The display module is used to display the processing results in the form of voice and / or text.

14. A human-computer interaction device, characterized in that, The device includes a processor and a memory for storing a program executable by the processor, and the processor for reading the program from the memory and performing the steps of the method according to any one of claims 1-12.