Robotic conversation method, apparatus, storage medium, and electronic device
By calling a program plugin in the target application to obtain text content and operation information, sending it to the model service device to generate and display the response result, the problem of low efficiency in robot dialogue is solved, and a convenient and efficient dialogue process is achieved.
Patent Information
- Application Number
- CN202310636271.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2043-05-31
AI Technical Summary
Current technologies for robot dialogue are inefficient, requiring users to manually organize their language and send it to the chatbot, resulting in inefficiency and wasted time.
By calling the program plugin in the target application, the text content and operation information are obtained and sent to the model service device to generate a response result. The response content is received and displayed, and the program plugin supports the target object to establish a robot dialogue with the model service device.
It improves the efficiency of robot dialogue, enabling users to easily communicate with model service devices, reducing manual operation steps, and enhancing the convenience and efficiency of dialogue.
Smart Images

Figure CN117076619B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a robot dialogue method, apparatus, storage medium, and electronic device. Background Technology
[0002] Currently, model-based chatbots (such as large-scale or small-scale chatbots) are widely used in various fields, including customer service, education, entertainment, and healthcare. Large-scale chatbots are artificial intelligence systems that use large language models to engage in natural conversations with humans. These models are typically trained using deep learning techniques to simulate human language and thought processes, enabling them to answer various types of questions. It should be understood that compared to small-scale chatbots, large-scale chatbots possess higher intelligence and a wider range of applications. Furthermore, large-scale chatbots usually require multiple models to collaborate to achieve better dialogue effectiveness and intelligence. However, in existing technologies, when an individual needs to converse with a chatbot, they typically need to organize the text to be discussed and send it to the chatbot to obtain a response. This process is inefficient and time-consuming. Therefore, how to facilitate convenient chatbot dialogue and improve dialogue efficiency has become a research hotspot. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a robot dialogue method, apparatus, storage medium, and electronic device to solve the problem of low efficiency in robot dialogue; correspondingly, embodiments of the present invention can facilitate robot dialogue through program plug-ins, which can effectively improve dialogue efficiency.
[0004] According to one aspect of the present invention, a robot dialogue method is provided, the method comprising:
[0005] The program plugin in the target application is invoked to obtain the text content determined by the target object in the target application, and to obtain the target operation information determined by the target object in the target application. The program plugin is used to support the target object to establish robot dialogue with the model service device on the target application.
[0006] The text content and the target operation information are sent to the model service device, so that the model service device generates a text response result based on the text content and the target operation information. The text response result includes the target response content executed according to the target operation information.
[0007] Receive the target response content returned by the model service device and display the target response content.
[0008] According to another aspect of the present invention, a robot dialogue device is provided, the device comprising:
[0009] The acquisition unit is used to call the program plugin in the target application to acquire the text content determined by the target object in the target application, and to acquire the target operation information determined by the target object in the target application. The program plugin is used to support the target object to establish robot dialogue with the model service device on the target application.
[0010] The processing unit is configured to send the text content and the target operation information to the model service device, so that the model service device generates a text response result based on the text content and the target operation information, wherein the text response result includes the target response content executed according to the target operation information;
[0011] The processing unit is also configured to receive the target response content returned by the model service device and display the target response content.
[0012] According to another aspect of the present invention, an electronic device is provided, the electronic device including a processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the methods mentioned above.
[0013] According to another aspect of the present invention, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods mentioned above.
[0014] This invention allows for the invocation of a program plugin within a target application to obtain the text content and target operation information determined by the target object within the application. The text content and target operation information are then sent to a model service device, enabling the model service device to generate a text response result based on the text content and target operation information. This response result includes the target response content executed according to the target operation information. The program plugin supports the target object establishing a robot dialogue with the model service device within the target application. The device can then receive and display the target response content returned by the model service device. Therefore, this invention allows for convenient robot dialogue via a program plugin, effectively improving dialogue efficiency. Attached Figure Description
[0015] Further details, features, and advantages of the invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0016] Figure 1 A flowchart illustrating a robot dialogue method according to an exemplary embodiment of the present invention is shown;
[0017] Figure 2a A schematic diagram of a response content display area according to an exemplary embodiment of the present invention is shown;
[0018] Figure 2b A schematic diagram of another response content display area according to an exemplary embodiment of the present invention is shown;
[0019] Figure 3 A flowchart illustrating another robot dialogue method according to an exemplary embodiment of the present invention is shown;
[0020] Figure 4a A schematic diagram of a text selection operation according to an exemplary embodiment of the present invention is shown;
[0021] Figure 4b A schematic diagram of a text input area according to an exemplary embodiment of the present invention is shown;
[0022] Figure 5 A schematic diagram of an operation display area according to an exemplary embodiment of the present invention is shown;
[0023] Figure 6 A flowchart illustrating yet another robot dialogue method according to an exemplary embodiment of the present invention is shown;
[0024] Figure 7 A schematic block diagram of a robot dialogue device according to an exemplary embodiment of the present invention is shown;
[0025] Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0026] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0027] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0028] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0029] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0031] It should be noted that the executing entity of the robot dialogue method provided in this embodiment of the invention can be one or more electronic devices, or it can be a target application in an electronic device; this invention does not limit this. The electronic device can be a terminal (i.e., a client) including the target application or a server. Therefore, when the executing entity includes multiple electronic devices, and these multiple electronic devices include at least one terminal and at least one server, the robot dialogue method provided in this embodiment of the invention can be jointly executed by the terminal and the server. Accordingly, the terminal mentioned herein can include, but is not limited to: smartphones, tablets, laptops, desktop computers, smartwatches, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The server mentioned herein can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc.
[0032] Based on the above description, embodiments of the present invention propose a robot dialogue method. This robot dialogue method can be executed by an electronic device (terminal or server) including a target application as mentioned above; or, the robot dialogue method can be executed jointly by a terminal and a server; or, the robot dialogue method can be executed by the target application in the electronic device. For ease of explanation, the following description will use the execution of the robot dialogue method by an electronic device as an example; such as... Figure 1 As shown, the robot dialogue method may include the following steps S101-S103:
[0033] S101, call the program plugin in the target application to obtain the text content determined by the target object in the target application, and obtain the target operation information determined by the target object in the target application. The program plugin is used to support the target object to establish robot dialogue with the model service device on the target application.
[0034] The target application can be a target browser (such as Internet Explorer or Chrome browser), in which case the program plugin is a browser plugin applied to the target browser; a browser plugin is a customizable plugin used in the target browser that can enhance its functionality. Optionally, the target application can also be an app (third-party application), in which case the program plugin is an app plugin applied to the app, and so on; this invention does not limit this.
[0035] It should be noted that the aforementioned model service device can be ChatGPT (a natural language processing tool driven by artificial intelligence technology), Wenxin Yiyan (a large language model), Chatman (a chatbot trained based on a large language model), etc.; this invention does not limit it in this regard. It should be understood that the model service device can provide model service functionality; for example, when the model service device is Chatman, the model service provided by the model service device is a platform service encapsulated by Chatman, which can provide chatbot capabilities to the access party (such as the target application) through API (Application Program Interface).
[0036] S102, the text content and target operation information are sent to the model service device so that the model service device generates a text response result based on the text content and target operation information. The text response result includes the target response content executed according to the target operation information.
[0037] The aforementioned target operation information can be any operation information from a set of operation information, including but not limited to: explanation, translation, summary, and extension, etc.; this invention does not limit this. For example, when the target operation information is explanation, the model service device can explain the text content, and the target response content is the explanation result of the text content being interpreted according to the explanation; when the target operation information is translation, the model service device can translate the text content, and the target response content is the translation result of the text content being translated according to the translation, etc.
[0038] It should be understood that when the target operation information is translation, the target operation information can also carry the translation type (such as Chinese to English or English to Chinese). That is to say, when the target object determines that the target operation information is translation, the corresponding translation type can also be determined, so that the model service device can translate the text content according to the corresponding translation type.
[0039] S103: Receive the target response content returned by the model service device and display the target response content.
[0040] Specifically, after receiving the target response content returned by the model service device, the electronic device can display the target response content within the response content display area. This response content display area can be a display area determined according to the display area where the text content is located (e.g., a display area whose distance from the display area where the text content is located is equal to a first preset distance threshold, or a display area whose distance from the display area where the text content is located is less than the first preset distance threshold, etc.), or any display area in the display interface, or a display area in a program plugin, etc.; this invention does not limit this. It should be noted that the first preset distance threshold can be set based on experience or based on actual needs; this invention does not limit this.
[0041] For example, such as Figure 2a As shown, taking display area 201 as the response content display area as an example, display area 201 is a display area outside the program plugin, such as the display area determined according to the display area where the text content is located, or any display area in the display interface. When the electronic device needs to display the target response content, it can display display area 201 in the display interface, thereby displaying the target response content in display area 201.
[0042] For example, such as Figure 2bAs shown, taking display area 203 as the response content display area as an example, display area 203 is the display area in the program plug-in. When the electronic device needs to display the target response content, it can display the program plug-in in the display interface and display the target response content in the display area 203 included in the program plug-in.
[0043] It should be understood that, Figure 2a and Figure 2b The examples provided exemplarily illustrate the response content display area, but the present invention does not limit this; for example, the response content display area may also include a minimize button; or, for example, the response content display area may not include, for example... Figure 2a Instead of using the cancel button 202 shown, clicking on a blank area cancels the display of the response content display area, and so on.
[0044] This invention allows for the invocation of a program plugin within a target application to obtain the text content and target operation information determined by the target object within the application. The text content and target operation information are then sent to a model service device, enabling the model service device to generate a text response result based on the text content and target operation information. This response result includes the target response content executed according to the target operation information. The program plugin supports the target object establishing a robot dialogue with the model service device within the target application. The device can then receive and display the target response content returned by the model service device. Therefore, this invention allows for convenient robot dialogue via a program plugin, effectively improving dialogue efficiency.
[0045] Based on the above description, this embodiment of the invention also proposes a more specific robot dialogue method. Accordingly, this robot dialogue method can be executed by the electronic device (terminal or server) including the target application mentioned above; or, the robot dialogue method can be executed jointly by the terminal and the server; or, the robot dialogue method can be executed by the target application in the electronic device. For ease of explanation, the following description will use the execution of the robot dialogue method by an electronic device as an example; please refer to [link to previous text]. Figure 3 The robot dialogue method may include the following steps S301-S305:
[0046] S301, call the program plugin in the target application to obtain the text content determined by the target object in the target application, and obtain the target operation information determined by the target object in the target application. The program plugin is used to support the target object to establish robot dialogue with the model service device on the target application.
[0047] Specifically, when acquiring text content determined by the target object within the target application, the electronic device can, upon detecting a text selection instruction from the target object within the target application, use the content selected by the text selection instruction as the text content determined by the target object within the target application. Alternatively, the electronic device can display a text input area of the program plugin within the target application, and upon detecting a text editing instruction from the target object on the text input area, use the content indicated by the text editing instruction as the text content. Or, the electronic device can, upon detecting a voice input instruction from the target object for the program plugin, use the content indicated by the voice input instruction as the text content, and so on.
[0048] It should be noted that when a target object performs a text selection operation within a target application, the electronic device can detect the text selection instruction within the target application. The content selected by this text selection instruction is the content chosen by the target object through the text selection operation performed within the target application, such as... Figure 4a As shown; the text selection operation can refer to cursor-based sliding operations, or continuous clicking operations on a line or paragraph of text, etc., and this invention does not limit it. In this embodiment of the invention, when the target application is a target browser and the program plugin is a browser plugin, the program plugin can inject a listener function for onmouseup (a built-in event of the webpage). Then, whenever the target object performs a text selection operation in the target application, the program plugin can receive the event. Based on this, the electronic device can obtain the text content through the target application's built-in API and getSelection (a function to obtain the selected content).
[0049] Optionally, when a text selection instruction is detected and the program plugin is displayed in the target application, the electronic device can add the content selected by the text selection instruction to the program plugin and display the content selected by the text selection instruction in the program plugin.
[0050] Correspondingly, when the target object performs a copy editing operation on the text input area, the electronic device can detect the copy editing instruction of the target object on the text input area. At this time, the content indicated by the copy editing instruction is: the content edited by the target object through the copy editing operation; such as Figure 4b As shown, the target object can perform text editing operations on the text input area 401 to obtain text content (i.e., the content in the text input area 401). The text editing operation can refer to input via keyboard or handwriting input, etc., and this invention does not limit it to these methods.
[0051] Similarly, when a target object performs a voice input operation on a program plugin, the electronic device can detect the voice input command from the target object to the program plugin. The content indicated by the voice input command is the content entered by the target object through the voice input operation performed on the program plugin. It should be noted that the program plugin may include a voice input button. Therefore, when performing a voice input operation, the target object can click the voice input button and record voice input to the program plugin after the click until the voice input button is clicked again, thus achieving the voice input operation; or, the target object can long-press the voice input button and record voice input to the program plugin while holding the button down, until the long-press operation ends, thus achieving the voice input operation, and so on. This invention does not limit the specific implementation process of the voice input operation.
[0052] Furthermore, when acquiring target operation information determined by the target object in the target application, the electronic device can, upon detecting a text selection instruction from the target object in the target application, determine the operation display area according to the display area where the text content is located, and display the operation information set in the operation display area, such as... Figure 5 As shown. Based on this, when the electronic device detects a first operation selection instruction from the target object to the operation information set, it uses the operation information selected by the first operation selection instruction as the target operation information determined by the target object in the target application.
[0053] It should be noted that the operation display area can be a display area whose distance from the display area where the text content is located is less than or equal to a second preset distance threshold. In this case, if the operation information set includes commonly used operation information configured for the target object (such as translation or explanation), the electronic device can display a pop-up window of the commonly used operation information configured for the target object next to the text content. Optionally, the operation display area can also be a display area whose distance from the display area where the text content is located is greater than the second preset distance threshold, and so on. This invention does not limit this. The second preset distance threshold can be set based on experience or according to actual needs, and this invention does not limit this. In other embodiments, the operation display area can also be any display area in the display interface, and so on.
[0054] In a specific implementation, when the target object performs a first operation selection operation on the operation information set, the electronic device can detect the target object's first operation selection instruction on the operation information set. At this time, the operation information selected by the first operation selection instruction is: the operation information selected by the target object in the first operation selection operation. The first operation selection operation can refer to a click operation on any operation information in the operation information set, or it can refer to a long press operation on any operation in the operation information set, etc., and this invention does not limit it in this way.
[0055] Alternatively, the program plugin may include a set of operation information to be selected, such as Figure 4b As shown; based on this, when obtaining the target operation information determined by the target object in the target application, the electronic device can display a program plugin in the target application, and when it detects the second operation selection instruction of the target object for the operation information set in the program plugin, it takes the operation information indicated by the second operation selection instruction as the target operation information. Specifically, when the target object performs the second operation selection operation for the operation information set in the program plugin, the electronic device can detect the second operation selection instruction of the target object for the operation information set in the program plugin. At this time, the operation information indicated by the second operation selection instruction is: the operation information selected by the target object in the second operation selection operation; wherein, the second operation selection operation can refer to a click operation for any operation information in the operation information set, or it can refer to a long press operation for any operation information in the operation information set, etc., and the present invention does not limit it in this way.
[0056] Accordingly, when acquiring the target operation information determined by the target object in the target application, the electronic device can also use the operation information indicated by the operation input instruction for the program plugin as the target operation information when it detects the operation input instruction for the target object. Specifically, when the target object performs an operation information input operation for the program plugin, the electronic device can detect the operation input instruction for the program plugin. At this time, the operation information indicated by the operation input instruction is: the operation information input by the target object. The operation information input operation can refer to input via keyboard or voice input, etc., and this invention does not limit it.
[0057] S302, the text content and target operation information are sent to the model service device so that the model service device can generate a text response result based on the text content and target operation information. The text response result includes the target response content executed according to the target operation information.
[0058] In practical implementation, electronic devices can send text content and target operation information to the front-end service device based on the target session identifier corresponding to the text content. This allows the front-end service device to then send the text content and target operation information to the model service device according to the target session identifier. Figure 6As shown. The front-end service device is used to: manage the correspondence between each object in at least one object and each session in at least one session, based on the session identifier. The front-end service device serves as a relay device for the model service device in robot dialogue, and the target object is any one of the at least one objects. In this case, after receiving the text content and target operation information sent by the front-end service device, the model service device can generate a text response result based on the text content and target operation information.
[0059] It should be noted that the front-end service device can be a terminal, a server, etc., and this invention does not limit it. This front-end service device can call the API of the model service device to provide chatbot capabilities to the target application and manage common user (i.e., object) operations; that is, common operation information set by any object will be stored in the front-end service device so that the front-end service device can manage the corresponding content. Correspondingly, the session identifier can be a numeric identifier (i.e., session number (sessionId), such as 1 or 2) or a letter identifier (such as a or b), and this invention does not limit it.
[0060] In this embodiment of the invention, before sending the text content and target operation information to the front-end service device based on the target session identifier corresponding to the text content, the electronic device can search for the target session identifier. If the target session identifier is found, the above-mentioned sending of the text content and target operation information to the front-end service device based on the target session identifier corresponding to the text content is executed. If the target session identifier is not found, a session creation command is sent to the front-end service device, so that the model service device, after receiving the initialized session creation command sent by the front-end service device, creates the target session indicated by the target session identifier. The initialized session creation command carries the target session identifier. The model service device also receives the target session identifier returned by the front-end service device. Optionally, the front-end service device can initialize the session creation command and send the target session identifier to the electronic device after obtaining the initialized session creation command; the front-end service device can also create the target session on the model service device and send the target session identifier to the electronic device after receiving the target session identifier returned by the model service device, etc. The invention does not limit this.
[0061] Furthermore, the text response result may also include a response identifier, which is used to indicate the target response content; wherein, the response identifier may be a number identifier or a letter identifier, and the present invention does not limit it.
[0062] S303, Receive the response identifier returned by the model service device.
[0063] In the specific implementation, after the model service device generates a response identifier, it can return the response identifier to the foreground service device. The electronic device can then receive the response identifier returned by the foreground service device, thus achieving the reception of the response identifier returned by the model service device. Figure 6 As shown.
[0064] S304. Obtain the first generated result from the target response content from the model service device based on the response identifier, and display the first generated result. The first generated result includes M characters from the target response content, where M is a positive integer.
[0065] In specific implementation, the electronic device can query the answer from the front-end service device based on the response identifier, which in turn allows the front-end service device to query the answer from the model service device based on the response identifier. This enables the electronic device to obtain the first generated result from the target response content from the model service device based on the response identifier. Correspondingly, upon obtaining the first generated result, the model service device can return the first generated result to the front-end service device. Based on this, the electronic device can receive the first generated result returned by the front-end service device, such as... Figure 6 As shown.
[0066] S305, determine the response status identifier corresponding to the target response content. If the response status identifier is an incomplete response identifier, continue to obtain the second generated result in the target response content from the model service device based on the response identifier until the response status identifier is updated to a completed response identifier, so as to receive the target response content returned by the model service device and display the target response content. The second generated result includes N characters in the target response content, where N is a positive integer.
[0067] The response status identifier can be either an "incomplete response" identifier or a "complete response" identifier. Optionally, the response status identifier can be a number or a letter, etc., and this invention does not limit this. It should be understood that the "incomplete response" identifier and the "complete response" identifier are different. For example, when the response status identifier is a number, the "complete response" identifier can be 1, while the "incomplete response" identifier can be 0, and so on.
[0068] It should be noted that the second generated result may or may not include the first generated result (i.e., the second generated result may be any generated result in the target response content other than the first generated result), and this invention does not limit this. It should be understood that if the electronic device receives the second generated result and the response status indicator is still "not finished answering," the electronic device may display the second generated result and continue to obtain other generated results from the target response content from the model service device, etc.
[0069] Furthermore, after the target response content has been generated by the model service device, the electronic device can receive a completed response flag from the model service device and update the response status flag accordingly, thus updating the response status flag to the completed response flag. Specifically, the model service device can return the completed response flag to the foreground service device, and the electronic device can then receive the completed response flag returned by the foreground service device, thereby enabling the electronic device to receive the completed response flag sent by the model service device.
[0070] Optionally, before the target response content is fully generated by the model service device, the electronic device can also receive an incomplete response flag sent by the model service device. That is, the model service device can return the incomplete response flag to the foreground service device, and the electronic device can receive the incomplete response flag returned by the foreground service device, thus enabling the electronic device to receive the incomplete response flag returned by the model service device. Alternatively, when acquiring the target response content, the electronic device can generate an incomplete response flag corresponding to the target response content to initialize the response status flag, until the received completed response flag is used to update the response status flag.
[0071] This invention embodiment can call a program plugin in the target application to obtain the text content determined by the target object in the target application, and after obtaining the target operation information determined by the target object in the target application, send the text content and target operation information to the model service device. The model service device then generates a text response result based on the text content and target operation information. The text response result includes the target response content executed according to the target operation information and a response identifier. Then, the response identifier returned by the model service device can be received, and a first generated result from the target response content can be obtained from the model service device based on the response identifier. The first generated result includes M characters from the target response content, where M is a positive integer. Further, a response status identifier corresponding to the target response content can be determined. If the response status identifier is "not completed," a second generated result from the target response content can be obtained from the model service device based on the response identifier until the response status identifier is updated to "completed," thereby achieving the reception and display of the target response content returned by the model service device. The second generated result includes N characters from the target response content, where N is a positive integer. As can be seen, the embodiments of the present invention can acquire generated results in real time and display received results, thereby improving display efficiency and increasing user engagement. Furthermore, the embodiments of the present invention can directly apply chatbot capabilities to target applications via program plugins. Target users can select and directly manipulate the names and sentences requiring model service device assistance within the target application to obtain target responses. In other words, the embodiments of the present invention enable target users to better and more conveniently read and understand the content within the target application, such as reading complex web documents, interpreting profound technical terms, and translating various foreign languages.
[0072] Based on the description of the relevant embodiments of the above-mentioned robot dialogue method, this invention also proposes a robot dialogue device, which can be a computer program (including program code) running in an electronic device; such as Figure 7 As shown, the robot dialogue device may include an acquisition unit 701 and a processing unit 702. The robot dialogue device can perform... Figure 1 or Figure 3 The robot dialogue method shown, i.e., the robot dialogue device can operate the above-mentioned unit:
[0073] The acquisition unit 701 is used to call the program plugin in the target application to acquire the text content determined by the target object in the target application, and to acquire the target operation information determined by the target object in the target application. The program plugin is used to support the target object to establish robot dialogue with the model service device on the target application.
[0074] Processing unit 702 is configured to send the text content and the target operation information to the model service device, so that the model service device generates a text response result based on the text content and the target operation information, wherein the text response result includes the target response content executed according to the target operation information;
[0075] The processing unit 702 is further configured to receive the target response content returned by the model service device and display the target response content.
[0076] In one embodiment, when the processing unit 702 sends the text content and the target operation information to the model service device, it may specifically be used to:
[0077] Based on the target session identifier corresponding to the text content, the text content and the target operation information are sent to the front-end service device, so that the front-end service device sends the text content and the target operation information to the model service device according to the target session identifier;
[0078] The front-end service device is used to: manage the correspondence between each object in at least one object and each session in at least one session according to the session identifier, and the front-end service device is: a relay device for the model service device to conduct robot dialogue, and the target object is any one of the at least one objects.
[0079] In another embodiment, the processing unit 702 can also be used for:
[0080] Locate the target session identifier;
[0081] If the target session identifier is found, then the target session identifier corresponding to the text content is used to send the text content and the target operation information to the front-end service device.
[0082] If the target session identifier is not found, a create session instruction is sent to the front-end service device, so that the model service device creates the target session indicated by the target session identifier after receiving the initialized create session instruction sent by the front-end service device. The initialized create session instruction carries the target session identifier; and the model service device receives the target session identifier returned by the front-end service device.
[0083] In another embodiment, the text response result further includes a response identifier, which is used to indicate the target response content; the processing unit 702 can also be used for:
[0084] Receive the response identifier returned by the model service device;
[0085] When the processing unit 702 receives the target response content returned by the model service device and displays the target response content, it may specifically be used to:
[0086] Based on the response identifier, a first generated result is obtained from the target response content from the model service device, and the first generated result is displayed. The first generated result includes M characters from the target response content, where M is a positive integer.
[0087] The response status identifier corresponding to the target response content is determined. If the response status identifier is an "not answered" identifier, the second generated result in the target response content is obtained from the model service device based on the response identifier until the response status identifier is updated to an "answered" identifier, so as to receive the target response content returned by the model service device and display the target response content. The second generated result includes N characters in the target response content, where N is a positive integer.
[0088] In another embodiment, the processing unit 702 can also be used for:
[0089] After the target response content is generated by the model service device, the system receives a completed response flag sent by the model service device.
[0090] The response status identifier is updated using the completed response identifier.
[0091] In another embodiment, when acquiring the text content determined by the target object in the target application, the acquisition unit 701 may specifically be used to:
[0092] When a text selection instruction is detected from the target object in the target application, the content selected by the text selection instruction is taken as the text content determined by the target object in the target application; or,
[0093] The text input area of the program plugin is displayed in the target application. When a text editing instruction is detected on the text input area by the target object, the content indicated by the text editing instruction is used as the text content; or,
[0094] When a voice input command from the target object to the program plugin is detected, the content indicated by the voice input command is used as the text content.
[0095] In another embodiment, the program plugin includes a set of operation information to be selected; when the acquisition unit 701 acquires the target operation information determined by the target object in the target application, it may specifically be used to:
[0096] When a text selection instruction from the target object in the target application is detected, an operation display area is determined according to the display area where the text content is located, and the operation information set is displayed in the operation display area; and, when a first operation selection instruction from the target object on the operation information set is detected, the operation information selected by the first operation selection instruction is taken as the target operation information determined by the target object in the target application; or...
[0097] The program plugin is displayed in the target application, and when a second operation selection instruction is detected from the target object for the operation information set in the program plugin, the operation information indicated by the second operation selection instruction is used as the target operation information; or...
[0098] When an operation input instruction from the target object to the program plugin is detected, the operation information indicated by the operation input instruction is taken as the target operation information.
[0099] In another embodiment, the target application is a target browser, and the program plugin is a browser plugin applied to the target browser.
[0100] According to one embodiment of the present invention, Figure 1 or Figure 3 Each step involved in the method shown can be derived from... Figure 7 The dialogue is performed by individual units within the robot dialogue device shown. For example, Figure 1 Step S101 shown can be performed by Figure 7 The acquisition unit 701 shown is executed, and steps S102 and S103 can both be performed by... Figure 7 The processing unit 702 shown executes this. For example, Figure 3 Step S301 shown can be performed by Figure 7 The acquisition unit 701 shown is executed, and steps S302-S305 can all be performed by... Figure 7 The processing unit 702 shown executes, etc.
[0101] According to another embodiment of the present invention, Figure 7Each unit in the illustrated robot dialogue device can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of the present invention. The above-mentioned units are divided based on logical functions. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of the present invention, any robot dialogue device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0102] According to another embodiment of the present invention, the following can be performed by running on a general-purpose electronic device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 1 or Figure 3 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 7 The robot dialogue device shown herein, and the robot dialogue method for implementing embodiments of the present invention, are described. The computer program may be recorded on, for example, a computer storage medium, loaded onto the aforementioned electronic device via the computer storage medium, and run therein.
[0103] This invention allows for the invocation of a program plugin within a target application to obtain the text content and target operation information determined by the target object within the application. The text content and target operation information are then sent to a model service device, enabling the model service device to generate a text response result based on the text content and target operation information. This response result includes the target response content executed according to the target operation information. The program plugin supports the target object establishing a robot dialogue with the model service device within the target application. The device can then receive and display the target response content returned by the model service device. Therefore, this invention allows for convenient robot dialogue via a program plugin, effectively improving dialogue efficiency.
[0104] Based on the description of the method and apparatus embodiments above, an exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method according to an embodiment of the present invention.
[0105] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0106] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0107] refer to Figure 8 The present invention will now be described in the form of a structural block diagram of an electronic device 800 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0108] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0109] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. Input unit 806 can be any type of device capable of inputting information to electronic device 800. Input unit 806 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 808 may include, but is not limited to, disks and optical discs. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0110] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the robot dialogue method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the robot dialogue method by any other suitable means (e.g., by means of firmware).
[0111] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0112] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0113] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0114] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0115] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0116] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0117] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A robot dialogue method, characterized in that, include: The program plugin in the target application is invoked to obtain the text content determined by the target object in the target application, and to obtain the target operation information determined by the target object in the target application. The program plugin is used to support the target object to establish robot dialogue with the model service device on the target application. The text content and the target operation information are sent to the model service device, so that the model service device generates a text response result based on the text content and the target operation information. The text response result includes the target response content executed by the text content according to the target operation information. The text response result also includes a response identifier, which is used to indicate the target response content. Receive the response identifier returned by the model service device; Based on the response identifier, the first generated result in the target response content is obtained from the model service device, and the first generated result is displayed. The first generated result includes M characters in the target response content, where M is a positive integer. The response status identifier corresponding to the target response content is determined. If the response status identifier is an "not answered" identifier, the second generated result in the target response content is obtained from the model service device based on the response identifier until the response status identifier is updated to an "answered" identifier, so as to receive the target response content returned by the model service device and display the target response content. The second generated result includes N characters in the target response content, where N is a positive integer.
2. The method according to claim 1, characterized in that, Sending the text content and the target operation information to the model service device includes: Based on the target session identifier corresponding to the text content, the text content and the target operation information are sent to the front-end service device, so that the front-end service device sends the text content and the target operation information to the model service device according to the target session identifier; The front-end service device is used to: manage the correspondence between each object in at least one object and each session in at least one session according to the session identifier, and the front-end service device is: a relay device for the model service device to conduct robot dialogue, and the target object is any one of the at least one objects.
3. The method according to claim 2, characterized in that, The method further includes: Locate the target session identifier; If the target session identifier is found, then the target session identifier corresponding to the text content is used to send the text content and the target operation information to the front-end service device. If the target session identifier is not found, a create session instruction is sent to the front-end service device, so that the model service device creates the target session indicated by the target session identifier after receiving the initialized create session instruction sent by the front-end service device. The initialized create session instruction carries the target session identifier; and the model service device receives the target session identifier returned by the front-end service device.
4. The method according to claim 1, characterized in that, The method further includes: After the target response content is generated by the model service device, the system receives a completed response flag sent by the model service device. The response status identifier is updated using the completed response identifier.
5. The method according to any one of claims 1-3, characterized in that, The step of obtaining the text content of the target object as determined in the target application includes: When a text selection instruction is detected from the target object in the target application, the content selected by the text selection instruction is taken as the text content determined by the target object in the target application; or, The text input area of the program plugin is displayed in the target application. When a text editing instruction is detected on the text input area by the target object, the content indicated by the text editing instruction is used as the text content; or, When a voice input command from the target object to the program plugin is detected, the content indicated by the voice input command is used as the text content.
6. The method according to any one of claims 1-3, characterized in that, The program plugin includes a set of operation information to be selected; The step of obtaining the target operation information determined by the target object in the target application includes: When a text selection instruction from the target object in the target application is detected, an operation display area is determined according to the display area where the text content is located, and the operation information set is displayed in the operation display area; and, when a first operation selection instruction from the target object on the operation information set is detected, the operation information selected by the first operation selection instruction is taken as the target operation information determined by the target object in the target application; or... The program plugin is displayed in the target application, and when a second operation selection instruction is detected from the target object for the operation information set in the program plugin, the operation information indicated by the second operation selection instruction is used as the target operation information; or... When an operation input instruction from the target object to the program plugin is detected, the operation information indicated by the operation input instruction is taken as the target operation information.
7. The method according to any one of claims 1-3, characterized in that, The target application is the target browser, and the program plugin is a browser plugin applied to the target browser.
8. A robot dialogue device, characterized in that, The device includes: The acquisition unit is used to call the program plugin in the target application to acquire the text content determined by the target object in the target application, and to acquire the target operation information determined by the target object in the target application. The program plugin is used to support the target object to establish robot dialogue with the model service device on the target application. The processing unit is configured to send the text content and the target operation information to the model service device, so that the model service device generates a text response result based on the text content and the target operation information. The text response result includes the target response content executed by the text content according to the target operation information. The text response result also includes a response identifier, which is used to indicate the target response content. The processing unit is further configured to receive the response identifier returned by the model service device; obtain a first generation result from the target response content based on the response identifier from the model service device, and display the first generation result, wherein the first generation result includes M characters in the target response content, where M is a positive integer; determine the response status identifier corresponding to the target response content; if the response status identifier is an incomplete response identifier, then continue to obtain a second generation result from the target response content based on the response identifier until the response status identifier is updated to a completed response identifier, so as to receive the target response content returned by the model service device and display the target response content, wherein the second generation result includes N characters in the target response content, where N is a positive integer.
9. An electronic device, characterized in that, include: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice interaction response method and device, storage medium and electronic equipment
CN114268696A
Session information interaction method and device and storage medium
CN114529304A
Robot conversation method and device
CN115470800A