Information processing method, information interaction method, and system, device and medium

By using language models to identify intentions and extract parameters in digital offline conference rooms, the problems of slow response speed and inconvenient operation of conference software are solved, fast and convenient voice control is achieved, and user experience is improved.

WO2025139661A1PCT designated stage expired Publication Date: 2025-07-03BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/136784
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-12-04
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the server response of the digital offline conference room is slow, the voice commands are complex, resulting in poor user experience, and the control conference software requires carrying a terminal device or close to the controller or touch screen, which is inconvenient to operate.

Method used

By obtaining natural language information, using the language model to perform intent recognition and parameter extraction, it is split into two subtasks: intent recognition and parameter extraction, first identify the task type and then extract parameter information, reducing the resource consumption of language model processing and improving response speed.

Benefits of technology

It effectively reduces the interface delay of the information processing process, improves the response speed and user experience of the conference software, and users can control the conference software through voice interaction without carrying a terminal device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136784_03072025_PF_FP_ABST
    Figure CN2024136784_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are an information processing method, an information interaction method, and a system, a device and a medium. The information processing method comprises: acquiring first information that represents a natural language, wherein the first information is used for instructing the execution of a target task; using a language model to perform intent recognition on the first information, so as to determine the task type of the target task; using the language model to perform parameter extraction on the first information, so as to obtain parameter information matching the task type; and on the basis of the task type and the parameter information, executing the target task.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information interaction method, system, device and medium

[0001] This application claims priority to Chinese Patent Application No. 202311812645.0 filed on December 26, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] Embodiments of the present disclosure relate to an information processing method, an information interaction method, a system, an electronic device, and a computer-readable storage medium. Background Art

[0003] With the continuous advancement of computer technology, digital offline conference rooms have emerged. These conference rooms are equipped with hardware such as a host computer, speakers, microphones, and cameras, eliminating the need for users to debug their equipment. Furthermore, the host computer is equipped with conferencing software, allowing users to create and sign in to meetings efficiently.

[0004] When users attend meetings in digital offline conference rooms, they can control the conference software through controller hardware, terminal devices, or touch screens. However, in these methods, the conference software server has difficulty achieving fast and efficient responses. Summary of the Invention

[0005] The present disclosure provides an information processing method. This method can reduce interface latency during information processing and improve response speed. The present disclosure also provides an information interaction method, system, electronic device, computer-readable storage medium, and computer program product corresponding to the above method.

[0006] In a first aspect, the present disclosure provides an information processing method, the method comprising:

[0007] Acquire first information representing natural language, where the first information is used to instruct execution of a target task;

[0008] Using a language model to perform intent recognition on the first information to determine a task type of the target task;

[0009] Using the language model to extract parameters from the first information to obtain parameter information matching the task type;

[0010] The target task is executed according to the task type and the parameter information.

[0011] In some possible implementations, the using a language model to perform intent recognition on the first information and determine the task type of the target task includes:

[0012] generating first prompt information according to the first information and a first prompt template, wherein the first prompt template includes attribute information of multiple task types;

[0013] The first prompt information is sent to a language model, and the task type of the target task returned by the language model is received.

[0014] In some possible implementations, sending the first prompt information to a language model and receiving the task type of the target task returned by the language model includes:

[0015] Sending the first prompt information to a language model so that the language model determines the task type of the target task, and determines the target character corresponding to the task type of the target task based on a mapping relationship between the task type and the character;

[0016] The target character returned by the language model is received.

[0017] In some possible implementations, sending the first prompt information to a language model and receiving the task type of the target task returned by the language model includes:

[0018] The first prompt information is sent to a language model, so that the language model determines whether the task type associated with the first information belongs to the multiple task types based on the attribute information of the multiple task types, and in response to the task type associated with the first information belonging to the multiple task types, determines the task type of the target task and returns the task type of the target task.

[0019] In some possible implementations, extracting parameters from the first information using the language model to obtain parameter information matching the task type includes:

[0020] generating second prompt information according to the first information and a second prompt template, wherein the second prompt template includes at least one parameter item matching the task type;

[0021] The second prompt information is sent to the language model, and parameter information corresponding to the at least one parameter item returned by the language model is received.

[0022] In some possible implementations, obtaining first information representing the natural language includes:

[0023] Obtain user input information and historical conversation information for the current conversation round;

[0024] In response to the historical dialogue information of the current dialogue round being empty, determining first information according to the user input information;

[0025] In response to the historical conversation information of the current conversation round being not empty, first information is determined according to the user's input information and the historical conversation information of the current conversation round.

[0026] In some possible implementations, the first information includes user input information and historical conversation information of the current conversation turn, the historical conversation information of the current conversation turn includes information about the first task, and using a language model to perform intent recognition on the first information and determine the task type of the target task includes:

[0027] A language model is used to identify the intent of the first information, so that the language model determines whether the first information is used to confirm the first task. In response to the first information being used to confirm the first task and the confirmation result of the first information for the first task is yes, the task type of the first task is used as the task type of the target task.

[0028] In some possible implementations, obtaining first information representing the natural language includes:

[0029] In response to a user's voice input request, obtaining a voice instruction input by the user;

[0030] Perform voice recognition on the voice instruction to determine first information representing natural language.

[0031] In a second aspect, the present disclosure provides an information processing method, the method comprising:

[0032] Acquire first information representing natural language, where the first information is used to instruct execution of a target task related to the meeting;

[0033] Using a language model to perform intent recognition on the first information, and determining a task type of the target task related to the meeting;

[0034] Using the language model to extract parameters from the first information to obtain parameter information matching the task type;

[0035] The target task related to the meeting is executed according to the task type and the parameter information.

[0036] In some possible implementations, the method further includes:

[0037] Generating an execution result after executing the target task related to the meeting;

[0038] The execution result is displayed in the client deployed on the conference room device.

[0039] In some possible implementations, the task type of the target task related to the meeting includes at least one of the following:

[0040] Create a meeting, sign in to a meeting, invite others to join a meeting, generate meeting minutes, leave a meeting, and control meeting room equipment.

[0041] In some possible implementations, the task type is creating a meeting, and the parameter information matching the task type includes at least one of the following: meeting member information, meeting duration information, and meeting identification information;

[0042] The task type is signing in or leaving a meeting, and the parameter information matching the task type includes meeting identification information;

[0043] The task type is to invite others to join a meeting, and the parameter information matching the task type includes at least one of the following: meeting identification information and information of members to be invited;

[0044] The task type is generating meeting minutes, and the parameter information matching the task type includes at least one of the following: meeting identification information, and information about members receiving meeting minutes;

[0045] The task type is controlling conference room equipment, and parameter information matching the task type includes at least one of the following: conference room equipment information and control parameter information.

[0046] In a third aspect, the present disclosure provides an information interaction method, the method comprising:

[0047] receiving first information representing natural language, where the first information is used to instruct execution of a target task related to the meeting;

[0048] After the server executes a target task related to the conference, receiving an execution result sent by the server, the execution result including a task type of the target task and parameter information matching the task type;

[0049] The execution result is presented.

[0050] In a fourth aspect, the present disclosure provides an information processing system, the system comprising:

[0051] an acquisition module, configured to acquire first information representing a natural language, wherein the first information is used to instruct execution of a target task;

[0052] an identification module, configured to use a language model to perform intent recognition on the first information and determine a task type of the target task;

[0053] an extraction module, configured to extract parameters from the first information using the language model to obtain parameter information matching the task type;

[0054] An execution module is used to execute the target task according to the task type and the parameter information.

[0055] In some possible implementations, the identification module is specifically configured to:

[0056] generating first prompt information according to the first information and a first prompt template, wherein the first prompt template includes attribute information of multiple task types;

[0057] The first prompt information is sent to a language model, and the task type of the target task returned by the language model is received.

[0058] In some possible implementations, the identification module is specifically configured to:

[0059] Sending the first prompt information to a language model so that the language model determines the task type of the target task, and determines the target character corresponding to the task type of the target task based on a mapping relationship between the task type and the character;

[0060] The target character returned by the language model is received.

[0061] In some possible implementations, the identification module is specifically configured to:

[0062] The first prompt information is sent to a language model, so that the language model determines whether the task type associated with the first information belongs to the multiple task types based on the attribute information of the multiple task types, and in response to the task type associated with the first information belonging to the multiple task types, determines the task type of the target task and returns the task type of the target task.

[0063] In some possible implementations, the extraction module is specifically configured to:

[0064] generating second prompt information according to the first information and a second prompt template, wherein the second prompt template includes at least one parameter item matching the task type;

[0065] The second prompt information is sent to the language model, and parameter information corresponding to the at least one parameter item returned by the language model is received.

[0066] In some possible implementations, the acquisition module is specifically configured to:

[0067] Obtain user input information and historical conversation information for the current conversation round;

[0068] In response to the historical dialogue information of the current dialogue round being empty, determining first information according to the user input information;

[0069] In response to the historical conversation information of the current conversation round being not empty, first information is determined according to the user's input information and the historical conversation information of the current conversation round.

[0070] In some possible implementations, the first information includes user input information and historical conversation information of a current conversation round, the historical conversation information of the current conversation round includes information about the first task, and the recognition module is specifically configured to:

[0071] A language model is used to identify the intent of the first information, so that the language model determines whether the first information is used to confirm the first task. In response to the first information being used to confirm the first task and the confirmation result of the first information for the first task is yes, the task type of the first task is used as the task type of the target task.

[0072] In some possible implementations, the acquisition module is specifically configured to:

[0073] In response to a user's voice input request, obtaining a voice instruction input by the user;

[0074] Perform voice recognition on the voice instruction to determine first information representing natural language.

[0075] In a fifth aspect, the present disclosure provides an electronic device, comprising a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory to cause the electronic device to perform the method according to the first aspect or any implementation of the first aspect, the second aspect or any implementation of the second aspect, or the third aspect.

[0076] In a sixth aspect, the present disclosure provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct an electronic device to execute the above-mentioned first aspect or any one of the implementations of the first aspect, or the second aspect or any one of the implementations of the second aspect, or the method in the third aspect.

[0077] In the seventh aspect, the present disclosure provides a computer program product comprising instructions, which, when run on an electronic device, enables the electronic device to execute the above-mentioned first aspect or any one of the implementations of the first aspect, or the second aspect or any one of the implementations of the second aspect, or the method in the third aspect.

[0078] Based on the implementations provided in the above aspects, the present disclosure can be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the technical methods of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments.

[0080] FIG1 is a flow chart of an information processing method provided by an embodiment of the present disclosure;

[0081] FIG2 is a flow chart of another information processing method provided by an embodiment of the present disclosure;

[0082] FIG3 is a flow chart of another information processing method provided by an embodiment of the present disclosure;

[0083] FIG4 is a flow chart of a voice conversation according to an embodiment of the present disclosure;

[0084] FIG5 is a flow chart of a voice module state change process according to an embodiment of the present disclosure;

[0085] FIG6 is a schematic diagram of an information interaction process provided by an embodiment of the present disclosure;

[0086] FIG7 is a schematic diagram of the structure of an information processing system provided by an embodiment of the present disclosure;

[0087] FIG8 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0088] The terms "first" and "second" in the embodiments of the present disclosure are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, the terms "first" and "second" may explicitly or implicitly include one or more of the features.

[0089] First, some technical terms involved in the embodiments of the present disclosure are introduced.

[0090] Deploying hardware such as a host computer, speakers, microphones, and cameras in offline meeting rooms, along with installing conferencing software on the host computer, has become a mature digital solution for offline meeting rooms. Users can quickly join meetings in digital offline meeting rooms without having to debug their equipment. Furthermore, users can use the conferencing software to create and sign in to meetings, connecting online and offline meetings for efficient meetings.

[0091] In a digital offline conference room, users can control the conference software in different ways. For example, a matching controller hardware device can be deployed on the desktop of the offline conference room, and the user can control the conference software by operating the controller. For another example, the user can use a terminal device (such as a mobile phone) to scan a QR code to open the Hypertext 5.0 (HTML5.0, H5) page of the conference software and control the conference software through the H5 page. For another example, the offline conference room can be equipped with a touch screen, and the user can control the conference software through the touch screen.

[0092] However, the above-mentioned methods of controlling conference software either require users to carry terminal devices for control or require users to walk near the controller or touch screen for control, both of which are inconvenient. Therefore, in the present disclosure, a digital offline conference room can support users to control the conference software through voice interaction.

[0093] Specifically, users can enter voice commands through the conferencing software client, and the server performs corresponding operations based on the voice commands, such as creating a schedule, signing in to a meeting, and ending a meeting. In this way, voice interaction can overcome the limitations of operating distance, eliminating the need for users to carry terminal devices to offline meeting rooms, and improving the convenience of user control of conferencing software.

[0094] However, in the above-mentioned voice interaction method of controlling the conference software, since the voice commands are relatively complex and the server's response speed is slow, the high response delay can reduce the user experience in the digital offline conference room.

[0095] In view of this, the present disclosure provides an information processing method. The method first obtains first information representing natural language, wherein the first information is used to indicate the execution of a target task, then uses a language model to identify the intent of the first information to determine the task type of the target task, and then uses the language model to extract parameters from the first information to obtain parameter information that matches the task type. The target task is then executed based on the task type and parameter information.

[0096] This method splits intent recognition and parameter extraction into two subtasks: first, using a language model to identify the task type, and then extracting parameter information. This allows the language model to process only the simpler subtasks during a natural language task, while also enabling targeted parameter extraction based on the target task. This effectively reduces the resources consumed by the language model during natural language task processing, lowers interface latency during information processing, and improves response speed.

[0097] To facilitate understanding of the technical solutions provided by the embodiments of the present disclosure, they will be described below with reference to the accompanying drawings.

[0098] Referring to FIG. 1 , a flow chart of an information processing method provided by an embodiment of the present disclosure is shown. The method can be applied to a server and specifically includes:

[0099] S101: Acquire first information representing a natural language.

[0100] The first information generally refers to information input by the user and used to express a specific instruction. In the embodiment of the present disclosure, the first information represents natural language. In other words, the user can input the first information by inputting natural language.

[0101] In addition, the first information may be used to indicate the execution of a target task. The target task may be different in different scenarios. For example, in a meeting scenario, the target task may be related to the meeting. For another example, in a word processing scenario, the target task may be related to word processing needs.

[0102] In some possible implementations, the user may input information through a voice dialogue. In a specific implementation, in response to the user's voice input request, a voice instruction input by the user is obtained, voice recognition is performed on the voice instruction, and the first information representing the natural language is determined.

[0103] In some embodiments, a voice input request can be associated with a wake-up word that is used to start audio capture. In other words, when the user speaks the wake-up word, a voice input request is generated. In response to the voice input request, an audio capture tool can be used to capture the user's voice command input, such as an audio capture tool based on real-time communication (RTC) to capture high-quality voice commands.

[0104] After completing the collection of the voice command, the server can perform voice recognition on the voice command and convert the voice command into text information (i.e., the first information). For example, a voice recognition tool based on automatic speech recognition (ASR) technology can be used to perform voice recognition to obtain the first information with high recognition accuracy.

[0105] In a human-computer interaction, a conversation turn can be understood as a conversation within a topic. A conversation turn consists of a user input and a server response to that input, i.e., a "question-answer" conversation. In the disclosed embodiments, conversation turns can be distinguished by wake-up words, meaning each time a user says the wake-up word, it is considered the start of a new conversation turn.

[0106] Taking into account that a conversation round may include one or more conversation rounds, in order to facilitate accurate information processing by the subsequent language model, in an embodiment of the present disclosure, the user's input information and the historical conversation information of the current conversation round can be obtained. In response to the historical conversation information of the current conversation round being empty, the first information is determined based on the user's input information. In response to the historical conversation information of the current conversation round being not empty, the first information is determined based on the user's input information and the historical conversation information of the current conversation round.

[0107] It is understood that if there is no historical conversation information for the current conversation turn, it indicates that the user's input information is the content of the first conversation turn. In this case, the first information may only include the user's input information. If there is historical conversation information for the current conversation turn, it indicates that the user's input information is not the content of the first conversation turn. In this case, the first information may include the user's input information and the content of multiple conversation turns in the current conversation turn. In this way, in the scenario of multiple conversation turns, the subsequent natural language task processing of the speech model on the first information can be combined with historical conversation information, thereby improving the accuracy of natural language task processing.

[0108] S102: Using a language model to perform intent recognition on the first information, and determining a task type of the target task.

[0109] The language model can be a model with natural language processing capabilities and capable of handling natural language tasks, such as a deep learning model trained on text data. Intent recognition can be understood as the language model analyzing the first information to identify the intended intent of the first information, representing natural language, i.e., the task type of the target task indicated by the first information.

[0110] In specific implementation, the server can generate first prompt information based on the first information and the first prompt template, where the first prompt template includes attribute information of multiple task types, and then send the first prompt information to the language model and receive the task type of the target task returned by the language model.

[0111] In the disclosed embodiment, intent recognition is achieved using a prompt-based learning approach. The first prompt template includes attribute information of multiple task types, where the attribute information of the task type can be natural language information describing the task type from multiple aspects, such as descriptive information of the task type and parameter items matching the task type. By generating a first prompt message (prompt) from the first prompt template and the first information, the language model can analyze the first information based on the attribute information of the multiple task types, thereby determining the task type of the target task and achieving intent recognition.

[0112] In some embodiments, to improve the accuracy of intent recognition, the language model may first perform fuzzy intent recognition, and then perform detailed intent recognition. Specifically, the server may send the first prompt information to the language model, so that the language model can determine whether the task type associated with the first information belongs to multiple task types based on the attribute information of multiple task types. In response to the task type associated with the first information belonging to multiple task types, the language model determines the task type of the target task and returns the task type of the target task.

[0113] In other words, because the first prompt template includes attribute information of multiple task types, the language model can understand the task types that the server can handle through the first prompt template. Furthermore, through the content in the first prompt information that guides the language model to first perform fuzzy range intent recognition and then perform detailed intent recognition, the language model, based on the guidance of the first prompt information, determines whether the server can handle the task type associated with the first information (i.e., the task type of the target task). When the server can handle the task type associated with the first information, it further specifically determines the task type of the target task. In this way, the recognition accuracy of the language model is improved.

[0114] Implementing intent recognition through prompt learning effectively improves the scalability of language models. As the types of tasks a server can handle increase, simply add the task type's attributes to the first prompt template. This first prompt informs the language model that the server can handle the new task type, eliminating the need for extensive model training and debugging.

[0115] Furthermore, in some embodiments, considering that errors may occur in the language model's intent recognition, the server can also provide secondary confirmation to the user, that is, present the intent recognition result to the user and ask the user whether the intent recognition result is correct. Since the user can perform secondary confirmation on the intent recognition result sent by the server (for example, reply confirmation, incorrect or exit, etc.), the current round of dialogue can include multiple rounds of dialogue. At this time, the language model can first determine whether the first information is used for secondary confirmation, and then determine the task type of the target task based on the judgment result.

[0116] Specifically, the first information may include the user's input information and historical conversation information from the current conversation round, where the historical conversation information from the current conversation round includes information about the first task (i.e., the intent recognition result returned by the language model). The language model is used to identify the intent of the first information, so that the language model can determine whether the first information is used to confirm the first task. In response to the first information being used to confirm the first task and the confirmation result of the first information for the first task is yes, the task type of the first task is used as the task type of the target task.

[0117] It is understandable that, in response to the first information not being used to confirm the first task, the server may determine the task type of the target task according to the original process.

[0118] In other words, for multi-round dialogue scenarios, the historical dialogue information includes the intent recognition results returned by the language model. If the first information is not used for secondary confirmation of the intent recognition result, the language model can analyze the first information for intent recognition and return the task type of the target task. If the first information is used for secondary confirmation of the intent recognition result, and the result of the secondary confirmation is yes, the language model can use the intent recognition result (i.e., the task type of the first task) directly as the task type of the target task. In this way, the language model first determines whether the first information is used for secondary confirmation, and then determines the task type of the target task based on the judgment result, which can further improve the accuracy of intent recognition.

[0119] The language model consumes a certain amount of resources (also called tokens) when processing natural language tasks. The number of tokens consumed by the language model is related to the number of characters. That is, the more characters in the processing results returned by the language model, the more tokens are consumed. To reduce the number of tokens consumed by the language model, in the embodiments of the present disclosure, a mapping relationship between tasks and characters can be pre-established, thereby simplifying the results returned by the language model after intent recognition into characters.

[0120] In specific implementation, the first prompt information can be generated based on the first information and the first prompt template, and the first prompt information can be sent to the language model, so that the language model determines the task type of the target task, and determines the target character corresponding to the task type of the target task based on the mapping relationship between the task type and the character, and returns the target character.

[0121] In the disclosed embodiment, a mapping relationship between task types and characters is pre-established. After the language model performs intent recognition and obtains the task type of the target task, the target character corresponding to the task type is returned based on this mapping relationship. This simplifies the output of the language model to characters, reducing the number of tokens consumed by the language model.

[0122] In addition, for the secondary confirmation scenario, considering that the user may not confirm the intent recognition result presented by the server, but instead reply with other content (for example, to exit the current conversation), in the embodiment of the present disclosure, when the first information includes the user's input information and the historical conversation information of the current conversation round (i.e., there are multiple conversation rounds), multiple intent recognition tasks can be called in parallel. In other words, the language model can concurrently determine whether the first information is a new instruction, whether it is for secondary confirmation, or whether it is to exit the current conversation. In this way, while allowing users to exit the current conversation, it further reduces latency and improves response speed.

[0123] S103: Utilize the language model to extract parameters from the first information to obtain parameter information matching the task type.

[0124] After the language model completes intent recognition, the first information can be further subjected to parameter extraction to obtain parameter information that matches the task type. The parameter information that matches the task type can be parameter information required to perform the target task, such as time information, personnel information, etc.

[0125] In specific implementation, the server can generate second prompt information based on the first information and the second prompt template, wherein the second prompt template includes at least one parameter item matching the task type, and then send the second prompt information to the language model, and receive parameter information corresponding to the at least one parameter item returned by the language model.

[0126] In the disclosed embodiments, parameter extraction is achieved using a prompt-based learning approach. Because the second prompt template includes at least one parameter item that matches the task type, the second prompt information is generated using the first information and the second prompt template. The language model analyzes the second prompt information and, based on the parameter item that matches the target task type, extracts the parameter information corresponding to the parameter item from the first information, thereby completing parameter extraction.

[0127] By splitting intent recognition and parameter extraction into two subtasks, the language model can perform parameter extraction on the first information after the task type of the target task has been determined. In this way, parameter information can be extracted based on the known parameter items that match the task type of the target task, avoiding mutual interference between parameter information of the same type under different intentions, and improving the accuracy of parameter extraction by the language model.

[0128] In some embodiments, considering that the process of parameter extraction for some target task types is relatively simple, the server can also perform parameter extraction directly by running back-end code instead of using a language model. Specifically, for some simpler task types, developers can pre-write code for parameter extraction. When the language model performs intent recognition and determines that the target task type matches the aforementioned simpler task type, the server can directly run the back-end code for parameter extraction without using a language model, thereby further reducing the token consumption of the language model and interface latency.

[0129] S104: Execute the target task according to the task type and parameter information.

[0130] After completing intent recognition and parameter extraction, the server can execute the target task. In specific implementation, the server can execute the target task by calling the target task's application programming interface (API).

[0131] Based on the above description, the embodiments of the present disclosure provide an information processing method. The method first obtains first information representing natural language, wherein the first information is used to indicate the execution of a target task, then uses a language model to perform intent recognition on the first information to determine the task type of the target task, then uses the language model to perform parameter extraction on the first information to obtain parameter information that matches the task type, and then executes the target task based on the task type and parameter information.

[0132] This method splits intent recognition and parameter extraction into two subtasks: first, using a language model to identify the task type, and then extracting parameter information. This allows the language model to process only the simpler subtasks during a natural language task, while also enabling targeted parameter extraction based on the target task. This effectively reduces the resources consumed by the language model during natural language task processing, lowers interface latency during information processing, and improves response speed.

[0133] As can be seen from the foregoing, the information processing method provided by the embodiments of the present disclosure can be applied to conference scenarios. Referring to FIG2 , which illustrates a flow chart of another information processing method, in some embodiments, the method can be applied to a server, and a client corresponding to the server can be deployed in a conference room device. The method includes:

[0134] S201: Acquire first information representing a natural language.

[0135] The first information may be used to instruct the execution of target tasks related to the meeting, such as meeting creation, meeting sign-in, and meeting minutes generation.

[0136] S202: Use a language model to perform intent recognition on the first information and determine a task type of a target task related to the meeting.

[0137] In a meeting scenario, the task types of target tasks related to a meeting may include at least one of the following: creating a meeting, signing in to a meeting, inviting others to join a meeting, generating meeting minutes, leaving a meeting, and controlling meeting room equipment.

[0138] Similar to the previous content, in the process of intent recognition based on prompt learning, the language model can first complete the fuzzy range of intent recognition, and then subdivide and identify the specific task type. Specifically, the first prompt template can include attribute information of multiple task types related to the meeting. The language model analyzes the first information in combination with the first prompt template to determine whether the task type associated with the first information belongs to the task type related to the meeting that can be handled by the server. For example, when the task type associated with the first information is to answer general knowledge questions, the task type does not belong to the task type that the server can handle. When the task type associated with the first information is to control conference room equipment, the task type belongs to the task type that the server can handle. Then, when the task type associated with the first information belongs to the multiple task types included in the first prompt template, the language model can perform specific intent recognition to determine the task type of the target task, such as lowering the volume of the conference room audio.

[0139] The following is an explanation of the secondary confirmation scenario. In an offline meeting room, the first information in the first round of conversation can be "Please sign in for me". The intent of the first information is recognized through the language model. The server can present "Do you want to sign in for the meeting?" in the client to provide a secondary confirmation to the user. In the second round of conversation, the user's input information can be "Yes". The historical conversation information of the current conversation round can include "Please sign in for me" and "Do you want to sign in for the meeting?". The historical conversation information of the current conversation round includes the task type of the first task (i.e., meeting sign-in). At this time, the first information includes the above-mentioned user input information and the above-mentioned historical conversation information.

[0140] For the above-mentioned multiple rounds of dialogue, the language model can identify the intent of the first information, determine that the first information is used to confirm the first task, and the confirmation result of the first information for the first task is yes. Therefore, the task type of the target task is meeting sign-in, and the intention identification is completed.

[0141] The following describes the process by which a language model returns target characters. In an offline meeting room, the first message in the first round of conversation might be "Sign me in." In this case, if no mapping between task types and characters has been established, the language model, after identifying the intent of the first message, might return "Tool:Check_In_Meeting_Room,Tool Input:{"user_confirm":"FALSE","schedule_id":""}" to the server. After establishing a mapping between tasks and characters, based on the mapping of Check_In_Meeting_Room = X, the language model can return "X" to the server, reducing token consumption.

[0142] Similarly, in the double-confirmation scenario, the first message includes the user's input and historical conversation information for the current conversation turn. The user's input can be "Confirm." In this case, if there's no pre-established mapping between task types and characters, the language model can recognize the intent of the first message and return "Tool:Check_In_Meeting_Room,Tool Input:{"user_confirm":"TRUE","schedule_id":"XXXXXXXXXXXXXXXXXXXX"}" to the server. After establishing a mapping between task types and characters, based on the mapping relationship of TRUE = T, the language model can return "T" to the server, reducing token consumption.

[0143] S203: Utilize the language model to extract parameters from the first information to obtain parameter information matching the task type.

[0144] When the task type is to create a meeting, the parameter information matching the task type may include at least one of the following: meeting member information, meeting duration information, and meeting identification information.

[0145] When the task type is signing in or leaving a meeting, the parameter information matching the task type may include meeting identification information.

[0146] When the task type is to invite others to join a meeting, the parameter information matching the task type may include at least one of the following: meeting identification information and information of members to be invited.

[0147] When the task type is generating meeting minutes, the parameter information matching the task type may include at least one of the following: meeting identification information and information about members receiving the meeting minutes.

[0148] When the task type is controlling conference room equipment, the parameter information matching the task type may include at least one of the following: conference room equipment information and control parameter information.

[0149] S204: Execute target tasks related to the meeting according to the task type and parameter information.

[0150] After the server uses the language model to extract parameter information and executes the target task related to the meeting, the server can generate the execution result after executing the target task related to the meeting and display the execution result in the client of the conference room device.

[0151] In some embodiments, developers can pre-configure the execution results corresponding to target tasks. For example, if the target task is to sign in to meeting A, the pre-configured execution results could be "Sign in Successfully" or "Sign in Failed." In this case, the server can match the execution results based on the execution status of the target task and display the results directly on the client.

[0152] In other embodiments, the language model can generate execution results based on the execution status of the target task and return the execution results to the server, thereby leveraging the natural language processing capabilities of the language model to generate execution results that meet the requirements of the target task. For example, if the target task is to create a one-hour meeting, the language model can generate a card message based on the meeting creation status, and then present the card message to the user on the client.

[0153] Furthermore, given that users can input information through voice conversations, the client can also inform users of the execution results through voice broadcasts. Specifically, this can be achieved by converting the execution results into voice based on text-to-speech (TTS) technology, and playing the corresponding voice to the user, thus completing voice interaction.

[0154] It should be noted that, unless there is any contradiction, the method described in the embodiment of the present disclosure can be combined with the method described in the above embodiment, and the resulting solution is also covered by the protection scope of the embodiment of the present disclosure.

[0155] The information processing method provided by the embodiment of the present disclosure will be introduced below in conjunction with specific scenarios.

[0156] During an offline meeting in an offline conference room, the conference software can provide users with the function of generating meeting minutes for the offline meeting. The information processing process in this scenario will be specifically described below with reference to FIG3.

[0157] Referring to the flowchart of another information processing method shown in Figure 3, after the user speaks the wake-up word, the voice module deployed on the conference room equipment can begin to receive the user's voice command, such as "Help me generate meeting minutes," and perform voice recognition on the voice command to obtain the user's input information. Next, the system checks whether the historical conversation information for the current conversation round is empty. If so, the user's input information is used as the first information. If not, the user's input information and the historical conversation information are used as the first information.

[0158] For single-round conversation scenarios, the language model is used to identify intent and determine the target task's task type. For multi-round conversation scenarios, the language model is used to determine whether the first information is used to confirm the first task (i.e., whether the first information is used for secondary confirmation). If so, the task type of the first task is determined as the target task's task type. If not, the target task's task type is determined. In this case, the target task's task type could be generating meeting minutes.

[0159] Next, a language model is used to extract parameters from the first information to obtain parameter information that matches the task type. In this case, the parameter information that matches the task type may be meeting identification information. Based on the task type and parameter information, the target task is executed, for example, calling a meeting minutes generation tool to generate meeting minutes.

[0160] After executing the target task, the server can generate the execution results and display them on the client, such as displaying "Meeting minutes generated" or displaying a link to the generated meeting minutes document. Furthermore, the server can convert the execution results into speech and broadcast the corresponding speech, thereby informing the user of the execution results through voice interaction. The language model can then determine whether the target task has been completed. If so, the information processing process ends. If not, the client can continue to receive audio and obtain the user's voice instructions.

[0161] In the embodiment of the present disclosure, the user can input information through voice dialogue, and the voice dialogue process will be introduced below.

[0162] Referring to the flow diagram of a voice conversation shown in Figure 4, the voice module is deployed in the conference room equipment of the offline conference room. After the user wakes up the voice module, the voice module obtains the user's voice command and sends the voice command to the server. The server sends the voice command to the voice / text conversion module. The voice / text conversion module performs voice recognition, obtains the user's input information, and returns the user's input information to the server.

[0163] The server then returns the user's input to the voice module, allowing it to present the user's input. The server can also determine the first message based on whether the historical conversation information for the current conversation turn is empty, and send the first message to the language model. The language model performs intent recognition and parameter extraction on the first message, returning the task type and parameter information to the server. Based on the task type and parameter information, the server executes the target task and generates an execution result.

[0164] The server sends the execution result to the user (i.e., the client) to inform them of the execution result. Simultaneously, the server sends the execution result to the speech / text conversion module, which converts the execution result into speech and returns the corresponding speech to the server. The server then sends the speech corresponding to the execution result to the speech module, which plays the speech corresponding to the execution result, completing the voice conversation.

[0165] In the above voice dialogue process, the voice module has different states. Referring to the flow diagram of the voice module state change shown in Figure 5, the initial state of the voice module is standby. After receiving the wake-up word, the state of the voice module changes to wakeup. At the same time, the voice module can send a start request to the voice / text conversion module so that the voice / text conversion module can subsequently perform voice recognition.

[0166] When the speech / text conversion module is activated, the speech module changes to a listening state. In the listening state, the speech module can receive voice commands and send the voice commands to the speech / text conversion module for speech recognition. If the speech / text conversion module is not successfully activated, the speech module changes to a standby state.

[0167] Then, when the end identifier is received, the state of the voice module changes to a waiting state (thinking). The end identifier is used to indicate that the input of the voice command has ended. For example, the end identifier can be the isFinal identifier sent by the server, or the end end point received locally by the voice module. In the waiting state, the server will determine the first information and send the first information to the language model for intent recognition and parameter extraction. If the end identifier is not received within the set time, the state of the voice module changes to a standby state.

[0168] Upon receiving the execution result from the server, the voice module enters the feedback state. If the execution result is not received within the set time, the voice module enters the standby state. Subsequently, if the target task indicated by the first message requires secondary confirmation, the voice module enters the feedback listening state and again determines whether an end indicator has been received. If an end indicator is received, the voice module enters the waiting state. If an end indicator is not received within the set time, the voice module enters the finish state. If the target task indicated by the first message does not require secondary confirmation, the voice module enters the finish state.

[0169] In addition, for some scenarios where the voice dialogue function needs to be disabled, the state of the voice module can be changed to disabled. In this way, the user interface (UI) corresponding to the state of the voice module can be presented on the client side, improving the user experience.

[0170] The present disclosure also provides an information interaction method. Referring to FIG6 , which is a flow chart of an information interaction method, in some embodiments, the method may be applied to a client, which may be deployed in a conference room device. The method includes:

[0171] S601: Receive first information representing a natural language;

[0172] S602: After the server executes the target task related to the conference, the server receives the execution result sent by the server;

[0173] S603: Present the execution result.

[0174] It should be noted that, unless there is any contradiction, the method described in the embodiment of the present disclosure can be combined with the method described in the above embodiment, and the resulting solution is also covered by the protection scope of the embodiment of the present disclosure.

[0175] The above describes in detail the information processing method provided by the embodiment of the present disclosure in conjunction with Figures 1 to 5. The following describes the system and device provided by the embodiment of the present disclosure in conjunction with the accompanying drawings.

[0176] Referring to the structural diagram of the information processing system shown in FIG7 , the system 70 includes:

[0177] An acquisition module 701 is configured to acquire first information representing a natural language, where the first information is used to instruct execution of a target task;

[0178] An identification module 702 is configured to use a language model to perform intent recognition on the first information and determine a task type of the target task;

[0179] An extraction module 703 is configured to extract parameters from the first information using the language model to obtain parameter information matching the task type;

[0180] The execution module 704 is configured to execute the target task according to the task type and the parameter information.

[0181] In some possible implementations, the identification module 702 is specifically configured to:

[0182] generating first prompt information according to the first information and a first prompt template, wherein the first prompt template includes attribute information of multiple task types;

[0183] The first prompt information is sent to a language model, and the task type of the target task returned by the language model is received.

[0184] In some possible implementations, the identification module 702 is specifically configured to:

[0185] Sending the first prompt information to a language model so that the language model determines the task type of the target task, and determines the target character corresponding to the task type of the target task based on a mapping relationship between the task type and the character;

[0186] The target character returned by the language model is received.

[0187] In some possible implementations, the identification module 702 is specifically configured to:

[0188] The first prompt information is sent to a language model, so that the language model determines whether the task type associated with the first information belongs to the multiple task types based on the attribute information of the multiple task types, and in response to the task type associated with the first information belonging to the multiple task types, determines the task type of the target task and returns the task type of the target task.

[0189] In some possible implementations, the extraction module 703 is specifically configured to:

[0190] generating second prompt information according to the first information and a second prompt template, wherein the second prompt template includes at least one parameter item matching the task type;

[0191] The second prompt information is sent to the language model, and parameter information corresponding to the at least one parameter item returned by the language model is received.

[0192] In some possible implementations, the obtaining module 701 is specifically configured to:

[0193] Obtain user input information and historical conversation information for the current conversation round;

[0194] In response to the historical dialogue information of the current dialogue round being empty, determining first information according to the user input information;

[0195] In response to the historical conversation information of the current conversation round being not empty, first information is determined according to the user's input information and the historical conversation information of the current conversation round.

[0196] In some possible implementations, the first information includes user input information and historical conversation information of the current conversation round, the historical conversation information of the current conversation round includes information about the first task, and the identification module 702 is specifically configured to:

[0197] A language model is used to identify the intent of the first information, so that the language model determines whether the first information is used to confirm the first task. In response to the first information being used to confirm the first task and the confirmation result of the first information for the first task is yes, the task type of the first task is used as the task type of the target task.

[0198] In some possible implementations, the obtaining module 701 is specifically configured to:

[0199] In response to a user's voice input request, obtaining a voice instruction input by the user;

[0200] Perform voice recognition on the voice instruction to determine first information representing natural language.

[0201] According to the embodiment of the present disclosure, the information processing system 70 may correspond to executing the method described in the embodiment of the present disclosure, and the above-mentioned and other operations and / or functions of each module / unit of the information processing system 70 are respectively for implementing the corresponding processes of each method in the embodiment shown in Figure 1. For the sake of brevity, they will not be repeated here.

[0202] The present disclosure also provides an electronic device, which is specifically configured to implement the functions of the information processing system 70 in the embodiment shown in FIG. 7 .

[0203] FIG8 is a schematic diagram of the structure of an electronic device 800. As shown in FIG8, the electronic device 800 includes a bus 801, a processor 802, a communication interface 803, and a memory 804. The processor 802, the memory 804, and the communication interface 803 communicate with each other via the bus 801.

[0204] Bus 801 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, FIG8 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.

[0205] The processor 802 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0206] The communication interface 803 is used for external communication. For example, the communication interface 803 can be used for communication with a terminal.

[0207] The memory 804 may include volatile memory, such as random access memory (RAM), or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0208] The memory 804 stores executable codes, and the processor 802 executes the executable codes to perform the aforementioned information processing method.

[0209] Specifically, when implementing the embodiment shown in FIG7 , and when each module or unit of the information processing system 70 described in the embodiment of FIG7 is implemented by software, the software or program code required to execute the functions of each module / unit in FIG7 may be partially or completely stored in the memory 804. The processor 802 executes the program code corresponding to each unit stored in the memory 804 to perform the aforementioned information processing method.

[0210] The present disclosure also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned information processing method applied to the information processing system 70.

[0211] The present disclosure also provides a computer program product comprising one or more computer instructions that, when loaded and executed on a computing device, fully or partially generate the process or function described in the present disclosure.

[0212] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0213] When the computer program product is executed by a computer, the computer performs any of the aforementioned information processing methods. The computer program product may be a software installation package, and when any of the aforementioned information processing methods is needed, the computer program product may be downloaded and executed on the computer.

[0214] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to the various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0215] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.

[0216] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0217] In the context of the embodiments of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0218] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0219] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0220] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0221] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0222] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present disclosure. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is intended to be construed in the widest manner consistent with the principles and novel features disclosed herein.

Claims

1. An information processing method, comprising: Obtaining first information representing natural language, where the first information is used to indicate the execution of a target task; Performing intent recognition on the first information by using a language model to determine the task type of the target task; Performing parameter extraction on the first information by using the language model to obtain parameter information matching the task type; Executing the target task according to the task type and the parameter information.

2. The information processing method according to claim 1, wherein, The performing intent recognition on the first information by using a language model to determine the task type of the target task includes: Generating first prompt information according to the first information and a first prompt template, where the first prompt template includes attribute information of multiple task types; Sending the first prompt information to the language model and receiving the task type of the target task returned by the language model.

3. The information processing method according to claim 2, wherein, The sending the first prompt information to the language model and receiving the task type of the target task returned by the language model includes: Sending the first prompt information to the language model, so that the language model determines the task type of the target task and determines a target character corresponding to the task type of the target task according to the mapping relationship between the task type and the character; and Receiving the target character returned by the language model.

4. The information processing method according to claim 2, wherein, The sending the first prompt information to the language model and receiving the task type of the target task returned by the language model includes: Sending the first prompt information to the language model, so that the language model determines whether the task type associated with the first information belongs to the multiple task types according to the attribute information of the multiple task types, and in response to the task type associated with the first information belonging to the multiple task types, determines the task type of the target task and returns the task type of the target task.

5. The information processing method according to claim 1, wherein, The performing parameter extraction on the first information by using the language model to obtain parameter information matching the task type includes: Generating second prompt information according to the first information and a second prompt template, where the second prompt template includes at least one parameter item matching the task type; Sending the second prompt information to the language model and receiving the parameter information corresponding to the at least one parameter item returned by the language model.

6. The information processing method according to any one of claims 1 to 5, wherein, The obtaining first information representing natural language includes: Obtaining the user's input information and the historical conversation information of the current conversation turn; In response to the historical conversation information of the current conversation turn being empty, determining the first information according to the user's input information; In response to the historical conversation information of the current conversation turn not being empty, determining the first information according to the user's input information and the historical conversation information of the current conversation turn.

7. The information processing method according to any one of claims 1 to 6, wherein, The first information includes the user's input information and the historical conversation information of the current conversation turn, and the historical conversation information of the current conversation turn includes information of a first task. The performing intent recognition on the first information by using a language model to determine the task type of the target task includes: Use a language model to perform intent recognition on the first information, so that the language model determines whether the first information is used to confirm the first task. In response to the first information being used to confirm the first task and the confirmation result of the first information for the first task being yes, use the task type of the first task as the task type of the target task.

8. The information processing method according to any one of claims 1 to 7, wherein, The obtaining of the first information representing natural language includes: In response to a voice input request from the user, obtain the voice command input by the user; Perform speech recognition on the voice command to determine the first information representing natural language.

9. An information processing method, including: Obtain first information representing natural language, where the first information is used to indicate the execution of a target task related to a meeting; Use a language model to perform intent recognition on the first information to determine the task type of the target task related to the meeting; Use the language model to extract parameters from the first information to obtain parameter information matching the task type; Execute the target task related to the meeting according to the task type and the parameter information.

10. The information processing method according to claim 9, further including: Generate an execution result after executing the target task related to the meeting; In the client deployed on the conference room device, display the execution result.

11. The information processing method according to claim 9 or 10, wherein The task type of the target task related to the meeting includes at least one of the following: Create a meeting, sign in to a meeting, invite others to join a meeting, generate a meeting minutes, leave a meeting, control conference room equipment.

12. The information processing method according to claim 11, wherein, When the task type is to create a meeting, the parameter information matching the task type includes at least one of the following: meeting member information, meeting duration information, meeting identification information; When the task type is to sign in to or leave a meeting, the parameter information matching the task type includes meeting identification information; When the task type is to invite others to join a meeting, the parameter information matching the task type includes at least one of the following: meeting identification information, member information to be invited; When the task type is to generate a meeting minutes, the parameter information matching the task type includes at least one of the following: meeting identification information, meeting minutes receiving member information; When the task type is to control conference room equipment, the parameter information matching the task type includes at least one of the following: conference room equipment information, control parameter information.

13. An information interaction method, including: Receive first information representing natural language, where the first information is used to indicate the execution of a target task related to a meeting; After the target task related to the meeting is executed on the server, receive the execution result sent by the server, where the execution result includes the task type of the target task and the parameter information matching the task type; Present the execution result.

14. An information processing system, including: An obtaining module, configured to obtain first information representing natural language, where the first information is used to indicate the execution of a target task; An identification module, configured to use a language model to perform intent recognition on the first information to determine the task type of the target task; An extraction module, configured to extract parameters from the first information by using the language model to obtain parameter information matching the task type; An execution module, configured to execute the target task according to the task type and the parameter information.

15. An electronic device, comprising a processor and a memory; Among them, The processor is configured to execute instructions stored in the memory, so that the electronic device executes the information processing method according to any one of claims 1 to 8, or claims 9 to 12, or claim 13.

16. A computer-readable storage medium, comprising instructions, wherein, The instructions instruct the electronic device to execute the information processing method according to any one of claims 1 to 8, or claims 9 to 12, or claim 13.

Citation Information

Patent Citations

  • Questioning type analysis node generation method and system and storage medium

    CN112270189A

  • Task-driven system based on natural language

    CN113515616A

  • Large model plug-in calling method and device, equipment and medium

    CN117112064A

  • Large model plug-in calling method and device, equipment and medium

    CN117112065A

Cited By

  • Multi-agent information processing method and device based on large model, electronic equipment, storage medium and program product

    CN120763321A

  • Task processing method and robot

    CN121187741A