Automatic voice conference management and control method and system based on large language model, and medium
Through the speech automation conference management method of large language model, the limitations of voice control conference management in the existing technology are solved, the full coverage automation of conference operations is realized, and the convenience and efficiency of conference management are improved.
Patent Information
- Application Number
- CN202510451717.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-08
AI Technical Summary
The existing conference management software has limited support for voice automation control, making it difficult for users to complete the creation, query, delete and modify meetings through voice commands, which affects the flexibility and convenience of conference management.
The voice automation conference management method based on the large language model is adopted, voice commands are received through the microphone, and the conference management and control commands are identified, time elements are extracted and converted by the large language model. Combined with multi-level judgment and processing mechanisms, the automated execution of conference operations and result feedback are realized.
It realizes full coverage automation of conference management, improves the convenience and efficiency of conference management, and ensures the accuracy of instruction recognition and time processing.
Smart Images

Figure CN120279916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of meeting management, and particularly to a method, system and medium for voice - automated meeting control based on a large language model. Background Art
[0002] With the popularization of meeting management software, users can perform operations such as creating, querying, deleting, and modifying meetings through a graphical interface. These software usually provide functions such as schedule arrangement, participant management, and meeting room reservation, making meeting management more convenient. However, in some cases, users may not be able to conveniently use the graphical interface for operations, such as when driving or when their hands are occupied.
[0003] Current meeting management software has limited support for voice - automated meeting control. It is difficult for users to complete operations such as creating, querying, deleting, and modifying meetings through voice commands. This limitation is particularly obvious in scenarios where users cannot use the graphical interface, affecting the flexibility and convenience of meeting management. In addition, existing voice control systems often have difficulty accurately understanding and executing complex meeting management instructions, resulting in a poor user experience. Summary of the Invention
[0004] In view of the above - mentioned technical problems, the present invention provides a method for voice - automated meeting control based on a large language model, which realizes the comprehensive coverage of meeting control instructions under voice control and improves the convenience and efficiency of meeting management.
[0005] In the first aspect of the present invention, there is provided a method for voice - automated meeting control based on a large language model, including: Receiving a voice command sent by a user through a microphone; Using a large language model to determine whether the voice command is a meeting - control - related command; If the voice command is a meeting - control - related command, then performing time - element extraction and meeting - type judgment; Converting the extracted time elements into a standard time format; According to the meeting type, using a large language model to perform corresponding meeting operations; Invoking a corresponding meeting interface to execute the meeting operation; and Announcing the operation result to the user through voice synthesis.
[0006] In a possible implementation manner, the step of using a large language model to determine whether the voice command is a meeting - control - related command includes: Using a classification model based on the BERT structure to perform a preliminary domain - falling judgment; For user queries that may be meeting commands, using a large language model to perform a secondary discrimination.
[0007] In a possible implementation, the time element extraction adopts the Global-Pointer algorithm.
[0008] In a possible implementation, the conversion of the extracted time elements into the standard time format includes: Identifying different time expressions using pre-defined regular expressions; Converting the time expressions into the standard time format; Using a large language model to judge the Chinese expression of the converted time and the extracted time elements.
[0009] In a possible implementation, the meeting operations include meeting reservation, meeting query, meeting deletion, and meeting modification.
[0010] In a possible implementation, the method further includes: When performing the meeting operations, retrieving the user's meeting list; Using a large language model to screen the meeting list in combination with the converted time elements.
[0011] In a possible implementation, the method further includes: Fine-tuning the large language model using supervised fine-tuning and reinforcement learning data.
[0012] In a possible implementation, the method further includes: During the inference process, adopting a dynamic decoding mechanism, giving the keys of the JSON according to the rules, and only allowing the model to infer the values of the JSON to ensure that the output is in a parsable JSON format.
[0013] In a second aspect of the present invention, there is provided a speech automated meeting control system based on a large language model, including: A speech receiving module, configured to receive a speech command issued by the user through a microphone; An instruction processing module, configured to use a large language model to judge whether the speech command is a meeting control related command. If the speech command is a meeting control related command, time element extraction and meeting type judgment are performed; A time conversion module, configured to convert the extracted time elements into the standard time format; A meeting operation module, configured to perform corresponding meeting operations using a large language model according to the meeting type, and call the corresponding meeting interface to execute the meeting operations; and A result broadcast module, configured to broadcast the operation result to the user through speech synthesis.
[0014] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is run by a computer, it executes the method described in the first aspect of the embodiments of the present invention.
[0015] In summary, compared with the prior art, the present invention has at least one of the following beneficial technical effects: By using a large language model to intelligently identify and process user voice commands, it realizes full coverage of control commands such as meeting reservation, query, deletion, and modification, greatly improving the automation level and convenience of meeting management; By adopting a multi-level discrimination and processing mechanism, the accuracy of command recognition is improved; Through the intelligent extraction and conversion of time elements, the accuracy of time processing is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 FIG. is a flowchart of an embodiment of a voice automation meeting control method based on a large language model of the present invention.
[0017] Figure 2 FIG. is a flowchart of an embodiment of a voice automation meeting control method based on a large language model of the present invention.
[0018] Figure 3 FIG. is a schematic structural diagram of an embodiment of a voice automation meeting control system based on a large language model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] It should be understood that the terms "first", "second", "third", etc. in the claims, specifications, and drawings of the present disclosure are used to distinguish different objects, rather than to describe a specific order. The terms "including" and "comprising" used in the specifications and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. It should also be understood that the terms used in the specifications of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure.
[0021] Referring to Figure 1 , the embodiments of the present invention disclose a voice automation meeting control method based on a large language model, including the following steps.
[0022] S1. Receive the voice command issued by the user through the microphone. Specifically, the microphone connected to the user terminal receives the user's voice input in real time. When the user issues a voice command related to the meeting, the microphone collects and transmits the voice signal to the backend processing module, and then the audio signal is converted into the corresponding text content through the Automatic Speech Recognition (ASR) engine, laying the foundation for the large language model to identify the meeting command type and extract its elements.
[0023] S2. Use the large language model to determine whether the voice command is a meeting control related command. The large language model can understand and analyze the semantics of the user's voice command to determine whether the command is related to meeting control.
[0024] Specifically, after receiving the converted text command, first use a lightweight classification model based on the BERT structure to make a preliminary judgment on the text, screening out the commands that may belong to the meeting control field to reduce the computational pressure of subsequent processing; for the commands preliminarily judged to be possibly relevant, further submit them to the large language model for fine-grained recognition. Combining with the preset few-shot examples, through prompt, guide the model to determine whether the command is a specific meeting control command, such as a meeting reservation, deletion, query, or modification command, so as to improve the accuracy and recall rate of recognition.
[0025] S3. If the judgment result is a meeting control related command, then perform time element extraction and meeting type judgment. Time element extraction can identify the time information contained in the voice command, such as the meeting date, time, etc. Meeting type judgment can determine the specific meeting operation type that the user wants to perform.
[0026] Specifically, when the large language model determines that the voice command is a meeting control related command, first use the Global-Pointer algorithm to extract the time elements in the command text. This algorithm has the ability to efficiently and accurately identify time phrases (such as "3 o'clock tomorrow afternoon"); at the same time, the large language model understands the command semantics according to the set prompt task, and judges the type of meeting operation from it, such as whether it is a meeting reservation, query, deletion, or modification. The context information and historical interaction content can also be combined in this process to improve the accuracy of meeting type judgment.
[0027] S4. Convert the extracted time elements into the standard time format.
[0028] Specifically, process the Chinese time elements extracted in step S3. First, use a word segmentation tool to split the time phrase into basic phrases, and then match and identify the time expression by combining a preset regular expression template to convert the time in natural language form (such as "next Monday morning") into a standard time format (such as "2025-04-14 09:00"); for periodic time expressions, such as "every Wednesday afternoon", specific regular rules can also be set for parsing. Finally, to ensure the accuracy of the time conversion result, the large language model can also be called to compare and verify the original Chinese expression and the conversion result, and if there are errors, the model will perform intelligent correction.
[0029] S5. According to the type of the meeting, use the large language model to perform corresponding meeting operations.
[0030] Specifically, according to the identified type of meeting operation, call the large language model to execute the corresponding meeting operation task. If it is a reservation meeting, further extract elements such as the meeting theme, time, and participants from the instruction and organize them into a structured format; if it is a query, deletion, or modification of a meeting, first combine the converted time elements to screen out the target meeting from the user's meeting list, and then perform corresponding processing according to the operation type, such as obtaining the meeting details, generating a deletion instruction, or extracting modification parameters. Throughout the process, by setting the system prompt and few-shot examples to guide the large language model to accurately complete various tasks, and supplemented by SFT and DPO fine-tuning methods to improve the execution effect of the model in the meeting scenario.
[0031] S6. Call the corresponding meeting interface to execute the meeting operation.
[0032] Specifically, after the large language model completes the extraction of meeting elements and the generation of instructions, call the corresponding meeting management interface according to the operation type for actual execution. For example, call the create meeting interface to submit reservation information, or call the delete interface to delete the screened meeting, or call the update interface to modify the content of the target meeting. When calling the interface, the structured parameters output by the model (such as standard time, meeting ID, participants, etc.) are directly used as the input for interface call to ensure accurate operation, so as to achieve seamless connection from voice instructions to actual meeting operations.
[0033] S7. Broadcast the operation result to the user through text-to-speech synthesis.
[0034] Specifically, after the meeting operation interface is executed, corresponding feedback content is generated according to the result returned by the interface, and the natural language form of the broadcast text is organized through text generation technology. Subsequently, the text is converted into a voice signal by calling the text-to-speech (TTS) algorithm and is broadcast to the user in real time through a speaker or a terminal device. For example, content such as "Your meeting has been successfully created" or "No meeting meeting the conditions was found for deletion" can be broadcast to achieve instant voice feedback on the operation result and improve the naturalness and efficiency of user interaction.
[0035] The method for voice automated meeting control based on a large language model in the embodiments of the present invention constructs a complete voice interaction closed loop through steps S1 to S7, realizing the full-process automated processing from voice command reception, intent recognition, element extraction, format conversion, to meeting operation execution and voice feedback. This method fully combines the advantages of the large language model in natural language understanding and information extraction, and cooperates with the auxiliary model and the rule engine to improve efficiency and accuracy, enabling users to perform meeting reservation, query, modification, and deletion operations through natural language without manual operation, significantly improving the convenience and intelligent level of human-computer interaction. The overall meeting control accuracy rate reaches more than 90%, and it has good application prospects and promotion value.
[0036] Further, as an implementation manner of the present invention, in step S2, using the large language model to determine whether the voice command is a meeting control related command includes: Performing a preliminary domain determination using a classification model based on the BERT structure; For user queries that may be meeting commands, using the large language model for secondary discrimination.
[0037] Specifically, after receiving the text command, the text is first input into a pre-trained classification model based on the BERT structure for semantic recognition and intent classification. This model has the ability to judge whether the input content belongs to the field of meeting control, that is, the domain determination ability, by learning a large number of labeled training data, and can quickly filter out obviously irrelevant non-meeting commands, thus effectively reducing the inference burden of the subsequent large language model and improving the response speed and resource utilization efficiency of the overall system.
[0038] When the classification model based on BERT preliminarily determines that the user query may be related to meeting control, the text is further input into the large language model, combined with the pre-designed system prompt and few-shot examples, to guide the model to deeply understand the semantics of the command to identify whether it is a specific meeting operation command, such as reserving, querying, deleting, or modifying a meeting. This step relies on the powerful semantic understanding and generalization ability of the large language model to effectively make up for the deficiency of the lightweight model in judgment ability under boundary conditions and complex expressions, thereby improving the accuracy and robustness of meeting command recognition.
[0039] Further, as an implementation manner of the present invention, the time element extraction adopts the Global-Pointer algorithm.
[0040] Specifically, after recognizing that the user voice instruction is a meeting-related instruction, the corresponding text is input into a sequence labeling model based on the Global-Pointer architecture to extract the time elements therein. This algorithm can effectively capture the start and end positions of time phrases in the text by constructing a global pointer mechanism, and has improved both the extraction accuracy and inference efficiency compared with the traditional BERT+CRF method. The training data is generated by a large language model combined with few-shot capabilities and a preset prompt, covering various Chinese time expressions, such as "3:00 pm next Monday" or "every Wednesday morning", etc., to ensure that the model still has high accuracy and strong generalization ability when facing diverse natural language expressions.
[0041] Further, as an implementation manner of the present invention, in step S4, converting the extracted time elements into a standard time format, including: Identifying different time expressions using pre-defined regular expressions; Converting the time expression into a standard time format; Using a large language model to judge the Chinese expression of the converted time and the extracted time elements.
[0042] Specifically, after completing the extraction of time elements, for the extracted Chinese time phrases, the text is tokenized, and a set of pre-defined regular expression rules are applied to perform pattern matching and recognition on various time expressions. These regular rules cover common time expression forms, such as "3:00 pm today", "every Monday morning", "the 15th of next month", etc., and can accurately extract information such as time units (such as days, weeks, months), time ranges, periodic characteristics, etc., providing a structured basis for subsequent time format standardization and ensuring adaptability to diverse natural language time descriptions.
[0043] After identifying the time expression, extract the key components in the time according to the regular expression matching results, such as date, week, time period, periodicity, etc. information, and perform context parsing and logical inference in combination with the current system time, such as parsing "next Wednesday morning" into the specific format of "2025-04-11 09:00". For relative time, fuzzy expressions or periodic times, they are completed and calculated through built-in time parsing functions, and finally all time information is uniformly converted into a structured time field that conforms to the standard time format (such as the ISO 8601 standard) for subsequent call by the meeting operation process.
[0044] After completing the conversion of the time expression to the standard time format, the conversion result, together with the original Chinese time expression and the extracted time elements, is input into the large language model. By setting prompts, the model is guided to judge the semantic consistency of the conversion result. The model will compare whether the original Chinese expression and the standard time format are logically and semantically consistent to determine whether the time conversion is accurate. If the model identifies deviations or ambiguities in the conversion, it will automatically give correction suggestions or directly modify it to a more semantically consistent standard time based on the context, thereby improving the robustness and accuracy of time processing.
[0045] Furthermore, as an implementation manner of the present invention, the method further includes: When performing the meeting operation, retrieve the user's meeting list; Use the large language model to screen the meeting list in combination with the converted time elements.
[0046] Specifically, when the user issues a voice command to query, delete, or modify a meeting, first retrieve the user's meeting list through the interface and obtain structured information including meeting time, title, participants, etc. Subsequently, input the converted standard time elements and the meeting list into the large language model, and set the system prompt to guide the model to screen the meeting according to time matching, semantic association, or context description, and identify the target meeting that best meets the user's intention. This method makes full use of the large language model's ability in multi-dimensional information understanding and screening judgment, improving the meeting matching accuracy and interaction experience under complex or ambiguous instructions.
[0047] Furthermore, as an implementation manner of the present invention, the method further includes: Fine-tune the large language model using supervised fine-tuning and reinforcement learning data.
[0048] Specifically, to improve the large language model's instruction understanding and task execution capabilities in the meeting control scenario, high-quality supervised fine-tuning (SFT) data covering various types of instructions such as reservation, query, deletion, and modification, and reinforcement learning (such as DPO) data based on human preference feedback are constructed to perform targeted fine-tuning on the large language model. The SFT stage is mainly used to train the model to accurately extract instruction elements and generate structured outputs in a standardized manner; the DPO stage guides the model to preferentially select results that better meet user expectations among multiple reasonable outputs. Through the joint training of these two types of data, the model shows higher accuracy, stability, and response consistency when processing actual meeting instructions, significantly enhancing the practicality and robustness of the system.
[0049] Furthermore, as an implementation manner of the present invention, the method further includes: During the inference process, a dynamic decoding mechanism is adopted. According to the rules, given the keys of the JSON, only let the model infer the values of the JSON to ensure that the output is a parsable JSON format.
[0050] Specifically, to ensure that the large language model outputs a parsable JSON format that strictly conforms to the expected structure during the inference process, when constructing the prompt, the keys in the JSON object are clearly specified in advance, and by accessing dynamic decoding frameworks such as sglang, the generation scope of the model is real-time constrained during the inference process. Only allow the model to generate corresponding values under the specified keys, and prohibit modifying or generating content with non-preset structures. This mechanism effectively avoids the problems of chaotic output structure or missing fields of the model, and ensures that each inference result conforms to the format specifications required by the subsequent meeting operation interface.
[0051] In summary, a specific implementation example of the voice automated meeting control method based on a large language model described in the present invention is as Figure 2 shown. The entire process includes four core stages: voice command reception and recognition, meeting intention judgment and element extraction, meeting operation execution, and result feedback, which reflects how the present invention realizes voice intelligent control of the entire life cycle of a meeting by means of a large language model (LLM).
[0052] First, the user issues a natural language command through the microphone to start the process and enter the first judgment node "meeting command". In this link, the voice signal is first converted into text content, and a classification model based on the BERT structure is used to make a preliminary intention judgment on the text to identify whether it belongs to the meeting control field command. If the judgment is a non-meeting-related command, the process ends; if it is a meeting command, it further enters the "meeting type" judgment node.
[0053] At the "meeting type" judgment node, identify the user's specific intention, whether it is to make a meeting reservation, or to query, delete, or modify. For all operation types, it is necessary to first extract the time element from the voice and convert it into a standard time format through a predefined regular expression combined with Chinese time processing logic. In this process, a large language model is also introduced to judge the consistency between the conversion result and the original Chinese expression to ensure the accuracy and reliability of the time information.
[0054] If the user's command is "reserve a meeting", then call the large language model to extract all the elements required for the reservation (such as the theme, time, participants, etc.), and after organizing them in a structured JSON format, complete the meeting creation operation by calling the "create meeting interface".
[0055] If the user instruction is "query, delete, modify", first obtain the full meeting list of the user through the interface, and input the meeting list and time elements into the large language model to filter out the target meetings that meet the conditions. If it is a "query" operation, the large language model will generate a natural language reply based on the filtering results and call the TTS module for voice broadcast; if it is "delete", the delete interface will be called to execute the meeting deletion; if it is "modify", the large language model also needs to further extract the modification parameters, such as the new time or content, and call the modification interface to complete the update.
[0056] Throughout the process, the large language model is also directionally optimized through methods such as supervised fine-tuning (SFT) and preference optimization (DPO) to make it more adaptable to the meeting scenario tasks; at the same time, a dynamic decoding mechanism is used in the inference stage to lock the key names of the JSON structure in advance, and only allow the model to generate key-value content, and cooperate with the json repair tool to ensure the structural parsability, enhancing the stability and practicality of the system.
[0057] Refer to Figure 3 , this embodiment of the present invention also discloses a voice automation meeting control system based on a large language model, including a voice receiving module 1, an instruction processing module 2, a time conversion module 3, a meeting operation module 4, and a result broadcast module 5.
[0058] The voice receiving module 1 is used to receive the voice instruction issued by the user through the microphone.
[0059] Specifically, the voice receiving module 1 continuously monitors the user's voice input through the microphone device connected to the user terminal, and collects it in real time when a voice signal is detected. The collected audio data is first pre-processed, such as noise reduction, echo cancellation and other pre-processing operations, to improve the quality of the voice signal; then the processed audio data is transmitted to the automatic speech recognition engine (ASR) for speech-to-text conversion, generating structured natural language text content. This text serves as the basic input for the subsequent module processing, providing accurate semantic information support for instruction recognition and task execution. The voice receiving module 1 can adapt to various types of microphone hardware, support multi-language recognition and far-field pickup, ensuring that the system can stably capture user instructions in various interaction scenarios.
[0060] The instruction processing module 2 is used to use the large language model to judge whether the voice instruction is a meeting control related instruction. If the voice instruction is a meeting control related instruction, time element extraction and meeting type judgment are performed.
[0061] Specifically, after receiving the text instruction output by the voice receiving module 1, the instruction processing module 2 first uses a lightweight classification model based on the BERT structure to perform a preliminary intention classification judgment on the text, which is used to quickly screen whether it belongs to the meeting control-related instructions and improve the system response efficiency. For the text judged to be possibly a meeting instruction, it is further input into the large language model for in-depth semantic analysis, and the specific meeting operation type (such as reservation, query, deletion, or modification) is judged by guiding the model through the preset system prompt and few-shot examples. At the same time, the instruction processing module 2 also calls the large language model to accurately extract the time elements in the instruction in combination with the Global-Pointer algorithm, providing key information support for subsequent time conversion and meeting operations to ensure the accuracy and integrity of semantic recognition.
[0062] The time conversion module 3 is used to convert the extracted time elements into the standard time format.
[0063] Specifically, after receiving the original Chinese time elements output by the instruction processing module 2, the time conversion module 3 first performs word segmentation on it, and applies a predefined set of regular expression templates to identify different types of time expressions, such as "3:00 pm tomorrow" or "every Friday morning", etc., and extracts specific structural information such as time units, values, and periodicity. Then, in combination with the current system time context, the relative time or fuzzy expression is converted into the standard time format (such as "2025-04-11T15:00:00") through the time parsing function, and the large language model is further called to judge the semantic consistency of the conversion result to ensure that the converted time is consistent with the original user expression meaning. If a logical conflict or deviation is identified, the time conversion module 3 can also automatically correct and revise it, so as to achieve highly robust and accurate time standardization processing.
[0064] The meeting operation module 4 is used to perform corresponding meeting operations using the large language model according to the meeting type, and call the corresponding meeting interface to execute the meeting operation.
[0065] Specifically, after the meeting operation module 4 obtains the standardized time information and meeting type, it calls the large language model to further process the instructions, extracts the required key parameters for different operation types (such as reservation, query, deletion, modification), such as meeting theme, participants, meeting ID, change fields, etc. For meeting reservation, the meeting operation module 4 organizes the extracted elements into structured data and calls the create meeting interface to complete the meeting generation; for query, deletion, and modification operations, the meeting operation module 4 first retrieves the user's meeting list and filters the meeting list in combination with the large language model to locate the target meeting, and then calls the corresponding interfaces (such as query interface, deletion interface, modification interface) according to different operation types to execute specific tasks. Throughout the process, the large language model combines the inference capabilities fine-tuned by SFT and DPO to ensure high accuracy and intelligent judgment capabilities in semantic understanding, data extraction, and meeting screening, thus realizing the automated execution of the entire meeting operation process.
[0066] The result announcement module 5 is used to announce the operation result to the user through voice synthesis.
[0067] Specifically, after the meeting operation module 4 completes the corresponding operation, the result announcement module 5 automatically generates the corresponding natural language feedback text according to the execution result returned by the interface, such as "The meeting has been successfully created", "No meeting meeting the conditions was found for deletion", etc. The system then calls the text-to-speech (TTS) engine to convert the text content into clear and natural voice signals and broadcasts them to the user in real time through the terminal device speaker or headphones. The result announcement module 5 can adjust the broadcast speed, tone, and voice type according to the user's preferences to improve the friendliness and personalized experience of voice interaction, ensuring that the user can clearly understand the meeting operation result without viewing the interface, thus realizing true voice closed-loop control.
[0068] The voice automated meeting control system based on the large language model provided by the present invention integrates voice recognition, large language model inference, structured data extraction, and interface calling capabilities, and can realize the full process automation from the reception of user voice instructions, intention recognition, element extraction, time conversion, to meeting operation execution and result announcement, greatly improving the intelligence and interaction convenience of meeting management. Users do not need to rely on traditional graphical interfaces or manual input, and can complete operations such as meeting reservation, query, deletion, or modification only through natural language voice, which is particularly suitable for scenarios such as traveling, driving, or other situations where manual operation is not convenient. The system uses the collaborative judgment of the BERT model and the large language model, combines the regular and semantic verification time conversion mechanism, and dynamic decoding to ensure the accuracy of the output format, making the overall system have good response speed and operation stability while ensuring intelligence, and has broad application prospects and promotion value.
[0069] An embodiment of the present invention also discloses a readable storage medium.
[0070] A readable storage medium stores a computer program which, when executed by a processor, implements the steps of the method for voice automated conference control based on a large language model according to any one of the above embodiments.
[0071] It can be understood that a computer-readable storage medium may include: any entity or device capable of carrying a computer program, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), and a software distribution medium, etc. The computer program includes computer program code. The computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. A computer-readable storage medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), and a software distribution medium, etc.
[0072] In some embodiments of the present invention, an electronic device may include a controller or a processor. The controller is a single-chip microcomputer chip integrating a processor, a memory, a communication module, etc. The processor may refer to the processor included in the controller. The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0073] Any process or method description shown in a flowchart or described in other ways herein may be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present invention belong.
[0074] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0075] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice automated conference control method based on a large language model, characterized in that, It includes: Receiving a voice command issued by the user through the microphone; Using a large language model to determine whether the voice command is a meeting control related command; If the voice command is a meeting control related command, perform time element extraction and meeting type judgment; Convert the extracted time elements into a standard time format; According to the meeting type, use a large language model to perform corresponding meeting operations; Call the corresponding meeting interface to execute the meeting operation; and Announce the operation result to the user through voice synthesis.
2. The voice automation conference control method according to claim 1, wherein The using a large language model to determine whether the voice command is a meeting control related command includes: Using a classification model based on the BERT structure to perform a preliminary domain determination; For user queries that may be meeting commands, use a large language model for secondary discrimination.
3. The voice automation conference control method according to claim 2, wherein The time element extraction adopts the Global-Pointer algorithm.
4. The voice automated conference control method according to claim 1, wherein The converting the extracted time elements into a standard time format includes: Using a predefined regular expression to identify different time expressions; Convert the time expression into a standard time format; Use a large language model to judge the Chinese expression of the converted time and the extracted time elements.
5. The voice automation conference control method according to claim 1, characterized in that, The meeting operations include meeting reservation, meeting query, meeting deletion, and meeting modification.
6. The voice automation conference control method according to claim 5, characterized in that, It also includes: When performing the meeting operation, retrieve the user's meeting list; Use a large language model to screen the meeting list in combination with the converted time elements.
7. The voice automation conference control method according to claim 1, wherein It also includes: Fine-tuning the large language model using supervised fine-tuning and reinforcement learning data.
8. The voice automation conference control method according to claim 1, wherein It also includes: Adopt a dynamic decoding mechanism during the inference process, give the keys of the JSON according to the rules, and only let the model infer the values of the JSON to ensure that the output is in a parsable JSON format.
9. A voice automated meeting control system based on a large language model, characterized in that, The system includes: A voice receiving module for receiving a voice command issued by the user through the microphone; An instruction processing module for using a large language model to determine whether the voice command is a meeting control related command, and if the voice command is a meeting control related command, perform time element extraction and meeting type judgment; A time conversion module for converting the extracted time elements into a standard time format; A meeting operation module for performing corresponding meeting operations according to the meeting type using a large language model, and calling the corresponding meeting interface to execute the meeting operation; and A result announcement module for announcing the operation result to the user through voice synthesis.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is run by the computer, it executes the voice automated meeting control method according to any one of claims 1 to 8.