LLM-based communication system intelligent voice control method and system

By converting and identifying voice commands based on LLM, the problem of inconvenience in operation of the PBX system is solved, and voice control and efficient command recognition are realized.

CN120580992APending Publication Date: 2025-09-02XIAMEN XINGZONG DIGITAL TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510770815.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The operation of existing PBX systems requires physical buttons, GUI or traditional IVR systems, resulting in inconvenience in operation.

Method used

The large language model based on LLM is used to identify and process voice instructions, convert them into text data and generate call instructions to realize voice control.

Benefits of technology

Operation is achieved through voice commands, which solves the problem of inconvenient key operation and improves the efficiency of command recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580992A_ABST
    Figure CN120580992A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an LLM-based communication system intelligent voice control method and system. The method comprises the steps of obtaining a first voice instruction sent by a user; converting the first voice instruction into first text data; performing semantic recognition on the first text data through a preset large language model to obtain first semantic information; and generating a first calling instruction according to the first semantic information, so that the PBX system executes the first calling instruction. After a first voice instruction sent by a user is received, the voice instruction is converted into text data, then semantic information corresponding to the text data is analyzed through a preset large language model, and then a first calling instruction is generated according to the first semantic information, so that a PBX system executes the first calling instruction. Therefore, operation can be performed through voice, the problem that key operation is inconvenient is solved, recognition can be performed through the preset large language model, and the instruction recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information technology, and in particular to an LLM-based intelligent voice control method and system for a communication system. Background Art

[0002] Currently, PBXs (Private Branch Exchanges) are widely used within companies and other environments. However, the inventors have discovered that while current PBX systems offer a variety of functions, such as extension management, call control, call routing, IVR (Interactive Voice Response), conferencing, voicemail, and call recording, executing these functions often requires users to use physical buttons, a graphical user interface (GUI), preset keyboard shortcuts, or traditional IVR systems, resulting in operational inconvenience. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a method and system for intelligent voice control of a communication system based on LLM to solve the problem of inconvenient PBX button operation. The specific technical solution is as follows:

[0004] In a first aspect of an embodiment of the present application, a method for intelligent voice control of a communication system based on LLM is provided, the method comprising:

[0005] Obtaining a first voice command issued by a user;

[0006] Converting the first voice instruction into first text data;

[0007] Performing semantic recognition on the first text data using a preset large language model to obtain first semantic information;

[0008] A first call instruction is generated according to the first semantic information, so that the PBX system executes the first call instruction, wherein the first call instruction corresponds to one or more operations.

[0009] In a possible implementation, generating a first call instruction according to the first semantic information includes:

[0010] Determining whether the operation corresponding to the first semantic information is a key operation by using multiple preset key operation types and / or preset evaluation logic;

[0011] If it is determined that the operation corresponding to the semantic information is a key operation, the first call instruction is generated according to the first semantic information.

[0012] In a possible implementation, determining whether the operation corresponding to the first semantic information is a key operation based on multiple preset key operation types includes:

[0013] Matching multiple preset key operation types with the first semantic information;

[0014] If there is a match, the operation corresponding to the first semantic information is determined to be a key operation.

[0015] In a possible implementation, if it is determined that the operation corresponding to the semantic information is a key operation, generating the first call instruction according to the first semantic information includes:

[0016] If the operation corresponding to the semantic information is determined to be a key operation, a secondary confirmation request is sent;

[0017] After receiving a confirmation instruction for the secondary confirmation request, the first calling instruction is generated according to the first semantic information.

[0018] In one possible implementation, the steps of obtaining a first voice command issued by a user; converting the first voice command into first text data; and performing semantic recognition on the first text data using a preset large language model to obtain first semantic information include:

[0019] receiving a voice deletion instruction for target content;

[0020] Converting the voice deletion instruction into deletion text data;

[0021] The deleted text data is semantically recognized by using a preset large language model to obtain deletion semantic information.

[0022] In a possible implementation, the PBX system executing the first calling instruction includes:

[0023] Receive a retracement call instruction for the target operation;

[0024] Identifying one or more corresponding target operations according to the callback call instruction;

[0025] Determining whether the corresponding one or more target operations are reversible operations;

[0026] If a reversible operation is determined, the corresponding one or more target operations are performed.

[0027] In a possible implementation manner, after the PBX system executes the first call instruction, the method further includes:

[0028] Obtaining an execution result of the first calling instruction;

[0029] Generate and send a prompt voice according to the execution result.

[0030] A second aspect of the embodiments of the present application provides an LLM-based intelligent voice control system for a communication system, the system comprising:

[0031] A first instruction acquisition module, configured to acquire a first voice instruction issued by a user;

[0032] A first instruction conversion module, configured to convert the first voice instruction into first text data;

[0033] a first semantic recognition module, configured to perform semantic recognition on the first text data using a preset large language model to obtain first semantic information;

[0034] The first instruction execution module is configured to generate a first call instruction according to the first semantic information, so as to enable the PBX system to execute the first call instruction, wherein the first call instruction corresponds to one or more operations.

[0035] In a possible implementation, the first instruction execution module includes:

[0036] an operation judgment submodule, configured to judge whether the operation corresponding to the first semantic information is a key operation based on a plurality of preset key operation types and / or preset evaluation logics;

[0037] The instruction execution submodule is used to generate the first call instruction according to the first semantic information if it is determined that the operation corresponding to the semantic information is a key operation.

[0038] In a possible implementation, the operation determination submodule is specifically configured to match a plurality of preset key operation types with the first semantic information; if a match occurs, determining that the operation corresponding to the first semantic information is a key operation.

[0039] In one possible implementation, the instruction execution submodule is specifically configured to send a secondary confirmation request if it is determined that the operation corresponding to the semantic information is a critical operation; and after receiving a confirmation instruction for the secondary confirmation request, generate the first call instruction according to the first semantic information.

[0040] In one possible implementation, the first instruction acquisition module is specifically used to receive a voice deletion instruction for target content; the first instruction conversion module is specifically used to convert the voice deletion instruction into deletion text data; and the first semantic recognition module is specifically used to perform semantic recognition on the deletion text data through a preset large language model to obtain deletion semantic information.

[0041] In one possible embodiment, the first instruction execution module is specifically used to receive a rollback call instruction for a target operation; identify the corresponding one or more target operations based on the rollback call instruction; determine whether the corresponding one or more target operations are reversible operations; if it is determined to be a reversible operation, execute the corresponding one or more target operations.

[0042] In a possible implementation, the system further includes:

[0043] The result prompt module is used to obtain the execution result of the first calling instruction; generate and send a prompt voice according to the execution result.

[0044] Another aspect of the present application provides an electronic device, including:

[0045] Memory for storing computer programs;

[0046] The processor is configured to implement any of the above-mentioned LLM-based intelligent voice control methods for a communication system when executing a program stored in the memory.

[0047] In another aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, it implements any of the above-mentioned LLM-based communication system intelligent voice control methods.

[0048] In another aspect of the embodiments of the present application, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-mentioned LLM-based intelligent voice control methods for a communication system.

[0049] Beneficial effects of the embodiments of the present application:

[0050] The embodiment of the present application provides a method and system for intelligent voice control of a communication system based on LLM, the method comprising: obtaining a first voice instruction sent by a user; converting the first voice instruction into first text data; performing semantic recognition on the first text data through a preset large language model to obtain first semantic information; generating a first call instruction based on the first semantic information, so that the PBX system executes the first call instruction, wherein the first call instruction corresponds to one or more operations. It can be seen that through the method of the embodiment of the present application, after receiving the first voice instruction sent by the user, the voice instruction can be converted into text data, and then the semantic information corresponding to the text data can be parsed through a preset large language model, and then the first call instruction can be generated based on the first semantic information, so that the PBX system executes the first call instruction, thereby not only realizing operation through voice and solving the problem of inconvenience of key operation, but also improving the recognition efficiency of the instruction through recognition by the preset large language model.

[0051] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0053] Figure 1 A flowchart of an LLM-based intelligent voice control method for a communication system provided in an embodiment of the present application;

[0054] Figure 2 A flowchart of voice command processing provided in an embodiment of the present application;

[0055] Figure 3 A flowchart of a delete instruction provided in an embodiment of the present application;

[0056] Figure 4 A flowchart of voice command processing including optional secondary confirmation provided in an embodiment of the present application;

[0057] Figure 5 A flowchart of a retrieval instruction provided in an embodiment of the present application;

[0058] Figure 6 A flowchart of a retracement instruction processing provided in an embodiment of the present application;

[0059] Figure 7A schematic diagram of a system architecture provided in an embodiment of the present application;

[0060] Figure 8 A schematic diagram of the structure of an intelligent voice control system for a communication system based on LLM provided in an embodiment of the present application;

[0061] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0063] In the first aspect of the embodiment of the present application, a communication system intelligent voice control method based on LLM is first provided. Figure 1 , Figure 1 A flow chart of an LLM-based intelligent voice control method for a communication system provided in an embodiment of the present application, the method comprising:

[0064] Step S11, obtaining a first voice command issued by a user;

[0065] Step S12, converting the first voice instruction into first text data;

[0066] Step S13, performing semantic recognition on the first text data using a preset large language model to obtain first semantic information;

[0067] Step S14: Generate a first call instruction according to the first semantic information, so that the PBX system executes the first call instruction, wherein the first call instruction corresponds to one or more operations.

[0068] Corresponding to step S11 above, the method of the embodiment of the present application is applied to a PBX (Private Branch Exchange) system and can be implemented through the system. Specifically, obtaining the first voice command issued by the user can be achieved through a terminal device of the system. In one embodiment, the system may include a voice receiving module, such as a microphone, to receive the voice command sent by the user. For example, when the user says "hang up extension 1000", the command can be received by the receiving module.

[0069] Corresponding to the above step S12, when converting the first voice instruction into the first text data, it can be recognized by the ASR (Automatic Speech Recognition) in the system. Specifically, it can be achieved by a hidden Markov model, a deep neural network or a hybrid network. Through this step, the voice instruction can be converted into text data, which facilitates further recognition and judgment based on the text data. For example, after receiving the voice instruction "Hang up extension 1000", the corresponding text "Hang up extension 1000" can be obtained through conversion.

[0070] Corresponding to the above-mentioned step S13, in the present application, semantic recognition is performed on the first text data through a preset large language model, which can be processed by the LLM (Large Language Model) core processing module in the PBX system, and semantic recognition can be performed in combination with the context of the first text data, thereby preventing semantic recognition errors. For example, after the text is converted, a conversion error may occur, such as converting the semantic meaning corresponding to hang up to "hang up", and in the embodiment of the present application, semantic recognition of the first text data through a preset large language model can be performed in combination with the context, thereby identifying that the correct voice should be hang up, which is convenient for subsequent operations. At the same time, in actual use, the first semantic information can include the operation intention and the corresponding parameters. For example, for "hang up extension 1000", the semantic recognition of the first text data through the preset large language model can obtain the operation intention as "hang up the call", and the corresponding parameter is "target extension number: 1000".

[0071] Corresponding to step S14 above, when a first call instruction is generated based on the first semantic information, it can be processed by the PBX system's operation instruction generation and execution module, causing the PBX system to execute the first call instruction. The first call instruction corresponds to one or more operations. In actual use, the number of operations is related to the content of the call instruction, and the one or more operations corresponding to a call instruction are consistent and logically consistent, without any contradictions or conflicts. For example, a command for shutting down a pass may include three operations.

[0072] It can be seen that through the method of the embodiment of the present application, after receiving the first voice command sent by the user, the voice command can be converted into text data, and then the semantic information corresponding to the text data is parsed through the preset large language model, and then a first call instruction is generated according to the first semantic information, so that the PBX system executes the first call instruction. In this way, not only can operation be performed through voice to solve the problem of inconvenient key operation, but also recognition can be performed through the preset large language model to improve the recognition efficiency of the instruction.

[0073] In one possible implementation, generating a first call instruction based on the first semantic information includes: determining whether the operation corresponding to the first semantic information is a critical operation using multiple preset key operation types and / or preset evaluation logic; and if the operation corresponding to the semantic information is determined to be a critical operation, generating the first call instruction based on the first semantic information. Specifically, when generating the call instructions based on the semantic instructions, a determination may be made as to whether the instructions are critical operations. Key operations herein may include operations with higher risks, while non-key operations may include operations with lower risks, such as querying status or controlling a single call. Specifically, when determining whether the operation corresponding to the first semantic information is a critical operation using multiple preset key operation types and / or preset evaluation logic, the multiple preset key operation types may include multiple pre-set key operation types. By determining whether the first semantic information is one of the multiple preset key operation types, if so, determining the operation corresponding to the semantic information is a critical operation; otherwise, determining it is not a critical operation. Determining whether the operation corresponding to the first semantic information is a critical operation using the preset evaluation logic can be accomplished through logical analysis, such as word semantics analysis, sentence structure analysis, and conjunction analysis. For example, when analyzing the sentence structure, the semantic information can be judged based on the order of the identified subject, predicate, object, or attributive, adverbial, and complement to determine whether it meets the correct order requirements, and thus determine whether it is a critical operation based on the judgment result. In one example, the processing of instructions containing conditional logic: for instructions such as "When extension A is busy, forward the incoming call to extension B", the large language model core processing module is responsible for understanding the conditions and corresponding operations. The PBX operation instruction generation and execution module converts this parsed logic into a series of API (Application Programming Interface) calls to configure the PBX call routing or forwarding rules. Therefore, if the operation corresponding to the semantic information is determined to be a critical operation, the first call instruction is generated according to the first semantic information.

[0074] In one possible implementation, the method of determining whether the operation corresponding to the first semantic information is a key operation through a plurality of preset key operation types includes: matching a plurality of preset key operation types with the first semantic information; if they match, determining that the operation corresponding to the first semantic information is a key operation. Specifically, the plurality of preset key operation types may include: defining deletion operations, batch modification operations, operations involving system core configuration, etc. When matching a plurality of preset key operation types with the first semantic information, whether there is a match can be determined by determining whether the operation corresponding to the first semantic information is one of a plurality of preset key operation types, and if so, determining that there is a match.

[0075] In one possible implementation, after the PBX system executes the first call instruction, the method further includes: obtaining the execution result of the first call instruction; generating and sending a prompt voice according to the execution result. Specifically, the execution result (success or failure information) can be returned to the user feedback module, and then TTS (speech synthesis) is performed and broadcast through the user feedback module, such as, "The call of extension 1000 has been successfully hung up", or the operation result is displayed on the corresponding user interface, or the corresponding prompt is briefly displayed on the interface (such as the lower right corner). In an example, the voice instruction processing flow (taking the instruction "Hang up extension 1000" as an example, see Figure 2 (Voice command processing flow diagram): 1. The user issues the voice command "Hang up extension 1000" through the voice input interface module. 2. The ASR (Analog Speech Recognition) module converts the voice command into text: "Hang up extension 1000." 3. The LLM core processing module receives the text, parses the semantics, identifies the action intent as "hang up the call," and extracts the parameter "target extension number: 1000." 4. The LLM core processing module's Operation Criticality Assessment Unit determines whether the "Hang up the call" action is a critical action. In this scenario, assume that "hanging up a single extension call" is not defined as a critical action by default. 5. The structured intent and parameters are passed to the PBX action command generation and execution module. 6. The intent-action mapping unit converts this intent and parameters into an API call to the PBX system to hang up extension 1000. 7. The PBX system interface module receives the command and sends it to the PBX entity for execution. The action and related information are recorded in the operation log. 8. The PBX system executes the hang-up action and returns the result (success or failure). 9. The execution result is transmitted to the User Feedback Module via the PBX Operation Command Generation and Execution Module. 10. The User Feedback Module announces "Successfully hung up the call at extension 1000" via TTS (Text-to-Speech) or displays the operation result on the corresponding user interface. It may also briefly display a prompt "Undo last operation" in the interface (e.g., in the lower right corner).

[0076] In a possible implementation, if the operation corresponding to the semantic information is determined to be a key operation, the first call instruction is generated according to the first semantic information, see Figure 3 , Figure 3 A schematic diagram of a secondary confirmation process provided in an embodiment of the present application, wherein the method further includes:

[0077] Step S31: If the operation corresponding to the semantic information is determined to be a key operation, a secondary confirmation request is sent;

[0078] Step S32: After receiving the confirmation instruction for the secondary confirmation request, generate the first call instruction according to the first semantic information.

[0079] As mentioned above, critical operations are often operations with higher risks. Therefore, after determining that the current operation is a critical operation, that is, a high-risk operation, these operations can be reconfirmed to prevent irreversible consequences after the operation. Specifically, the secondary confirmation request can be sent to the user through voice, display, and other methods. For example, play to the user by voice: "Please confirm whether to delete ***." When receiving the confirmation instruction for the secondary confirmation request, it can also be received by voice. For example, when the user voice reply is detected: "Confirm", the first call instruction is generated according to the first semantic information. In actual use, the secondary confirmation request can also be sent to the user by display, and the user can also confirm by pressing a designated button, etc.

[0080] In one embodiment, when the received instruction is a delete instruction, the steps of obtaining a first voice instruction issued by a user, converting the first voice instruction into first text data, and performing semantic recognition on the first text data using a preset large language model to obtain first semantic information include: receiving a voice instruction for deleting target content; converting the voice instruction into deleting text data; and performing semantic recognition on the deleting text data using a preset large language model to obtain deletion semantic information. Specifically, receiving the voice instruction for deleting target content can be performed, as described in the above embodiment, via a microphone in a PBX system, for example. Converting the voice instruction into deleting text data can be performed by a text conversion module. Semantic recognition of the deleting text data using a preset large language model to obtain deletion semantic information can be performed. Specifically, semantic recognition can be performed using the LLM in conjunction with context to prevent recognition errors. In conjunction with the secondary confirmation described in the above embodiment, if the deletion semantic information is determined to be a critical operation, a secondary confirmation request is sent to facilitate user reconfirmation. Because deleting partial content during normal use may render related records inaccessible, the present application provides a method for requesting user reconfirmation before deleting content to prevent misoperation and enhance security. In an embodiment of the present application, receiving a confirmation instruction for the secondary confirmation request can be implemented in a variety of ways, such as receiving a confirmation instruction sent by the user through voice, or receiving a confirmation instruction sent by the user through a key or other method. In one example, the sending user's secondary confirmation request can be sent by voice or display, such as playing a voice "Please confirm whether to delete ***", and when the user replies with a voice "Confirm", the corresponding content is deleted. For another example, "Please confirm whether to delete ***" is displayed through the display module, and then confirmation can be made by pressing a key or touching, that is, a confirmation instruction for the user's secondary confirmation request, thereby deleting.

[0081] In an example, the voice command processing flow with optional secondary confirmation (taking the command "delete all call recordings" as an example, see Figure 4 (Schematic diagram of the command processing flow with secondary confirmation): 1. The user issues the voice command "Delete all call recordings" through the voice input interface module. 2. The voice recognition module converts the voice command into text: "Delete all call recordings." 3. The large language model core processing module receives the text, parses the semantics, and identifies the intended operation as "Delete all recordings," with no specific parameters or with the parameter "All." 4. The operation criticality assessment unit of the LLM core processing module, based on a pre-set rule base, determines "Delete all call recordings" as a critical operation and marks it as requiring secondary confirmation. 5. The user feedback module receives the secondary confirmation request and issues a confirmation prompt to the user, for example, via text-to-text transmission: "You have requested to delete all call recordings. This operation is irreversible. Do you confirm this? Please answer 'Yes' or 'No.'" or displaying a confirmation dialog box on the user interface. 6. The system waits for the user to confirm the request via voice or user interface. If the user confirms (e.g., by answering "Yes" or clicking the "Confirm" button), the LLM core processing module passes the confirmed intent to the PBX operation command generation and execution module. Subsequent steps are the same as steps 6-10 in point B: the deletion operation is executed and the results are reported. If the user denies the operation (e.g., answers "No" or clicks the "Cancel" button) or fails to respond within the preset time, the operation is aborted. The user feedback module notifies the user that "Operation Cancelled" or displays a corresponding prompt. The abort event is also recorded in the operation log. In actual use, the key operation rule library of the double confirmation mechanism can be configured and managed by the system administrator, allowing administrators to customize which types of PBX operations require double confirmation according to enterprise policies or set different confirmation trigger thresholds. Users can also enable or disable double confirmation prompts for certain common but sensitive operations in their personal settings.

[0082] In a possible implementation, the PBX system executes the first call instruction, see Figure 5 , the method further comprises:

[0083] Step S51, receiving a call back instruction for a target operation;

[0084] Step S52: identifying one or more corresponding target operations according to the callback call instruction;

[0085] Step S53, determining whether the corresponding one or more target operations are reversible operations;

[0086] Step S54: If it is determined to be a reversible operation, the corresponding one or more target operations are executed.

[0087] In an embodiment of the present application, the user can use the recall instruction to recall the corresponding operation. Specifically, the recall instruction can be sent by the user through voice or other means. The recall instruction can include one or more corresponding target operation information, so that the corresponding one or more target operations can be identified through the recall instruction. For example, the recall instruction is "recall the previous ***", and the corresponding target operation can be identified as the previous operation according to the recall instruction. In actual use, since the user generally recalls the previous one or more operations when recalling, when it is not clear which operation it is, it can be defaulted to the one or more operations corresponding to the previous call instruction. Because some operations have been executed and cannot be recalled, or are not recallable, such as the previous operation is to hang up the phone, if it has been executed, it cannot be recalled, while other instructions can be recalled, such as query operations. Therefore, in the present application, it is determined whether the target operation is a reversible operation, and the target operation is recalled when it is determined to be a reversible operation. In an example, the low-cost, one-click instruction execution recall mechanism corresponding to the present application can be found in Figure 6, including: 1. User-initiated withdrawal: After the previous voice command is executed, if the user believes that the LLM recognition is wrong or the operation result is unexpected, he or she can initiate a withdrawal request through a short voice command (such as "undo", "cancel just now", "wrong, go back") or click the "undo last operation" prompt button provided by the user feedback module on the interface. 2. Withdrawal intention recognition and operation positioning: The withdrawal intention recognition and confirmation unit of the LLM core processing module receives and parses the user's voice withdrawal command or interface withdrawal signal. The operation log recording unit of the context information management module is queried to locate the most recent PBX operation triggered and executed by the user through a voice command and its related detailed information in the record. 3. Reversibility judgment and reverse instruction generation: The reversible operation judgment and retraction instruction execution unit of the PBX operation instruction generation and execution module determines whether the operation can be safely and effectively undone based on the type of the located operation and the predefined PBX operation reversibility rules; (1) If it is judged to be a reversible operation, the unit automatically generates one or more reverse PBX instructions required to execute the operation; (2) If it is judged to be an irreversible operation, the judgment result is notified to the LLM core processing module. 4. Execution retraction or feedback infeasibility: For reversible operations, the generated reverse PBX instructions are sent to the PBX entity through the PBX system interface module for execution. For irreversible operations, the LLM core processing module explains to the user through the user feedback module that the specific operation cannot be directly "undone with one click". 5. User confirmation and feedback: The user feedback module clearly informs the user of the execution result of the retraction operation. 6. Log update: The instructions, execution process and results of this retraction operation are all recorded in detail in the operation log of the context information management module. In actual use, the operation log recording unit can support recording the latest N operation histories and allow users to selectively retract. During actual use, if the corresponding target instruction cannot be withdrawn, suggestions for compensation or processing can also be sent to the user.

[0088] To illustrate the solution of the embodiment of the present application, the following is described in conjunction with the corresponding system architecture and workflow. Figure 7, including: system architecture, including: 1. Voice input interface module. 2. Voice recognition module. 3. Large language model core processing module, instruction parsing unit, intention recognition unit, context understanding unit, dialogue management and clarification unit, complex instruction decomposition unit, withdrawal intention recognition and confirmation unit, operation criticality assessment unit: After identifying the user's intention and parameters, this unit determines whether the PBX operation corresponding to the intention is a critical operation based on the preset rule base (for example, defining deletion operations, batch modification operations, and operations involving system core configuration as critical operations) or dynamic evaluation logic. If it is a critical operation, the operation is marked as requiring secondary confirmation. 4. Context information management module, including: PBX status interface; user data storage; conversation history and operation log recording unit: records each interactive conversation triggered by voice commands in detail and in an orderly manner, including: the text representation of the original voice, the final text recognized by ASR, the intent and parameters parsed by LLM, the specific PBX instructions generated accordingly, the timestamp of the instruction execution, the status snapshot of PBX-related entities before execution (for rollback-configured operations), the PBX execution status feedback and final results, and whether the user has reconfirmed key operations. 5. PBX operation instruction generation and execution module, including: intent-operation mapping unit, parameter adaptation and verification unit, instruction execution control unit, reversible operation judgment and instruction rollback execution unit, and result processing unit. 6. PBX system API Gateway (interface module). 7. User feedback module, including: an information integration unit, which integrates LLM clarification questions, PBX execution results, operation confirmation prompts (including reconfirmation requests for key operations), and the results of undo operations; a feedback generation unit, which converts information into natural language text or other forms, such as generating clear reconfirmation queries; and a TTS / UI (output interface), which announces feedback via speech synthesis or displays it through a graphical user interface. The user interface can be designed to briefly display a clickable "Undo Last Operation" button in a specific area (such as the lower right corner of the screen) after executing a voice command. For operations requiring reconfirmation, a clear confirmation dialog box or voice prompt will pop up.

[0089] As can be seen, the methods of the embodiments of the present application enable deep semantic understanding and fine-grained execution of PBX voice commands based on a large language model. This not only establishes a context management mechanism within the PBX voice control system, supporting the large language model's ability to resolve ambiguous commands and determine the user's true intent, but also utilizes the large language model to intelligently decompose and logically plan complex user voice commands, converting them into action sequences or configuration rules executable by the PBX system. This also provides a unified natural language voice interaction interface and implementation method covering a wide range of PBX functions, as well as a PBX voice control system with clarifying interaction and intelligent fault tolerance. A low-cost user-triggered reversal mechanism for recent operations is implemented, including operation history recording, operation reversibility determination, and automatic generation and execution of reverse commands. Furthermore, an optional secondary confirmation mechanism can be implemented within the large language model-based PBX voice control system. This mechanism includes determining the risk level of the corresponding operation via the operation criticality assessment unit within the large language model core processing module, and automatically requesting explicit confirmation from the user through the user feedback module before executing a high-risk operation. The PBX operation is only executed after receiving a positive response.

[0090] In a second aspect of the embodiment of the present application, an intelligent voice control system for a communication system based on LLM is provided. Figure 8 , Figure 8 A schematic diagram of the structure of an intelligent voice control system for a communication system based on LLM provided in an embodiment of the present application, the system includes:

[0091] A first instruction acquisition module 801 is used to acquire a first voice instruction issued by a user;

[0092] A first instruction conversion module 802 is configured to convert the first voice instruction into first text data;

[0093] A first semantic recognition module 803 is configured to perform semantic recognition on the first text data using a preset large language model to obtain first semantic information;

[0094] The first instruction execution module 804 is configured to generate a first call instruction according to the first semantic information, so as to enable the PBX system to execute the first call instruction, wherein the first call instruction corresponds to one or more operations.

[0095] In a possible implementation, the first instruction execution module includes:

[0096] an operation judgment submodule, configured to judge whether the operation corresponding to the first semantic information is a key operation based on a plurality of preset key operation types and / or preset evaluation logics;

[0097] The instruction execution submodule is used to generate the first call instruction according to the first semantic information if it is determined that the operation corresponding to the semantic information is a key operation.

[0098] In a possible implementation, the operation determination submodule is specifically configured to match a plurality of preset key operation types with the first semantic information; if a match occurs, determining that the operation corresponding to the first semantic information is a key operation.

[0099] In one possible implementation, the instruction execution submodule is specifically configured to send a secondary confirmation request if it is determined that the operation corresponding to the semantic information is a critical operation; and after receiving a confirmation instruction for the secondary confirmation request, generate the first call instruction according to the first semantic information.

[0100] In one possible implementation, the first instruction acquisition module is specifically used to receive a voice deletion instruction for target content; the first instruction conversion module is specifically used to convert the voice deletion instruction into deletion text data; and the first semantic recognition module is specifically used to perform semantic recognition on the deletion text data through a preset large language model to obtain deletion semantic information.

[0101] In one possible embodiment, the first instruction execution module is specifically used to receive a rollback call instruction for a target operation; identify the corresponding one or more target operations based on the rollback call instruction; determine whether the corresponding one or more target operations are reversible operations; if it is determined to be a reversible operation, execute the corresponding one or more target operations.

[0102] In a possible implementation, the system further includes:

[0103] The result prompt module is used to obtain the execution result of the first calling instruction; generate and send a prompt voice according to the execution result.

[0104] It can be seen that through the system of the embodiment of the present application, after receiving the first voice command sent by the user, the voice command can be converted into text data, and then the semantic information corresponding to the text data can be parsed through the preset large language model. Then, a first call instruction is generated according to the first semantic information, so that the PBX system executes the first call instruction. In this way, not only can operation be performed through voice to solve the problem of inconvenient key operation, but also recognition can be performed through the preset large language model to improve the recognition efficiency of the instruction.

[0105] The present application also provides an electronic device, such as Figure 9 Shown, including:

[0106] Memory 901, used for storing computer programs;

[0107] The processor 902 is configured to execute the program stored in the memory 901 and implement the following steps:

[0108] Obtaining a first voice command issued by a user;

[0109] Converting the first voice instruction into first text data;

[0110] Performing semantic recognition on the first text data using a preset large language model to obtain first semantic information;

[0111] A first call instruction is generated according to the first semantic information, so that the PBX system executes the first call instruction, wherein the first call instruction corresponds to one or more operations.

[0112] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0113] The communication interface is used for communication between the above electronic device and other devices.

[0114] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk memory. Alternatively, the memory may be at least one storage device located away from the processor.

[0115] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0116] In another embodiment provided in the present application, a computer-readable storage medium is also provided, which stores a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned LLM-based communication system intelligent voice control methods are implemented.

[0117] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any of the LLM-based intelligent voice control methods for a communication system in the above embodiments.

[0118] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a solid-state drive (SSD).

[0119] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0120] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system, electronic device, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For related portions, refer to the descriptions of the method embodiments.

[0121] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.

Claims

1. An intelligent voice control method for a communication system based on LLM, characterized in that: The method comprises: Obtaining a first voice command issued by a user; Converting the first voice instruction into first text data; Performing semantic recognition on the first text data using a preset large language model to obtain first semantic information; A first call instruction is generated according to the first semantic information, so that the PBX system executes the first call instruction, wherein the first call instruction corresponds to one or more operations.

2. The method according to claim 1, characterized in that Generating a first call instruction according to the first semantic information includes: Determining whether the operation corresponding to the first semantic information is a key operation by using multiple preset key operation types and / or preset evaluation logic; If it is determined that the operation corresponding to the semantic information is a key operation, the first call instruction is generated according to the first semantic information.

3. The method according to claim 2, characterized in that The determining whether the operation corresponding to the first semantic information is a key operation based on a plurality of preset key operation types includes: Matching multiple preset key operation types with the first semantic information; If there is a match, the operation corresponding to the first semantic information is determined to be a key operation.

4. The method according to claim 2, characterized in that If it is determined that the operation corresponding to the semantic information is a key operation, generating the first call instruction according to the first semantic information includes: If the operation corresponding to the semantic information is determined to be a key operation, a secondary confirmation request is sent; After receiving a confirmation instruction for the secondary confirmation request, the first calling instruction is generated according to the first semantic information.

5. The method according to claim 1, characterized in that The method of obtaining a first voice command issued by a user; converting the first voice command into first text data; and performing semantic recognition on the first text data using a preset large language model to obtain first semantic information includes: receiving a voice deletion instruction for target content; Converting the voice deletion instruction into deletion text data; The deleted text data is semantically recognized by using a preset large language model to obtain deletion semantic information.

6. The method according to claim 1, characterized in that The PBX system executes the first calling instruction, including: Receive a retracement call instruction for the target operation; Identifying one or more corresponding target operations according to the callback call instruction; Determining whether the corresponding one or more target operations are reversible operations; If a reversible operation is determined, the corresponding one or more target operations are performed.

7. The method according to claim 1, characterized in that After the PBX system executes the first call instruction, the method further includes: Obtaining an execution result of the first calling instruction; Generate and send a prompt voice according to the execution result.

8. An intelligent voice control system for a communication system based on LLM, characterized in that: The system comprises: A first instruction acquisition module, configured to acquire a first voice instruction issued by a user; A first instruction conversion module, configured to convert the first voice instruction into first text data; a first semantic recognition module, configured to perform semantic recognition on the first text data using a preset large language model to obtain first semantic information; The first instruction execution module is configured to generate a first call instruction according to the first semantic information, so as to enable the PBX system to execute the first call instruction, wherein the first call instruction corresponds to one or more operations.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Speech command processing method, apparatus and system

    CN106992001A

  • Articulated naturality web business executing method and device

    CN109256126A

  • Voice interaction method and device

    CN110265016A

  • Vehicle-mounted voice instruction withdrawing method and device

    CN116884406A

  • Equipment interaction method based on intelligent voice recognition, equipment and medium

    CN117524218A