Voice control operation method, device and equipment

Through the voice control operation method, users' voice requests on the SaaS platform are analyzed and interface operation instructions are generated, solving the problem of low backend operation efficiency of multi-tenant SaaS platform and improving user experience and system stability.

CN120220673APending Publication Date: 2025-06-27BEIJING BAILONG MAYUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510327539.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the multi-tenant SaaS platform, background operations rely on graphical user interfaces, resulting in users needing to have certain technical background and operation skills, and the interaction efficiency through keyboard or mouse is low, affecting the user experience.

Method used

A method of voice control operation is provided, by obtaining the interface information of the current graphical user interface and the user's voice operation request, parsing the voice request, generating interface operation instructions, and executing these instructions to complete the operation of the user's intention.

Benefits of technology

It improves the user's operation convenience and efficiency on the SaaS platform, lowers the operation threshold, improves the user experience, and automatically recognizes and analyzes voice operation requests, avoids misoperation and errors, and improves the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220673A_ABST
    Figure CN120220673A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of computers, and discloses a voice control operation method, device and equipment, and the method comprises the steps: obtaining the interface information of a current graphical user interface, and a voice operation request input by a user based on the current graphical user interface; analyzing the voice operation request based on the interface information to obtain a voice analysis result; generating an interface operation instruction based on the voice analysis result; and based on the interface operation instruction, performing interface operation on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface, and displaying an operation result. By applying the technical scheme of the invention, the efficiency and the reliability of the control operation on the graphical user interface can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and particularly to methods, devices, and equipment for voice control operations. Background Art

[0002] In a multi-tenant SaaS platform, background operations usually rely on a graphical user interface (GUI). These methods require users to have a certain technical background and operation skills, and need to use a keyboard or mouse for interaction. However, in cases where frequent or urgent operations are required, interacting through a keyboard or mouse may reduce the efficiency of user operations, thus affecting the user experience of background operations. Summary of the Invention

[0003] In view of the above problems, embodiments of the present invention provide a method, device, and equipment for voice control operations.

[0004] According to one aspect of the embodiments of the present invention, there is provided a method for voice control operations, which is applied to a SaaS platform, and multiple graphical user interfaces are configured on the SaaS platform. The method includes: obtaining interface information of the current graphical user interface and a voice operation request input by a user based on the current graphical user interface; parsing the voice operation request based on the interface information to obtain a voice parsing result; generating an interface operation instruction based on the voice parsing result; and performing an interface operation on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface based on the interface operation instruction, and displaying the operation result. Through the above process, the operation convenience and efficiency of the user on the SaaS platform can be greatly improved. The user can complete complex operations only through voice instructions without the need for cumbersome clicks or keyboard inputs, which reduces the operation threshold and improves the user experience. In addition, by automatically identifying and parsing the user's voice operation request, errors caused by misoperations or improper operations are avoided, and the stability and reliability of the system are improved.

[0005] In an alternative embodiment, parsing the voice operation request based on the interface information to obtain a voice parsing result includes:

[0006] Based on the interface information, obtaining the component identifiers of each operation control component on the current graphical user interface;

[0007] Identifying keywords in the voice operation request and matching the keywords with the component identifiers to identify component operation instructions associated with the component identifiers;

[0008] If the matching is successful, obtaining a voice parsing result based on the component operation instruction and the component identifier;

[0009] If the matching fails, obtain the voice operation request again, or prompt the user to enter the correct voice input.

[0010] In an alternative embodiment, generating an interface operation instruction based on the voice parsing result includes:

[0011] Determine the corresponding interface operation action based on the component operation instructions corresponding to each component identifier in the voice parsing result;

[0012] Generate an interface operation instruction according to the interface operation action and the layout information of the current graphical user interface.

[0013] In an alternative embodiment, determining the corresponding interface operation action based on the component operation instructions corresponding to each component identifier in the voice parsing result includes:

[0014] Obtain the order and dependency relationship of the appearance of each component identifier in the voice operation request;

[0015] Analyze the execution priority of each component operation instruction according to the order and dependency relationship;

[0016] Sort the component operation instructions according to the execution priority;

[0017] Match the sorted component operation instructions with the calibrated operation instructions stored in the interface operation action library to obtain an instruction matching result;

[0018] Determine the interface operation action corresponding to each component operation instruction based on the instruction matching result.

[0019] In an alternative embodiment, generating an interface operation instruction according to the interface operation action and the layout information of the current graphical user interface includes:

[0020] Identify the target operation control component associated with the interface operation action in the current graphical user interface;

[0021] Determine the position information of the target operation control component in the current graphical user interface according to the layout information;

[0022] Construct an interface operation instruction based on the position information and the interface operation action. The interface operation instruction includes the operation type and operation parameters of the target operation control component to control the current graphical user interface to perform the corresponding operation.

[0023] In an alternative embodiment, generating an interface operation instruction based on the voice parsing result further includes:

[0024] Obtain the target instruction extraction model;

[0025] Extract instructions from the speech parsing result based on the target instruction extraction model, and generate interface operation instructions based on the instruction extraction result.

[0026] In an alternative embodiment, obtaining the target instruction extraction model includes:

[0027] Obtain training sample data and corresponding sample labels, where the sample labels are used to represent the interface operation instructions corresponding to the speech parsing result;

[0028] Input the training sample data and the corresponding sample labels into the initial instruction extraction model to train the initial instruction extraction model and obtain an instruction extraction training model;

[0029] Test the instruction extraction training model based on the test data to obtain a test result;

[0030] If the test result meets the preset conditions, use the instruction extraction training model as the target instruction extraction model;

[0031] If the test result does not meet the preset conditions, adjust the parameters of the initial instruction extraction model and retrain the initial instruction extraction model until an instruction extraction training model that meets the preset conditions is obtained, and use the instruction extraction training model that meets the conditions as the target instruction extraction model.

[0032] In an alternative embodiment, based on the interface operation instructions, perform interface operations on the current graphical user interface and / or other graphical user interfaces and display the operation results, including:

[0033] Obtain the instruction confirmation result feedback by the user for the interface operation instructions;

[0034] If the instruction confirmation result is to confirm execution, perform the interface operation corresponding to the interface operation instructions, and obtain the status information of the current graphical user interface and / or other graphical user interfaces in real time, so as to update and display the content of the current graphical user interface and / or other graphical user interfaces based on the status information;

[0035] If the instruction confirmation result is to cancel execution, do not execute the interface operation instructions and keep the status of the current graphical user interface and / or other graphical user interfaces unchanged.

[0036] According to another aspect of the embodiments of the present invention, a device for voice control operations is provided, which is applied to a SaaS platform. Multiple graphical user interfaces are configured on the SaaS platform, including: an information acquisition module for acquiring the interface information of the current graphical user interface and the voice operation request input by the user based on the current graphical user interface; a voice parsing module for parsing the voice operation request based on the interface information to obtain a voice parsing result; an instruction generation module for generating an interface operation instruction based on the voice parsing result; and an interface operation module for performing interface operations on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface based on the interface operation instruction and displaying the operation result. Through the above modules, the operation convenience and efficiency of the user on the SaaS platform can be greatly improved. The user can complete complex operations only through voice instructions without the need for cumbersome clicks or keyboard inputs, which reduces the operation threshold and improves the user experience. In addition, by automatically identifying and parsing the user's voice operation request, errors caused by misoperations or improper operations are avoided, and the stability and reliability of the system are improved.

[0037] According to another aspect of the embodiments of the present invention, a computer device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete mutual communication through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations of the aforementioned method for voice control operations.

[0038] According to yet another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which at least one executable instruction is stored, and the executable instruction causes a computer device / device to execute the operations of the aforementioned method for voice control operations.

[0039] According to yet another aspect of the embodiments of the present invention, a computer program product is provided, including computer instructions for causing a computer to execute the operations of the method for voice control operations in the first aspect or any corresponding embodiment thereof.

[0040] The above description is only an overview of the technical solutions of the embodiments of the present invention. In order to be able to understand the technical means of the embodiments of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features, and advantages of the embodiments of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically given below. Description of the Drawings

[0041] The drawings are only used to illustrate the embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0042] Figure 1The figure shows a schematic flowchart of a method for voice control operation provided by the present invention;

[0043] Figure 2 The figure shows another schematic flowchart of a method for voice control operation provided by the present invention;

[0044] Figure 3 The figure shows a schematic structural diagram of a device for voice control operation provided by the present invention;

[0045] Figure 4 The figure shows a schematic structural diagram of a computer device provided by the present invention. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] Figure 1 The figure shows a flowchart of the first embodiment of a method for voice control operation of the present invention. As Figure 1 shown, applied to the SaaS platform, where multiple graphical user interfaces are configured on the SaaS platform, the method includes the following steps:

[0048] Step 110, obtaining the interface information of the current graphical user interface and the voice operation request input by the user based on the current graphical user interface.

[0049] Among them, the interface information includes but is not limited to the layout of the current graphical user interface, the positions of the function buttons, and the text or image content displayed on the interface. The voice operation request refers to an instruction for operating the graphical user interface issued by the user through a voice input device, such as a microphone. By obtaining the interface information and the voice operation request, the SaaS platform can accurately understand the user's intention and perform corresponding operations according to the user's voice operation request.

[0050] For example, if a user wishes to operate the graphical user interface by the voice command "Open File Manager". In this case, the SaaS platform will first parse the interface information to determine whether there is a function button named "File Manager" in the current graphical user interface and its location. Then, the SaaS platform will simulate the operation of the user clicking on this function button according to the voice operation request, so as to open the File Manager. In this way, the user can quickly complete various operations through voice commands without manually operating the graphical user interface, greatly improving the operation efficiency and convenience.

[0051] Step 120, parse the voice operation request based on the interface information to obtain a voice parsing result.

[0052] Among them, the voice parsing result includes but is not limited to the identified operation object, operation action, and possible operation parameters. For example, in the voice operation request "Open File Manager", the operation object is "File Manager" and the operation action is "Open". If the voice operation request also includes a specific file path or file name, these will also be parsed as operation parameters. After obtaining the voice parsing result, the SaaS platform will, based on this result and in combination with the interface information, simulate the actual operation of the user and execute the corresponding function. This process ensures that the SaaS platform can accurately respond to the user's voice commands and achieve the operation effect expected by the user.

[0053] In some alternative embodiments, when parsing the voice operation request based on the interface information to obtain a voice parsing result, the component identifiers of each operation control component on the current graphical user interface can be obtained based on the interface information; the keywords in the voice operation request are identified and the keywords are matched with the component identifiers to identify the component operation instructions associated with the component identifiers; if the matching is successful, the voice parsing result is obtained based on the component operation instructions and the component identifiers; if the matching fails, the voice operation request is obtained again, or the user is prompted to enter correct voice input.

[0054] Specifically, during the matching process, first, based on the interface information of the current graphical user interface, determine the layout and component information of the current graphical user interface, and construct a component identification library. This component identification library contains the unique identifiers of all operable control components in the current graphical user interface, as well as their corresponding function descriptions. Subsequently, using natural language processing technology, extract keywords from the voice operation request. These keywords are usually a direct manifestation of the user's intention, such as "open", "close", "edit", etc. When the keywords are successfully matched with the component identifiers, further parse out the specific component operation instructions. This instruction details the specific operation that the user hopes to perform on the operation control component, such as "open a certain folder in the file manager". If the matching fails, that is, the user's intention cannot be accurately recognized from the voice operation request, or the user's intention does not match the operation control components on the current graphical user interface, then one of two strategies will be adopted: either re-obtain the voice operation request to obtain a clearer or more correct instruction; or prompt the user to make a correct voice input through voice or graphical interface means to ensure flexible response in the face of ambiguous or incorrect voice instructions and improve the user experience.

[0055] In some alternative embodiments, based on the interface information, the component identifiers of each operation control component on the current graphical user interface and the corresponding component attribute information can also be obtained. The component identifier can be the unique identifier of the component, which is used to distinguish different components on the interface. The component attribute information describes the characteristics of the component, such as type, position, size, status, etc. By obtaining this information, the SaaS platform can more accurately understand the interface layout and component functions, providing a basis for subsequent parsing of the operation object and operation action in the voice operation request. For example, if the voice instruction is "click button 1", the SaaS platform needs to first determine whether there is a component labeled "button 1" on the interface and understand the position and clickable status of this component, so as to accurately simulate the user's click operation. Secondly, identify the keywords in the voice operation request and match the keywords with the component identifiers to identify the component operation instructions associated with the component identifiers; if the keywords match the component identifiers successfully, determine the specific operation instructions according to the attribute information of the component. For example, if the keyword is "open" and the matched component is the "open" button in a file selection dialog box, the operation instruction is to trigger the click event of this button to open the file selection dialog box. In this way, the SaaS platform can flexibly convert the user's voice instructions into various specific interface operations to meet the user's needs in different scenarios. At the same time, this matching process also has a certain degree of fault tolerance. Even if there are minor deviations or ambiguities in the user's voice instructions, the SaaS platform can make reasonable inferences and corrections based on the context information and component attribute information to ensure the accuracy and efficiency of the operation. For example, when the user misreads "click the submit button" as "click submit" due to accent or speech rate problems, the SaaS platform can understand the user's intention to perform the submission operation based on the context and automatically match the correct "submit" button for clicking. In addition, if the user does not clearly specify the operation object, the SaaS platform can also intelligently predict and execute the most likely operation according to the current focus position on the interface or user habits. This intelligent fault tolerance mechanism not only improves the user experience but also significantly enhances the practicality and flexibility of the SaaS platform. Moreover, if the keyword fails to match the component identifier, the user's voice operation request can be parsed again, or the user can be informed through the interface display or sound that the input voice instruction is incorrect, guiding the user to re-enter the correct voice operation request, thereby improving the user experience. For example, when the user's voice instruction cannot be recognized, a voice prompt can be given: "Sorry, I didn't understand your instruction. Please re-enter." Or, corresponding error prompt information can be displayed on the interface to guide the user to perform the correct operation. At the same time, the user's incorrect input can also be recorded for subsequent data analysis to optimize the voice recognition algorithm and improve the recognition accuracy of the system.

[0056] In some alternative embodiments, when parsing a voice operation request based on interface information to obtain a voice parsing result, component identifiers of each operation control component on the current graphical user interface may also be obtained based on the interface information; keywords in the voice operation request are identified, and the keywords are matched with the component identifiers to identify voice instructions associated with the component identifiers; if the match is successful, natural language conversion is performed on the voice instructions based on the semantic space constructed based on the current graphical user interface to obtain a voice parsing result; if the match fails, the voice instructions are parsed based on general semantic rules to obtain the closest voice parsing result, and the user is prompted to confirm or correct it.

[0057] In some alternative embodiments, when performing natural language conversion on voice instructions based on the semantic space constructed based on the current graphical user interface to obtain a voice parsing result, a dynamic semantic space may be constructed based on information such as the layout, type, attributes of the operation control components on the current graphical user interface, and the user's historical operation behaviors. This semantic space can understand and map the relationship between user voice instructions and graphical user interface elements. For example, if the user voice instruction is "click on that red button", the system can identify through the semantic space that "red" corresponds to a certain red button component on the graphical user interface, so as to accurately perform the click operation. In this way, the system can more intelligently parse and execute user voice instructions, improving the accuracy and efficiency of operations.

[0058] Furthermore, during the process of natural language conversion after a successful match, based on the context information of the current graphical user interface, such as the currently selected component, the opened window, or the active area, etc., it is ensured that the converted instruction can accurately reflect the user's true intention. This conversion process is not limited to simple keyword replacement, but also includes the understanding and processing of implicit logical relationships in voice instructions. For example, for an instruction like "insert picture B into document A", the relationship between "document A" and "picture B" needs to be understood and correctly converted into an operation instruction for the corresponding components on the graphical user interface. In addition, in the case of a failed match, when adopting the strategy of parsing based on general semantic rules, the system will as much as possible infer the operation that the user may want to perform based on the user's voice instruction and the context information of the current graphical user interface, and provide the closest parsing result for the user to confirm or correct. This process not only improves the fault tolerance of the system, but also enables the user to interact with the system more naturally and smoothly.

[0059] In some alternative embodiments, it is also possible to predict the user's likely next operation by analyzing the user's historical operation habits, and pre-load relevant components or functions in advance to reduce the user's waiting time. For example, if the user often opens the file manager immediately after using a certain application, then when the user starts the application, the relevant components of the file manager will be pre-loaded to ensure that when the user issues an instruction to open the file manager, it can respond quickly.

[0060] Furthermore, in order to improve the recognition accuracy of voice commands and the user experience, it is also possible to preprocess the user's voice input. The preprocessing steps may include noise cancellation, voice enhancement, and voice segmentation, etc., to ensure that the system can clearly capture the user's voice commands. In addition, the user can also set specific voice commands for commonly used operation control components according to their own usage habits, thereby further improving the operation efficiency and convenience.

[0061] Step 130, generate an interface operation instruction based on the voice parsing result.

[0062] Among them, the above interface operation instruction can be executed through the interface control module of the operating system, that is, the interface control module performs corresponding interface operations according to the instruction content. For example, if the voice parsing result is to open a certain application, the interface operation instruction will be an instruction to start the application; if the voice parsing result is to close the current window, the interface operation instruction will be an instruction to close the current window. In this way, the user can directly control the interface operations of the device through voice commands, greatly improving the operation efficiency and convenience. At the same time, since the interface operation instruction is generated based on the voice parsing result, it can also effectively avoid the misoperation problems that may be caused by traditional manual operations.

[0063] In some alternative embodiments, when generating an interface operation instruction based on the voice parsing result, it is possible to determine the corresponding interface operation action based on the component operation instructions corresponding to the component identifiers in the voice parsing result; and generate an interface operation instruction according to the interface operation action and the layout information of the current graphical user interface.

[0064] In some alternative embodiments, when determining the corresponding interface operation actions based on the component operation instructions corresponding to each component identifier in the voice parsing result, the order and dependency of each component identifier appearing in the voice operation request may be obtained first; according to the order and dependency, the execution priorities of each component operation instruction may be analyzed; the component operation instructions may be sorted according to the execution priorities; the sorted component operation instructions may be matched with the calibrated operation instructions stored in the interface operation action library to obtain an instruction matching result; based on the instruction matching result, the interface operation actions corresponding to each component operation instruction may be determined. For example, if the voice operation request contains multiple component operation instructions and there are dependencies between these instructions, such as the need to select a certain file first and then perform an open operation, then the component operation instruction for selecting the file needs to be executed first. By obtaining the order and dependency of these component identifiers appearing in the voice operation request, the execution priorities of each component operation instruction can be determined, so as to ensure that the generation order of the interface operation instructions is correct, avoid operation failures or incorrect operations caused by incorrect execution orders, and further improve the accuracy and reliability of voice control and enhance the user experience.

[0065] After the execution priorities are determined, based on the conflict situation between the component operation instructions, critical or urgent component operation instructions may be preferentially executed, or the user may be prompted to make a manual selection. For example, when multiple component operation instructions need to be executed simultaneously but there are conflicts in their interface operations, a decision can be made according to a preset conflict resolution strategy to preferentially execute critical or urgent operation instructions, or prompt the user to make a manual selection. The conflict resolution strategy can be formulated based on various factors such as the importance, urgency, and user habits of the component operation instructions. For example, if a certain component operation instruction is related to system security or data integrity, it may be set to be preferentially executed; if a certain component operation instruction is frequently used by the user, then this instruction can be preferentially executed by default to improve efficiency. At the same time, when a conflict is detected and an automatic decision cannot be made, the user can be informed through voice prompts, screen displays, etc., so that the user can manually select the instruction to be executed. Such a design not only ensures the automated processing ability of the SaaS platform but also takes into account the user's participation and control, making the entire voice control operation more intelligent and user-friendly. In addition, for some complex interface operation actions, the user can also add or modify the mapping relationship between specific operation instructions and actions in the interface operation action library according to their own needs, thereby further enhancing the flexibility and adaptability of the operation and improving the overall user experience.

[0066] In some alternative embodiments, when generating an interface operation instruction based on an interface operation action and the layout information of the current graphical user interface, the target operation control component associated with the interface operation action in the current graphical user interface may be identified first; according to the layout information, the position information of the target operation control component in the current graphical user interface is determined; based on the position information and the interface operation action, an interface operation instruction is constructed, and the interface operation instruction includes the operation type and operation parameters for the target operation control component to control the current graphical user interface to perform corresponding operations.

[0067] Specifically, during the process of constructing the interface operation instruction, it can also be constructed based on the status information of the operation control component. For example, when a button component may be in an available, disabled, or hidden state, corresponding operation instructions can be determined according to this status information. If the button is in a disabled or hidden state, no operation instruction for this button will be generated to avoid invalid operations or confusing the user. In addition, to enhance the robustness and stability of the SaaS platform, error checking and verification can also be performed when generating the interface operation instruction, including the legality verification of the operation control component, the rationality check of the operation type and operation parameters, etc. Only the interface operation instructions that pass these checks and verifications will be finally executed by the system to ensure that the generated interface operation instructions not only conform to the user's operation intention but also can effectively and accurately control the current graphical user interface to perform corresponding operations, thereby providing a more fluent and stable operation experience for the user.

[0068] Step 140, based on the interface operation instruction, perform an interface operation on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface, and display the operation result.

[0069] Among them, during the interface operation process, corresponding operations, such as clicking, dragging, and entering text, can be performed on the corresponding graphical user interface elements (including operation control components) according to the type and content of the interface operation instruction. At the same time, the display content of the graphical user interface is updated in real time to reflect the result of the operation. For example, if the user requests to open a file through the interface operation instruction, the content of the file is displayed in the graphical user interface; if the user requests to modify the value of a certain setting item, the display value of this setting item is updated in the graphical user interface, so that the user can interact with the system through the interface operation instruction to implement various functional operations and obtain immediate feedback results.

[0070] In some alternative embodiments, when performing an interface operation on the current graphical user interface and / or other graphical user interfaces based on an interface operation instruction and displaying the operation result, the instruction confirmation result feedback by the user for the interface operation instruction may be obtained first; if the instruction confirmation result is to confirm execution, the interface operation corresponding to the interface operation instruction is executed, and the status information of the current graphical user interface and / or other graphical user interfaces is obtained in real time, so as to update and display the content of the current graphical user interface and / or other graphical user interfaces based on the status information; if the instruction confirmation result is to cancel execution, the interface operation instruction is not executed, and the status of the current graphical user interface and / or other graphical user interfaces remains unchanged, so as to further improve the accuracy of the interface operation and the user's operation experience. Before the user confirms the execution of the interface operation instruction, no operation is performed, thus avoiding interface chaos or errors caused by misoperation or uncertain operation intentions. At the same time, obtaining and updating the status information of the graphical user interface in real time can ensure that the interface content seen by the user is always the latest and most accurate, thereby enhancing the user's trust and operation confidence in the SaaS platform. In addition, if the user chooses to cancel the execution of the interface operation instruction, the SaaS platform will also respond immediately and keep the status of the current interface unchanged, avoiding unnecessary interface changes and user confusion.

[0071] For example, the interface operation instruction may include clicking on the current graphical user interface, submitting form data, navigating from the current graphical user interface to other graphical user interfaces associated with the current graphical user interface, performing creator operations on resources, etc., and the interface operation instruction may be displayed in the target area of the current graphical user interface in JSON format.

[0072] The method for voice control operation according to the embodiments of the present invention includes obtaining the interface information of the current graphical user interface and the voice operation request input by the user based on the current graphical user interface; parsing the voice operation request based on the interface information to obtain a voice parsing result; generating an interface operation instruction based on the voice parsing result; and performing an interface operation on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface based on the interface operation instruction and displaying the operation result, which can greatly improve the operation convenience and efficiency of the user on the SaaS platform. The user can complete complex operations only by voice instructions without the need for cumbersome clicking or keyboard input, reducing the operation threshold and improving the user experience. In addition, by automatically identifying and parsing the user's voice operation request, errors caused by misoperation or improper operation are avoided, and the stability and reliability of the system are improved.

[0073] Figure 2 The flowchart of another embodiment of the method for voice control operation according to the present invention is shown. As Figure 2As shown, it is applied to a SaaS platform, and multiple graphical user interfaces are configured on the SaaS platform. The method includes the following steps:

[0074] Step 210, obtain the interface information of the current graphical user interface and the voice operation request input by the user based on the current graphical user interface.

[0075] For details, please refer to Figure 1 Step 110 of the embodiment shown, which will not be elaborated here.

[0076] Step 220, parse the voice operation request based on the interface information to obtain a voice parsing result.

[0077] For details, please refer to Figure 1 Step 120 of the embodiment shown, which will not be elaborated here.

[0078] Step 230, generate an interface operation instruction based on the voice parsing result.

[0079] Specifically, the above step 230 includes:

[0080] Step 2301, obtain a target instruction extraction model.

[0081] Among them, the above target instruction extraction model is a pre-trained model for extracting specific instructions corresponding to the target operation from the voice parsing result.

[0082] In some optional implementation manners, when obtaining the target instruction extraction model, training sample data and corresponding sample labels can be obtained, and the sample labels are used to represent the interface operation instructions corresponding to the voice parsing result; input the training sample data and the corresponding sample labels into an initial instruction extraction model to train the initial instruction extraction model to obtain an instruction extraction training model; test the instruction extraction training model based on test data to obtain a test result; if the test result meets the preset conditions, use the instruction extraction training model as the target instruction extraction model; if the test result does not meet the preset conditions, adjust the parameters of the initial instruction extraction model and re-train the initial instruction extraction model until an instruction extraction training model that meets the preset conditions is obtained, and use the instruction extraction training model that meets the conditions as the target instruction extraction model.

[0083] During the training process, the preset conditions may include but are not limited to evaluation metrics such as the accuracy, recall rate, and F1 score of the model reaching preset thresholds. In addition, in order to improve the generalization ability of the model, technical means such as cross-validation and data augmentation can also be adopted. Once the target instruction extraction model is determined, it can be applied to the actual SaaS platform to automatically extract the corresponding interface operation instructions according to the user's voice operation request, thus realizing intelligent operation control, which not only improves the user's operation efficiency but also reduces the operation complexity and enhances the user experience. At the same time, due to the adoption of machine learning technical means, this method can also continuously learn and optimize with the user's use, further improving its performance and accuracy.

[0084] In addition, the target instruction extraction model can adopt deep learning algorithms, such as convolutional neural network CNN, recurrent neural network RNN, or its variants (such as long short-term memory network LSTM), etc. Through the learning of a large number of voice operation requests and corresponding interface operation instructions, it can accurately understand the voice parsing results and extract the target instructions.

[0085] Furthermore, the training sample data and corresponding sample labels can be stored in a schema structure. For example:

[0086]

[0087]

[0088]

[0089] Step 2302, based on the target instruction extraction model, extract instructions from the voice parsing results to generate interface operation instructions based on the instruction extraction results.

[0090] In specific implementation, first, the user's voice operation request is input into the pre-trained speech recognition model to obtain the voice parsing result. Subsequently, the voice parsing result is input into the target instruction extraction model, and the model will extract the corresponding interface operation instructions from the voice parsing result according to the learned knowledge. To ensure the accuracy and feasibility of the instructions, the extracted interface operation instructions can also be verified and optimized to meet the requirements of the actual application scenario. Finally, the generated interface operation instructions are sent to the corresponding execution module to realize intelligent operation control. The whole process is efficient and accurate, greatly enhancing the user's operation experience and efficiency.

[0091] Step 240, based on the interface operation instructions, perform interface operations on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface, and display the operation results.

[0092] For details, please refer to Figure 1Step 140 of the illustrated embodiment will not be elaborated here.

[0093] In summary, the method for voice control operation in the embodiment of the present invention obtains the interface information of the current graphical user interface and the voice operation request input by the user based on the current graphical user interface; parses the voice operation request based on the interface information to obtain a voice parsing result; generates an interface operation instruction based on the voice parsing result; and performs an interface operation on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface based on the interface operation instruction, and displays the operation result, which can greatly improve the operation convenience and efficiency of the user on the SaaS platform. The user can complete complex operations only through voice instructions without the need for cumbersome clicks or keyboard inputs, reducing the operation threshold and improving the user experience. In addition, by automatically identifying and parsing the user's voice operation request, errors caused by misoperation or improper operation are avoided, and the stability and reliability of the system are improved.

[0094] Figure 3 The structural schematic diagram of an embodiment of a device for voice control operation of the present invention is shown. As Figure 3 shown, applied to the SaaS platform, where multiple graphical user interfaces are configured on the SaaS platform, the device includes:

[0095] An information acquisition module 310, configured to acquire the interface information of the current graphical user interface and the voice operation request input by the user based on the current graphical user interface;

[0096] A voice parsing module 320, configured to parse the voice operation request based on the interface information to obtain a voice parsing result;

[0097] An instruction generation module 330, configured to generate an interface operation instruction based on the voice parsing result;

[0098] An interface operation module 340, configured to perform an interface operation on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface based on the interface operation instruction, and display the operation result.

[0099] In an alternative embodiment, the voice parsing module 320 includes:

[0100] An identifier acquisition sub-module, configured to acquire the component identifiers of each operation control component on the current graphical user interface based on the interface information;

[0101] A data matching sub-module, configured to identify the keywords in the voice operation request and match the keywords with the component identifiers to identify the component operation instructions associated with the component identifiers;

[0102] A result acquisition sub-module, configured to, if the matching is successful, obtain a voice parsing result based on the component operation instruction and the component identifier;

[0103] A failure handling sub-module, configured to, if the matching fails, re-obtain the voice operation request or prompt the user to perform correct voice input.

[0104] In an optional implementation manner, the instruction generation module 330 includes:

[0105] An action determination sub-module, configured to determine corresponding interface operation actions based on the component operation instructions corresponding to the component identifiers in the voice parsing result;

[0106] An instruction generation sub-module, configured to generate an interface operation instruction according to the interface operation action and the layout information of the current graphical user interface.

[0107] In an optional implementation manner, the action determination sub-module includes:

[0108] A relationship acquisition unit, configured to acquire the order and dependency relationship of the appearance of each component identifier in the voice operation request;

[0109] A priority analysis unit, configured to analyze the execution priorities of the component operation instructions according to the order and dependency relationship;

[0110] An instruction sorting unit, configured to sort the component operation instructions according to the execution priorities;

[0111] An instruction matching unit, configured to match the sorted component operation instructions with the calibrated operation instructions stored in the interface operation action library to obtain an instruction matching result;

[0112] An action determination unit, configured to determine the interface operation actions corresponding to the component operation instructions based on the instruction matching result.

[0113] In some optional implementation manners, the instruction generation sub-module includes:

[0114] A component recognition unit, configured to recognize a target operation control component associated with the interface operation action in the current graphical user interface;

[0115] A position determination unit, configured to determine the position information of the target operation control component in the current graphical user interface according to the layout information;

[0116] An instruction construction unit, configured to construct an interface operation instruction based on the position information and the interface operation action, where the interface operation instruction includes an operation type and operation parameters for the target operation control component to control the current graphical user interface to perform corresponding operations.

[0117] In some alternative embodiments, the instruction generation module 330 further includes:

[0118] a model acquisition sub-module, configured to acquire a target instruction extraction model;

[0119] an instruction extraction sub-module, configured to extract instructions from the speech parsing result based on the target instruction extraction model, so as to generate an interface operation instruction based on the instruction extraction result.

[0120] In some alternative embodiments, the above-mentioned model acquisition sub-module is specifically configured to acquire training sample data and corresponding sample labels, where the sample labels are used to represent the interface operation instructions corresponding to the speech parsing result; input the training sample data and the corresponding sample labels into an initial instruction extraction model to train the initial instruction extraction model to obtain an instruction extraction training model; test the instruction extraction training model based on test data to obtain a test result; if the test result meets a preset condition, use the instruction extraction training model as the target instruction extraction model;

[0121] if the test result does not meet the preset condition, adjust the parameters of the initial instruction extraction model, and re-train the initial instruction extraction model until an instruction extraction training model that meets the preset condition is obtained, and use the instruction extraction training model that meets the condition as the target instruction extraction model.

[0122] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding method embodiments above, and will not be repeated here.

[0123] Through the above device and its components, the technical solution provided by the embodiments of the present invention has the following advantages:

[0124] The voice control operation device according to the embodiments of the present invention obtains the interface information of the current graphical user interface and the voice operation request input by the user based on the current graphical user interface; parses the voice operation request based on the interface information to obtain a voice parsing result; generates an interface operation instruction based on the voice parsing result; based on the interface operation instruction, performs an interface operation on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface, and displays the operation result, which can greatly improve the operation convenience and efficiency of the user on the SaaS platform. The user can complete complex operations only through voice instructions without the need for cumbersome clicks or keyboard inputs, which reduces the operation threshold and improves the user experience. In addition, by automatically identifying and parsing the user's voice operation request, errors caused by misoperation or improper operation are avoided, and the stability and reliability of the system are improved.

[0125] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention, asFigure 4 As shown, the computer device includes: one or more processors 410, a memory 420, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 4 In FIG., a processor 410 is taken as an example.

[0126] The processor 410 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 410 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.

[0127] Among them, the memory 420 stores instructions executable by at least one processor 410, so that at least one processor 410 executes the method shown in the above embodiments.

[0128] The memory 420 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device presented by a kind of mini-program landing page, etc. In addition, the memory 420 can include a high-speed random access memory and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 420 can optionally include a memory remotely set relative to the processor 410, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a server cluster, a mobile communication network, and combinations thereof.

[0129] The memory 420 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 420 can also include a combination of the above types of memories.

[0130] The computer device further includes a communication interface 430 for the computer device to communicate with other devices or communication networks.

[0131] An embodiment of the present invention also provides a computer-readable storage medium storing at least one executable instruction. When the executable instruction runs on a computer device / a device for voice control operations, it causes the computer device / a device for voice control operations to execute the method for voice control operations in any of the above method embodiments.

[0132] An embodiment of the present invention also provides a computer program product including computer instructions for causing a computer to execute the method for voice control operations in the first aspect above or any corresponding implementation manner thereof.

[0133] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. In addition, embodiments of the present invention are not directed to any particular programming language.

[0134] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that embodiments of the present invention may be practiced without these specific details. Similarly, in order to streamline the present invention and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present invention, the various features of the embodiments of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. Among them, the claims following the specific implementation manners are hereby expressly incorporated into the specific implementation manners, where each claim itself serves as a separate embodiment of the present invention.

[0135] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive.

[0136] It should be noted that the above embodiments are illustrative of the present invention rather than restrictive thereof, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A method for voice control operation, characterized in that: Applied to a SaaS platform, where a plurality of graphical user interfaces are configured on the SaaS platform, the method comprises: Acquire interface information of a current graphical user interface and a voice operation request input by a user based on the current graphical user interface; Parsing the voice operation request based on the interface information to obtain a voice analysis result; Generate interface operation instructions based on the voice analysis result; Based on the interface operation instruction, an interface operation is performed on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface, and the operation result is displayed.

2. The method according to claim 1, characterized in that The parsing the voice operation request based on the interface information to obtain a voice parsing result includes: Based on the interface information, obtaining component identifications of each operation control component on the current graphical user interface; identifying a keyword in the voice operation request, and matching the keyword with the component identifier to identify a component operation instruction associated with the component identifier; If the match is successful, the speech analysis result is obtained based on the component operation instruction and the component identifier; If the match fails, the voice operation request is acquired again, or the user is prompted to perform a correct voice input.

3. The method according to claim 1, characterized in that The generating of the interface operation instruction based on the speech analysis result includes: Determine the corresponding interface operation action based on the component operation instruction corresponding to each component identifier in the speech analysis result; The interface operation instruction is generated according to the interface operation action and the layout information of the current graphical user interface.

4. The method according to claim 3, characterized in that The determining the corresponding interface operation action based on the component operation instruction corresponding to each component identifier in the speech analysis result includes: Obtaining the order and dependency relationship of each component identifier in the voice operation request; Analyzing the execution priority of each component operation instruction according to the sequence and dependency relationship; sorting the component operation instructions according to the execution priority; Matching the sorted component operation instructions with the calibration operation instructions stored in the interface operation action library to obtain an instruction matching result; Based on the instruction matching result, the interface operation action corresponding to each of the component operation instructions is determined.

5. The method according to claim 3, characterized in that: The generating the interface operation instruction according to the interface operation action and the layout information of the current graphical user interface includes: Identifying a target operation control component associated with the interface operation action in the current graphical user interface; Determining, according to the layout information, position information of the target operation control component in the current graphical user interface; Based on the position information and the interface operation action, the interface operation instruction is constructed, and the interface operation instruction includes the operation type and operation parameters of the target operation control component to control the current graphical user interface to perform a corresponding operation.

6. The method according to claim 1, characterized in that The generating of the interface operation instruction based on the speech analysis result also includes: Obtain a target instruction extraction model; The speech analysis result is subjected to instruction extraction based on the target instruction extraction model, so as to generate an interface operation instruction based on the instruction extraction result.

7. The method according to claim 5, characterized in that The acquiring target instruction extraction model comprises: Acquire training sample data and corresponding sample labels, where the sample labels are used to represent interface operation instructions corresponding to speech analysis results; Inputting the training sample data and the corresponding sample labels into an initial instruction extraction model to train the initial instruction extraction model to obtain an instruction extraction training model; Testing the instruction extraction training model based on the test data to obtain a test result; If the test result meets the preset conditions, the instruction extraction training model is used as the target instruction extraction model; If the test result does not meet the preset conditions, the parameters of the initial instruction extraction model are adjusted, and the initial instruction extraction model is retrained until an instruction extraction training model that meets the preset conditions is obtained, and the instruction extraction training model that meets the conditions is used as the target instruction extraction model.

8. The method according to claim 1, characterized in that The performing interface operations on the current graphical user interface and / or other graphical user interfaces based on the interface operation instructions and displaying the operation results includes: Obtaining a confirmation result of a command fed back by a user in response to the interface operation command; If the instruction confirmation result is confirmation execution, then the interface operation corresponding to the interface operation instruction is executed, and the status information of the current graphical user interface and / or other graphical user interfaces is acquired in real time, so as to update and display the content of the current graphical user interface and / or other graphical user interfaces based on the status information; If the instruction confirmation result is to cancel the execution, the interface operation instruction will not be executed, and the status of the current graphical user interface and / or other graphical user interfaces will remain unchanged.

9. A device for voice control operation, characterized in that: Applied to a SaaS platform, where a plurality of graphical user interfaces are configured on the SaaS platform, the device comprises: An information acquisition module, used to acquire interface information of a current graphical user interface and a voice operation request input by a user based on the current graphical user interface; A voice analysis module, used to analyze the voice operation request based on the interface information to obtain a voice analysis result; An instruction generation module, used to generate interface operation instructions based on the speech analysis result; The interface operation module is used to perform interface operations on the current graphical user interface and / or other graphical user interfaces associated with the current graphical user interface based on the interface operation instructions, and display the operation results.

10. A computer device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operation of the method for voice control operation as described in any one of claims 1-8.