Method and device for intelligently generating service form through voice, and medium
Through the method of intelligently generating business forms for voice, the inefficiency of manually filling out forms and the difficulty of speech recognition technology in dealing with complex business logic is solved, and automated data entry and verification is realized, reducing error rate and operational complexity.
Patent Information
- Application Number
- CN202510064625.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, it is inefficient and prone to errors in manual filling of forms, and the existing voice recognition technology is difficult to meet the needs of complex business logic processing and generate business forms.
The method of intelligently generating business forms through voice includes obtaining recordings and converting them into audio files, using ASR to convert audio files into text files, and streaming them to the terminal. Extract the feature parameters in the text for verification, and build a business form based on the preset prompt words, function call parameters and historical session records.
It realizes the automatic processing of large amounts of data entry and verification work, reduces the manual labor burden of employees, reduces the error rate caused by human negligence, simplifies the operation process, and improves the operation efficiency and satisfaction of users.
Smart Images

Figure CN119988583A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to a method, device and medium for generating a business form through voice intelligence. Background Art
[0002] In the traditional business processing process, corporate employees often need to manually complete various types of form filling tasks, such as business trip applications, meeting minutes, work daily reports, etc. This method is not only time-consuming and laborious, but also often leads to data entry errors.
[0003] In recent years, with the development of speech recognition technology and natural language processing technology, especially the application of deep learning technology and large-scale language models, it has become feasible to directly create business forms through voice commands. However, most products and services on the market are limited to simple command execution or information retrieval, and fail to fully cope with complex business scenarios and generate business forms.
[0004] Through the above analysis, the problems and defects of the prior art are as follows:
[0005] Manually filling out forms in the prior art is inefficient and error-prone, and existing speech recognition technology is difficult to meet the needs of complex business logic processing and generate business forms. Summary of the invention
[0006] The embodiments of the present application provide a method, device and medium for generating business forms through voice intelligence, which can solve the problem of manually filling out forms in the prior art being inefficient and error-prone, and the problem that the existing voice recognition technology is difficult to meet the needs of complex business logic processing and generate business forms.
[0007] In the first aspect, the embodiment of the present application provides a method for generating a business form through voice intelligence. The method includes: obtaining a recording and converting the recording to obtain an audio file; converting the audio file into a text file through ASR, and sending the text file to the terminal in a streaming manner, wherein the text file includes a first text; extracting a first feature parameter in the first text, and performing parameter verification on the feature parameter; if the verification passes, supplementing the context of the first text according to preset prompt words, function call parameters and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; extracting a second feature parameter of the second text, and constructing a business form according to the second feature parameter.
[0008] In one implementation of the present application, a first feature parameter in the first text is extracted, and a parameter check is performed on the feature parameter, specifically including: matching the form type according to the first feature parameter; judging whether there are missing feature parameters according to the mapping relationship between the form type and the required items, and judging whether there are erroneous logical parameters through a supervised learning algorithm.
[0009] In one implementation of the present application, after extracting the first characteristic parameter in the first text and verifying the characteristic parameter, the method also includes: if the parameter verification fails, sending the missing characteristic parameter or the erroneous logical parameter to the terminal to prompt the user to supplement the missing characteristic parameter or correct the erroneous logical parameter; after the missing characteristic parameter has not been supplemented for more than a preset time, supplementing it with the backup data to obtain the supplemented first text and output it; in the case of a logical conflict, replacing the erroneous logical parameter to obtain the corrected first text, and simultaneously outputting the first text before correction and the first text after correction.
[0010] In one implementation of the present application, the context of the first text in the text file is supplemented according to preset prompt words, function call parameters and historical session records to obtain a second text, specifically including: segmenting the first text to obtain text segmentation, wherein the text segmentation includes a root node, a qualifier, a subject and an object; traversing the text segmentation and prompt words, and calculating the co-occurrence frequency of the text segmentation, historical records and prompt words; removing prompt words with a co-occurrence frequency of 0, converting the text segmentation into a vector, and obtaining a text segmentation vector; calculating the similarity between the text segmentation vector and the prompt word vector, and calculating the semantic similarity and dependency relationship between the remaining prompt words and the text segmentation; extracting the dependency relationship and prompt words above the preset threshold, and supplementing the prompt words to the first text to obtain the second text.
[0011] In one implementation of the present application, after obtaining the second text, the method also includes: mapping the second characteristic parameter in the second text to the corresponding field of the form; executing the form verification logic, storing the verified business form in the database, and generating a unique form identifier.
[0012] In one implementation of the present application, before obtaining the recording and converting the recording to an audio file, the method also includes: obtaining the input text and output form of the historical record, using a machine learning algorithm to extract features from the input text and output form to obtain prompt words, required items, and verification logic; and updating the prompt words, required items, and verification logic based on user feedback and historical records.
[0013] In one implementation of the present application, the method further includes: allocating a unique session identifier to the user based on the user identity authentication information; and applying a form template in the history record according to the user's history record.
[0014] In one implementation of the present application, the method further includes: when the user selects encryption or desensitization processing, encrypting or replacing the corresponding text; in the case of encryption, providing an encryption key for the business form.
[0015] In a second aspect, an embodiment of the present application also provides a device for generating a business form through voice intelligence, the device comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: obtain a recording, and convert the recording to obtain an audio file; convert the audio file into a text file through ASR, and send the text file to a terminal in a streaming manner, wherein the text file comprises a first text; extract a first feature parameter from the first text, and perform parameter verification on the feature parameter; if the verification passes, supplement the context of the first text according to preset prompt words, function call parameters and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; extract a second feature parameter of the second text, and construct a business form according to the second feature parameter.
[0016] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for generating a business form through voice intelligence, storing computer executable instructions, wherein the computer executable instructions are configured to: obtain a recording, and convert the recording to obtain an audio file; convert the audio file into a text file through ASR, and send the text file to a terminal in a streaming manner, wherein the text file includes a first text; extract a first feature parameter from the first text, and perform parameter verification on the feature parameter; if the verification passes, supplement the context of the first text according to preset prompt words, function call parameters, and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; extract a second feature parameter of the second text, and construct a business form according to the second feature parameter.
[0017] The embodiments of the present application provide a method, device and medium for intelligently generating business forms through voice, which introduce intelligent forms and voice recognition technology, and can automatically process the entry and verification of large amounts of data, greatly reducing the manual labor burden of employees and significantly reducing the error rate caused by human negligence; users only need to use voice or simple operations to complete the filling and submission of forms, which greatly simplifies the operation process; the system can provide instant feedback on user input, prompt errors or missing items, and improve user operation efficiency and satisfaction; electronic forms also use data encryption technology to ensure the security of data transmission and storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1A flow chart of a method for generating a business form through voice intelligence provided in an embodiment of the present application;
[0020] Figure 2 An overall logic diagram of generating a business form through voice intelligence provided by an embodiment of the present application;
[0021] Figure 3 A schematic diagram of missing logic parameters for generating a business form through voice intelligence provided in an embodiment of the present application;
[0022] Figure 4 A schematic diagram of a business form for intelligently generating a business form through voice provided in an embodiment of the present application;
[0023] Figure 5 A schematic diagram of the internal structure of a device for intelligently generating a business form through voice provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0025] The embodiments of the present application provide a method, device and medium for generating business forms through voice intelligence, which solves the problem of manually filling out forms in the prior art being inefficient and error-prone, and the problem that the existing voice recognition technology is difficult to meet the needs of complex business logic processing and generate business forms.
[0026] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0027] Figure 1 A flow chart of a method for generating a business form through voice intelligence provided by an embodiment of the present application. Figure 1 As shown, a method for generating a business form through voice intelligence provided by an embodiment of the present application specifically includes the following steps:
[0028] Step 10: Get the recording and convert it into an audio file;
[0029] In this step, if Figure 2As shown, an assistant web page can be developed, using the WEB (World Wide Web) browser recording API (Application Programming Interface), or the assistant web page can be integrated into the mobile or PC software, and the recording interface provided by the software can be accessed to realize the recording function. Users can collect voice commands through the built-in or external microphone of the terminal device (such as a smart phone, tablet computer, etc.).
[0030] Step 20: converting the audio file into a text file through ASR, and sending the text file to the terminal in a streaming manner, wherein the text file includes the first text;
[0031] In this step, the terminal device encodes the collected sound signal and converts it into an audio file format suitable for transmission, such as WAV (Windows Media Audio), MP3, etc., and uploads the audio file to ASR (Automatic Speech Recognition, Automatic Speech Recognition) system, the ASR system will pre-process the uploaded audio files, including noise reduction, echo removal, sound enhancement, etc., to improve the recognition accuracy. In specific implementation, first extract audio features from the audio file, and predict the character or phoneme probability of each time step based on these audio features. During the prediction process, the audio features of each time step can be analyzed by the language model of the deep neural network, and a probability distribution can be output to represent the phoneme or subword that may correspond to the time step and its probability. The predicted phoneme or subword probability distribution needs to be further processed by the decoder. The decoder will consider the constraints of the language model to find the most grammatical text sequence; the language model is a model trained based on a large amount of text data, which can evaluate the linguistic rationality and probability of a text sequence. During the decoding process, the language model will score the candidate text sequences to help the decoder select the optimal recognition result.
[0032] Furthermore, it supports function calls and streaming returns, and instantly displays the converted text on the user interface to improve user experience and interaction efficiency.
[0033] Step 30: extracting a first characteristic parameter from the first text, and performing parameter verification on the characteristic parameter;
[0034] As an optional embodiment, extracting the first characteristic parameter in the first text and performing parameter verification on the characteristic parameter may specifically include: Step 301: matching the form type according to the first characteristic parameter; Step 302: judging whether there are missing characteristic parameters according to the mapping relationship between the form type and the required items, and judging whether there are erroneous logical parameters through a supervised learning algorithm.
[0035] In this step, form types can include business trip applications, meeting minutes, reimbursement details, etc. Because there are a large number of historical records for reference, supervised learning is used in the early stage to judge the user's intention based on the characteristic parameters in the voice; further, the required items are checked, and for types such as strings and numbers, their lengths or values are checked to see if they are within a reasonable range. For example, email addresses, phone numbers, dates and times should follow a specific format; ensure that all required fields are filled in; according to business rules, check whether there are logical conflicts or situations that do not conform to business logic between parameters, such as whether the number of business trip days conflicts with the departure and return dates.
[0036] As an optional embodiment, after extracting the first characteristic parameter in the first text and verifying the characteristic parameter, the method may further include: Step 303: if the parameter verification fails, sending the missing characteristic parameter or the erroneous logical parameter to the terminal to prompt the user to supplement the missing characteristic parameter or correct the erroneous logical parameter; Step 304: after the missing characteristic parameter has not been supplemented for more than a preset time, using the backup data to supplement it, obtaining the supplemented first text and outputting it; Step 305: in the case of a logical conflict, replacing the erroneous logical parameter to obtain the corrected first text, and simultaneously outputting the first text before correction and the first text after correction.
[0037] In this step, when a problem is found, it is first sent to the customer, such as Figure 3 As shown, if the user makes corrections or supplements, the work continues on the corrected content; the above entire interactive process can be repeated, and voice or text can be used for correction again; considering that the user is in a meeting and the recorded content is the speaker's speech, there is no way to correct it at the time, so the backup data can be used for supplementation, or the system can make self-modifications based on historical records, etc., but the supplemented and modified content is shown in another color to remind the user of the problem here.
[0038] Furthermore, you can introduce an A / B testing mechanism to compare the effects of different form design schemes and select the best one; you can also send two schemes to users at the same time.
[0039] Step 40: If the verification is passed, the context of the first text is supplemented according to the preset prompt words, function call parameters and historical session records to obtain a second text, wherein the prompt words include prompt word vectors;
[0040] In this step, the prompt words can be converted into vector mode in advance to improve calculation efficiency.
[0041] As an optional embodiment, according to preset prompt words, function call parameters and historical session records, the context of the first text in the text file is supplemented to obtain the second text, which may specifically include: step 401: segmenting the first text to obtain text segmentation, wherein the text segmentation includes a root node, a qualifier, a subject and an object; step 402: traversing the text segmentation and prompt words, and calculating the co-occurrence frequency of the text segmentation, the historical records and the prompt words; step 403: removing the prompt words with a co-occurrence frequency of 0, converting the text segmentation into a vector, and obtaining a text segmentation vector; step 404: calculating the similarity between the text segmentation vector and the prompt word vector, and calculating the semantic similarity and dependency relationship between the remaining prompt words and the text segmentation;
[0042] In this step, the text segmentation and the remaining prompt words are converted into vector forms respectively. By calculating the similarity between these vectors, such as cosine similarity, the semantic proximity between the text segmentation and the prompt word can be quantified. R = ∑t∈Tmaxd∈D(α·freq(t,d)+β·sim(t,d)) where: freq(t,d) is the frequency of co-occurrence of prompt word t and segmentation d in the text, sim(t,d) is the semantic similarity between prompt word t and segmentation d, which can be obtained by calculating cosine similarity using word vectors, α and β are weight parameters for adjusting the importance of frequency and semantic similarity, which can be determined according to actual conditions during specific implementation.
[0043] Step 405: extract the dependency relationship and prompt words above a preset threshold, and add the prompt words to the first text to obtain a second text.
[0044] As an optional embodiment, after obtaining the second text, the method may also include: Step 406: mapping the second characteristic parameter in the second text to the corresponding field of the form; Step 407: executing the form verification logic, storing the verified business form in the database, and generating a unique form identifier.
[0045] In this step, verify the logic again to see if there are any inconsistencies and generate a unique form identifier for subsequent query and modification.
[0046] Step 50: extract the second characteristic parameter of the second text, and construct a business form according to the second characteristic parameter.
[0047] In this step, if Figure 4As shown, the created form is displayed to the user in the form of a card through the user terminal device, and the user can confirm the form content or provide feedback; the form data after the user's confirmation is stored in the server, and the API interface of the relevant business system is called for processing; the backend service performs subsequent processing on the success and failure of the call respectively; the backend service returns the call status to the user terminal device through preset friendly prompt words and encapsulated structured form data. Through voice recognition and language processing, the user's voice instructions are converted into specific business forms to improve the work efficiency within the enterprise and reduce errors in manual operations.
[0048] As an optional embodiment, before obtaining the recording and converting the recording to an audio file, the method may also include: Step 01: obtaining the input text and output form of the historical record, using a machine learning algorithm to extract features from the input text and the output form to obtain prompt words, required items and verification logic; Step 02: updating the prompt words, required items and verification logic based on user feedback and historical records.
[0049] In this step, assuming that there is a form for collecting personal information of employees, which includes the age and year of employment of the employees, the logical rule may be: current year - hire_year>=18 (ensuring that the employee is at least 18 years old when he / she is employed).
[0050] As an optional embodiment, the method may further include: allocating a unique session identifier to the user based on the user identity authentication information; and applying a form template in the history record according to the user's history record.
[0051] In this step, a unique session identifier is assigned to the user to track and manage all operations and interaction data of the user during the form generation process.
[0052] As an optional embodiment, the method may further include: if the user selects encryption or desensitization processing, encrypting or replacing the corresponding text; in the case of encryption, providing an encryption key for the business form.
[0053] In this step, for example, specific characters or virtual values can be used to replace the real values in the sensitive data, such as replacing the mobile phone number with a unified placeholder; the sensitive data is no longer useful by truncating, encrypting, hiding, etc. the field data value, such as using special characters (*) to replace part of the numbers in the ID number; the order of elements in the sensitive data is disrupted to make it lose its original meaning; the sensitive data is encrypted by encryption keys and algorithms so that only those with decryption keys can restore the original data, while maintaining the logical consistency of the ciphertext and the original data; for numerical data, the mean is first calculated, and then the desensitized values are randomly distributed around the mean to keep the total data unchanged.
[0054] In summary, the embodiments of the present application provide a method, device and medium for generating business forms through voice intelligence, including voice collection, voice-to-text conversion, text analysis and business intent recognition, form creation and parameter verification, form display and user feedback, data persistence and result processing, etc., which improves the efficiency and accuracy of business form creation, reduces the need for manual intervention, and provides a flexible interaction method so that users can interact with the system through natural language.
[0055] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, the embodiment of this application also provides a device for generating a business form through voice intelligence, and its structure is as follows: Figure 5 shown.
[0056] Figure 5 A schematic diagram of the internal structure of a device for generating a business form through voice intelligence provided in an embodiment of the present application. Figure 5 As shown, the device includes:
[0057] at least one processor 501;
[0058] and, a memory 502 communicatively connected to the at least one processor;
[0059] Among them, the memory 502 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 501 so that at least one processor 501 can: obtain a recording and convert the recording to an audio file; convert the audio file into a text file through ASR, and send the text file to the terminal in a streaming manner, wherein the text file includes a first text; extract a first feature parameter in the first text, and perform parameter verification on the feature parameter; if the verification passes, supplement the context of the first text according to preset prompt words, function call parameters and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; extract a second feature parameter of the second text, and construct a business form according to the second feature parameter.
[0060] Some embodiments of the present application provide corresponding Figure 1A non-volatile computer storage medium for generating a business form through voice intelligence stores computer executable instructions, wherein the computer executable instructions are configured to: obtain a recording, and convert the recording to obtain an audio file; convert the audio file into a text file through ASR, and send the text file to a terminal in a streaming manner, wherein the text file includes a first text; extract a first feature parameter from the first text, and perform parameter verification on the feature parameter; if the verification passes, supplement the context of the first text according to preset prompt words, function call parameters and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; extract a second feature parameter of the second text, and construct a business form according to the second feature parameter.
[0061] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the IoT device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0062] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.
[0063] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0064] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0065] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0067] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0068] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0069] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0070] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0071] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for generating a business form through voice intelligence, characterized in that: The method comprises: Obtaining a recording, and converting the recording to obtain an audio file; Converting the audio file into a text file through ASR, and sending the text file to a terminal in a streaming manner, wherein the text file includes a first text; Extracting a first characteristic parameter from the first text, and performing parameter verification on the characteristic parameter; If the verification is passed, the context of the first text is supplemented according to the preset prompt word, function call parameters and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; A second characteristic parameter of the second text is extracted, and a business form is constructed according to the second characteristic parameter.
2. A method for generating a business form through voice intelligence according to claim 1, characterized in that: Extracting a first characteristic parameter from the first text and performing parameter verification on the characteristic parameter specifically includes: Matching the form type according to the first characteristic parameter; According to the mapping relationship between the form type and the required items, it is determined whether there are missing feature parameters, and it is determined whether there are erroneous logical parameters through a supervised learning algorithm.
3. The method for generating a business form through voice intelligence according to claim 2, characterized in that: After extracting the first characteristic parameter from the first text and verifying the characteristic parameter, the method further includes: If the parameter check fails, the missing characteristic parameters or the erroneous logic parameters are sent to the terminal to prompt the user to supplement the missing characteristic parameters or correct the erroneous logic parameters; When the missing feature parameters are not supplemented for a preset time, the backup data is used to supplement them, and the supplemented first text is obtained and output; In case of a logic conflict, the erroneous logic parameter is replaced to obtain a corrected first text, and the first text before correction and the first text after correction are output simultaneously.
4. The method for generating a business form through voice intelligence according to claim 1, characterized in that: According to the preset prompt words, function call parameters and historical session records, the context of the first text in the text file is supplemented to obtain the second text, which specifically includes: Performing word segmentation on the first text to obtain text word segmentation, wherein the text word segmentation includes a root node, a qualifier, a subject, and an object; Traversing the text segmentation words and prompt words, and calculating the co-occurrence frequency of the text segmentation words, historical records and prompt words; Removing the prompt words with a co-occurrence frequency of 0, converting the text segmentation into a vector, and obtaining a text segmentation vector; Calculate the similarity between the text segmentation vector and the prompt word vector, and calculate the semantic similarity and dependency relationship between the remaining prompt words and the text segmentation; The dependency relationship and prompt words above a preset threshold are extracted, and the prompt words are added to the first text to obtain a second text.
5. The method for generating a business form through voice intelligence according to claim 4, characterized in that: After obtaining the second text, the method further includes: Mapping the second characteristic parameter in the second text to the corresponding field of the form; The form validation logic is executed, the business form that passes the validation is stored in the database, and a unique form identifier is generated.
6. The method for generating a business form through voice intelligence according to claim 1, characterized in that: Before obtaining the recording and converting the recording to obtain an audio file, the method further includes: Obtain input text and output form of the historical records, and use a machine learning algorithm to extract features from the input text and output form to obtain prompt words, required items, and verification logic; Update the prompt words, required items and verification logic based on user feedback and historical records.
7. The method for generating a business form through voice intelligence according to claim 1, characterized in that: The method further comprises: Based on the user authentication information, a unique session identifier is assigned to the user; Based on the user's history, apply a form template from the history.
8. The method for generating a business form through voice intelligence according to claim 1, characterized in that: The method further comprises: If the user chooses encryption or desensitization, the corresponding text will be encrypted or replaced; In case of encryption, an encryption key is provided to the business form.
9. A device for generating business forms through voice intelligence, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtaining a recording, and converting the recording to obtain an audio file; Converting the audio file into a text file through ASR, and sending the text file to a terminal in a streaming manner, wherein the text file includes a first text; Extracting a first characteristic parameter from the first text, and performing parameter verification on the characteristic parameter; If the verification is passed, the context of the first text is supplemented according to the preset prompt word, function call parameters and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; A second characteristic parameter of the second text is extracted, and a business form is constructed according to the second characteristic parameter.
10. A non-volatile computer storage medium for generating a business form through voice intelligence, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Obtaining a recording, and converting the recording to obtain an audio file; Converting the audio file into a text file through ASR, and sending the text file to a terminal in a streaming manner, wherein the text file includes a first text; Extracting a first characteristic parameter from the first text, and performing parameter verification on the characteristic parameter; If the verification is passed, the context of the first text is supplemented according to the preset prompt word, function call parameters and historical session records to obtain a second text, wherein the prompt word includes a prompt word vector; A second characteristic parameter of the second text is extracted, and a business form is constructed according to the second characteristic parameter.
Citation Information
Patent Citations
Method and apparatus for generating business document according to speech
CN107274889A
Text matching algorithm serving intelligent question and answer system
CN112988970A
Information input method and device, equipment and storage medium
CN113448963A
Business form filling method and device, electronic equipment and storage medium
CN115509485A