Testament text generation method, electronic equipment and computer readable medium

Through the voice input and recognition technology on the system side, combined with regional information and intention recognition model, the problem of time-consuming and high error rate of will text generation is solved, real-time generation and efficient review of will documents are realized.

CN120354835APending Publication Date: 2025-07-22FUAI (WUHAN) TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202311715570.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing testament text generation method takes a long time, the testator's accent leads to many document errors, and the audio recognition accuracy is low, which affects the efficiency of will document generation.

Method used

Through the system response to user operations, send audio of will problem, perform voice text recognition, display reply audio text, extract key fields and fill in the will text template, and use the voice text recognition model and voice intention recognition model that matches regional information to improve recognition accuracy.

Benefits of technology

Real-time generation, review and correction of will documents is realized, and the efficiency and accuracy of will text generation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354835A_ABST
    Figure CN120354835A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a testament text generation method, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the following steps: in response to detection of a selection operation of a user acting on a voice input button in a testament generation page, for each testament question preset in a system, executing the following processing steps: sending a question audio corresponding to the testament question to a user side of the user so as to enable the user side to play the question audio; responding to the detected reply audio corresponding to the question audio sent by the user side, performing voice text recognition on the reply audio to obtain a reply audio text, and displaying the reply audio text in the testament generation page; extracting each key field in the reply audio text to obtain a key field text; and filling each key field text into a preset testament text template to obtain a testament text. According to the embodiment, the testament document can be formed in real time, auditing and error correction can be carried out in time, and the testament text generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to a will text generation method, an electronic device, and a computer-readable medium. Background Art

[0002] Currently, for the formulation of a will text, the commonly adopted method is as follows: a professional (lawyer) consults the testator face to face and records to form a will document; or records with a voice recorder, and after forming the will text, the testator then conducts confirmation and review.

[0003] However, the above method usually has the following technical problems:

[0004] First, after the will text is formed and then confirmed and reviewed by the testator, when an error occurs, the will text needs to be modified again, which is time-consuming.

[0005] Second, when the testator has an accent, it is difficult for professionals to accurately form a will document, resulting in many errors in the formed will document and affecting the generation efficiency of the will document.

[0006] Third, the intention in the user's audio is not recognized. Although modal particles seem simple, due to the subtle differences in the context and pronunciation, the expressed intentions are different, resulting in a low accuracy rate of audio recognition.

[0007] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention

[0008] This summary of the disclosure is used to introduce concepts in a brief form, and these concepts will be described in detail in the subsequent detailed implementation section. This summary of the disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0009] Some embodiments of the present disclosure propose a will text generation method, an electronic device, and a computer-readable medium to solve one or more of the technical problems mentioned in the above background art section.

[0010] In a first aspect, some embodiments of the present disclosure provide a method for generating a will text, the method comprising: in response to detecting a selection operation by a user on a voice input button in a will generation page, for each will question preset in the system, performing the following processing steps: sending a question audio corresponding to the will question to the user terminal of the user, so that the user terminal plays the question audio; in response to detecting a reply audio corresponding to the question audio sent by the user terminal, performing speech text recognition on the reply audio to obtain a reply audio text, and displaying the reply audio text in the will generation page; in response to receiving a selection operation by the user on a confirmation control corresponding to the reply audio text, extracting each keyword field in the reply audio text to obtain a keyword field text; the system terminal fills each keyword field text into a preset will text template to obtain a will text.

[0011] In a second aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having stored thereon one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method described in any implementation manner of the first aspect.

[0012] In a third aspect, some embodiments of the present disclosure provide a computer-readable medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0013] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the will text generation method of some embodiments of the present disclosure, the generation time of the will document is shortened, and the will document can be generated at any time. Specifically, when an error occurs, it is necessary to modify the will text again. The reason for the long time consumption is that after the will text is formed, it needs to be confirmed and reviewed by the testator. When an error occurs, it is necessary to modify the will text again, which takes a long time. Based on this, in the will text generation method of some embodiments of the present disclosure, the system end responds to the detection of the user's selection operation on the voice input button in the will generation page, and for each will question preset in the system, the following processing steps are executed: First, send the question audio corresponding to the will question to the user's client, so that the client plays the question audio. Thus, the information of the testator can be automatically queried. Second, in response to the detection of the reply audio corresponding to the question audio sent by the client, perform speech text recognition on the reply audio to obtain the reply audio text, and display the reply audio text on the will generation page. Thus, the audio text of the testator can be recognized in time, so that the testator can review and correct the text in real time. Then, in response to receiving the user's selection operation on the confirmation control corresponding to the reply audio text, extract each keyword field in the reply audio text to obtain the keyword field text. Thus, after the testator determines the text, the keyword fields can be extracted, which is convenient for filling into the will text template. Finally, the system end fills each keyword field text into the preset will text template to obtain the will text. Thus, the will document can be formed in real time, and the review and correction can be carried out in time, improving the generation efficiency of the will text. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals indicate the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0015] Figure 1 is a flowchart of some embodiments of the will text generation method according to the present disclosure;

[0016] Figure 2 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0018] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0019] It should be noted that concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.

[0023] Figure 1 is a flowchart of some embodiments of a will text generation method according to the present disclosure. The flow 100 of some embodiments of the will text generation method according to the present disclosure is shown. The will text generation method includes the following steps:

[0024] Step 101, in response to detecting a selection operation by the user on the voice input button in the will generation page, for each will question preset in the system, perform the following processing steps:

[0025] Step 1011, send a question audio corresponding to the above will question to the user terminal of the above user, so that the above user terminal plays the above question audio.

[0026] In some embodiments, the above-mentioned system terminal can send a question audio corresponding to the above-mentioned will question to the user terminal of the above-mentioned user, so that the above-mentioned user terminal plays the above-mentioned question audio. The system terminal can refer to a terminal that has set various will question audios and automatically generates will texts. The will question audio can refer to an audio corresponding to a certain will question. For example, the will question audio can be the audio of "May I ask who are your heirs?". The will generation page can refer to a page constructed by the system terminal for generating will texts. The will generation page is displayed on the user terminal for the user to perform selection operations. The voice input button can refer to a control set on the will generation page. After the user clicks the voice input button, a voice interaction with the system terminal can be requested. The selection operation can be a click operation.

[0027] In practice, the user can click the voice input button on the will generation page of the user terminal. After that, after the system terminal detects the selection operation of the user on the voice input button on the will generation page, it can send a question audio corresponding to the above-mentioned will question to the user terminal of the above-mentioned user, so that the above-mentioned user terminal plays the above-mentioned question audio.

[0028] Step 1012, in response to detecting the reply audio corresponding to the above-mentioned question audio sent by the above-mentioned user terminal, perform voice text recognition on the above-mentioned reply audio to obtain a reply audio text, and display the above-mentioned reply audio text on the above-mentioned will generation page.

[0029] In some embodiments, the above-mentioned system terminal can, in response to detecting the reply audio corresponding to the above-mentioned question audio sent by the above-mentioned user terminal, perform voice text recognition on the above-mentioned reply audio to obtain a reply audio text, and display the above-mentioned reply audio text on the above-mentioned will generation page. The reply audio can refer to the user's reply voice to the question audio. For example, the Whisper model of OpenAI can be used to perform voice text recognition on the above-mentioned reply audio to obtain a reply audio text. After that, the reply audio text can be sent to the above-mentioned user terminal for display on the above-mentioned will generation page.

[0030] In practice, the above-mentioned system terminal can perform voice text recognition on the above-mentioned reply audio through the following steps:

[0031] The first step is to obtain the user information of the above-mentioned user, where the above-mentioned user information includes geographical information. The geographical information can represent the user's birth address or the current user's residence address. The geographical information can be set according to the address indicated by the user's actual accent.

[0032] Second step, select the speech text recognition model corresponding to the above regional information from the pre-trained speech text recognition model set as the target speech text recognition model. Each speech text recognition model corresponds to a regional information. The speech text recognition model can be a speech text recognition model pre-trained according to the accent corresponding to a certain regional information. For example, the speech text recognition model can be an RNN-T (Recurrent Neural Network Transducer) model. That is, the speech text recognition model indicating the above regional information can be selected from the pre-trained speech text recognition model set as the target speech text recognition model.

[0033] Third step, perform frame division on the above reply audio to generate a set of reply audio frames. The above reply audio can be frame-divided with a preset frame length to obtain a set of reply audio frames. The frame length of each reply audio frame in the set of reply audio frames except the last reply audio frame is equal to the above preset frame length. The frame length of the above last reply audio frame is less than or equal to the above preset frame length.

[0034] Fourth step, perform division processing on the above set of reply audio frames to generate a set of reply audio frame sequences. The above set of reply audio frames can be divided according to the audio frames corresponding to one character to generate a set of reply audio frame sequences. That is, the reply audio frame sequence can represent one character.

[0035] Step 5: Input the above reply audio frame sequence set into the above target speech text recognition model to determine the character node path graph corresponding to the above reply audio frame sequence set. Among them, the above character node path graph includes: each character node corresponding to the reply audio frame sequence in the above reply audio frame sequence set. There is a corresponding node label for each reply audio frame sequence. Each pair of character nodes is connected by a connection line. There is a corresponding grammar score between each pair of character nodes. There is a corresponding node score for each character node. The various reply audio frame sequences in the above reply audio frame sequence set have a sequence order. Here, the target speech text recognition model can refer to a model used to recognize each character that the reply audio frame sequence may correspond to, score each recognized character (grammar score), and score the connection relationship between every two characters (with a sequence order) (node score). The above target speech text recognition model is also used to output each word that the audio frame sequence set may correspond to. The higher the score, the higher the possibility that the audio frame sequence corresponds to the character with this score. For example, the target speech text recognition model can refer to a pre-trained convolutional neural network model or a recurrent neural network model. For example, there can be multiple recognized characters for a certain audio frame. For example, the characters "Li" and "Li". In practice, each recognized character and its score by the above target speech text recognition model can be sorted according to the order of each reply audio frame sequence. A character node is set between every two characters. Every two character nodes are connected by a connection line. Thus, a character node path graph is constructed. The connection line can represent the specific character and the pronunciation of the character. Blank nodes can be added to both the starting position and the ending position of each character node path in the character node path graph.

[0036] Step 6: Generate a target character node path graph according to the respective grammar scores and node scores corresponding to the above character node path graph.

[0037] Among them, Step 6 can include the following sub-steps:

[0038] The first sub-step: For each character node path in the above character node path graph, determine the sum of the respective grammar scores and node scores corresponding to the above character node path as the character node path score. In practice, first, determine each character node corresponding to the above character node path as a character node group. Secondly, determine the sum of the respective node scores corresponding to the above character node group as the total node score. Then, determine each character corresponding to the above character node path as a character group. Then, the sum of the respective grammar scores corresponding to the above character group can be determined as the total grammar score. Finally, the sum of the above total node score and the above total grammar score can be determined as the character node path score.

[0039] The second sub-step is to sort the scores of the determined paths of each text node in descending order to obtain a sequence of scores of text node paths.

[0040] The third sub-step is to sequentially select a preset number of scores of text node paths from the above sequence of scores of text node paths as an alternative sequence of scores of text node paths. Here, there is no limitation on the setting of the preset number. For example, the preset number can be 4.

[0041] The fourth sub-step is to determine the paths of each text node corresponding to the above alternative sequence of scores of text node paths as an alternative group of text node paths.

[0042] The fifth sub-step is to generate a target text node path graph according to the above alternative group of text node paths. In practice, each phoneme node path corresponding to the above alternative group of phoneme node paths in the above phoneme node path graph can be retained. That is, the phoneme node paths other than each phoneme node path corresponding to the above alternative group of phoneme node paths in the above phoneme node path graph are removed. Thus, an alternative phoneme node path graph is obtained.

[0043] The seventh step is to determine a target text node path according to the path selection information corresponding to the above target text node path graph submitted by the user. The above path selection information includes: a sequence of node labels. The sequence of node labels can represent a sequence composed of each node label selected by the user.

[0044] It should be noted that the target text node path graph can be displayed on the will generation page of the user side, and the user can select an alternative text node path as the path selection information. After the user's selection is completed, it can be sent to the system side.

[0045] Among them, the above seventh step may include the following sub-steps:

[0046] The first sub-step is to perform the following processing steps for each reply audio frame sequence in the above set of reply audio frame sequences:

[0047] 1. Determine the paths of each text node corresponding to the above reply audio frame sequence as an alternative group of text node paths.

[0048] 2. Determine the alternative text node corresponding to the target node label in the above alternative group of text node paths as the target text node. Among them, the above target node label is: the serial number corresponding to the node label in the above sequence of node labels and the node label with the same serial number as the above reply audio frame sequence in the above set of reply audio frame sequences, and the node label corresponding to the above target text node is the same as the above target node label.

[0049] The second sub-step is to determine the target text node paths based on the determined target text nodes. The alternative text node paths corresponding to the respective target text nodes in the above-mentioned target text node path diagram can be determined as the target text node paths.

[0050] The eighth step is to generate a response audio text corresponding to the above-mentioned response audio based on the above-mentioned target text node paths. The respective text combinations corresponding to the above-mentioned target text node paths can be combined as the response audio text of the above-mentioned response audio.

[0051] The above relevant content is an inventive point of the present disclosure, thereby solving the second technical problem mentioned in the background art, "resulting in a relatively large number of errors in the formed will document, affecting the generation efficiency of the will document." The factors that result in a relatively large number of errors in the formed will document and affect the generation efficiency of the will document are often as follows: when the testator has an accent, it is difficult for professionals to accurately form a will document, resulting in a relatively large number of errors in the formed will document and affecting the generation efficiency of the will document. If the above factors are solved, the effect of improving the generation efficiency of the will document can be achieved. To achieve this effect, first, the above-mentioned response audio frame set is divided to generate a response audio frame sequence set. Thereby, it is convenient to divide the respective audio frames corresponding to the pronunciation of one character together, facilitating subsequent recognition of the character corresponding to the audio. Secondly, the above-mentioned response audio frame sequence set is input into the above-mentioned target speech text recognition model to determine the text node path diagram corresponding to the above-mentioned response audio frame sequence set. Among them, the above-mentioned text node path diagram includes: the respective text nodes corresponding to the response audio frame sequences in the above-mentioned response audio frame sequence set, each response audio frame sequence has a corresponding node label, the connection lines are used to connect between every two text nodes, there is a corresponding grammar score between every two text nodes, each text node has a corresponding node score, and there is a sequence order among the respective response audio frame sequences in the above-mentioned response audio frame sequence set. Thereby, the recognition path of the audio can be gradually determined to prevent too many paths from causing inaccurate path selection. Then, a target text node path diagram is generated based on the respective grammar scores and node scores corresponding to the above-mentioned text node path diagram. Thereby, the text node path diagram can be initially simplified. Then, based on the path selection information corresponding to the above-mentioned target text node path diagram submitted by the user, the target text node path is determined. Thereby, a relatively accurate text node path can be selected according to the user's selection of the audio recognition result. Finally, a response audio text corresponding to the above-mentioned response audio is generated based on the above-mentioned target text node path. Thereby, the accuracy of audio recognition is improved. Thus, the errors in the formed will document are reduced, and the generation efficiency of the will document is improved.

[0052] Optionally, before selecting the speech text recognition model corresponding to the above regional information from the pre-trained speech text recognition model set as the target speech text recognition model, the above method further includes:

[0053] First step, for each piece of regional information, perform the following processing steps:

[0054] First processing step, obtain the audio training sample data corresponding to the above regional information. The audio training sample data may include the sample audio corresponding to the above regional information.

[0055] Second processing step, input the sample audio in the above audio training sample data into the speech recognition sub-network in the initial speech text recognition model to obtain the audio recognition result corresponding to the above sample audio. The speech recognition sub-network can be used to recognize the input sample audio to output the audio recognition result. The speech recognition sub-network can be a model using automatic speech recognition technology (ASR, Automatic Speech Recognition).

[0056] Third processing step, input the intermediate layer features output by the target network layer in the above speech recognition sub-network into the speech intent recognition sub-network in the above initial speech text recognition model to obtain the speech intent. Among them, the above target network layer is the network layer other than the input layer and the output layer. Which specific layer the target network layer is can be set according to the actual situation. As an example, the output of the N / 2 layer can be used as the intermediate layer features. N is the total number of layers of the speech recognition sub-network. It should be noted that using the N / 2 layer (using the ceiling calculation method) can balance efficiency and performance. As an example, the speech intent recognition sub-network can adopt the model structure of natural language processing (NLP, Natural Language Processing). The intermediate layer features are used as the input data of the speech intent recognition sub-network. The intermediate layer features not only contain the text information expressed by the audio, but also contain information such as intonation or semantics.

[0057] Fourth processing step, based on the above audio recognition result, the above speech intent, the text label and the speech intent label corresponding to the above sample audio in the above audio training sample data, train the above initial speech text recognition model to obtain the trained speech text recognition model.

[0058] Among them, the fourth processing step may include the following sub-steps:

[0059] First sub-step, according to the preset cross-entropy loss function, determine whether the sample loss value is less than or equal to the preset threshold.

[0060] Second sub-step, in response to determining that the above sample loss value is greater than the threshold, use the adaptive gradient optimization algorithm to adjust the network parameters of the initial speech text recognition model, and re-obtain the audio training sample data, and train the initial speech text recognition model again. Among them, the adaptive gradient optimization algorithm can be the Adaptive Moment Estimation (ADAM) optimization algorithm. In the update process of adjusting the network parameters of the initial speech text recognition model, the ADAM optimization algorithm can be used for gradient backpropagation.

[0061] The above related content, as an inventive point of the present disclosure, solves the third technical problem mentioned in the background art, "resulting in a low accuracy rate of audio recognition." The factors that cause a low accuracy rate of audio recognition are often as follows: the intention in the user's audio is not recognized. Although filler words seem simple, due to the subtle differences in the context and pronunciation, the expressed intentions are different. If the above factors are solved, the effect of improving the accuracy rate of audio recognition can be achieved. To achieve this effect, first, obtain the audio training sample data corresponding to the above regional information. Then, input the sample audio in the above audio training sample data into the speech recognition sub-network in the initial speech text recognition model to obtain the audio recognition result corresponding to the above sample audio. Then, input the intermediate layer features output by the target network layer in the above speech recognition sub-network into the speech intention recognition sub-network in the above initial speech text recognition model to obtain the speech intention. Among them, the above target network layer is the network layer other than the input layer and the output layer. Thus, when the speech recognition sub-network recognizes and analyzes the sample audio, the intermediate layer features output by its target network layer are used as input data and thus input into the speech intention recognition sub-network to recognize the intention expressed by the tone in the sample audio. Finally, based on the above audio recognition result, the above speech intention, the text label and the speech intention label corresponding to the above sample audio in the above audio training sample data, train the above initial speech text recognition model to obtain the trained speech text recognition model. Thus, by using the text information, intonation or semantic information, etc. contained in the intermediate layer features, the accuracy of intention prediction can be improved. Thus, the accuracy rate of audio recognition is improved.

[0062] Step 1013, in response to receiving the above selection operation of the user on the confirmation control corresponding to the above reply audio text, extract each keyword field in the above reply audio text to obtain the keyword field text.

[0063] In some embodiments, the system side may, in response to receiving a selection operation of the user on the confirmation control corresponding to the reply audio text, extract each keyword field in the reply audio text to obtain keyword field text. The confirmation control may be a control indicating that the reply audio text is correct. That is, the user can click the confirmation control corresponding to the reply audio text on the will generation page of the user side. Thus, the system side can extract each keyword field in the reply audio text according to a preset keyword field template to obtain keyword field text. For example, the keyword field template may include keyword fields such as user name, property type, property name, heir name, etc. That is, each keyword field in the reply audio text that matches the keyword fields in the keyword field template can be extracted to obtain keyword field text.

[0064] Step 102, the system side fills each keyword field text into a preset will text template to obtain a will text.

[0065] In some embodiments, the system side may fill each keyword field text into a preset will text template to obtain a will text. The will text template may be a preset template for filling in each keyword field related to the will. Such as, user name, property type, property name, heir name, etc.

[0066] Optionally, in response to detecting a selection operation of the user on the button to enter the chat room on the will generation page, the user side jumps to the chat room page.

[0067] In some embodiments, the user side may, in response to detecting a selection operation of the user on the button to enter the chat room on the will generation page, jump to the chat room page.

[0068] In an actual application scenario, the user can complete the operation of placing an order and making a payment through the platform; afterwards, an order will pop up on the user side, and the user can click the "enter chat room button" on the order page (will generation page) to enter the chat room page (voice chat room).

[0069] Optionally, the system side collects the user chat audio sequence and the target user chat audio sequence of the user side and the target user side on the chat room page.

[0070] In some embodiments, the system side collects the user chat audio sequence and the target user chat audio sequence of the user side and the target user side on the chat room page. The target user side may refer to the user side operated by a lawyer. The user chat audio may be the reply audio of the user. The target user chat audio may be the question audio proposed by the target user.

[0071] It should be noted that after the user enters the chat room page, the system will send a notification to the target user terminal to enable the lawyer to enter the chat room page. Thus, the user terminal can have an audio chat with the target user terminal on the above-mentioned chat room page. For example, the lawyer can send a question audio to the user on the chat room page, and then the user can reply with a response audio.

[0072] Optionally, the above system terminal performs speech text recognition on each user chat audio in the above user chat audio sequence to generate user chat audio text, obtaining a user chat audio text sequence.

[0073] In some embodiments, the above system terminal can perform speech text recognition on each user chat audio in the above user chat audio sequence to generate user chat audio text, obtaining a user chat audio text sequence. Here, the implementation manner of speech text recognition can refer to the implementation manner in step 1032, which will not be elaborated here.

[0074] Optionally, the above system terminal performs speech text recognition on each target user chat audio in the above target user chat audio sequence to generate target user chat audio text, obtaining a target user chat audio text sequence.

[0075] In some embodiments, the above system terminal can perform speech text recognition on each user chat audio in the above user chat audio sequence to generate user chat audio text, obtaining a user chat audio text sequence. Here, the implementation manner of speech text recognition can refer to the implementation manner in step 1032, which will not be elaborated here.

[0076] Optionally, for each target user chat audio text in the target user chat audio text sequence, the above system terminal performs the following processing steps:

[0077] First step, determine whether there is a target user chat audio text corresponding to the above target user chat audio text in the target user chat audio text sequence. That is, determine whether there is a target user chat audio text in the above target user chat audio text sequence whose expressed question is the same as the question expressed by the above target user chat audio text. For example, the question expressed by the above target user chat audio text is "Whether property one is inherited by the daughter". Another target user chat audio text expresses the question "Whether property one is inherited by the son", and these two questions belong to the same question.

[0078] In the second step, in response to determining whether there is a target user chat audio text corresponding to the above-mentioned target user chat audio text in the target user chat audio text sequence, the above-mentioned target user chat audio text and each corresponding target user chat audio text are constructed into a single question text. That is, the above-mentioned target user chat audio text and each corresponding target user chat audio text can be merged into a single question text.

[0079] In the third step, at least one user chat audio text corresponding to the above-mentioned single question text is selected from the above-mentioned user chat audio text sequence. That is, at least one user chat audio text is a text that replies to the question corresponding to the single question text.

[0080] In the fourth step, a will question reply text is constructed based on the above-mentioned single question text and the above-mentioned at least one user chat audio text. That is, the above-mentioned single question text and the above-mentioned at least one user chat audio text can be merged into a will question reply text.

[0081] Optionally, the above-mentioned system end generates a will text according to each will question reply text.

[0082] In some embodiments, the above-mentioned system end can generate a will text according to each will question reply text. That is, each keyword field in each will question reply text can be filled into the corresponding position in a preset will text template to obtain a will text.

[0083] Next, refer to Figure 2 , which shows a schematic structural diagram of an electronic device (for example, a computing device) 200 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 2 The electronic device shown is only an example and should not impose any limitation on the functions and usage scopes of the embodiments of the present disclosure.

[0084] As Figure 2As shown, the electronic device 200 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 201, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 202 or a program loaded from a storage device 208 into a random access memory (RAM) 203. In the RAM 203, various programs and task data required for the operation of the electronic device 200 are also stored. The processing device 201, the ROM 202, and the RAM 203 are connected to each other through a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.

[0085] Generally, the following devices may be connected to the I / O interface 205: an input device 206 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 207 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 208 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 209. The communication device 209 may allow the electronic device 200 to communicate with other devices wirelessly or wirelessly to exchange task data. Although Figure 2 an electronic device 200 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 2 Each block shown in may represent a device or, as needed, multiple devices.

[0086] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 209, or installed from the storage device 208, or installed from the ROM 202. When the computer program is executed by the processing device 201, the above functions defined in the methods of some embodiments of the present disclosure are executed.

[0087] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a task data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated task data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0088] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital task data in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0089] The above computer-readable medium may be included in the above electronic device; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: when the system side detects a selection operation by the user on the voice input button in the will generation page, for each will question preset in the system, perform the following processing steps: send a question audio corresponding to the will question to the user side of the user, so that the user side plays the question audio; in response to detecting the reply audio corresponding to the question audio sent by the user side, perform speech text recognition on the reply audio to obtain a reply audio text, and display the reply audio text on the will generation page; in response to receiving the selection operation by the user on the confirmation control corresponding to the reply audio text, extract each keyword field in the reply audio text to obtain a keyword field text; the system side fills each keyword field text into a preset will text template to obtain a will text.

[0090] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include product-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0092] The functions described above can be performed, at least in part, by one or more hardware logic components. By way of example, and not limitation, the types of hardware logic components that may be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0093] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for generating a will text, comprising: Upon detecting a selection operation by the user on the voice input button in the will generation page, the system end performs the following processing steps for each will question preset in the system: Sending a question audio corresponding to the will question to the user's client terminal so that the client terminal plays the question audio; In response to detecting the reply audio corresponding to the question audio sent by the client terminal, performing speech text recognition on the reply audio to obtain a reply audio text, and displaying the reply audio text in the will generation page; In response to receiving the selection operation by the user on the confirmation control corresponding to the reply audio text, extracting each keyword field in the reply audio text to obtain a keyword field text; The system end fills each keyword field text into a preset will text template to obtain a will text.

2. The method according to claim 1, wherein, The method further includes: The client terminal jumps to the chat room page in response to detecting the selection operation by the user on the enter chat room button in the will generation page; The system end collects the user chat audio sequence and the target user chat audio sequence of the user's client terminal and the target user's client terminal in the chat room page; The system end performs speech text recognition on each user chat audio in the user chat audio sequence to generate a user chat audio text, obtaining a user chat audio text sequence; The system end performs speech text recognition on each target user chat audio in the target user chat audio sequence to generate a target user chat audio text, obtaining a target user chat audio text sequence; For each target user chat audio text in the target user chat audio text sequence, the system end performs the following processing steps: Determining whether there is a target user chat audio text corresponding to the target user chat audio text in the target user chat audio text sequence; In response to determining whether there is a target user chat audio text corresponding to the target user chat audio text in the target user chat audio text sequence, constructing the target user chat audio text and the corresponding target user chat audio texts into a single question text; Selecting at least one user chat audio text corresponding to the single question text from the user chat audio text sequence; Constructing a will question reply text according to the single question text and the at least one user chat audio text; The system end generates a will text according to each will question reply text.

3. The method according to claim 1, wherein The performing speech text recognition on the reply audio to obtain a reply audio text includes: Obtaining the user information of the user, where the user information includes geographical information; Selecting a speech text recognition model corresponding to the geographical information from a pre-trained set of speech text recognition models as the target speech text recognition model, where each speech text recognition model corresponds to a geographical information; Performing frame splitting on the reply audio to generate a set of reply audio frames; Performing partitioning processing on the set of reply audio frames to generate a set of reply audio frame sequences; Input the set of reply audio frame sequences into the target speech text recognition model to determine the text node path graph corresponding to the set of reply audio frame sequences. Among them, the text node path graph includes: each text node corresponding to the reply audio frame sequence in the set of reply audio frame sequences, each reply audio frame sequence has a corresponding node label, two text nodes are connected by a connection line, there is a corresponding syntax score between every two text nodes, each text node has a corresponding node score, and there is a sequential order among the reply audio frame sequences in the set of reply audio frame sequences; Generate a target text node path graph according to the respective syntax scores and node scores corresponding to the text node path graph; Determine a target text node path according to the path selection information corresponding to the target text node path graph submitted by the user; Generate a reply audio text corresponding to the reply audio according to the target text node path.

4. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-3.

5. A computer-readable medium having a computer program stored thereon, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Voice recognition method and system

    CN109545218A

  • Voice recognition method and device, electronic equipment and storage medium

    CN110534109A

  • Conference summary generation method and device, computer equipment and storage medium

    CN111986677A

  • Insurance public estimation remote access summary method, device and equipment

    CN113283995A

  • Alarm audio recognition method and device, electronic equipment and computer medium

    CN116580701A