Intelligent system for establishing old-age self-written will based on multi-round voice recognition interaction
By constructing an intelligent will-making system for elderly people with multi-round voice recognition interaction, the problems of insufficient voice recognition accuracy and non-standard will content in elderly people's handwritten wills have been solved. The system realizes automatic recognition and dynamic completion of wills, ensuring the legality and validity of wills.
Patent Information
- Application Number
- CN202610706414.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-05-21
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, the accuracy of speech recognition for elderly people in handwritten will systems is insufficient, and they cannot adapt to the characteristics of elderly people such as mixed dialects, unclear speech, and slow speech speed. This results in the content of the will being non-standard and the logic being unclear. In addition, the existing systems lack multi-round interaction and dynamic completion mechanisms, which leads to a high risk of the will being invalid.
A smart system for drafting handwritten wills for the elderly based on multi-round speech recognition interaction is constructed, including modules for historical will data processing, elderly speech recognition and element extraction, will sentence reconstruction and element matching, confidence calculation and question generation, and multi-round interaction and dynamic updating. Through speech interaction, will sentence reconstruction and confidence calculation, the system ensures the completeness and logic of the will content.
It achieves voice adaptation for the elderly, automatically identifies the core elements of a will, lowers the threshold for making a will, ensures the completeness and logical clarity of the will content, effectively avoids the risk of invalid wills, and protects the property rights of the elderly and family harmony.
Smart Images

Figure CN122417019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice recognition interaction technology, and more specifically, to an intelligent system for drafting handwritten wills for the elderly based on multi-round voice recognition interaction. Background Technology
[0002] As my country's population continues to age, the legal needs of the elderly for wealth transfer and protection of personal rights are surging. As an important legal tool for the elderly to independently arrange their estate and maintain family harmony, the demand for handwritten wills is experiencing explosive growth.
[0003] Currently, manual services suffer from high costs, long appointment cycles, and significant time and space limitations. Furthermore, the elderly generally lack professional legal knowledge, making their wills prone to invalidity due to issues such as non-standard forms and unclear content. Judicial practice data shows that over 60% of handwritten wills have varying degrees of legal flaws. Simultaneously, existing digital systems using text-based interaction are extremely unfriendly to the elderly with declining vision and weak digital skills. General speech recognition technology has an accuracy rate of less than 70% in recognizing common speech characteristics of the elderly, such as mixed dialects, unclear pronunciation, slow speech, and frequent pauses, failing to accurately extract core information from the will. In addition, existing systems often use a one-time input model, which cannot adapt to the fragmented, disordered, and logically disjointed expressions of the elderly, easily leading to omissions and matching errors. Existing improvements are mostly limited to simply optimizing the basic accuracy of the speech recognition model, without combining the legal logic of will making to construct targeted semantic processing and dynamic interactive completion mechanisms, failing to effectively solve core problems such as unclear expression of intent and contradictory content. To mitigate these issues, a smart will-making system for the elderly based on multi-round speech recognition interaction is proposed. Summary of the Invention
[0004] The purpose of this invention is to provide an intelligent system for drafting handwritten wills by the elderly based on multi-round voice recognition interaction, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, an intelligent system for drafting handwritten wills for the elderly based on multi-round speech recognition interaction is provided, including a historical will data processing module, an elderly speech recognition and element extraction module, a will sentence reconstruction and element matching module, a confidence calculation and question generation module, and a multi-round interaction and dynamic update module. The historical will data processing module is used to collect historical will data, extract will-making frameworks from the historical will data, and divide the extracted will-making frameworks into will units and keywords corresponding to each will unit. The elderly speech recognition and element extraction module is used to record the speech data of elderly users, train an automatic speech recognition model to divide the speech data into will statements, and identify the recipient, received items and receiving details for each will statement. The module also matches the identified received items with the keywords of each will unit. The will statement restructuring and element matching module is used to restructure the matched will statements in each will unit according to the recording time and the recipient's position in the statement, and match exclusive receiving items and receiving details for each recipient according to the statement restructuring result; The confidence calculation and question generation module is used to calculate the distance between related sentences in the voice data, taking the receiver's sentence position as the center and combining the received items and receiving details. Based on the distance between related sentences, a confidence level is set for the corresponding received items and receiving details. At the same time, a confidence threshold is set, and questions are summarized for received items and receiving details that are below the confidence threshold to generate questions to be determined for each receiver. The multi-round interaction and dynamic update module is used to have the question to be determined interact with the elderly user through voice through an intelligent robot, and to update the original voice data corresponding to the question to be determined based on the latest voice data obtained from the interaction. At the same time, the confidence threshold is dynamically updated in sync with the updated voice data until there are no more questions to be determined, and then the handwritten will is completed.
[0006] As a further improvement to this technical solution, the historical will data processing module establishes a will database inside the intelligent robot, storing all handwritten wills that have been completed in the past in the database. At the same time, the intelligent robot collects legally effective handwritten will texts and notarized will texts from the network through the network module, and then summarizes them as historical will data. Furthermore, all historical will data has been anonymized, retaining only the structure and logical information of the wills, and containing no personal privacy information. All historical wills data are segmented into words, tagged with parts of speech and semantic roles, and the general logical structure of will making is extracted as a standard will making framework. The standard will-making framework is broken down into several independent but logically coherent sub-units. For each sub-unit, frequently occurring legal terms and everyday expressions commonly used by elderly users are extracted as corresponding keywords.
[0007] As a further improvement to this technical solution, the elderly voice recognition and element extraction module uses an intelligent robot to enable voice recording to record the voice data of elderly users. Synchronously record the timestamp of each segment of audio data; An automatic speech recognition model was established by collecting speech samples from elderly users with different dialects, accents, speaking speeds, and physical conditions. An elderly-specific speech dataset was constructed, and the automatic speech recognition model was trained based on the elderly-specific speech dataset to enable the automatic speech recognition model to have real-time speech recognition capabilities. Then, the trained automatic speech recognition model is used to segment the speech data into will statements.
[0008] As a further improvement to this technical solution, in the elderly speech recognition and element extraction module, the continuous speech stream is segmented based on the pause duration, semantic integrity and will unit keywords of the speech data to obtain multiple independent and semantically complete will statements, while meaningless interjections and repetitive statements are removed. Named entity recognition technology is used to extract names and kinship terms from the will as recipients, real estate, movable property, digital assets and property rights as received items, and descriptions of property shares, distribution conditions, delivery methods and custody methods as receiving details. The semantic similarity between the identified received items and the keywords of each will unit is calculated, and the received items are matched to the will unit with the highest semantic similarity.
[0009] As a further improvement to this technical solution, in the will statement reorganization and element matching module, within the same will unit, the will statements are initially sorted according to the chronological order of their recording time, and then clustered with the recipient as the core, all will statements that mention the same recipient are arranged together, and the statement that first fully mentions the recipient's name is taken as the starting sentence of the content corresponding to that recipient, and subsequent statements are arranged sequentially after the starting sentence. Perform semantic association analysis on the recombined statements, and match all received items and details mentioned in all statements after the starting sentence of the same recipient and before the starting sentence of the next recipient to the name of that recipient; For statements with contradictory content, both contradictory items are marked as high-risk items to be confirmed and directly pushed to the confidence calculation and question generation module.
[0010] As a further improvement to this technical solution, in the confidence calculation and question generation module, the time position of the starting statement containing the recipient's name is taken as the origin, and the time interval and semantic interval between all statements mentioning the recipient's corresponding received items or receiving details and the core starting statement are calculated. The weighted sum of the time interval and semantic interval is used as the final distance between related statements. The weighting coefficients for the time interval and the semantic interval total 1. The distance between related statements shows a linear negative correlation with confidence levels. The closer the related statements are, the higher the confidence level of the corresponding received items or received details; The greater the distance between the related statements, the lower the confidence level of the corresponding received item or received details.
[0011] As a further improvement to this technical solution, a confidence threshold is set in the confidence calculation and question generation module; Compare the confidence levels corresponding to the received items and receipt details with the confidence thresholds; If the confidence level is higher than the confidence threshold, it will not be set as a question statement; Conversely, if the confidence level is lower than the confidence threshold, it is set as a problem statement; All question statements are categorized and summarized according to the recipient, corresponding questions to be determined are generated for each question statement, and the questions to be determined are transformed into conversational question statements; Each question contains only one item to be confirmed.
[0012] As a further improvement to this technical solution, in the multi-round interaction and dynamic update module, the intelligent robot plays the spoken question of the question to be determined to the elderly user based on the voice broadcast function; Meanwhile, after each question to be determined is played, a fixed pause is taken as the voice interaction data collection phase, and the voice data collected from elderly users in this phase is used as the answer voice for that question to be determined. Based on the audio response, make a normal judgment on the question to be answered; If the current problem to be determined is correctly identified, the confidence threshold is adjusted based on the confidence level. The higher the confidence level, the smaller the adjustment to the confidence threshold; If the current question is judged incorrectly, the answer will be replaced with the corresponding position in the original voice data, and the corresponding will module will be updated. When the confidence level of all received items and details is higher than the current confidence level threshold, and the elderly user confirms the correctness of the judgment by answering the voice prompts, then the will is generated by merging the corresponding received items and details of all recipients.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This intelligent will-making system for the elderly, based on multi-turn speech recognition interaction, solves the core pain point of poor adaptability of general speech recognition to elderly voices by constructing a 100,000-hour-level elderly-specific speech dataset covering more than 30 mainstream dialects, different accents, and different physical conditions across 34 provincial-level administrative regions in China. The system adopts a pure voice interaction method throughout the entire process, eliminating the need for any complex text input or touch screen operation for the elderly. They can complete the will-making process simply through natural dialogue. It also supports breakpoint resume and context memory functions, automatically retaining the elderly's expressions and avoiding repeated questions, significantly lowering the threshold for will-making and enabling elderly people with different levels of education and digital abilities to complete the will-making process independently and conveniently.
[0014] 2. This intelligent will-making system for the elderly, based on multi-round voice recognition interaction, extracts a standard will-making framework through text segmentation, semantic role labeling, and hierarchical clustering algorithms, ensuring the legality of the will structure from the source. Simultaneously, through will sentence recombination, precise element matching, and confidence calculation mechanisms based on the distance between related sentences, it can automatically sort out the fragmented expressions of the elderly, accurately identify core elements such as the recipient, received items, and details of receipt, and automatically detect content contradictions and uncertainties. Through multi-round voice interaction, it provides targeted supplementation and confirmation, ensuring the will's content is complete, logically clear, and the expression of intent is genuine. This effectively avoids the risk of will invalidity due to formal defects or content errors, and truly protects the property rights of the elderly and family harmony. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the structure of the intelligent will-making system for the elderly based on multi-round voice recognition interaction according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Please see Figure 1 As shown, the purpose of this embodiment is to provide an intelligent system for drafting handwritten wills for the elderly based on multi-round speech recognition interaction, including a historical will data processing module, an elderly speech recognition and element extraction module, a will sentence reconstruction and element matching module, a confidence calculation and question generation module, and a multi-round interaction and dynamic update module. The historical will data processing module is used to collect historical will data, extract will-making frameworks from the historical will data, and divide the extracted will-making frameworks into will units and corresponding keywords for each will unit. In the historical will data processing module, a will database is established inside the intelligent robot to store all handwritten wills that have been completed in history. At the same time, the intelligent robot collects legally effective handwritten will texts and notarized will texts through the network module, and then summarizes them as historical will data. The historical will data collection program is initiated, automatically connecting to the local will database of the intelligent robot, exporting the plain text of all handwritten wills that have been completed and confirmed as valid through blockchain judicial evidence preservation, and synchronously recording the timestamp and evidence preservation number of each will. At the same time, it connects to two official authoritative data sources, namely the China Judgments Online website and the Ministry of Justice's Notarized Will Publicity Platform, through the network module to collect publicly available effective handwritten and notarized will texts. Furthermore, all historical will data has been anonymized, retaining only the structure and logical information of the wills, and containing no personal privacy information. Using a legally-adjusted ERNIE-NER model, all personal privacy entities are automatically identified and marked, including names, ID numbers, addresses, contact information, bank account numbers, and specific property addresses. Irreversible privacy deletion is then performed, clearing all marked privacy entities and retaining only non-identifiable information such as the logical structure of the will, property types, distribution methods, and expression habits.
[0018] All historical wills data are segmented into words, tagged with parts of speech and semantic roles, and the general logical structure of will making is extracted as a standard will making framework. The Sentence-BERT pre-trained model is used to convert each structured semantically annotated will text into a 768-dimensional fixed-length semantic vector. A hierarchical clustering algorithm is used to cluster all the will semantic vectors, and wills with similar structures are grouped into the same cluster. Semantic parsing is performed on the central vector of each cluster to extract the logical structure nodes and their order in which all wills in the cluster are shared. The common structure of all clusters is combined to generate the final standard will-making framework, which is required to include five sub-units: basic information of the testator, property list, property distribution, special arrangements, and confirmation of will wishes. The standard will-making framework is broken down into several independent but logically coherent sub-units. For each sub-unit, frequently occurring legal terms and everyday expressions commonly used by elderly users are extracted as corresponding keywords.
[0019] For each subunit, word frequency statistics are performed on all desensitized will texts to calculate the frequency and distribution of each word in the unit. The TF-IDF algorithm is used to extract the Top 50 core keywords for each subunit and automatically distinguish between legal terms and common expressions used by elderly users. Establish a two-way mapping table between legal terminology and colloquial expressions to achieve automatic conversion of expressions such as "real estate ↔ house / shop / homestead" and "bank deposit ↔ passbook / money in bank card". Then, store the keywords of each unit and the two-way mapping table into the global keyword library.
[0020] The elderly speech recognition and element extraction module is used to record the speech data of elderly users, train the automatic speech recognition model to divide the speech data into will statements, and identify the recipient, received items and receiving details for each will statement. It also matches the identified received items with the keywords of each will unit. In the elderly speech recognition and element extraction module, the voice recording function is activated by an intelligent robot to record the voice data of elderly users; After receiving the voice command to start making a will, the intelligent robot automatically starts the full-duplex voice recording function, using a 16kHz sampling rate and 16-bit mono PCM encoding format to record the elderly user's continuous voice stream. Simultaneously, a high-precision timestamp recording module is started to assign a unique absolute timestamp and relative timestamp to each 10ms voice segment, ensuring that the timing of the voice segments is traceable.
[0021] Synchronously record the timestamp of each segment of audio data; An automatic speech recognition model was established by collecting speech samples from elderly users with different dialects, accents, speaking speeds, and physical conditions. An elderly-specific speech dataset was constructed, and the automatic speech recognition model was trained based on the elderly-specific speech dataset to enable the automatic speech recognition model to have real-time speech recognition capabilities. We constructed a voice dataset specifically for the elderly, collecting voice samples from elderly users in 34 provincial-level administrative regions across the country, including 30+ mainstream dialects, 100+ local accents, a speech rate of 60-120 words / minute, and different physical conditions (wearing dentures, mild post-stroke sequelae, vocal cord aging), with a total effective sample size of ≥100,000 hours. The speech dataset is standardized and labeled, including text content, dialect type, accent features, speech rate, and speaker's physical condition. At the same time, invalid samples with background noise ≥60 decibels and unclear speech are automatically removed. The OpenAI WhisperLargev3 model was used as the base model, and fine-tuned training was performed using an elderly-specific speech dataset. The training was iterated for ≥50 rounds until the word error rate (WER) of the model on the independent test set was ≤8%. The trained ASR model specifically designed for the elderly is quantized into INT8 format and deployed to the local end of the intelligent robot to achieve offline real-time speech recognition. The inference latency for a single speech segment is ≤200ms. At the same time, a real-time speech inference pipeline is started, which inputs the locally cached continuous speech stream into the ASR model in segments and outputs preliminary text recognition results with accurate timestamps.
[0022] Then, the trained automatic speech recognition model is used to segment the speech data into will statements.
[0023] In the elderly speech recognition and element extraction module, the continuous speech stream is segmented based on the pause duration, semantic integrity and will unit keywords of the speech data to obtain multiple independent and semantically complete will statements, while meaningless interjections and repetitive sentences are removed. The initial text recognition results are preprocessed to automatically remove meaningless interjections (ah, oh, um, that, etc.), filler words, and completely repetitive sentences, resulting in a clean, continuous text stream. Simultaneously, a three-rule fusion segmentation method is employed to calculate the comprehensive segmentation confidence of each candidate segmentation point, as shown in the following formula: ; in, The overall confidence score for candidate split points, with a value range of [0,1]. The confidence score for pause duration is 1 if the pause duration is ≥2 seconds, and 0 otherwise. The confidence score is set to 1 if the statement is semantically complete, and 0 otherwise. The confidence level is triggered by keywords; it is 1 when the core keyword of the will unit is identified, and 0 otherwise. , , The weighting coefficients for pause duration (0.6), semantic integrity (0.3), and keyword triggering (0.1) are respectively. The confidence threshold for segmentation is set to 0.7. When the overall confidence of a candidate segmentation point is ≥0.7, the statement is segmented at that position to obtain multiple independent candidate will statements. A second semantic integrity check is performed on each candidate will statement. The ERNIE 3.0 semantic model is used to determine whether the statement can independently express a complete meaning, and semantically incomplete statement fragments are eliminated. Named entity recognition technology is used to extract names and kinship terms from the will as recipients, real estate, movable property, digital assets and property rights as received items, and descriptions of property shares, distribution conditions, delivery methods and custody methods as receiving details. The ERNIE-Legal-NER model, finely tuned in the legal field, is used to perform named entity recognition and semantic role labeling on each will statement. At the same time, the names and kinship terms (such as eldest son, daughter, grandson) in the statement are identified as recipients, and a unique ID is assigned to each recipient, recording their position in the statement and the corresponding timestamp. Identify real estate (houses, shops, homesteads), movable property (vehicles, bank deposits, gold), digital assets, and property rights (equity, debt) in the statement as received items, and record the item name, type, and corresponding timestamp; The system identifies descriptions in statements regarding property shares (e.g., half, all, 300,000 yuan), distribution conditions (e.g., being 18 years of age or older, responsible for elderly care), delivery methods (e.g., delivery after death, delivery in installments), and custody methods (e.g., custody by the daughter), and binds them one-to-one with the corresponding received items. The semantic similarity between the identified received items and the keywords of each will unit is calculated, and the received items are matched to the will unit with the highest semantic similarity.
[0024] The system utilizes a global keyword library generated by the historical wills data processing module to obtain the core keyword list and pre-calculated semantic vectors for each will unit. The Sentence-BERT model is then used to transform each extracted received item into a 768-dimensional semantic vector. The cosine similarity between each vector and the keyword vectors of each will unit is calculated. The received item is then matched to the will unit with the highest cosine similarity: if the highest similarity is ≥0.7, the match is successful; if the highest similarity is <0.7, the received item is marked as pending classification and pushed to the multi-round interaction module for manual confirmation. The formula is as follows: ; in, This is the semantic similarity value, ranging from -1 to 1. The closer the value is to 1, the higher the semantic matching degree. The extracted 768-dimensional Sentence-BERT semantic vector of the received item. The 768-dimensional Sentence-BERT semantic vector of the core keywords of a certain will unit.
[0025] The will statement restructuring and element matching module is used to restructure the matched will statements in each will unit according to the recording time and the recipient's position in the statement. Based on the statement restructuring results, each recipient is matched with exclusive receiving items and receiving details. In the module of will statement reorganization and element matching, within the same will unit, the will statements are initially sorted according to the chronological order of their recording time. Then, the recipient is used as the core for clustering, and all will statements that mention the same recipient are arranged together. The statement that first fully mentions the recipient's name is taken as the starting sentence of the content corresponding to that recipient, and subsequent statements are arranged sequentially after the starting sentence. The system receives a structured sentence dataset output by the elderly speech recognition and element extraction module. All sentences are grouped and isolated according to will units (testator's basic information, property list, property distribution, etc.) to ensure that sentences across units are not confused. Then, the system extracts the starting absolute timestamp of each will sentence in each unit and sorts the sentences in ascending order of timestamps to restore the original expression order of the elderly user. Traverse all statements within a unit, extract the recipient ID, recipient entity type (full name / kinship title), and the entity's position in the statement from each statement, and filter out the statement that first fully mentions the recipient's name (rather than just mentioning the kinship title), and use it as the starting sentence for all corresponding content for that recipient. The exact matching clustering method based on receiver ID is used to group all statements that mention the same receiver ID in a unit into the same statement set. At the same time, for each receiver's statement set, the statements are arranged in ascending order according to the start timestamp of the statement and concatenated sequentially after the start base sentence of the receiver to form an independent receiver statement block. Arrange all receiver statement blocks according to the chronological order of their starting reference sentences, and generate a complete statement sequence after recombination within the unit.
[0026] Perform semantic association analysis on the recombined statements, and match all received items and details mentioned in all statements after the starting sentence of the same recipient and before the starting sentence of the next recipient to the name of that recipient; For statements with contradictory content, both contradictory items are marked as high-risk items to be confirmed and directly pushed to the confidence calculation and question generation module.
[0027] The predefined will content contradiction type library includes four types of contradictions: the same received item is distributed to multiple different recipients, the total share of the property of the same recipient exceeds 100%, the distribution conditions conflict with each other, and the statements before and after are completely contradictory.
[0028] The confidence calculation and question generation module is used to calculate the distance between related sentences in the voice data, taking the receiver's sentence position as the center and combining the received items and received details. Based on the distance between related sentences, confidence is set for the corresponding received items and received details. At the same time, a confidence threshold is set. For received items and received details that are below the confidence threshold, questions are summarized and questions to be determined are generated for each receiver. In the confidence calculation and question generation module, taking the time position of the starting statement containing the recipient's name as the origin, the time interval and semantic interval between all statements mentioning the recipient's corresponding received items or receiving details and the core starting statement are calculated, and the weighted sum of the time interval and semantic interval is used as the final distance between related statements. Receive complete structured data output by the will statement reconstruction and element matching module, including the unique identifier of all recipients, the core starting statement corresponding to each recipient, the complete content and time information of all related statements, and the binding relationship between received items and receipt details; For each recipient, a unique starting statement is located, which is the statement in which the recipient's name is first mentioned in full. The absolute time position of this statement is extracted as the time origin for all subsequent calculations. Then, each associated statement under the recipient's name is traversed, and the absolute time position of each associated statement is extracted one by one. The time interval between each statement and the time origin is calculated. Compare the semantic content of each related statement with the corresponding core starting statement one by one, calculate the degree of semantic difference between the two, and obtain the semantic interval of each related statement. At the same time, according to the preset time interval weight and semantic interval weight, the time interval and semantic interval of each statement are weighted and summed to obtain the final association distance of the related statement. The weighting coefficients for the time interval (0.6) and the semantic interval (0.4) total 1. The distance between related statements is linearly negatively correlated with confidence level, as shown in the following formula: ; in, This corresponds to the confidence level value for the received item / received details. This represents the distance between related statements.
[0029] The closer the related statements are, the higher the confidence level of the corresponding received items or received details; The greater the distance between the related statements, the lower the confidence level of the corresponding received item or received details.
[0030] In the confidence calculation and question generation module, a confidence threshold is set; Compare the confidence levels corresponding to the received items and receipt details with the confidence thresholds; If the confidence level is higher than the confidence threshold, it will not be set as a question statement; Conversely, if the confidence level is lower than the confidence threshold, it is set as a problem statement; All question statements are categorized and summarized according to the recipient, corresponding questions to be determined are generated for each question statement, and the questions to be determined are transformed into conversational question statements; For each questionable received item or receipt details, a separate question to be confirmed is generated, strictly adhering to the principle that one question corresponds to only one matter to be confirmed, and multiple questions are not merged into the same question. At the same time, all structured questions to be confirmed are uniformly converted into conversational questions that are easy for elderly users to understand, avoiding the use of any legal jargon and adopting expressions used in daily communication. Each question contains only one item to be confirmed.
[0031] The multi-round interaction and dynamic update module is used to have the questions to be determined interact with the elderly user through voice through the intelligent robot, and to update the original voice data corresponding to the questions to be determined based on the latest voice data obtained from the interaction. At the same time, the confidence threshold is dynamically updated in sync with the updated voice data until there are no more questions to be determined, and then the handwritten will is completed.
[0032] In the multi-round interaction and dynamic update module, the intelligent robot plays the spoken question of the question to be determined to the elderly user based on the voice broadcast function; Meanwhile, after each question to be determined is played, a fixed pause is taken as the voice interaction data collection phase, and the voice data collected from elderly users in this phase is used as the answer voice for that question to be determined. The questions are prioritized according to their impact on the legal validity of the will, with questions about the distribution of core assets and the identity of the recipients listed first, and questions about minor details listed last. An orderly queue of interactive questions is generated, and the voice broadcast module of the intelligent robot is invoked, using an elderly-friendly tone and a moderate speaking speed to play the first conversational question in the queue order. After the question is played, it automatically pauses for 5 seconds as a dedicated voice interaction collection phase. At the same time, the full-duplex voice recording function is activated to collect only the user's voice within this time period as the answer to the corresponding question. The collected answer voice is then converted into text in real time to obtain the preliminary answer text content. If the answer is semantically ambiguous or cannot be recognized, the question will be played again automatically once; if it is still invalid after repetition, it will be marked as pending further processing. Based on the audio response, make a normal judgment on the question to be answered; If the current problem to be determined is correctly identified, the confidence threshold is adjusted based on the confidence level. The higher the confidence level, the smaller the adjustment to the confidence threshold; If the current question is judged incorrectly, the answer will be replaced with the corresponding position in the original voice data, and the corresponding will module will be updated. When the confidence level of all received items and details is higher than the current confidence level threshold, and the elderly user confirms the correctness of the judgment by answering the voice prompts, then the will is generated by merging the corresponding received items and details of all recipients.
[0033] After completing the entire processing flow for a single question, remove the question from the interactive question queue, continue playing and process the next question to be determined. After every 5 questions are completed, perform a global confidence check and iterate through the latest confidence values of all received items and receiving details. If there are still elements with confidence levels lower than the current threshold, a new question to be determined is automatically generated and added to the interaction queue, and the interaction process continues to loop. At the same time, when the confidence levels of all received items and received details are higher than the current threshold, the final confirmation voice is played, the user's answer to the final confirmation is collected, and the user's voice is identified and determined to determine whether the user has clearly confirmed that all the content is their true intention. After final confirmation, all recipients' exclusive received items and receiving details data are merged, and a complete handwritten will text is generated according to the standard will structure. The generated will text, along with the entire process interaction data of this will, is then pushed to the blockchain for evidence storage.
[0034] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent system for drafting handwritten wills for the elderly based on multi-turn voice recognition interaction, characterized in that: It includes modules for processing historical wills data, elderly speech recognition and element extraction, will sentence reconstruction and element matching, confidence calculation and question generation, and multi-round interaction and dynamic update. The historical will data processing module is used to collect historical will data, extract will-making frameworks from the historical will data, and divide the extracted will-making frameworks into will units and keywords corresponding to each will unit. The elderly speech recognition and element extraction module is used to record the speech data of elderly users, train an automatic speech recognition model to divide the speech data into will statements, and identify the recipient, received items and receiving details for each will statement. The module also matches the identified received items with the keywords of each will unit. The will statement restructuring and element matching module is used to restructure the matched will statements in each will unit according to the recording time and the recipient's position in the statement, and match exclusive receiving items and receiving details for each recipient according to the statement restructuring result; The confidence calculation and question generation module is used to calculate the distance between related sentences in the voice data, taking the receiver's sentence position as the center and combining the received items and receiving details. Based on the distance between related sentences, a confidence level is set for the corresponding received items and receiving details. At the same time, a confidence threshold is set, and questions are summarized for received items and receiving details that are below the confidence threshold to generate questions to be determined for each receiver. The multi-round interaction and dynamic update module is used to have the question to be determined interact with the elderly user through voice through an intelligent robot, and to update the original voice data corresponding to the question to be determined based on the latest voice data obtained from the interaction. At the same time, the confidence threshold is dynamically updated in sync with the updated voice data until there are no more questions to be determined, and then the handwritten will is completed.
2. The intelligent system for making handwritten wills for the elderly based on multi-turn speech recognition interaction as described in claim 1, characterized in that: In the historical will data processing module, a will database is established inside the intelligent robot to store all handwritten wills that have been completed in the past. At the same time, the intelligent robot collects legally effective handwritten will texts and notarized will texts through the network module, and then summarizes them as historical will data. Furthermore, all historical will data has been anonymized, retaining only the structure and logical information of the wills, and containing no personal privacy information. All historical wills data are segmented into words, tagged with parts of speech and semantic roles, and the general logical structure of will making is extracted as a standard will making framework. The standard will-making framework is broken down into several independent but logically coherent sub-units. For each sub-unit, frequently occurring legal terms and everyday expressions commonly used by elderly users are extracted as corresponding keywords.
3. The intelligent system for drafting handwritten wills for the elderly based on multi-turn speech recognition interaction as described in claim 1, characterized in that: In the elderly speech recognition and element extraction module, the voice recording function is activated by an intelligent robot to record the voice data of elderly users. Synchronously record the timestamp of each segment of audio data; An automatic speech recognition model was established by collecting speech samples from elderly users with different dialects, accents, speaking speeds, and physical conditions. An elderly-specific speech dataset was constructed, and the automatic speech recognition model was trained based on the elderly-specific speech dataset to enable the automatic speech recognition model to have real-time speech recognition capabilities. Then, the trained automatic speech recognition model is used to segment the speech data into will statements.
4. The intelligent system for drafting handwritten wills for the elderly based on multi-turn speech recognition interaction as described in claim 1, characterized in that: In the elderly speech recognition and element extraction module, the continuous speech stream is segmented based on the pause duration, semantic integrity and will unit keywords of the speech data to obtain multiple independent and semantically complete will statements, while meaningless interjections and repetitive statements are removed. Named entity recognition technology is used to extract names and kinship terms from the will as recipients, real estate, movable property, digital assets and property rights as received items, and descriptions of property shares, distribution conditions, delivery methods and custody methods as receiving details. The semantic similarity between the identified received items and the keywords of each will unit is calculated, and the received items are matched to the will unit with the highest semantic similarity.
5. The intelligent system for drafting handwritten wills for the elderly based on multi-turn speech recognition interaction as described in claim 1, characterized in that: In the will statement reorganization and element matching module, within the same will unit, the will statements are initially sorted according to the chronological order of their recording time. Then, the recipient is used as the core for clustering, and all will statements that mention the same recipient are arranged together. The statement that first fully mentions the recipient's name is taken as the starting sentence of the content corresponding to that recipient, and subsequent statements are arranged sequentially after the starting sentence. Perform semantic association analysis on the recombined statements, and match all received items and details mentioned in all statements after the starting sentence of the same recipient and before the starting sentence of the next recipient to the name of that recipient; For statements with contradictory content, both contradictory items are marked as high-risk items to be confirmed and directly pushed to the confidence calculation and question generation module.
6. The intelligent system for drafting handwritten wills for the elderly based on multi-turn speech recognition interaction as described in claim 1, characterized in that: In the confidence calculation and question generation module, taking the time position of the starting statement containing the recipient's name as the origin, the time interval and semantic interval between all statements mentioning the recipient's corresponding received items or receiving details and the core starting statement are calculated, and the weighted sum of the time interval and semantic interval is used as the final distance between related statements. The weighting coefficients for the time interval and the semantic interval total 1. The distance between related statements shows a linear negative correlation with confidence levels. The closer the related statements are, the higher the confidence level of the corresponding received item or received details; The greater the distance between the related statements, the lower the confidence level of the corresponding received item or received details.
7. The intelligent system for making handwritten wills for the elderly based on multi-turn speech recognition interaction as described in claim 1, characterized in that: In the confidence calculation and question generation module, a confidence threshold is set; Compare the confidence levels corresponding to the received items and receipt details with the confidence thresholds; If the confidence level is higher than the confidence threshold, it will not be set as a question statement; Conversely, if the confidence level is lower than the confidence threshold, it is set as a problem statement; All question statements are categorized and summarized according to the recipient, corresponding questions to be determined are generated for each question statement, and the questions to be determined are transformed into conversational question statements; Each question contains only one item to be confirmed.
8. The intelligent system for making handwritten wills for the elderly based on multi-turn speech recognition interaction as described in claim 1, characterized in that: In the multi-round interaction and dynamic update module, the intelligent robot plays the spoken question of the question to be determined to the elderly user based on the voice broadcast function. Meanwhile, after each question to be determined is played, a fixed pause is taken as the voice interaction data collection phase, and the voice data collected from elderly users in this phase is used as the answer voice for that question to be determined. Based on the audio response, make a normal judgment on the question to be answered; If the current problem to be determined is correctly identified, the confidence threshold is adjusted based on the confidence level. The higher the confidence level, the smaller the adjustment to the confidence threshold; If the current question is judged incorrectly, the answer will be replaced with the corresponding position in the original voice data, and the corresponding will module will be updated. When the confidence level of all received items and details is higher than the current confidence level threshold, and the elderly user confirms the correctness of the judgment by answering the voice prompts, then the will is generated by merging the corresponding received items and details of all recipients.