Voice data processing method and processing system

By fine-tuning the Transformer model to adapt to health management scenarios, the problem of insufficient accuracy of speech recognition and information extraction models in chronic disease management has been solved, achieving higher accuracy in speech data conversion and text data extraction, and providing detailed health management records and guidance.

CN121768698APending Publication Date: 2026-03-31SHENZHEN SIBIONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing speech recognition and information extraction models have limited accuracy in chronic disease management, especially in blood glucose management, and are unable to effectively identify disease and drug names.

Method used

A pre-trained model based on the Transformer model is used to fine-tune speech and text data, which are then used for speech recognition and information extraction, respectively, to adapt to health management scenarios and improve the accuracy of speech data conversion and text data information extraction.

Benefits of technology

It improves the accuracy of speech recognition and information extraction, and can record behavioral activities in detail during the health management process, providing users with accurate reference information and guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768698A_ABST
    Figure CN121768698A_ABST
Patent Text Reader

Abstract

The invention discloses a voice data processing method and processing system. The processing method comprises the following steps: receiving voice data of a target object; the voice data is automatically converted into first text data by utilizing a first model, the first model is obtained by finely adjusting a first pre-training model based on a Transform model through a voice data set, and the voice data set comes from a health management process; target information of the first text data is extracted through a second model to obtain second text data, the second text data comprises entity information, and the entity information is related to behavior activities of the target object in the health management process and comprises entity name information, entity quantity and time information; the second model is obtained by finely adjusting a second pre-training model based on a Transform model through a text data set, and the text data set comes from a health management process; and ranking the entity information based on the temporal information. Therefore, the speech recognition accuracy and the information extraction accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of data processing technology, and specifically to a method and system for processing voice data. Background Technology

[0002] A continuous glucose monitoring (CGM) device is used to monitor the body's glucose concentration in real time. Typically, a CGM is used in conjunction with an application installed on a mobile smart device (such as a smartphone). The application displays the glucose concentration data obtained by the CGM in real time, allowing users (such as diabetic patients) to intuitively and conveniently obtain their own glucose concentration data. Furthermore, to facilitate user management of their glucose levels, the application usually includes a form-filling and check-in function. Users can manually enter text to fill out forms and record their diet, medication, exercise, and other information, providing reference information for health management.

[0003] In existing technologies, to facilitate form filling and attendance tracking for certain special users (such as illiterate users, users with poor vision, or patients with diabetic retinopathy), mobile smart terminal applications can generally use speech-to-text models (i.e., speech recognition models) to automatically convert the user's voice input into text content. Furthermore, the application can also use information extraction models to extract key information from the text content to complete the attendance tracking operation.

[0004] However, in the field of chronic disease management, represented by blood glucose management, the accuracy of general speech recognition models, which are trained using public data, in recognizing speech such as disease and drug names needs to be improved. In addition, the accuracy of general information extraction models, which are also trained using public data, in extracting key information such as disease and drug names also needs to be improved. Summary of the Invention

[0005] This disclosure is made in view of the above-mentioned situation, and its purpose is to provide a method and system for processing speech data that can improve the accuracy of speech recognition and information extraction.

[0006] Therefore, the first aspect of this disclosure provides a method for processing voice data, which is a method for processing voice data generated during health management, comprising: receiving voice data of a target object; automatically converting the voice data into first text data using a first model, wherein the first model is obtained by fine-tuning a first pre-trained model based on a Transformer model using a voice dataset, the voice dataset originating from the health management process; extracting target information from the first text data using a second model to obtain second text data, wherein the second text data includes entity information, wherein the entity information is related to the target object's behavior and activities during the health management process and includes entity name information, entity quantity, and time information, wherein the second model is obtained by fine-tuning a second pre-trained model based on a Transformer model using a text dataset, the text dataset originating from the health management process; and sorting the entity information based on the time information to obtain an entity information sequence.

[0007] In the first aspect of this disclosure, since the first model is obtained by fine-tuning a first pre-trained model based on a Transformer model using a speech dataset of health management processes, the first model can adapt to the application scenarios of health management, improving its generalization ability to speech data of health management processes, thereby improving the accuracy of speech recognition for health management processes. Furthermore, since the second model is obtained by fine-tuning a second pre-trained model based on a Transformer model using a text dataset of health management processes, the second model can adapt to the application scenarios of health management, improving its generalization ability to text data of health management processes, thereby improving the accuracy of information extraction from the text data of health management processes. In this scenario, when the target audience uses voice to check in for health management, the first model can improve the accuracy of automatically converting voice data into first text data, while the second model can improve the accuracy of extracting target information from the first text data to obtain second text data. Since the entity information is related to the target audience's behavior during the health management process, sorting the entity information based on time information allows for the classification and sorting of the specific content of the behavior activities during the health management process according to the chronological order of the behavior activities. This enables clear, detailed, and accurate recording of the behavior activities during the health management process, providing accurate reference information for the target audience's health management.

[0008] Furthermore, in the voice data processing method according to the first aspect of this disclosure, optionally, the entity information further includes entity categories, which include at least one of diet, medication, exercise, and single-point analyte detection. In this case, since the entity information is related to the target object's behavioral activities during the health management process, by including at least one of diet, medication, exercise, and single-point analyte detection in the entity categories, the entity information can more comprehensively cover the content of the target object's behavioral activities during the health management process, providing comprehensive and accurate reference information for the target object's health management.

[0009] Furthermore, in the voice data processing method according to the first aspect of this disclosure, optionally, it further includes: acquiring analyte concentration data within the target object's body; acquiring analyte changes within a preset time period after the occurrence of the behavioral activity represented by the second text data, based on the second text data and the analyte concentration data; and generating guidance suggestions for health management based on the analyte changes. In this case, by combining the second text data with the analyte concentration data to acquire analyte changes, it is easier to analyze the reasons for fluctuations in the analyte concentration data of the target object within a preset time period after the occurrence of the relevant behavioral activity, thereby providing the target object with practical guidance suggestions for health management and optimizing the target object's self-management level.

[0010] Furthermore, in the voice data processing method according to the first aspect of this disclosure, optionally, it further includes: acquiring analyte concentration data within the target object's body; displaying a trend graph of the analyte concentration data changing over time; and marking the entity information in the trend graph based on the time information in the entity information. In this case, by displaying a trend graph of the analyte concentration data changing over time, the target object can intuitively and clearly view the trend of changes in the analyte concentration data after behavioral activities occur during health management. Additionally, by marking the entity information in the second text data in the trend graph, the entity information can be matched with the analyte concentration data, thereby facilitating accurate analysis of the reasons for fluctuations in the analyte concentration data of the target object within a preset time period after the relevant behavioral activities occur.

[0011] Furthermore, in the voice data processing method according to the first aspect of this disclosure, optionally, the voice data is obtained by receiving a check-in event from the target object. In this case, since the voice data involved in the data processing method comes from the check-in event of the target object, the target object only needs to check in by voice to generate second text data to record the specific content of the behavioral activities during the health management process, thereby providing convenience for the target object (especially for example, illiterate users, users with poor vision, or patients with diabetic retinopathy, etc.) to conduct health management through check-in.

[0012] Furthermore, in the speech data processing method according to the first aspect of this disclosure, optionally, the first pre-trained model is a Whisper model, and / or the second pre-trained model is a UIE model. In this case, since the Whisper model has good generalization performance in the field of speech recognition, fine-tuning the Whisper model using a speech dataset of health management processes can improve the Whisper model's generalization ability in the field of health management, thereby improving the Whisper model's accuracy in speech recognition of health management processes. Additionally, since the UIE model has good generalization performance in the field of text information extraction, fine-tuning the UIE model using a text dataset of health management processes can improve the UIE model's generalization ability in the field of health management, thereby improving the UIE model's accuracy in extracting information from health management processes.

[0013] Furthermore, in the speech data processing method according to the first aspect of this disclosure, optionally, the first pre-trained model is a multi-task prediction model, wherein the multi-task prediction includes predicting whether the speech data is a human voice, predicting the language used in the speech data, predicting the start and end times of the human voice in the speech data, and predicting the text of the human voice in the speech data. This can accelerate the training and convergence speed of the first pre-trained model and improve its generalization ability.

[0014] Furthermore, in the speech data processing method according to the first aspect of this disclosure, optionally, the total duration of the speech dataset is not less than 12 hours, and / or the amount of data in the text dataset is not less than 70,000 entries. In this case, compared to a large-scale speech dataset, the cost of collecting and managing a 12-hour speech dataset is lower, the computational resource requirements for training the first pre-trained model are also lower, and it is sufficient to help the first pre-trained model learn the general characteristics of speech data in the health management field. Similarly, compared to a large-scale text dataset, the cost of collecting and managing a text dataset with 70,000 entries is lower, the computational resource requirements for training the second pre-trained model are also lower, and it is sufficient to help the second pre-trained model learn the general characteristics of text data in the health management field.

[0015] Furthermore, in the speech data processing method according to the first aspect of this disclosure, optionally, the loss function of the first pre-trained model is: L 10 =L 12 +L 14 +L 16 +L 18 , where L 10 L represents the loss function of the first pre-trained model. 12Used to verify whether there is loss of human voice in the voice data, L 14 L is used to verify the language loss used in the speech data. 16 Used to verify the loss of human voice start and end times in speech data, L 18 The loss function used to verify the text of human voice in the speech data; and / or the loss function of the second pre-trained model is: L 20 =L 22 +L 24 +L 26 , where L 20 L represents the loss function of the second pre-trained model. 22 L represents the loss of positive and negative sample pairs in text structure. 24 L represents the loss generated from structured data. 26 This indicates the loss of retrospective semantics.

[0016] In addition, a second aspect of this disclosure provides a voice data processing system, including at least one processing circuit configured to perform the processing method involved in the first aspect of this disclosure.

[0017] According to this disclosure, a method and system for processing speech data that can improve the accuracy of speech recognition and information extraction can be provided. Attached Figure Description

[0018] This disclosure will now be explained in further detail by way of example only with reference to the accompanying drawings.

[0019] Figure 1 This is a schematic diagram illustrating an application scenario of the data processing method involved in the examples of this disclosure.

[0020] Figure 2 This is a flowchart illustrating the data processing method involved in the example of this disclosure.

[0021] Figure 3 This is a schematic diagram illustrating the conversion of speech data into first text data as described in the example of this disclosure.

[0022] Figure 4 This is a schematic diagram illustrating the process of obtaining the first model as described in this disclosure example.

[0023] Figure 5 This is a schematic diagram illustrating information extraction from first text data as described in this disclosure example.

[0024] Figure 6 This is a schematic diagram illustrating the process of obtaining the second model as described in this disclosure example.

[0025] Figure 7This is a schematic diagram illustrating the display of second text data in a timeline manner as described in this disclosure example.

[0026] Figure 8 This is a flowchart illustrating a first embodiment of the data processing method involved in the present disclosure.

[0027] Figure 9 This is a flowchart illustrating a second embodiment of the data processing method involved in the present disclosure.

[0028] Figure 10 This is a graph showing the trend of analyte concentration data over time as described in this disclosure example.

[0029] Figure 11 This is a flowchart illustrating the data processing methods involved in the examples of this disclosure. Detailed Implementation

[0030] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0031] It should be noted that the terms "first," "second," "third," and "fourth," etc., in this disclosure, claims, and the aforementioned drawings are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. In the following description, the same reference numerals are used for the same parts, and repeated descriptions are omitted. Additionally, the drawings are merely schematic diagrams, and the scale of the dimensions of the parts or the shape of the parts may differ from the actual figures.

[0032] This disclosure relates to a method for processing voice data, specifically a method for processing voice data generated during health management. Using the processing method disclosed herein to process voice data generated during health management can improve the accuracy of voice recognition and information extraction.

[0033] In some examples, the methods for processing speech data can be simply referred to as data processing methods. In other examples, data processing methods may also be called data transformation methods, data extraction methods, or information extraction methods, etc.

[0034] In some examples, the object generating the voice data can be referred to as the target object. In some examples, data processing methods can be used to process the voice data generated by the target object during health management.

[0035] In some examples, the target audience can be a living organism. In some examples, the target audience can be a patient with diabetes, hypopituitarism, thyroid disease, or obesity. In some examples, the target audience can also be a person in good health. In some examples, the target audience can be anyone who needs health management, regardless of age, gender, race, health condition, etc. Additionally, in some examples, the target audience can also be referred to as a user.

[0036] In some examples, health management can refer to the systematic and continuous monitoring and recording of a target individual's behavior, medication management, and physical condition to help the target individual understand their own health status, thereby implementing health interventions to prevent diseases (such as chronic diseases) or slow the progression of diseases, and ultimately improve the target individual's health status and quality of life.

[0037] In some examples, health management may include chronic disease management. In some examples, chronic diseases may include diabetes, such as type 1 or type 2 diabetes. In some examples, chronic diseases may also include hypertension, asthma, and cardiovascular diseases. In still other examples, health management may refer to chronic disease management.

[0038] In some examples, chronic disease management can refer to the target individual's real-time and continuous monitoring of their physical condition and the continuous recording of their diet, medication, and exercise, thereby providing guidance or behavioral recommendations for the target individual's health monitoring.

[0039] In some examples, health management may include check-in events. In some examples, check-in events can represent a target individual recording their diet, medication use, exercise, and analyte concentrations in their body. For example, the target individual may record their diet, medication use, exercise, and analyte concentrations via voice or text.

[0040] The following description, using diabetic patients as the target group and in conjunction with the accompanying drawings, explains the data processing methods involved in this disclosure.

[0041] Figure 1 This is a schematic diagram illustrating an application scenario of the data processing method involved in the examples of this disclosure. Figure 2 This is a flowchart illustrating the data processing method involved in the example of this disclosure.

[0042] The data processing methods disclosed herein can be applied to, for example... Figure 1 In the scenario shown. See also some examples. Figure 1 The first model 10 can input the voice data of the target object into the first model 10, which can convert the voice data into first text data. The second model 12 can receive the first text data and extract the target information from the first text data to obtain the second text data.

[0043] See in some examples Figure 2 The data processing method may include: receiving voice data of the target object (step S110), automatically converting the voice data into first text data using the first model 10 (step S120), extracting target information from the first text data using the second model 12 to obtain second text data (step S130), and sorting the entity information based on time information (step S140).

[0044] In some examples, in step S110, voice data of the target object may be received. In some examples, voice data generated by the target object during health management may be received. In some examples, voice data of the target object used for health management may be received.

[0045] In some examples, in step S110, the voice data can be obtained by receiving a check-in event from the target object. That is, the voice data involved in the data processing method can come from the check-in event of the target object. In this case, the target object only needs to check in by voice to generate second text data to record the specific content of the behavioral activities in the health management process, thereby providing convenience for the target object (especially illiterate users, users with poor vision, or patients with diabetic retinopathy, etc.) to carry out health management through check-in.

[0046] In some examples, voice data can be obtained by receiving check-in events from the target object for health management.

[0047] In some examples, in step S110, the voice data may be voice data recorded by the target object during health management, including information on diet, medication, exercise, and the concentration of analytes in the body.

[0048] Figure 3 This is a schematic diagram illustrating the conversion of speech data into first text data as described in the example of this disclosure.

[0049] In some examples, in step S120, the first model 10 can be used to automatically convert the speech data into the first text data.

[0050] In some examples, the first text data can be the transcription result of the speech data in step S110. In some examples, the first text data and the speech data in step S110 can have a mapping relationship, and the first text data can reflect the content of the speech data. For example, in Figure 3 In the example shown, the voice data is: "I have taken two metformin tablets now", and the corresponding first text data is: "I have taken two metformin tablets now".

[0051] In some examples, the first model 10 may be obtained by fine-tuning a first pre-trained model based on a Transformer model. The Transformer model may include an encoder and a decoder (i.e., the Transformer model may be a model that includes an encoder and a decoder architecture).

[0052] In some examples, within the field of machine learning, fine-tuning can refer to the process of further training an already pre-trained model for a specific task or dataset. Additionally, the purpose of fine-tuning can be to better adapt a pre-trained model to new tasks.

[0053] In some examples, the first model 10 can be obtained by fine-tuning a first pre-trained Transformer-based model using a speech dataset, which may originate from a health management process. That is, the first model 10 can be obtained by fine-tuning a first pre-trained Transformer-based model using a speech dataset related to the health management process. In this case, since the first model 10 is obtained by fine-tuning a first pre-trained Transformer-based model using a speech dataset from the health management process, the first model 10 can adapt to the application scenario of health management, improving its generalization ability to speech data from the health management process, thereby improving the speech recognition accuracy for the health management process.

[0054] In some examples, the voice dataset can originate from the voice data used by the target object when performing check-in events during health management. In other examples, voice data, such as that of diabetic patients checking in, can be collected from video websites to obtain the voice dataset.

[0055] In some examples, the total duration of the speech dataset can be no less than 12 hours. In this case, compared to large-scale speech datasets, the cost of collecting and managing a 12-hour speech dataset is lower, the computational resource requirements for training the first pre-trained model are also less, and it is sufficient to help the first pre-trained model learn the general characteristics of speech data in the health management domain. For example, the total duration of the speech dataset can be 12 hours or more.

[0056] In some examples, the first pre-trained model based on the Transformer model (hereinafter referred to as the first pre-trained model) can be a pre-trained Automatic Speech Recognition (ASR) model.

[0057] In some examples, the first pre-trained model can be a model pre-trained using a public speech dataset. That is, the pre-training data for the first pre-trained model can come from a public speech dataset. For example, a public speech dataset can include speech data from multiple fields such as healthcare, finance, law, and news.

[0058] In some examples, the first pre-trained model can be a multi-task prediction model. In some examples, multi-task prediction may include predicting whether the speech data is a human voice, predicting the language used in the speech data, predicting the start and end times of the human voice in the speech data, and predicting the text of the human voice in the speech data. This can accelerate the training and convergence speed of the first pre-trained model and improve its generalization ability.

[0059] In some examples, the first pre-trained model can be a WeNet model, an AISHELL model, or a Whisper model. In some examples, preferably, the first pre-trained model can be a Whisper model. In this case, since the Whisper model has good generalization performance in the field of speech recognition, fine-tuning the Whisper model using a speech dataset of health management processes can improve the Whisper model's generalization ability in the field of health management, thereby improving the Whisper model's accuracy in speech recognition of health management processes.

[0060] In some examples, the Whisper model can perform multiple prediction tasks simultaneously through encoding and decoding, such as converting speech data into text data, translating other languages ​​into English, and identifying whether there is a human voice.

[0061] It should be noted that in the field of chronic disease management, represented by blood glucose management, English speech data (such as CGM, DKA, HbA1c) is more likely to appear; that is, the speech data involved in this disclosure can include English speech. After researching various ASR models, the inventors found that the WeNet model, AISHELL model, and Whisper model each have their own characteristics. For example, the WeNet model performs well in Chinese speech recognition but poorly in English spelling (such as CGM, HbA1c); the AISHELL model performs well in English spelling but slightly worse in Chinese speech recognition; the generalization performance of both the WeNet and AISHELL models is inferior to that of the Whisper model. The inventors prioritized the Whisper model as the first pre-trained model after comprehensively considering the application scenario, generalization ability, and model performance.

[0062] Figure 4 This is a schematic diagram illustrating the process of obtaining the first model 10 as described in this disclosure example.

[0063] See in some examples Figure 4 Obtaining the first model 10 may include: cleaning and labeling the speech dataset to obtain a labeled speech dataset (step S122); dividing the labeled speech dataset into a first training dataset, a first validation dataset, and a first test dataset (step S124); and fine-tuning the first pre-trained model based on the Transformer model using the first training dataset and the first validation dataset (step S126).

[0064] In some examples, in step S122, the speech dataset can be cleaned first, and then the speech dataset can be labeled to obtain a labeled speech dataset.

[0065] In some examples, data cleaning of the speech dataset may include removing background noise and silent portions of the speech dataset. Additionally, in step S122, the duration of the labeled speech dataset after data cleaning can be 12 hours.

[0066] In some examples, in step S124, the labeled speech dataset can be divided according to a first preset ratio to obtain a first training dataset, a first validation dataset, and a first test dataset. In some examples, the labeled speech dataset can be divided in a 4:1:1 ratio. For example, for a 12-hour labeled speech dataset, it can be separated into an 8-hour first training set, a 2-hour first validation set, and a 2-hour first test set.

[0067] In some examples, in step S126, a first pre-trained model based on the Transformer model can be fine-tuned using a first training dataset and a first validation dataset to obtain a first model 10.

[0068] In some examples, the loss function of the first pre-trained model may include losses for verifying whether the current speech data contains human voices, losses for verifying the language of the current speech data, losses for verifying the timestamp of the current speech data, and losses for verifying that the current speech data recognizes text. The loss function of the first pre-trained model can be expressed as:

[0069] L 10 =L 12 +L 14 +L 16 +L 18 ,…………Formula (1)

[0070] Among them, L 10 L represents the loss function of the first pre-trained model. 12 Used to verify whether there is loss of human voice in the voice data, L 14 L is used to verify the language loss used in the speech data. 16 Used to verify the loss of human voice start and end times in speech data, L 18 Used to verify the loss of text in speech data containing human voices.

[0071] In some examples, the formula for calculating whether there is a loss of human voice in the speech data can be:

[0072]

[0073] Among them, L 12 This indicates whether the speech data has lost human voice input, where y represents the true label (where 1 represents human voice and 0 represents background noise). This indicates the probability that the predicted speech data contains human voices.

[0074] Furthermore, the formula for calculating the language loss used in the speech data can be:

[0075]

[0076] Among them, L 14 The loss represents the language used in the speech data, where N represents the number of language categories, and y represents the language loss. i One-Hot encoding representing real language (if the real language is class i, y i =1, others y i (for 0), This represents the predicted probability for the i-th language class.

[0077] Furthermore, the formula for calculating the loss of human voice start and end times in speech data can be:

[0078]

[0079] Among them, L 16 t represents the loss due to the start and end times of human voices in the speech data. start and t end These represent the start and end times of the actual human voice. and These represent the predicted start and end times of the human voice.

[0080] Furthermore, the formula for calculating the loss of human voice text in speech data can be:

[0081]

[0082] Among them, L 18 This represents the loss of human voice text in the speech data, where T represents the length of the target text, and y represents the loss. t P(y) represents the actual character of the target text at time step t. t |x) indicates that at time step t, the character y is... i The predicted probability.

[0083] In some examples, the loss function of the first pre-trained model can be expressed as:

[0084] L 10 =λ1L 12 +λ2L 14 +λ3L 16 +λ4L 18 ,………Formula (6)

[0085] Wherein, λ1, λ2, λ3, and λ4 represent the weight hyperparameters of the corresponding loss, which can be adjusted according to the actual needs of the task to balance the importance of different tasks.

[0086] It should be noted that in the loss function of the first pre-trained model, the speech data can be any speech data input to the first pre-trained model. For example, during pre-training, the speech data can be the speech data from the pre-training data. Furthermore, during application, the speech data can be the currently received speech data.

[0087] In some examples, in step S126, the Whisper model can be fine-tuned using the first training dataset and the first validation dataset to obtain the first model 10. In this case, because the Whisper model has excellent generalization performance, after being trained on a speech dataset in the health management field, the Whisper model can adapt to the application scenario of health management, improve the generalization ability of speech data in the health management process, and thus improve the speech recognition accuracy of the health management process.

[0088] In some examples, after step S126, the generalization ability of the first model 10 can be tested using the first test dataset.

[0089] Figure 5 This is a schematic diagram illustrating information extraction from first text data as described in this disclosure example.

[0090] See back Figure 2 In step S130, the target information in the first text data can be extracted using the second model 12 to obtain the second text data.

[0091] In some examples, the target information in the first text data can be information related to health management. For example, the target information could be information related to diet, medication, exercise, or analyte testing in the first text data. Alternatively, the target information could be information related to the names of foods, medications, exercises, or analyte concentrations in the first text data. Furthermore, the target information can also be referred to as key information.

[0092] In some examples, the second text data can be used to record the specific content of behavioral activities during the health management process. In some examples, the second text data can be structured data. For example, in... Figure 5 In the example shown, the first text data could be: "I have now taken two metformin tablets," and the second text data could be: Figure 5 The structured data shown.

[0093] In some examples, the second text data may include entity information. This entity information may be related to the target object's behavioral activities during the health management process.

[0094] In some examples, the target's behavioral activities during the health management process may include activities such as diet, medication, exercise, and analyte testing.

[0095] For example, in Figure 5 In the example shown, entity information may include “Drug: Metformin”, “Quantity: Two tablets”, and “Dose date: 2024-05-11 17:50”.

[0096] In some examples, analyte detection can include continuous analyte detection and single-point analyte detection. Single-point analyte detection can involve using a single-point analyte collection device to detect bodily fluids (e.g., blood) that have left the target object. In some examples, the single-point analyte collection device can be a finger-prick blood glucose meter. For instance, a blood sample can be obtained by pricking the fingertip with a disposable needle, using a test strip or pipette, and then using a finger-prick blood glucose meter to detect the blood sample and obtain analyte concentration data in the blood sample.

[0097] In some examples, entity information may include entity name information, entity quantity, and time information. Entity name information can be a name used to uniquely identify an entity. Entity quantity can be a measure of the entity. That is, entity quantity can be used to characterize a quantity associated with the entity name information. For example, in... Figure 5 In the example shown, the entity name is "metformin" and the quantity is "two tablets". Another example is that the entity name could be "running" and the quantity could be "1 hour".

[0098] In some examples, entity name information may include at least one of the following: the food consumed by the target object during health management, the medication taken, the name of the exercise, and the analyte detected at a single point. Additionally, entity quantity may include at least one of the following: the quantity or weight of food, the amount of medication, the intensity or duration of exercise, and the concentration data of the analyte.

[0099] In some examples, time information can be correlated with the timing of the target subject's behavioral activities during the health management process.

[0100] In some examples, time information can be extracted from the first text data, and the time information of the second text data can be determined based on this time information. In some examples, time information can be extracted from the first text data, and the current time of the second model 12 can be obtained from the time information of the second text data based on this time information. In some examples, the time information may include terms such as now, current, immediate, just now, a moment ago, at this time, and so on.

[0101] For example, if the first text data is "I have taken two metformin tablets now", and the information describing the time in the first text data is "now", then the information "now" can be extracted, and the current time of the model can be obtained from the second model 12 as the time information of the second text data.

[0102] In some examples, time information can be extracted from the first text data and used as the time information for the second text data. In other examples, time information can be extracted from the first text data, and then converted (e.g., processed to a specified format) before being used as the time information for the second text data.

[0103] For example, if the first text data is "I ate two bowls of white porridge at 8:30 in the morning", the information describing the time in the first text data is "8:30 in the morning". Then, the information "8:30 in the morning" can be extracted and processed into "08:30" as the time information of the second text data.

[0104] In some examples, in response to the lack of time information in the first text data, the time when the first text data is input into the second model 12 can be used as the time information for the second text data.

[0105] For example, if the first text data is "I ate two bowls of white porridge", and there is no information describing the time in the first text data, then the time when the first text data is input into the second model 12 can be recorded and used as the time information of the second text data.

[0106] In some examples, entity information may also include entity categories. In some examples, entity categories may include at least one of diet, medication, exercise, and single-point analyte detection. In this case, since the entity information is related to the target object's behavioral activities during health management, including at least one of diet, medication, exercise, and single-point analyte detection in the entity categories allows the entity information to comprehensively cover the target object's behavioral activities during health management, providing comprehensive and accurate reference information for the target object's health management.

[0107] In some examples, the second model 12 can be used to extract the first text data into structured data (e.g., Figure 5 (The structured data shown). Additionally, structured data can include entity name information and entity quantity. For example, structured data can include information such as food name, weight, and quantity. As another example, structured data can include information such as drug name and quantity. In some examples, structured data can also include time information (see...). Figure 5 ).

[0108] In some examples, the second model 12 may be obtained by fine-tuning a second pre-trained model based on a Transformer model. Additionally, the Transformer model may include an encoder and a decoder (i.e., the Transformer model may be a model containing an encoder and decoder architecture).

[0109] In some examples, the second model 12 can be obtained by fine-tuning a second pre-trained Transformer model using a text dataset, which may originate from the health management process. That is, the second model 12 can be obtained by fine-tuning a second pre-trained Transformer model using a text dataset related to the health management process. In this case, since the second model 12 is obtained by fine-tuning a second pre-trained Transformer model using a text dataset from the health management process, the second model 12 can adapt to the application scenario of health management, improving its generalization ability to text data from the health management process, thereby increasing the accuracy of information extraction from the text data of the health management process.

[0110] In some examples, the text dataset can originate from the text data used by the target object when performing check-in events during health management. In other examples, text data, such as that of diabetic patients checking in, can be collected from video websites to obtain the text dataset.

[0111] In some examples, the text dataset can contain at least 70,000 entries. In this case, compared to large-scale text datasets, collecting and managing a text dataset of 70,000 entries is less costly, requires fewer computational resources for training the second pre-trained model, and is sufficient to help the second pre-trained model learn the general characteristics of text data in the health management domain. For example, the text dataset can contain 70,000 or more entries.

[0112] In some examples, the second pre-trained model based on the Transformer model (hereinafter referred to as the second pre-trained model) can be a pre-trained information extraction (IE) model.

[0113] In some examples, the second pre-trained model can be a model pre-trained using a public text dataset. That is, the pre-training data for the second pre-trained model can come from a public text dataset. For example, a public text dataset can include text data from multiple fields such as medicine, finance, law, and news.

[0114] In some examples, the second pre-trained model can perform multi-task prediction. Furthermore, multi-task prediction can include entity recognition of text data, relation extraction from text data, event extraction from text data, and sentiment polarity determination of text data.

[0115] In some examples, the second pre-trained model can be a Universal Information Extraction (UIE) model. In this case, since the UIE model has good generalization performance in the field of text information extraction, fine-tuning the UIE model using a text dataset of health management processes can improve the UIE model's generalization ability in the health management field, thereby improving the accuracy of the UIE model in extracting information from the health management process.

[0116] In some examples, the UIE model can perform multiple prediction tasks simultaneously through encoding and decoding, such as extracting entity information, relational information, temporal information, and sentiment information from text data at the same time.

[0117] Figure 6 This is a schematic diagram illustrating the process of obtaining the second model 12 as described in this disclosure example.

[0118] See in some examples Figure 6 Obtaining the second model 12 may include: cleaning and labeling the text dataset to obtain a labeled text dataset (step S132); dividing the labeled text dataset into a second training dataset, a second validation dataset, and a second test dataset (step S134); and fine-tuning the second pre-trained model based on the Transformer model using the second training dataset and the second validation dataset (step S136).

[0119] In some examples, in step S132, the text dataset can be cleaned first, and then the text data can be labeled to obtain a labeled text dataset.

[0120] In some examples, data cleaning of a text dataset may include: removing special characters and punctuation marks from the text dataset, correcting spelling errors in the text dataset, and identifying and handling duplicate content in the text dataset.

[0121] In some examples, in step S134, the labeled text dataset can be divided according to a second preset ratio to obtain a second training dataset, a second validation dataset, and a second test dataset. In some examples, the labeled text dataset can be divided in a 5:1:1 ratio. For example, for a labeled speech dataset containing 70,000 text data points, it can be separated into a second training set containing 50,000 text data points, a second validation set containing 10,000 text data points, and a second test set containing 10,000 text data points.

[0122] In some examples, in step S136, a second pre-trained model based on the Transformer model can be fine-tuned using a second training dataset and a second validation dataset to obtain a second model 12.

[0123] In some examples, the loss function of the second pre-trained model may include the loss from text-structured positive and negative sample pairs, the loss from structured data generation, and the loss from reviewing semantics. The loss function of the second pre-trained model can be expressed as:

[0124] L 20 =L 22 +L 24 +L 26 ,…………Formula (7)

[0125] Among them, L 20 L represents the loss function of the second pre-trained model. 22 L represents the loss of positive and negative sample pairs in text structure. 24 L represents the loss generated from structured data. 26 This indicates the loss of retrospective semantics.

[0126] In some examples, the loss of positive and negative sample pairs from text structure can represent the loss of positive and negative sample pairs from text to structure. Furthermore, positive samples can refer to correctly extracted structured data from the text, while negative samples can refer to noise or randomly generated invalid structures.

[0127] In some examples, the loss formula for positive and negative sample pairs in text structuring can be:

[0128]

[0129] Among them, L 22 The loss P(y) represents the loss of positive and negative sample pairs in the text structure. + |x) represents the correct structure y generated from given text x. + The probability, P(y) - |x) represents the structure y that generates negative samples given text x. - The probability of.

[0130] In some examples, the second pre-trained model can learn the ability to accurately map structured information from the input text using a loss of positive and negative sample pairs of text structure.

[0131] In some examples, the formula for calculating the loss from structured data generation can be:

[0132]

[0133] Among them, L 24Denote the loss of structure data generation, T denote the length of the generated sequence, and y t denote the t-th element in the target structure data sequence, and P θ (y t |y<t,x) denotes the probability that the model parameters θ of the second pre-trained model generate the target sequence y t under the condition x (input text). Additionally, in Equation (9), the matching degree between the generated structure and the target structure (such as entity, relationship, or event label) can be calculated through log-likelihood estimation.

[0134] In some examples, the second pre-trained model can learn to generate a reasonable structure according to the instructions and prompts through the loss of structure data generation. Additionally, the loss of structure data generation can measure the rationality of the generated structure by the matching degree between the generated sequence and the target sequence. For example, the higher the log-likelihood value of the generated target structure data and the real structure data at each time step, the smaller the loss of structure data generation.

[0135] In some examples, the calculation formula for the loss of retrospected semantics can be:

[0136] L 26 = KL(P text ||P gen ), ………… Equation (10)

[0137] where, L 26 denotes the loss of retrospected semantics, P text denotes the probability distribution of the semantic representation of the input text, P gen denotes the probability distribution of the generated structured data, and KL(P text ||P gen ) denotes the KL divergence between the input text distribution and the generated structure distribution.

[0138] Additionally, the definition of KL divergence can be

[0139]

[0140] In some examples, Equation (11) can be used to measure the difference between the generated structure distribution P gen and the original text distribution P text .

[0141] In some examples, the second pre-trained model can reduce the forgetting of the above prompt by the generated structure through the loss of retrospected semantics. For example, by minimizing the KL divergence, the distribution of the generated structured data can be made as close as possible to the distribution of the original text, thereby reducing the model's forgetting of the semantics of the original text.

[0142] In some examples, the loss function of the second pre-trained model can be expressed as:

[0143] L 20 =λ5L 22 +λ6L 24 +λ7L 26 ,…………Formula (12)

[0144] Wherein, λ5, λ6, and λ7 represent the weight hyperparameters of the corresponding loss, which can be adjusted according to the actual needs of the task to balance the importance of different tasks.

[0145] In some examples, in step S136, the UIE model can be fine-tuned using a second training dataset and a second validation dataset to obtain a second model 12. In this case, because the UIE model has excellent generalization performance, after being trained on a text dataset in the field of health management, the UIE model can adapt to the application scenarios of health management, improve the UIE model's ability to generalize to text data in the health management process, and thus improve the accuracy of information extraction from text data in the health management process.

[0146] In some examples, after step S136, the generalization ability of the second model 12 can be tested using a second test dataset.

[0147] In some examples, after obtaining the first model 10 and the second model 12, the first model 10 and the second model 12 can be deployed in an application for use by the target object. Additionally, the application may include at least one of a webpage, a mini-program, and an app.

[0148] Figure 7 This is a schematic diagram illustrating the display of second text data in a timeline manner as described in this disclosure example.

[0149] See back Figure 2 In step S140, entity information can be sorted based on time information to obtain an entity information sequence. In this case, since the entity information is related to the target object's behavior and activities during the health management process, sorting the entity information based on time information allows for the classification and sorting of the specific content of the behavior and activities during the health management process according to the chronological order of their occurrence. This enables clear, detailed, and accurate recording of the behavior and activities during the health management process, providing accurate reference information for the target object's health management.

[0150] In some examples, the entity information sequence can be a sequence of multiple entity information items arranged in chronological order. In other examples, the entity information sequence can be obtained by sorting multiple entity information items corresponding to multiple received voice data items based on time information.

[0151] In some examples, multiple entity information can be sorted according to the chronological order of multiple behavioral activities of the target object during the health management process to obtain an entity information sequence.

[0152] In some examples, information about multiple entities can be sorted in the form of a timeline.

[0153] In some examples, the second text data can be displayed as a check-in table or a timeline. In other words, the second text data can be displayed as a check-in table or a timeline.

[0154] In some examples, displaying the second text data in a timeline format can mean arranging multiple entity information items chronologically from top to bottom. This allows the target audience to view the second text data intuitively and conveniently. Figure 7 This illustration shows the content of the second text data displayed in a timeline format.

[0155] In some examples, displaying the second text data in the form of a check-in table can refer to sorting multiple entity information by time information and then displaying it in a list.

[0156] In this disclosure, when the target object performs health management by checking in via voice, the first model 10 can improve the accuracy of automatically converting voice data into first text data, and the second model 12 can improve the accuracy of extracting target information from the first text data to obtain second text data. When health management is carried out based on the second text data, accurate reference information can be provided to the target object.

[0157] Figure 8 This is a flowchart illustrating a first embodiment of the data processing method involved in the present disclosure.

[0158] In some examples, the analytes involved in this disclosure may be one or more of glucose, acetylcholine, amylase, bilirubin, cholesterol, human chorionic gonadotropin, creatine kinase, creatine, creatine anhydride, DNA, fructosamine, glutamine, growth hormone, hormones, ketone bodies, lactate, oxygen, peroxides, prostate-specific antigen, prothrombin, RNA, thyroid-stimulating hormone, or troponin.

[0159] The following description uses glucose as an example to further illustrate the data processing method involved in this disclosure. It should be noted that, for other analytes, those skilled in the art can easily adapt the data processing method used for glucose to other analytes with minor modifications.

[0160] See in some examples Figure 8The data processing method may further include: acquiring analyte concentration data within the target object's body (step S210); acquiring analyte changes within a preset time period after the behavioral activity represented by the second text data occurred (step S220); and generating guidance suggestions for health management based on the analyte changes (step S230). In this case, by combining the second text data with the analyte concentration data to acquire analyte changes, it is easier to analyze the reasons for fluctuations in the analyte concentration data of the target object within a preset time period after the relevant behavioral activity occurred, thereby providing the target object with practical guidance suggestions for health management and optimizing the target object's self-management level.

[0161] In some examples, in step S210, analyte concentration data within the target body can be acquired. For example, glucose concentration data within the target body can be acquired using a continuous glucose monitoring (CGM) device (see [link to CGM]). Figure 3 For example, glucose concentration data in a target subject can be obtained through a single-point analyte collection device (such as a fingertip blood glucose meter).

[0162] See in some examples Figure 8 In step S220, the changes in analyte within a preset time period after the behavioral activity represented by the second text data can be obtained based on the second text data and the analyte concentration data. In some examples, the changes in analyte within a preset time period after the target object performs a check-in event can be obtained based on the second text data and the analyte concentration data. Specifically, the changes in analyte within a preset time period after delaying the time information in the second text data can be obtained based on the entity information in the second text data and the analyte concentration data.

[0163] In some examples, analyte variation can include whether analyte concentration data exceeds a target range within a preset time period. Furthermore, the preset time period can be determined based on the entity category. Additionally, the target range can be determined based on the analyte category.

[0164] For example, based on the food consumed by the target subject during eating (entity name information), the quantity of food (entity quantity), and glucose concentration data, it can be determined whether the glucose concentration data within two hours (a preset time period) after the target subject's meal exceeds the target range. Here, the target range for glucose concentration can be 4.4-7.8 mmol / L.

[0165] Additionally, in some examples, the analyte variation may include whether the rate of change in analyte concentration exceeds a preset value within a preset time period. Furthermore, the preset value can be determined based on the type of analyte.

[0166] In some examples, in step S230, guidance recommendations for health management can be generated based on changes in the analyte.

[0167] In some examples, in step S230, in response to the analyte concentration data exceeding the target range, a first guidance recommendation for health management can be generated, which can be used to suggest the target subject's next improvement activities.

[0168] For example, if a target's glucose concentration exceeds the target range within two hours after breakfast, it may be because they ate an extra bowl of plain porridge, which contains rice dextrin, causing a rapid increase in glucose concentration beyond the target range. In this case, the first guidance suggestion can be generated: suggest that the target replace the plain porridge with foods with a lower glycemic index or eat only half a bowl of plain porridge.

[0169] For example, if the target's glucose concentration exceeds the target range within two hours after exercise, it may be because the exercise duration or intensity is too high, causing the glucose concentration to drop rapidly beyond the target range. In this case, the first guidance suggestion can be generated: suggest that the target shorten the exercise duration or reduce the exercise intensity.

[0170] In some examples, in step S230, in response to the analyte concentration data being within the target range, a second guidance recommendation for health management can be generated, which can be used to encourage the target subject to maintain behavioral activity.

[0171] For example, if the target subject's glucose concentration data remains within the target range for two hours after breakfast, a second guidance suggestion can be generated: encourage the target subject to maintain this breakfast routine.

[0172] Figure 9 This is a flowchart illustrating a second embodiment of the data processing method involved in the present disclosure. Figure 10 This is a graph showing the trend of analyte concentration data over time as described in this disclosure example.

[0173] See in some examples Figure 9 The data processing method may further include: acquiring analyte concentration data within the target object (step S310); displaying a trend graph of analyte concentration data changing over time (step S320); and marking entity information in the trend graph based on time information (step S330).

[0174] In some examples, in step S310, analyte concentration data within the target organism can be obtained. See the relevant description in step S210 for details.

[0175] In some examples, in step S320, a trend graph of the analyte concentration data over time can be displayed (see [reference]). Figure 10 In this context, displaying a trend graph of analyte concentration data over time allows the target population to intuitively and clearly view the changing trends of analyte concentration data after behavioral activities occur during health management. For example, the trend graph of analyte concentration data over time can be displayed on a mobile smart terminal (e.g., a mobile phone) application (e.g., an app).

[0176] In some examples, trend graphs may include curves showing the fluctuation of analyte concentration data over time. Figure 10 The diagram illustrates the fluctuation curve of glucose concentration data over time.

[0177] In some examples, in step S330, entity information can be labeled in the trend graph based on time information in the entity information. That is, entity information in the second text data can be labeled in the trend graph based on time information in the second text data. In this case, by labeling entity information in the second text data in the trend graph, the entity information can be matched with analyte concentration data, thereby facilitating accurate analysis of the reasons for fluctuations in analyte concentration data within a preset time period after the occurrence of relevant behavioral activities of the target object.

[0178] See in some examples Figure 10 Geometric pattern A can be marked on the fluctuation curve along the time axis of the trend chart based on the coordinates corresponding to the time information in the second text data. Geometric pattern A can be configured to open upon triggering (e.g., a single click or double click), and can display entity information after being opened. For example, in... Figure 10 In the example shown, when geometric pattern A is opened, it can display entity information: "09:00 Breakfast: Two bowls of white porridge, one salted duck egg".

[0179] Figure 11 This is a flowchart illustrating the data processing methods involved in the examples of this disclosure.

[0180] The execution flow of the data processing method involved in this disclosure will now be described in conjunction with the first model 10 and the second model 12.

[0181] See Figure 11The data processing method disclosed herein may include: inputting voice data of the target object into the first model 10 (step S410); the first model 10 converting the voice data into first text data (step S420); the second model 12 receiving the first text data and extracting target information from the first text data to obtain second text data (step S430); displaying the second text data in the form of a check-in table or timeline (step S440); and generating guidance recommendations for health management based on the second text data and analyte concentration data (step S450).

[0182] In addition, this disclosure also relates to a voice data processing system (also referred to as a processing system), which may include at least one processing circuit that can be configured to perform the data processing method disclosed herein.

[0183] While the present disclosure has been specifically described above in conjunction with the accompanying drawings and examples, it is to be understood that the foregoing description does not limit the present disclosure in any way. Those skilled in the art can make modifications and variations to the present disclosure as needed without departing from its essential spirit and scope, and all such modifications and variations shall fall within the scope of the present disclosure.

Claims

1. A method of processing voice data, which is a method for processing voice data generated in a health management process, characterized by, The method comprises: receiving voice data of a target object; automatically converting the voice data into first text data using a first model, the first model is obtained by fine-tuning a first pre-trained model based on a Transformer model using a voice data set, the voice data set is derived from a health management process; extracting target information in the first text data using a second model to obtain second text data, the second text data comprising entity information, wherein the entity information is related to the behavior activities of the target object in the health management process and comprises entity name information, entity quantity and time information, the second model is obtained by fine-tuning a second pre-trained model based on a Transformer model using a text data set, the text data set is derived from a health management process; and sorting the entity information based on the time information to obtain an entity information sequence.

2. The method of processing voice data according to claim 1, wherein the entity information further comprises an entity category, and the entity category comprises at least one of diet, medication, exercise and single-point analyte detection.

3. The method of processing voice data according to claim 1, further comprising: obtaining analyte concentration data in the target object; based on the second text data and the analyte concentration data, obtaining analyte changes in a preset time period after the behavior activities represented by the second text data; and generating a guidance suggestion for health management based on the analyte changes.

4. The method of processing voice data according to claim 1, further comprising: obtaining analyte concentration data in the target object; displaying a trend graph of the analyte concentration data over time; and based on the time information in the entity information, marking the entity information in the trend graph.

5. The method of processing voice data according to any one of claims 1 to 4, wherein the voice data is obtained by receiving a check-in event from the target object.

6. The method of processing voice data according to claim 1, wherein the first pre-trained model is a Whisper model, and / or the second pre-trained model is a UIE model.

7. The method of processing voice data according to claim 1, wherein the first pre-trained model is a multi-task prediction model, and the multi-task prediction comprises predicting whether the voice data is human voice, predicting the language used in the voice data, predicting the start time and end time of the human voice in the voice data, and predicting the text of the human voice in the voice data.

8. The method of processing voice data according to claim 1, wherein the total duration of the voice data set is not less than 12 hours, and / or the data amount of the text data set is not less than 70,000.

9. The method according to claim 1, wherein the loss function of the first pre-trained model is: L 10 = L 12 + L 14 + L 16 + L 18 , wherein, L 10 represents a loss function of the first pre-training model, L 12 a loss for verifying whether the voice data has human voice, L 14 a loss for verifying the language used by the voice data, L 16 a loss for verifying the start time and end time of the human voice in the voice data, L 18 a loss for verifying the text of the human voice in the voice data; and / or the loss function of the second pre-trained model is: L 20 = L 22 + L 24 + L 26 , wherein, L 20 represents the loss function of the second pre-training model, L 22 represents the loss of the positive and negative sample pair of text structuring, L 24 represents the loss of structure data generation, L 26 represents the loss of review semantics.

10. A system for processing voice data, characterized by comprising at least one processing circuitry configured to perform the processing method of any one of claims 1 to 9.