Data information generation method and device, computer device, and storage medium
Patent Information
- Application Number
- CN202210956026.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-10
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-08-10
AI Technical Summary
然而,然而穷举规则依赖于领域知识,并且依赖于外部知识图谱的构建,且对于不同领域的病患通常所书写的身体状态档案具有对应的书写规律,即不同的领域的医务人员具有不同的身体状态档案描述逻辑,因此采用穷举规则所生成的身体状态档案,内容较为生硬且描述逻辑不连贯
[0061]上述数据信息生成方法、装置、计算机设备、存储介质和计算机程序产品,先获取待处理数据,并基于待处理数据得到特征向量,再基于特征向量得到数据信息隐变量以及撰写风格隐变量,并基于数据信息隐变量得到预测数据信息,并基于撰写风格隐变量得到预测撰写风格,预测撰写风格用于描述预测数据信息的内容描述逻辑,从而根据预测数据信息以及预测撰写风格,生成待处理数据对应的目标数据信息。在考虑到预测数据信息的基础上,进一步考虑到预测撰写风格,因此能够提升数据信息生成的灵活性。其次,通过预测撰写风格描述预测数据信息的内容描述逻辑,因此根据预测数据信息以及预测撰写风格所生成的目标数据信息,能够保证目标数据信息的描述逻辑连贯性。
Smart Images

Figure CN115312151B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer and communication technology, and in particular to a data information generation method, apparatus, computer equipment, and storage medium. Background Technology
[0002] With the development of internet technology and the proliferation of various internet applications, in internet hospitals, specifically in scenarios like health status inquiries, after a patient inputs information describing their health status, the system first collects information from multiple dimensions of that description and then outputs a health status profile. This profile helps patients understand their own health and better articulate their condition, thus improving communication efficiency during consultations. Secondly, it also helps doctors understand a patient's basic situation before seeing them, reducing repetitive health status inquiries, avoiding missed questions, and ultimately improving the efficiency of health status inquiries.
[0003] Currently, physical condition inquiry systems primarily use rule-based methods to output physical condition profiles based on collected information. Specifically, this involves exhaustively enumerating all rules applicable to the patient's physical condition profile information (including relevant descriptions) and customizing corresponding rule templates to generate the patient's corresponding physical condition profile. However, exhaustive rule enumeration relies on domain knowledge and the construction of external knowledge graphs. Furthermore, physical condition profiles typically exhibit different writing patterns for patients from different domains; that is, medical personnel from different domains have different descriptive logics for physical condition profiles. Therefore, physical condition profiles generated using exhaustive rules tend to be rigid in content and lack coherent descriptive logic. Thus, improving the flexibility and coherence of descriptive logic in physical condition profile generation is a pressing issue that needs to be addressed. Summary of the Invention
[0004] Therefore, it is necessary to provide a data information generation method, apparatus, computer equipment, and storage medium that can improve the flexibility of data information generation and the logical coherence of description, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for generating data information. The method includes:
[0006] Acquire the data to be processed and obtain the feature vector based on the data to be processed;
[0007] Latent variables of data information and writing style are obtained based on feature vectors;
[0008] Predicted data information is obtained based on latent variables of data information, and predicted writing style is obtained based on latent variables of writing style. The predicted writing style is used to describe the content description logic of the predicted data information.
[0009] Based on the predicted data and the predicted writing style, the target data information corresponding to the data to be processed is generated.
[0010] In one embodiment, obtaining a feature vector based on the data to be processed includes:
[0011] The data to be processed is encoded by the encoding network in the data information generation model to obtain feature vectors. The data information generation model is obtained by adjusting the prior network in the initial data information generation model based on the first data information, the second data information, the first writing style, and the second writing style. The first data information is obtained based on the latent variables of the sample data, the first writing style is obtained based on the latent variables of the writing style of the sample data, and the latent variables of the data information and the writing style of the sample data are obtained based on the feature vectors of the sample data. The sample data is generated based on the second data information and the second writing style.
[0012] In one embodiment, latent variables of data information and latent variables of writing style are obtained based on feature vectors, including:
[0013] The prior network in the data information generation model is based on feature vectors to obtain latent variables of data information and writing style.
[0014] In one embodiment, the prior network includes a first prior network and a second prior network;
[0015] The prior network in the data information generation model, based on feature vectors, obtains latent variables of data information and writing style, including:
[0016] The first prior network calculates the prior distribution of the data information based on the feature vector, and obtains the first prior distribution corresponding to the data information.
[0017] The second prior network calculates the prior distribution of writing style based on feature vectors, and obtains the second prior distribution corresponding to the writing style.
[0018] Latent variables of data information are obtained from the first prior distribution, and latent variables of writing style are obtained from the second prior distribution.
[0019] In one embodiment, predictive data information is obtained based on latent variables of data information, and predictive writing style is obtained based on latent variables of writing style, including:
[0020] The classification network in the data information generation model predicts writing style based on latent variables of writing style.
[0021] The decoding network in the data information generation model obtains predicted data information based on the latent variables of the data information.
[0022] In one embodiment, the classification network in the data information generation model obtains a predicted writing style based on latent variables of writing style, including:
[0023] The classification network in the data information generation model is based on the latent variable of writing style to obtain the probability that the data to be processed belongs to each writing style;
[0024] Based on the probability that the data to be processed belongs to each writing style, the predicted writing style is determined.
[0025] In one embodiment, the training process of the data information generation model includes:
[0026] Based on the first data information, the second data information, and the first prior distribution, adjust the first prior network in the initial data information generation model;
[0027] Based on the first and second writing styles, as well as the second prior distribution, adjust the second prior network in the initial data information generation model.
[0028] In one embodiment, the training process of the data information generation model includes:
[0029] The first posterior network in the initial data information generation model calculates the posterior distribution based on the first data information and the second data information to obtain the first posterior distribution corresponding to the data information.
[0030] The second posterior network in the initial data information generation model calculates the posterior distribution based on the first writing style and the second writing style to obtain the second posterior distribution corresponding to the writing style.
[0031] Based on the first data information and the second data information, the first prior distribution and the first posterior distribution, adjust the first prior network in the initial data information generation model;
[0032] Based on the first and second writing styles, the second prior distribution, and the second posterior distribution, adjust the second prior network in the initial data information generation model.
[0033] In one embodiment, the training process of the data information generation model includes:
[0034] The decoding network in the initial data information generation model obtains the first data information of the sample data based on the latent variables of the sample data.
[0035] Based on the first data information and the second data information, the first prior distribution and the first posterior distribution, adjust the first prior network in the initial data information generation model, including:
[0036] Based on the first data information and the second data information, the first prior distribution and the first posterior distribution, adjust the first prior network and the decoding network in the initial data information generation model.
[0037] In one embodiment, predicting writing style based on latent writing style variables includes:
[0038] The classification network in the initial data information generation model obtains the predicted writing style of the sample data based on the latent variable of writing style in the sample data.
[0039] Based on the first and second writing styles, the second prior distribution, and the second posterior distribution, adjust the second prior network in the initial data information generation model, including:
[0040] Based on the first and second writing styles, the second prior distribution, and the second posterior distribution, adjust the second prior network and the classification network in the initial data information generation model.
[0041] Secondly, this application also provides a data information generation apparatus. The apparatus includes:
[0042] The acquisition module is used to acquire the data to be processed and obtain the feature vector based on the data to be processed;
[0043] The latent variable acquisition module is used to obtain latent variables of data information and writing style latent variables based on feature vectors;
[0044] The prediction module is used to obtain predicted data information based on latent variables of data information, and to obtain predicted writing style based on latent variables of writing style. The predicted writing style is used to describe the content description logic of the predicted data information.
[0045] The generation module is used to generate target data information corresponding to the data to be processed based on the predicted data information and the predicted writing style.
[0046] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0047] Acquire the data to be processed and obtain the feature vector based on the data to be processed;
[0048] Latent variables of data information and writing style are obtained based on feature vectors;
[0049] Predicted data information is obtained based on latent variables of data information, and predicted writing style is obtained based on latent variables of writing style. The predicted writing style is used to describe the content description logic of the predicted data information.
[0050] Based on the predicted data and the predicted writing style, the target data information corresponding to the data to be processed is generated.
[0051] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0052] Acquire the data to be processed and obtain the feature vector based on the data to be processed;
[0053] Latent variables of data information and writing style are obtained based on feature vectors;
[0054] Predicted data information is obtained based on latent variables of data information, and predicted writing style is obtained based on latent variables of writing style. The predicted writing style is used to describe the content description logic of the predicted data information.
[0055] Based on the predicted data and the predicted writing style, the target data information corresponding to the data to be processed is generated.
[0056] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0057] Acquire the data to be processed and obtain the feature vector based on the data to be processed;
[0058] Latent variables of data information and writing style are obtained based on feature vectors;
[0059] Predicted data information is obtained based on latent variables of data information, and predicted writing style is obtained based on latent variables of writing style. The predicted writing style is used to describe the content description logic of the predicted data information.
[0060] Based on the predicted data and the predicted writing style, the target data information corresponding to the data to be processed is generated.
[0061] The aforementioned data information generation method, apparatus, computer equipment, storage medium, and computer program product first acquire the data to be processed, and then obtain a feature vector based on the data. Next, they obtain latent variables for data information and latent variables for writing style based on the feature vectors. Based on the latent variables for data information, they obtain predicted data information, and based on the latent variables for writing style, they obtain a predicted writing style. The predicted writing style is used to describe the content description logic of the predicted data information. Therefore, based on the predicted data information and the predicted writing style, the target data information corresponding to the data to be processed is generated. By considering both the predicted data information and the predicted writing style, the flexibility of data information generation is improved. Furthermore, by describing the content description logic of the predicted data information through the predicted writing style, the logical coherence of the description of the target data information generated based on the predicted data information and the predicted writing style can be guaranteed. Attached Figure Description
[0062] Figure 1 This is an application environment diagram of a data information generation method in one embodiment;
[0063] Figure 2 This is a flowchart illustrating a data information generation method in one embodiment;
[0064] Figure 3 This is a schematic diagram of an embodiment of the data to be processed in one example;
[0065] Figure 4 This is a schematic diagram illustrating the decoupling of writing style and data information in one embodiment.
[0066] Figure 5 This is a schematic diagram of an embodiment of the target data information corresponding to the data to be processed in one embodiment;
[0067] Figure 6 This is a flowchart illustrating the process of obtaining a feature vector based on the data to be processed in one embodiment.
[0068] Figure 7 This is a schematic diagram of the structure of the coding network in one embodiment;
[0069] Figure 8 This is a partial flowchart illustrating the process of obtaining latent variables of data information and writing style based on feature vectors in one embodiment.
[0070] Figure 9 This is a partial flowchart illustrating the process of obtaining latent variables of data information and writing style based on feature vectors in another embodiment.
[0071] Figure 10This is a partial flowchart illustrating the process of obtaining predicted data information based on latent variables of data information and obtaining predicted writing style based on latent variables of writing style in one embodiment.
[0072] Figure 11 This is a schematic diagram of the decoding network structure in one embodiment;
[0073] Figure 12 This is a partial flowchart illustrating the process of obtaining predicted writing style in one embodiment;
[0074] Figure 13 This is a partial flowchart illustrating the training process of a data information generation model in one embodiment.
[0075] Figure 14 This is a partial flowchart illustrating the training process of the data information generation model in another embodiment;
[0076] Figure 15 This is a partial flowchart illustrating the training process of the data information generation model in another embodiment;
[0077] Figure 16 This is a partial flowchart illustrating the training process of the data information generation model in another embodiment;
[0078] Figure 17 This is a schematic diagram of the complete process of a data information generation method in one embodiment;
[0079] Figure 18 This is a structural block diagram of a data information generation device in one embodiment;
[0080] Figure 19 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0081] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0082] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Based on the cloud computing business model, cloud technology encompasses network technology, information technology, integration technology, management platform technology, and application technology. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.
[0083] The solutions provided in this application relate to Artificial Intelligence as a Service (AIaaS) in cloud technology. AIaaS is also commonly referred to as "AI as a Service." This is a mainstream service model for artificial intelligence platforms. Specifically, AIaaS platforms break down several common AI services and provide them as independent or packaged services in the cloud. This service model is similar to opening an AI-themed marketplace: all developers can access and use one or more AI services provided by the platform through API interfaces. Some experienced developers can also use the AI framework and AI infrastructure provided by the platform to deploy and maintain their own dedicated cloud AI services. The following embodiments illustrate this further:
[0084] The data information generation method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another server.
[0085] Specifically, taking server 104 as an example, before determining the data information, a data information generation model needs to be trained. Terminal 102 can send instructions to server 104 to train the model, or server 104 can directly start model training; this is not limited here. If target data information needs to be generated, server 104 can obtain the data to be processed from the data storage system, or obtain the data to be processed through communication with terminal 102; this is not limited here either. Based on this, server 104 calls the encoding network in the trained data information generation model to encode the data to be processed, obtaining feature vectors. The prior network in the data information generation model obtains latent variables for data information and latent variables for writing style based on the feature vectors. Then, predicted data information is obtained based on the latent variables for data information, and predicted writing style is obtained based on the latent variables for writing style. Finally, the target data information corresponding to the data to be processed is generated based on the predicted data information and predicted writing style.
[0086] Secondly, taking a high-computing-power terminal 102 as an example, before determining the data information, a data information generation model needs to be trained. The terminal 102 can train the model itself or obtain it through communication with the server 104; this is not limited here. Based on this, the terminal 102 obtains the data to be processed and encodes it using the encoding network in the data information generation model to obtain feature vectors. The prior network in the data information generation model then uses these feature vectors to obtain latent variables for data information and writing style. Based on these latent variables, it obtains predicted data information and a predicted writing style. Finally, based on the predicted data information and the predicted writing style, it generates the target data information corresponding to the data to be processed.
[0087] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and aircraft. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers. This invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0088] Furthermore, the data information generation method provided in this application embodiment can be applied to scenarios such as online real-time physical status inquiry and electronic physical status profile generation. Taking the scenario of online real-time physical status inquiry as an example, the physical status description information input by the patient and the response information input by the doctor in the online physical status inquiry can constitute data to be processed. Then, the data to be processed (i.e., the physical status description information input by the patient and the response information input by the doctor) is input into the data information generation model. The encoding network in the data information generation model encodes the data to be processed to obtain feature vectors. The prior network in the data information generation model obtains data information latent variables and writing style latent variables based on the feature vectors. Based on this, the classification network in the data information generation model obtains the predicted writing style based on the writing style latent variables, and the decoding network in the data information generation model obtains the predicted data information based on the data information latent variables. Thus, based on the predicted data information and the predicted writing style, the target data information corresponding to the data to be processed is generated. The target data information is the physical status profile report corresponding to the physical status description information input by the patient and the response information input by the doctor in the online physical status inquiry.
[0089] Secondly, taking the application in generating electronic health records as an example, the relevant information about the patient's input description of their health status is identified as the data to be processed. This data (i.e., the relevant information about the health status description) is then input into a data generation model. The encoding network in the model encodes the data to obtain feature vectors. Based on these feature vectors, the prior network in the model obtains latent variables for data information and writing style. Based on these latent variables, the classification network in the model predicts the writing style, and the decoding network obtains predicted data information. Thus, based on the predicted data information and the predicted writing style, the target data information corresponding to the data to be processed is generated. This target data information is the electronic health record corresponding to the relevant information about the patient's input description of their health status. It should be understood that the above example is only for understanding this solution and should not be construed as a limitation of this solution.
[0090] Based on this, in one embodiment, such as Figure 2 As shown, a data information generation method is provided, which can be applied to... Figure 1 Taking a terminal as an example, it can be understood that this method can also be applied to a server, and to a system that includes both a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0091] Step 202: Obtain the data to be processed and obtain the feature vector based on the data to be processed.
[0092] The data to be processed includes either text data or image data. Furthermore, the feature vector can be a text feature vector, an image feature vector, or a multimodal feature vector that includes both text and image feature vectors.
[0093] Specifically, the terminal can obtain data to be processed through real-time user input, local storage, or communication with a server. The method of obtaining the data is not limited here. For example, in a scenario involving the generation of electronic health records, the data to be processed would be information related to the patient's description of their health status. Secondly, in a scenario involving online health status inquiries, the data to be processed would be the patient's description of their health status and the doctor's response.
[0094] Based on this, the terminal encodes the data to be processed to obtain a feature vector, which is a low-dimensional vector. Furthermore, since the data to be processed includes either text data or image data, when the data is text data, the feature vector obtained based on the text data is specifically a text feature vector. Similarly, when the data is image data, the feature vector obtained based on the text data is specifically an image feature vector. And when the data includes both text and image data, the feature vector obtained based on both text and image data is specifically a multimodal feature vector including both text and image feature vectors. The specific feature vector needs to be determined based on the data to be processed; no limitation is made here.
[0095] To facilitate understanding, we will use the generation of electronic body status profiles in the intelligent body status inquiry process as an example for illustration. Figure 3 As shown, Figure 3 In Figure (A), the intelligent body status inquiry process typically begins with the patient providing basic information (e.g., information describing their body status 301, gender, and age 302). After obtaining this basic information, the intelligent body status inquiry system sequentially inquires about the patient's current body status description 303. Figure 3 In Figure (B), the intelligent body status inquiry system further inquires about the patient's past body status description information 304 and body status reflection information 305, and finally summarizes the data to be processed, which includes relevant information on body status description 301, gender and age information 302, current body status description information 303, past body status description information 304 and body status reflection information 305.
[0096] Step 204: Obtain latent variables of data information and writing style based on feature vectors.
[0097] Specifically, the terminal performs prior distribution calculation on the data information based on the feature vector, and performs prior distribution calculation on the writing style based on the feature vector, thereby determining the latent variables of the data information based on the prior distribution calculation results of the data information, and determining the latent variables of the writing style based on the prior distribution calculation results of the writing style.
[0098] Furthermore, the aforementioned data information refers to the specific data information included in the data to be processed. For example, if the data to be processed is the text data "coughing and feeling unwell," then the data information is the text data "coughing and feeling unwell." Secondly, if the data to be processed is image data and "blunt force injury," then the data information includes image data and "blunt force injury."
[0099] Step 206: Obtain predicted data information based on the latent variables of data information, and obtain predicted writing style based on the latent variables of writing style. The predicted writing style is used to describe the content description logic of the predicted data information.
[0100] The predicted data information includes either textual or image information. Secondly, the prediction writing style describes the content description logic of the predicted data information. For example, the writing style can include, but is not limited to, a preset rule writing style, a specialist physician writing style, and a medical staff writing style. A preset rule writing style could be: using preset punctuation marks and transitional text to describe each piece of information in the predicted data information. A specialist physician writing style could be: describing each piece of information in the predicted data information using professional language, and could also consider the writing styles of different specialist physicians, incorporating commonly used language and transitional phrases.
[0101] Specifically, the terminal obtains the predicted data information corresponding to the data to be processed based on the latent variables of the data information. Next, it classifies the writing styles corresponding to the data to be processed based on the latent variables of the style, then obtains the probability that the data to be processed belongs to each writing style, and finally determines the predicted writing style corresponding to the data to be processed based on the probability of the data to be processed belonging to each writing style. It should be understood that the predicted writing style can include one or more writing styles; this is not limited here.
[0102] To facilitate understanding of the decoupled architecture between writing style and data information, such as Figure 4As shown, the data to be processed 401 is first encoded to obtain feature vector 402. Then, based on feature vector 402, the prior distribution of the data information in the data to be processed is calculated, and the latent variable 4031 of the data information is determined based on the result of the prior distribution calculation of the data information. Similarly, the prior distribution of the writing style is calculated based on feature vector 402, and the latent variable 4032 of the data information is determined based on the result of the prior distribution calculation of the data information. Then, the predicted data information 4041 is obtained based on the latent variable 4031 of the data information, and the predicted writing style 4042 is obtained based on the latent variable 4032 of the writing style, so as to achieve the purpose of decoupling the writing style from the data information.
[0103] Step 208: Generate target data information corresponding to the data to be processed based on the predicted data information and the predicted writing style.
[0104] Specifically, the terminal describes the predicted data information and the predicted writing style obtained in step 208. Since the predicted writing style is used to describe the content description logic of the predicted data information, the terminal describes the content of the predicted data information according to the content description logic of the predicted writing style. This usually involves adding punctuation marks and transition text to the predicted data information, thereby generating the target data information corresponding to the data to be processed, and making the target data information more flexible in terms of content description logic.
[0105] Furthermore, if applied to the generation of electronic body status profiles, the data to be processed includes information related to the patient's input description of their body status, current body status description, past body status description, and body status reflection information. Then, using the data information generation method provided in this solution, an electronic body status profile corresponding to the relevant information such as the description of the body status, current body status description, past body status description, and body status reflection information is generated. In practical applications, the generated electronic body status profile can also be pushed to doctors, allowing them to understand the patient's body status before formally treating the patient, thereby improving medical efficiency.
[0106] Secondly, when applied to online health status inquiries, doctors can engage in real-time, multi-round interactions with patients, asking questions about symptoms and examinations. This allows doctors to predict the patient's illness based on the information gathered during the inquiry process, providing real-time information about the possible disease type and offering medication or examination suggestions. The terminal can also identify the information obtained from the real-time doctor-patient inquiries as data to be processed. After the patient completes the online health status inquiry, a health status profile summary corresponding to the information obtained from the real-time doctor-patient inquiries is generated and sent to both the doctor and the patient to facilitate subsequent offline medical visits and other diagnostic and treatment procedures.
[0107] To facilitate understanding of this solution, we will again consider the scenario of generating electronic body status profiles in the intelligent body status inquiry process, and based on... Figure 3 Taking the example shown, firstly, referring to Figure (3) again, we can obtain the data to be processed, including information related to body status description 301, gender and age information 302, current body status description information 303, past body status description information 304, and body status reflection information 305. Through the aforementioned method, we can obtain data such as... Figure 5 The electronic body status profile shown (i.e., the target data information corresponding to the data to be processed) includes relevant information 301 based on the description of the body status and body status data information 501 obtained from the collected information on "duration of cough occurrence", and the body status data information 501 specifically describes "the patient coughed for 1 day".
[0108] Similarly, the electronic health status file also includes current health status description data 502 obtained based on current health status description information 303, specifically describing "current health status description data 502: patient 1 day ago xxxxxxxxxxxx". It also includes past health status description data 503 obtained based on past health status description information 304, health status reflection data 504 obtained based on health status reflection information 305, and gender and age information 505, specifically describing "patient denies health status 1, denies health status 2, denies health status 3, denies health status 4"; and health status reflection data 504 specifically describing "patient denies health status reflection information 1, denies health status reflection information 2, and denies health status reflection information 3". It should be understood that the foregoing examples are for understanding this solution and should not be construed as limiting this solution.
[0109] The aforementioned data generation method considers both predicted data information and predicted writing style, thus enhancing the flexibility of data generation. Secondly, by describing the content description logic of the predicted data information through the predicted writing style, the logical coherence of the target data information generated based on the predicted data information and the predicted writing style can be guaranteed.
[0110] Figure 2 The illustrated embodiment describes the following: The terminal needs to encode the data to be processed to obtain a feature vector. The specific implementation method for using a model to encode the data to be processed to obtain the feature vector will be described below:
[0111] In one embodiment, such as Figure 6 As shown, the feature vector obtained based on the data to be processed includes:
[0112] Step 602: The data to be processed is encoded through the encoding network in the data information generation model to obtain feature vectors. The data information generation model is obtained by adjusting the prior network in the initial data information generation model based on the first data information and the second data information, as well as the first writing style and the second writing style. The first data information is obtained based on the latent variables of the data information of the sample data, the first writing style is obtained based on the latent variables of the writing style of the sample data, and the latent variables of the data information and the writing style of the sample data are obtained based on the feature vectors of the sample data. The sample data is generated based on the second data information and the second writing style.
[0113] Before encoding the data to be processed using a model, a data information generation model needs to be trained first. Based on this, the training method for the data information generation model includes: first, acquiring sample data, which is generated from second data information and a second writing style. The second data information comprises the specific data information included in the sample data, while the second writing style describes the content description logic of the second data information. Then, the sample data is input into the initial data information generation model, and the encoding network within the initial data information generation model encodes the sample data to obtain the corresponding feature vector.
[0114] Then, the first prior network in the model is generated from the initial data information. Based on the feature vectors corresponding to the sample data, the prior distribution of the sample data information is calculated, resulting in the prior distribution of the sample data information. The latent variables of the sample data information are then obtained from this prior distribution. Similarly, the second prior network in the model is generated from the initial data information. Based on the feature vectors corresponding to the sample data, the prior distribution of the writing style of the sample data is calculated, resulting in the prior distribution of the writing style of the sample data. The latent variables of the writing style of the sample data are then obtained from this prior distribution.
[0115] Then, the classification network in the initial data information generation model obtains the first writing style of the sample data based on the writing style latent variables of the sample data, and the decoding network in the initial data information generation model obtains the first data information of the sample data based on the data information latent variables of the sample data. Therefore, based on the first and second data information, as well as the first and second writing styles, the prior network in the initial data information generation model is adjusted to obtain the data information generation model.
[0116] Based on this, after acquiring the data to be processed, the terminal calls the data information generation model trained in the above manner, and specifically uses the encoding network in the data information generation model to encode the data to be processed to obtain feature vectors. Specifically, the encoding network is used to convert either text data or image data into feature vectors.
[0117] Specifically, the encoding network is an encoder using a self-attention mechanism, and more specifically, a simplified bidirectional auto-regressive encoder (BART). Based on this, for ease of understanding the encoding network, as follows... Figure 7 As shown, the encoding network is specifically composed of a self-attention mechanism layer 701 and a first normalization layer 702, followed by a cascaded forward fully connected network 703 and a second normalization layer 704. Therefore, when the data to be processed is input into the encoding network, the encoding network encodes the data to be processed through the self-attention mechanism layer 701, the first normalization layer 702, the fully connected network 703, and the second normalization layer 704, thereby converting the data to be processed into a feature vector.
[0118] In this embodiment, since the data information generation model is obtained through iterative training, the data to be processed is encoded through the encoding network in the data information generation model. This not only ensures the specific feasibility of the data information generation method, but also further guarantees the reliability of the obtained feature vector.
[0119] In one embodiment, such as Figure 8 As shown, latent variables of data information and writing style are obtained based on feature vectors, including:
[0120] Step 802: The prior network in the data information generation model obtains the latent variables of data information and the latent variables of writing style based on the feature vector.
[0121] Specifically, the terminal inputs the data to be processed into the data information generation model. The encoding network in the data information generation model first encodes the data to be processed to obtain feature vectors. Then, the encoding network inputs the feature vectors into the prior network. The prior network calculates the prior distribution of the feature vectors. Based on the obtained prior distribution calculation results, the latent variables of the data information and the latent variables of the writing style are determined.
[0122] The following section will explain in detail how to use prior networks to calculate prior distributions in order to determine latent variables of data information and writing style:
[0123] In one embodiment, such as Figure 9 As shown, the prior network includes a first prior network and a second prior network. Based on this, in step 802, the prior network in the data information generation model obtains the data information latent variables and writing style latent variables based on the feature vectors, including:
[0124] Step 902: The first prior network calculates the prior distribution of the data information based on the feature vector to obtain the first prior distribution corresponding to the data information.
[0125] The first prior network is used to calculate the prior distribution of the data. The prior distribution is a probability distribution that determines the cause based on historical patterns, and specifically, the first prior distribution is a Gaussian distribution. Secondly, the parameters for calculating the prior distribution of the data are obtained through knowledge distillation. Knowledge distillation refers to using the parameters for calculating the posterior distribution of the data to guide the training of the parameters for calculating the prior distribution, ensuring consistency between the parameters for calculating the prior and posterior distributions. This ensures that the error between the obtained posterior and prior distributions of the data is within a preset error range.
[0126] Specifically, since the prior network includes a first prior network and a second prior network, and based on step 802, the encoding network inputs the feature vector into the first prior network, and the first prior network calculates the prior distribution of the data information based on the feature vector, thereby obtaining the first prior distribution corresponding to the data information.
[0127] For ease of understanding, the aforementioned first prior distribution can refer to a Gaussian distribution where the first data information mapping vector has a mean and the second data information mapping vector has a variance. Based on this, see formulas (1) to (3):
[0128] μ c =W μc *h+b μc (1)
[0129] σ c =W σc *h+b σc (2)
[0130] P c =μ c +σ c ⊙∈ (3)
[0131] Where h is the feature vector, W μc W σc b μc and b σc For the parameters used in prior distribution calculation (i.e., the parameters of the first prior network), μ c Let σ be the first data information mapping vector. c P is the second data information mapping vector. c It is the first prior distribution.
[0132] Therefore, it can be seen that the first prior network, based on formula (1), maps the feature vector to the first data information mapping vector μ corresponding to the data information. c Secondly, the first prior network, based on formula (2), maps the feature vector to the second data information mapping vector σ corresponding to the data information. c Therefore, as shown in formula (3), the first prior distribution P c It is to map the first data information to the vector μ c The mean is the second data information mapping vector σ. c A Gaussian distribution with variance N(μ) c , σ c ).
[0133] Step 904: The second prior network calculates the prior distribution of writing style based on the feature vector to obtain the second prior distribution corresponding to the writing style.
[0134] The second prior network is used to calculate the prior distribution of writing style. The prior distribution is a probability distribution that determines the cause based on historical patterns, and specifically, the second prior distribution is a Gaussian distribution. Furthermore, the parameters for calculating the prior distribution of writing style are obtained through knowledge distillation. Knowledge distillation refers to using the parameters for calculating the posterior distribution of writing style during training to guide the training of the parameters for calculating the prior distribution of writing style, ensuring consistency between the parameters for calculating the prior and posterior distributions of writing style. This ensures that the error between the obtained posterior distribution and the obtained prior distribution of writing style is within a preset error range.
[0135] Specifically, since the prior network includes a first prior network and a second prior network, and based on step 802, the encoding network inputs the feature vector into the second prior network, and the second prior network calculates the prior distribution of the writing style based on the feature vector, thereby obtaining the second prior distribution corresponding to the writing style.
[0136] For ease of understanding, the aforementioned second prior distribution can refer to a Gaussian distribution where the first writing style mapping vector has a mean and the second writing style mapping vector has a variance. Based on this, see formulas (4) to (6):
[0137] μ s =W μs *h+b μs (4)
[0138] σ s =W σs *h+b σs (5)
[0139] P s =μ s +σ s ⊙∈ (6)
[0140] Where h is the feature vector, W μs W σs b μs and b σs For the parameters used in prior distribution calculation (i.e., the parameters of the second prior network), μ s For the first writing style mapping vector, σ s For the second writing style mapping vector, P s This is the second prior distribution.
[0141] Therefore, it can be seen that the second prior network, based on formula (4), maps the feature vector to the first writing style mapping vector μ corresponding to the writing style. s Secondly, the first prior network, based on formula (5), maps the feature vectors to the second writing style mapping vector σ corresponding to the writing style.s Therefore, as shown in formula (6), the second prior distribution P s It is the first writing style mapping vector μ s The mean, the second writing style mapping vector σ s A Gaussian distribution with variance N(μ) s , σ s ).
[0142] Step 906: Obtain latent variables of data information from the first prior distribution and latent variables of writing style from the second prior distribution.
[0143] Here, the first prior distribution can refer to a Gaussian distribution where the first data information mapping vector has a mean and the second data information mapping vector has a variance. The second prior distribution can refer to a Gaussian distribution where the first writing style mapping vector has a mean and the second writing style mapping vector has a variance.
[0144] Specifically, the terminal samples noise values from a standard normal distribution, and then transforms the first and second data information mapping vectors based on these noise values to obtain latent data variables. Similarly, the terminal samples noise values from a standard normal distribution, and then transforms the first and second writing style mapping vectors based on these noise values to obtain latent writing style variables. Here, the standard normal distribution refers to a normal distribution with a mean of 0 and a variance of 1, denoted as N(0, 1).
[0145] Specifically, noise values are obtained by sampling from a standard normal distribution, and then the latent variables of the data information are calculated using formula (7):
[0146] z c =μ C +σ C ×x (7)
[0147] Among them, z c μ is a latent variable for data information. C Let σ be the first data information mapping vector. C Let x be the second data information mapping vector, and let x be the noise value.
[0148] Similarly, sampling is performed from the standard normal distribution to obtain noise values, and then the writing style latent variable is calculated using formula (8):
[0149] z s =μ s +σ s ×x (8)
[0150] Among them, z s To write style latent variables, μ s For the first writing style mapping vector, σs Let x be the second writing style mapping vector, and x be the noise value.
[0151] In this embodiment, since the prior distribution is a probability distribution that determines the cause based on historical patterns, the prior distribution calculated through the prior distribution can obtain a generalized distribution that is closer to the true result. Based on this, noise values are obtained by sampling from the standard normal distribution. Based on the noise values, each mapping vector is transformed to obtain the latent variables of data information and writing style, thereby ensuring the differentiability of the data information generation model and thus ensuring the reliability and feasibility of the data information generation model.
[0152] In one embodiment, such as Figure 10 As shown, predicted data information is obtained based on latent variables of data information, and predicted writing style is obtained based on latent variables of writing style, including:
[0153] Step 1002: The classification network in the data information generation model obtains the predicted writing style based on the latent variable of writing style.
[0154] The classification network is used to assign writing style labels to the data to be processed, and each writing style label describes a writing style. The set of writing styles included in this embodiment includes: preset rule writing styles, specialist physician writing styles, medical staff writing styles, and other writing styles. The specialist physician writing styles can be further divided into specialist physician writing styles corresponding to multiple departments according to actual needs.
[0155] Specifically, to decouple writing style from data information, the data information generation model obtains the writing style latent variable from the second prior distribution through a second prior network. This second prior network then inputs the writing style latent variable into the classification network to ensure that it does not contain data information. The classification network then predicts the writing style based on this latent variable. The classification network primarily consists of linear layers. These linear layers receive the style latent variable as input and output the probability that the data to be processed belongs to each writing style. Thus, the predicted writing style is obtained based on the probability of the data belonging to each writing style.
[0156] Step 1004: The decoding network in the data information generation model obtains the predicted data information based on the latent variables of the data information.
[0157] Specifically, in order to decouple writing style from data information, the data information generation model obtains the latent variables of data information from the first prior distribution through the first prior network. The first prior network then inputs the latent variables of data information into the decoding network to ensure that the latent variables of data information do not contain writing style-related information. Thus, the decoding network obtains the predicted data information based on the latent variables of data information.
[0158] Preferably, in the specific implementation process, when obtaining feature vectors based on the data to be processed, the data information generation model performs dropout masking on the data to be processed. Therefore, the data information generation model encodes the masked data to be processed and then obtains the feature vectors corresponding to the masked data to be processed, until the latent variables of the data information are obtained. The specific implementation method is similar to the aforementioned embodiments and will not be repeated here. Based on this, since the data information generation model performs dropout masking on the data to be processed, the first prior network inputs the latent variables of the data information to the decoding network, and the decoding network outputs the data that has been masked out from the data to be processed based on the latent variables of the data information.
[0159] Based on this, the decoding network includes a masked self-attention mechanism layer, a third normalization layer, a fully connected network, a fourth normalization layer, and a combination of fully connected layers and normalization layers. For ease of understanding, as follows... Figure 11 As shown, the first prior network inputs the latent variables of the data information into the decoding network. The masking self-attention mechanism layer 1101 in the decoding network processes the latent variables of the data information and obtains the data-related information of the masked data. This information is then processed by the third normalization layer 1102, the fully connected network 1103, the fourth normalization layer 1104, and the fully connected and normalized layers 1105. The fully connected and normalized layers 1105 then output the data in the data to be processed that has been masked. Finally, the data information generation model specifically obtains the predicted data information based on the data in the data to be processed that has been masked and the data in the data to be processed that has not been masked.
[0160] In this embodiment, by inputting the latent variable of writing style into the classification network and the latent variable of data information into the decoding network, the writing style and data information are decoupled, thereby avoiding the interference of data information on the writing style prediction process and the interference of writing style on the data information prediction process, thus improving the accuracy and reliability of predicting writing style and predicting data information.
[0161] In one embodiment, such as Figure 12 As shown, the classification network in the data information generation model predicts writing style based on latent variables of writing style, including:
[0162] Step 1202: The classification network in the data information generation model obtains the probability that the data to be processed belongs to each writing style based on the latent variable of writing style.
[0163] The classification network mainly consists of linear layers.
[0164] Specifically, the data information generation model obtains the writing style latent variable from the second prior distribution through the second prior network. The second prior network then inputs the writing style latent variable into the classification network, and each linear layer in the classification network outputs the probability that the data to be processed belongs to each writing style.
[0165] For example, if each writing style specifically includes: the pre-defined rule writing style, the specialist doctor writing style, the medical staff writing style, and other writing styles, then the probability that the data to be processed belongs to each writing style includes: the probability that the data to be processed belongs to the pre-defined rule writing style, the probability that the data to be processed belongs to the specialist doctor writing style, the probability that the data to be processed belongs to the medical staff writing style, and the probability that the data to be processed belongs to other writing styles.
[0166] Step 1204: Determine the predicted writing style based on the probability that the data to be processed belongs to each writing style.
[0167] Specifically, the predicted writing style is the writing style with the highest probability among all writing styles to which the data to be processed belongs. In detail, the probabilities of the data to be processed belonging to each writing style are sorted from largest to smallest, and the writing style with the highest probability is selected.
[0168] For example, the probabilities of the data to be processed belonging to various writing styles include: the probability of the data belonging to the preset rule writing style, the probability of the data belonging to the specialist physician writing style, the probability of the data belonging to the medical staff writing style, and the probability of the data belonging to other writing styles. The probability of the data belonging to the preset rule writing style is 50%, the probability of the data belonging to the specialist physician writing style is 80%, the probability of the data belonging to the medical staff writing style is 60%, and the probability of the data belonging to other writing styles is 20%. Therefore, 80% can be determined as the highest probability, and the writing style corresponding to 80% is specifically the specialist physician writing style; that is, the predicted writing style is the specialist physician writing style.
[0169] In this embodiment, the probability of the data to be processed belonging to each writing style is determined by the probability of the highest value, so as to ensure the accuracy and reliability of the predicted writing style.
[0170] The methods for generating data information have been described in detail above. However, data information generation requires a trained data information generation model. Therefore, the training process of the data information generation model will be described in detail below:
[0171] In one embodiment, such as Figure 13 As shown, the training process of the data information generation model includes:
[0172] Step 1302: Adjust the first prior network in the initial data information generation model based on the first data information, the second data information, and the first prior distribution.
[0173] Specifically, based on the first data information, the second data information, and the first prior distribution, the first prior network in the initial data information generation model is adjusted to calculate the value of the first loss function. When the value of the first loss function does not reach the preset threshold or the training does not reach the maximum number of iterations, the gradient of the network parameters in the first prior network is calculated. Then, the network parameters in the first prior network are updated using a pre-set model optimization algorithm until the value of the preset loss function reaches the preset threshold or the training reaches the maximum number of iterations, thus obtaining the first prior network trained in the target latent variable model.
[0174] The model optimization algorithms include, but are not limited to, SGD (stochastic gradient descent), ADAM (adaptive moment estimation), conjugate gradient, momentum optimization, simulated annealing, and ant colony optimization. For example, the server calculates the value of a first loss function based on the first and second data information and a first prior distribution. It then calculates the gradient of the network parameters in the first prior network based on the value of the first loss function. Finally, it updates the network parameters in the first prior network using the SGD algorithm. When a preset training completion condition is met, the trained first prior network is obtained. This preset training completion condition could be that the value of the preset loss function reaches a preset threshold or the number of training iterations reaches the maximum number; this is not limited here.
[0175] Step 1304: Adjust the second prior network in the initial data information generation model according to the first writing style, the second writing style, and the second prior distribution.
[0176] Specifically, based on the first writing style, the second writing style, and the second prior distribution, the second prior network in the initial writing style generation model is adjusted to calculate the value of the second loss function. When the value of the second loss function does not reach a preset threshold or the training has not reached the maximum number of iterations, the gradient of the network parameters in the second prior network is calculated. Then, a pre-set model optimization algorithm is used to update the network parameters in the second prior network until the value of the preset loss function reaches a preset threshold or the training has reached the maximum number of iterations, thus obtaining the trained second prior network in the target latent variable model.
[0177] Similar to step 1302, the model optimization algorithm includes, but is not limited to, SGD, ADAM, conjugate gradient, momentum optimization, simulated annealing, and ant colony optimization. For example, the server calculates the value of the second loss function based on the first writing style, the second writing style, and the second prior distribution. It then calculates the gradient of the network parameters in the second prior network based on the value of the second loss function. Finally, it uses the ADAM algorithm to update the network parameters in the second prior network. When a preset training completion condition is met, the trained second prior network is obtained. This preset training completion condition can be that the value of the preset loss function reaches a preset threshold or the number of training iterations reaches the maximum number; it is not limited here.
[0178] In this embodiment, the network parameters of the first prior network and the second prior network are adjusted respectively to ensure that the writing style and data information are also decoupled during the model training process, so as to avoid mutual interference between the writing style and data information. This improves the reliability of the decoupled prediction of the data information generation model for the writing style and data information, and improves the efficiency and accuracy of data information generation.
[0179] In one embodiment, such as Figure 14 As shown, the training process of the data information generation model includes:
[0180] Step 1402: The first posterior network in the initial data information generation model calculates the posterior distribution based on the first data information and the second data information to obtain the first posterior distribution corresponding to the data information.
[0181] The first posterior network is used to calculate the posterior distribution of the data information. The posterior distribution refers to the probability distribution of the cause estimated based on the known result. Specifically, the first data information and the second data information are input into the first posterior network in the initial data information generation model. The first posterior network calculates the posterior distribution of the first data information and the second data information to obtain the first posterior distribution corresponding to the data information.
[0182] Step 1404: The second posterior network in the initial data information generation model calculates the posterior distribution based on the first writing style and the second writing style to obtain the second posterior distribution corresponding to the writing style.
[0183] The second posterior network is used to calculate the posterior distribution of writing styles. The posterior distribution refers to the probability distribution of the cause estimated based on the known result. Specifically, the first writing style and the second writing style are input into the second posterior network in the initial data information generation model. The second posterior network calculates the posterior distribution of the first writing style and the second writing style to obtain the first posterior distribution corresponding to the writing style.
[0184] Step 1406: Adjust the first prior network in the initial data information generation model based on the first data information, the second data information, the first prior distribution, and the first posterior distribution.
[0185] Specifically, a first posterior network is used as the teacher network to train the student network (i.e., the first prior network) to achieve knowledge transfer. Based on this, similar to the aforementioned embodiments, according to the first and second data information, the first prior distribution, and the first posterior distribution, the first prior network in the initial writing style generation model is adjusted to calculate the value of the first loss function. When the value of the first loss function does not reach a preset threshold or the training does not reach the maximum number of iterations, the gradient of the network parameters in the first prior network is calculated. Then, a pre-set model optimization algorithm is used to update the network parameters in the first prior network until the value of the preset loss function reaches a preset threshold or the training reaches the maximum number of iterations, thus obtaining the trained first prior network in the target latent variable model.
[0186] Therefore, considering the first posterior distribution, the network parameters of the first prior network in the initial data information generation model are adjusted using the first loss function. This allows the first prior distribution obtained through the first prior network in the trained data information generation model to approach the first posterior distribution. It should be understood that the network parameters in the first posterior network are already trained and do not require updating.
[0187] Step 1408: Adjust the second prior network in the initial data information generation model according to the first writing style and the second writing style, the second prior distribution and the second posterior distribution.
[0188] Specifically, a second posterior network is used as the teacher network to train the student network (i.e., the second prior network) to achieve knowledge transfer. Based on this, similar to the aforementioned embodiments, according to the first and second writing styles, the second prior distribution, and the second posterior distribution, the second prior network in the initial writing style generation model is adjusted to calculate the value of the second loss function. When the value of the second loss function does not reach a preset threshold or the training does not reach the maximum number of iterations, the gradient of the network parameters in the second prior network is calculated. Then, a pre-set model optimization algorithm is used to update the network parameters in the second prior network until the value of the preset loss function reaches a preset threshold or the training reaches the maximum number of iterations, thus obtaining the trained second prior network in the latent variable model of the training target.
[0189] Therefore, considering the second posterior distribution, the network parameters of the second prior network in the initial writing style generation model are adjusted using the second loss function. This allows the second prior distribution obtained through the second prior network in the trained writing style generation model to approach the second posterior distribution. It should be understood that the network parameters in the second posterior network are already trained and do not require updating.
[0190] In this embodiment, the first prior network is trained using the first posterior network as a pre-trained network, thereby ensuring that the first prior distribution and the first posterior distribution are consistent. Similarly, the second prior network is trained using the second posterior network as a pre-trained network, thereby ensuring that the second prior distribution and the second posterior distribution are consistent. This removes the latent variable exposure bias problem, enabling the data information generation model to generate highly accurate predicted data information and predicted writing style when deployed using both the first and second prior distributions, thus improving the accuracy of data information generation.
[0191] In one embodiment, such as Figure 15 As shown, the training process of the data information generation model includes:
[0192] Step 1502: The decoding network in the initial data information generation model obtains the first data information of the sample data based on the latent variables of the sample data.
[0193] The decoding network is used to decode the latent variables of the data to obtain the data information. It can use a neural network, such as an RNN (Recurrent Neural Network). Alternatively, a neural network symmetrical to the encoding network can be used directly as the decoding network. Preferably, the decoding network can use complex deep neural networks that combine attention mechanisms and transformers (a type of BERT model) for decoding.
[0194] Specifically, the latent variables of the sample data are input into the decoding network in the initial data information generation model for decoding to obtain the first data information of the sample data.
[0195] Based on this, step 1406, according to the first data information and the second data information, the first prior distribution and the first posterior distribution, adjusts the first prior network in the initial data information generation model, including:
[0196] Step 1504: Based on the first data information and the second data information, the first prior distribution and the first posterior distribution, adjust the first prior network and the decoding network in the initial data information generation model.
[0197] Specifically, a first loss value is calculated based on the first and second data information, the first prior distribution, and the first posterior distribution. This first loss value is used to measure the quality of the trained first prior network and decoding network. Then, the first prior network and decoding network in the initial data generation model are adjusted based on the first loss value. It should be understood that the network parameters in the first posterior network and encoding network are already trained and therefore do not need to be adjusted.
[0198] Furthermore, specifically according to formula (9), the first loss value is calculated based on the first data information, the second data information, the first prior distribution, and the first posterior distribution:
[0199] L 数据信息 =-E q(z|x) logp(x|z)+KL(q(z|x)||p(z)) (9)
[0200] Among them, L 数据信息 Let z be the first loss value, x be the second data information, q(z|x) be the first posterior distribution, p(z) be the first prior distribution, p(x|z) be the decoding network, and KL(q(z|x)||p(z)) be the loss information calculated using the KL divergence loss function between the first prior distribution and the first posterior distribution. q(z|x) logp(x|z) is the loss for reconstructing the input.
[0201] In this embodiment, encoding and decoding are performed using an encoding network and a decoding network, which improves the efficiency and accuracy of encoding and decoding. Based on this, during model training, a first loss value is calculated according to the first data information, the second data information, the first prior distribution, and the first posterior distribution. The network parameters in the first prior network and the decoding network are then adjusted, thereby obtaining a highly reliable data information generation model and ensuring the accuracy of the data information generation model.
[0202] In one embodiment, such as Figure 16 As shown, the predicted writing style is obtained based on latent variables of writing style, including:
[0203] Step 1602: The classification network in the initial data information generation model obtains the predicted writing style of the sample data based on the latent variable of writing style in the sample data.
[0204] The classification network mainly consists of linear layers.
[0205] Specifically, the latent variable of writing style is input into the classification network of the initial data information generation model. Each linear layer in the classification network outputs the probability that the data to be processed belongs to each writing style. The predicted writing style is the writing style corresponding to the highest probability among all writing styles. Specifically, the probabilities of the data to be processed belonging to each writing style are sorted from largest to smallest, and the writing style corresponding to the highest probability is selected. The method for obtaining the predicted writing style is similar to the previous embodiment and will not be repeated here.
[0206] Based on this, step 1408, according to the first writing style and the second writing style, the second prior distribution and the second posterior distribution, adjusts the second prior network in the initial data information generation model, including:
[0207] Step 1604: Adjust the second prior network and classification network in the initial data information generation model according to the first writing style and the second writing style, the second prior distribution and the second posterior distribution.
[0208] Specifically, a second loss value is calculated based on the first and second writing styles. This second loss value measures the quality of the trained classification network and is specifically the cross-entropy loss function value. Then, a third loss value is calculated based on the second prior and second posterior distributions. This third loss value measures the quality of the trained second prior network. Finally, based on the second and third loss values, the initial writing style and the initial data information generation model's second prior network and classification network are adjusted. It should be understood that the network parameters in the second posterior network and the encoding network are already trained and therefore do not require adjustment.
[0209] Furthermore, specifically according to formula (10), the second loss value is calculated based on the first writing style and the second writing style:
[0210]
[0211] Among them, L 撰写风格 The second loss value, y ik This is the second writing style.
[0212] In this embodiment, during model training, a second loss value is calculated based on the first and second writing styles, and a third loss value is calculated based on the second prior distribution and the second posterior distribution. This adjusts the network parameters in the second prior network and the classification network. While ensuring the decoupling of writing style from data information, a highly reliable data information generation model is obtained, thus guaranteeing the accuracy of the data information generation model.
[0213] Based on the foregoing embodiments, the complete process of the data information generation method will be described below, such as... Figure 17As shown, a data information generation method is provided, which can be applied to... Figure 1 Taking a terminal as an example, it can be understood that this method can also be applied to a server, and to a system that includes both a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0214] Step 1701: Obtain the data to be processed, and use the data information to generate the encoding network in the model to encode the data to be processed and obtain the feature vector.
[0215] The data to be processed includes either text data or image data. Furthermore, the feature vector can be a text feature vector, an image feature vector, or a multimodal feature vector that includes both text and image feature vectors.
[0216] Specifically, the terminal can obtain the data to be processed through data input by the user in real time, or it can obtain the data to be processed from local storage, or it can obtain the data to be processed through communication and interaction with the server. There are no restrictions on the method of obtaining the data to be processed here.
[0217] Based on this, after acquiring the data to be processed, the terminal invokes the data information generation model, and specifically uses the encoding network in the data information generation model to encode the data to be processed to obtain feature vectors. Specifically, the encoding network is used to convert either text data or image data into feature vectors.
[0218] Step 1702: The first prior network calculates the prior distribution of the data information based on the feature vector to obtain the first prior distribution corresponding to the data information.
[0219] The first prior network is used to calculate the prior distribution of the data. The prior distribution is a probability distribution that determines the cause based on historical patterns, and specifically, the first prior distribution is a Gaussian distribution. Secondly, the parameters for calculating the prior distribution of the data are obtained through knowledge distillation. Knowledge distillation refers to using the parameters for calculating the posterior distribution of the data to guide the training of the parameters for calculating the prior distribution, ensuring consistency between the parameters for calculating the prior and posterior distributions. This ensures that the error between the obtained posterior and prior distributions of the data is within a preset error range.
[0220] Specifically, the encoding network inputs the feature vector into the first prior network, which then calculates the prior distribution of the data information based on the feature vector, thereby obtaining the first prior distribution corresponding to the data information.
[0221] Step 1703: The second prior network calculates the prior distribution of writing style based on the feature vector to obtain the second prior distribution corresponding to the writing style.
[0222] The second prior network is used to calculate the prior distribution of writing style. The prior distribution is a probability distribution that determines the cause based on historical patterns, and specifically, the second prior distribution is a Gaussian distribution. Furthermore, the parameters for calculating the prior distribution of writing style are obtained through knowledge distillation. Knowledge distillation refers to using the parameters for calculating the posterior distribution of writing style during training to guide the training of the parameters for calculating the prior distribution of writing style, ensuring consistency between the parameters for calculating the prior and posterior distributions of writing style. This ensures that the error between the obtained posterior distribution and the obtained prior distribution of writing style is within a preset error range.
[0223] Specifically, the encoding network inputs the feature vector into the second prior network, which calculates the prior distribution of the writing style based on the feature vector, thereby obtaining the second prior distribution corresponding to the writing style.
[0224] Step 1704: Obtain the latent variables of data information from the first prior distribution and the latent variables of writing style from the second prior distribution.
[0225] Here, the first prior distribution can refer to a Gaussian distribution where the first data information mapping vector has a mean and the second data information mapping vector has a variance. The second prior distribution can refer to a Gaussian distribution where the first writing style mapping vector has a mean and the second writing style mapping vector has a variance.
[0226] Specifically, the terminal samples noise values from a standard normal distribution, and then transforms the first and second data information mapping vectors based on these noise values to obtain latent data variables. Similarly, the terminal samples noise values from a standard normal distribution, and then transforms the first and second writing style mapping vectors based on these noise values to obtain latent writing style variables. Here, the standard normal distribution refers to a normal distribution with a mean of 0 and a variance of 1, denoted as N(0, 1).
[0227] Step 1705: The classification network in the data information generation model obtains the predicted writing style based on the latent variable of writing style.
[0228] The classification network mainly consists of linear layers and is used to assign writing style labels to the data to be processed. Each writing style label describes a writing style.
[0229] Specifically, to decouple writing style from data information, the data information generation model obtains the writing style latent variable from the second prior distribution through the second prior network. This second prior network then inputs the writing style latent variable into the classification network to ensure that it does not contain data information. The classification network then predicts the writing style based on this latent variable. The classification network mainly consists of linear layers. These linear layers receive the style latent variable as input and output the probability that the data to be processed belongs to each writing style. The probabilities of the data belonging to each writing style are then sorted from largest to smallest, and the writing style corresponding to the highest probability is selected.
[0230] Step 1706: The decoding network in the data information generation model obtains the predicted data information based on the latent variables of the data information.
[0231] Specifically, in order to decouple writing style from data information, the data information generation model obtains the latent variables of data information from the first prior distribution through the first prior network. The first prior network then inputs the latent variables of data information into the decoding network to ensure that the latent variables of data information do not contain writing style-related information. Thus, the decoding network obtains the predicted data information based on the latent variables of data information.
[0232] Preferably, in the specific implementation process, when obtaining feature vectors based on the data to be processed, the data information generation model performs dropout masking on the data to be processed. Therefore, the data information generation model encodes the masked data to be processed and then obtains the feature vectors corresponding to the masked data to be processed, until the latent variables of the data information are obtained. The specific implementation method is similar to the aforementioned embodiments and will not be repeated here. Based on this, since the data information generation model performs dropout masking on the data to be processed, the first prior network inputs the latent variables of the data information to the decoding network, and the decoding network outputs the data that has been masked out in the data to be processed based on the latent variables of the data information.
[0233] Step 1707: Based on the predicted data information and the predicted writing style, generate the target data information corresponding to the data to be processed.
[0234] Specifically, the terminal describes the predicted data information and the predicted writing style based on the obtained predicted data information. Since the predicted writing style is used to describe the content description logic of the predicted data information, the terminal describes the content of the predicted data information according to the content description logic of the predicted writing style. This usually involves adding punctuation marks and transition text to the predicted data information, thereby generating the target data information corresponding to the data to be processed, and making the target data information more flexible in terms of content description logic.
[0235] It is understandable that the training process of the data information generation model, and Figure 17 The response embodiments of the steps shown have been described in detail in the foregoing embodiments, and therefore will not be repeated here.
[0236] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0237] Based on the same inventive concept, this application also provides a data information generation apparatus for implementing the data information generation method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more data information generation apparatus embodiments provided below can be found in the limitations of the data information generation method described above, and will not be repeated here.
[0238] In one embodiment, such as Figure 18 As shown, a data information generation device is provided, including: an acquisition module 1802, a latent variable acquisition module 1804, a prediction module 1806, and a generation module 1808, wherein:
[0239] The acquisition module 1802 is used to acquire the data to be processed and obtain the feature vector based on the data to be processed;
[0240] Latent variable acquisition module 1804 is used to obtain data information latent variables and writing style latent variables based on feature vectors;
[0241] The prediction module 1806 is used to obtain predicted data information based on the latent variables of data information and to obtain predicted writing style based on the latent variables of writing style. The predicted writing style is used to describe the content description logic of the predicted data information.
[0242] The generation module 1808 is used to generate target data information corresponding to the data to be processed based on the prediction data information and the prediction writing style.
[0243] In one embodiment, the acquisition module 1802 is further configured to encode the data to be processed through the encoding network in the data information generation model to obtain a feature vector. The data information generation model is obtained by adjusting the prior network in the initial data information generation model based on the first data information and the second data information, as well as the first writing style and the second writing style. The first data information is obtained based on the data information latent variables of the sample data, the first writing style is obtained based on the writing style latent variables of the sample data, and the data information latent variables and the writing style latent variables of the sample data are obtained based on the feature vector of the sample data. The sample data is generated based on the second data information and the second writing style.
[0244] In one embodiment, the latent variable acquisition module 1804 is also used to obtain data information latent variables and writing style latent variables based on feature vectors in the prior network of the data information generation model.
[0245] In one embodiment, the prior network includes a first prior network and a second prior network;
[0246] The latent variable acquisition module 1804 is also used for the first prior network to calculate the prior distribution of the data information based on the feature vector, and to obtain the first prior distribution corresponding to the data information; and the second prior network to calculate the prior distribution of the writing style based on the feature vector, and to obtain the second prior distribution corresponding to the writing style; and to obtain the latent variables of the data information from the first prior distribution, and the latent variables of the writing style from the second prior distribution.
[0247] In one embodiment, the prediction module 1806 is further configured to: ...
[0248] In one embodiment, the prediction module 1806 is further configured to: use the classification network in the data information generation model to obtain the probability that the data to be processed belongs to each writing style based on the latent variable of writing style; and determine the predicted writing style based on the probability that the data to be processed belongs to each writing style.
[0249] In one embodiment, the data information generation device further includes a model training module 1810;
[0250] The model training module 1810 is used to adjust the first prior network in the initial data information generation model according to the first data information, the second data information, and the first prior distribution; and to adjust the second prior network in the initial data information generation model according to the first writing style, the second writing style, and the second prior distribution.
[0251] In one embodiment, the model training module 1810 is further configured to: generate a first posterior network in the initial data information generation model, calculate a posterior distribution based on the first data information and the second data information to obtain a first posterior distribution corresponding to the data information; and generate a second posterior network in the initial data information generation model, calculate a posterior distribution based on the first writing style and the second writing style to obtain a second posterior distribution corresponding to the writing style; and adjust the first prior network in the initial data information generation model according to the first data information and the second data information, the first prior distribution and the first posterior distribution; and adjust the second prior network in the initial data information generation model according to the first writing style and the second writing style, the second prior distribution and the second posterior distribution.
[0252] In one embodiment, the model training module 1810 is further configured to generate the first data information of the sample data by the decoding network in the initial data information generation model based on the latent variables of the sample data; and adjust the first prior network and the decoding network in the initial data information generation model according to the first data information, the second data information, the first prior distribution, and the first posterior distribution.
[0253] In one embodiment, the model training module 1810 is further configured to: generate a classification network in the initial data information generation model based on the latent variable of writing style of the sample data to obtain the predicted writing style of the sample data; and adjust the second prior network and the classification network in the initial data information generation model according to the first writing style and the second writing style, the second prior distribution and the second posterior distribution.
[0254] Each module in the aforementioned data generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0255] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 19As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a data information generation method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0256] Those skilled in the art will understand that Figure 19 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0257] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0258] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above-described method embodiments.
[0259] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0260] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0261] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0262] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0263] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for generating data information, characterized in that, The method includes: The process involves acquiring data to be processed, masking the data using the encoding network in a data information generation model, and then encoding the masked data to obtain a feature vector. The data to be processed includes at least information related to the patient's input description of their physical state. The data information generation model is obtained by adjusting the prior network in an initial data information generation model based on first and second data information, as well as a first and second writing style. The first data information is obtained based on the latent variables of the sample data, the first writing style is obtained based on the latent variables of the writing style of the sample data, and both the latent variables of the sample data and the latent variables of the writing style are obtained based on the feature vector of the sample data. The sample data is generated based on the second data information and the second writing style. The first prior network in the data information generation model calculates the prior distribution of the data information based on the feature vector to obtain the first prior distribution corresponding to the data information. The first prior distribution refers to a Gaussian distribution where the first data information mapping vector has a mean and the second data information mapping vector has a variance. Noise values are sampled from the standard normal distribution. Based on the noise values, the first data information mapping vector and the second data information mapping vector are transformed to obtain latent variables of the data information. The latent variables of the data information are then input into the decoding network to ensure that the latent variables of the data information do not contain any writing style-related information. The second prior network in the data information generation model calculates the prior distribution of writing style based on the feature vector to obtain the second prior distribution corresponding to the writing style. The second prior distribution refers to a Gaussian distribution where the first writing style mapping vector has a mean and the second writing style mapping vector has a variance. The noise value is obtained by sampling from the standard normal distribution. The first writing style mapping vector and the second writing style mapping vector are transformed according to the noise to obtain the writing style latent variable. The writing style latent variable is then input into the classification network to ensure that the writing style latent variable does not contain data information. The decoding network in the data information generation model obtains the data that has been masked out in the predicted data to be processed based on the latent variables of the data information, and obtains the predicted data information based on the data that has been masked out in the predicted data to be processed and the data that has not been masked in the data to be processed. The classification network in the data information generation model classifies the writing style corresponding to the data to be processed based on the writing style latent variable to obtain the predicted writing style, so as to decouple the writing style from the data information; the predicted writing style is used to describe the content description logic of the predicted data information, and the writing style includes at least the preset rule writing style, the specialist doctor writing style and the medical staff writing style. According to the content description logic of the predicted writing style, punctuation marks and transition texts for the predicted data information are added to complete the content description of the predicted data information, generating target data information corresponding to the data to be processed, so that the content description logic of the target data information matches the predicted writing style; the target data information is the body status file corresponding to the relevant information of the body status description input by the patient.
2. The method according to claim 1, characterized in that, The classification network in the data information generation model classifies the writing style corresponding to the data to be processed based on the writing style latent variable to obtain the predicted writing style, including: The classification network in the data information generation model obtains the probability that the data to be processed belongs to each writing style based on the writing style latent variable. The predicted writing style is determined based on the probability that the data to be processed belongs to each writing style.
3. The method according to claim 1, wherein the training process of the data information generation model includes: Based on the first data information, the second data information, and the first prior distribution, adjust the first prior network in the initial data information generation model; Based on the first writing style and the second writing style, as well as the second prior distribution, adjust the second prior network in the initial data information generation model.
4. The method according to claim 1, characterized in that, The training process of the data information generation model includes: The first posterior network in the initial data information generation model calculates the posterior distribution based on the first data information and the second data information to obtain the first posterior distribution corresponding to the data information. The second posterior network in the initial data information generation model calculates the posterior distribution based on the first writing style and the second writing style to obtain the second posterior distribution corresponding to the writing style. Based on the first data information and the second data information, the first prior distribution and the first posterior distribution, adjust the first prior network in the initial data information generation model; Based on the first writing style and the second writing style, the second prior distribution and the second posterior distribution, the second prior network in the initial data information generation model is adjusted.
5. The method according to claim 4, characterized in that, The training process of the data information generation model includes: The decoding network in the initial data information generation model obtains the first data information of the sample data based on the latent variables of the sample data. The step of adjusting the first prior network in the initial data information generation model based on the first data information and the second data information, the first prior distribution, and the first posterior distribution includes: Based on the first data information and the second data information, the first prior distribution and the first posterior distribution, adjust the first prior network and the decoding network in the initial data information generation model.
6. The method according to claim 4, characterized in that, The process of predicting writing style based on the latent variables of writing style includes: The classification network in the initial data information generation model obtains the predicted writing style of the sample data based on the latent variable of writing style in the sample data. Based on the first writing style and the second writing style, the second prior distribution, and the second posterior distribution, the second prior network in the initial data information generation model is adjusted, including: Based on the first writing style and the second writing style, the second prior distribution and the second posterior distribution, adjust the second prior network and the classification network in the initial data information generation model.
7. A data information generation device, characterized in that, The device includes: An acquisition module is used to acquire data to be processed, and to perform masking processing on the data to be processed through the encoding network in the data information generation model, and to encode the masked data to obtain a feature vector; the data to be processed includes at least relevant information on the description of the patient's physical state; the data information generation model is obtained by adjusting the prior network in the initial data information generation model based on first data information and second data information, as well as first writing style and second writing style; the first data information is obtained based on the latent variables of the data information of the sample data, the first writing style is obtained based on the latent variables of the writing style of the sample data, and the latent variables of the data information and the writing style of the sample data are obtained based on the feature vector of the sample data; the sample data is generated based on the second data information and the second writing style. The latent variable acquisition module is used in the data information generation model to calculate the prior distribution of the data information based on the feature vector, thereby obtaining the first prior distribution corresponding to the data information. The first prior distribution refers to a Gaussian distribution where the first data information mapping vector has a mean and the second data information mapping vector has a variance. Noise values are sampled from a standard normal distribution. Based on the noise values, the first and second data information mapping vectors are transformed to obtain latent variables. These latent variables are then input into the decoding network to ensure that the latent variables do not contain any writing style. The data information generation model includes: a second prior network in the data information generation model; and a second prior network in the model that calculates the prior distribution of writing style based on the feature vector to obtain the second prior distribution corresponding to the writing style. The second prior distribution refers to a Gaussian distribution where the first writing style mapping vector has a mean and the second writing style mapping vector has a variance. Noise values are sampled from the standard normal distribution. The first and second writing style mapping vectors are then transformed according to the noise to obtain latent variables of the writing style. These latent variables are then input into the classification network to ensure that the latent variables of the writing style do not contain data information. The prediction module is used by the decoding network in the data information generation model to obtain the masked data in the data to be processed based on the latent variables of the data information, and to obtain the predicted data information based on the masked data and the unmasked data in the data to be processed; the classification network in the data information generation model classifies the writing style corresponding to the data to be processed based on the writing style latent variables to obtain the predicted writing style, so as to decouple the writing style from the data information; the predicted writing style is used to describe the content description logic of the predicted data information, and the writing style includes at least the preset rule writing style, the specialist doctor writing style, and the medical staff writing style; The generation module is used to add punctuation marks and transition text to the predicted data information according to the content description logic of the predicted writing style, so as to complete the content description of the predicted data information and generate the target data information corresponding to the data to be processed, so that the content description logic of the target data information matches the predicted writing style; the target data information is the body status file corresponding to the relevant information of the body status description input by the patient.
8. The apparatus according to claim 7, characterized in that, The prediction module is also used to enable the classification network in the data information generation model to obtain the probability that the data to be processed belongs to each writing style based on the writing style latent variable. The predicted writing style is determined based on the probability that the data to be processed belongs to each writing style.
9. The apparatus according to claim 7, wherein the data information generation apparatus further comprises a model training module; The model training module is used to adjust the first prior network in the initial data information generation model according to the first data information, the second data information, and the first prior distribution; and to adjust the second prior network in the initial data information generation model according to the first writing style, the second writing style, and the second prior distribution.
10. The apparatus according to claim 7, characterized in that, The model training module is also used to perform posterior distribution calculation based on the first data information and the second data information in the first posterior network of the initial data information generation model, and to obtain the first posterior distribution corresponding to the data information. The second posterior network in the initial data information generation model calculates the posterior distribution based on the first writing style and the second writing style to obtain the second posterior distribution corresponding to the writing style; the first prior network in the initial data information generation model is adjusted according to the first data information and the second data information, the first prior distribution and the first posterior distribution; the second prior network in the initial data information generation model is adjusted according to the first writing style and the second writing style, the second prior distribution and the second posterior distribution.
11. The apparatus according to claim 10, characterized in that, The model training module is further configured to use the decoding network in the initial data information generation model to obtain the first data information of the sample data based on the latent variables of the sample data; the step of adjusting the first prior network in the initial data information generation model according to the first data information and the second data information, the first prior distribution and the first posterior distribution includes: adjusting the first prior network and the decoding network in the initial data information generation model according to the first data information and the second data information, the first prior distribution and the first posterior distribution.
12. The apparatus according to claim 10, characterized in that, The model training module is further used to obtain the predicted writing style of the sample data based on the writing style latent variable of the sample data by the classification network in the initial data information generation model; and to adjust the second prior network and the classification network in the initial data information generation model according to the first writing style and the second writing style, the second prior distribution and the second posterior distribution.
13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.