A method and system for generating a medical, nursing and rehabilitation field business arrangement file based on a large model
Patent Information
- Application Number
- CN202410402435.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-04-03
AI Technical Summary
然而,现在的大语言模型在通用领域的表现很好,但是在医疗、养老、康复领域的表现不佳,需要通过适当的方法将通用大模型改造成专业大模型,在改造过程中如何融入领域知识并且快速、有效、低成本地改造是目前研究的重点
[0035]1、本发明引入大模型能力,可以自动完成电子化业务编排,相较于纯手动的人工编排方法,本发明方法能够降低智力成本,同时充分利用已有业务的复用和敏捷需求迭代能力,从而缩短了应用建设和维护的周期和成本,进一步提升了IT服务业务水平。
Smart Images

Figure CN118132518B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of business orchestration, and in particular to a method and system for generating business orchestration documents in the medical, elderly care and rehabilitation field based on a large model. Background Technology
[0002] As society ages further and seniors place increasing importance on their health, the need for integration among medical, rehabilitation, and elderly care services is becoming increasingly urgent. However, there is currently no method to provide a service arrangement plan that simultaneously involves all three areas for an individual senior citizen.
[0003] While general-purpose visual component-based business orchestration methods are relatively mature—for example, Chinese patent document CN113448547A discloses a visual task orchestration method, and Chinese patent document CN111522543A discloses a visual application component orchestration method and system—these methods remain relatively complex for professionals in the medical, elderly care, and rehabilitation fields. They require considerable training and cannot allow professionals to focus solely on business orchestration. Furthermore, general solutions typically only address a single business within a specific field, failing to meet the needs of multi-field integrated orchestration. Currently, there is no mature solution for integrated orchestration across these three fields that lowers the barrier to entry for experts in all three areas.
[0004] Large language models, or simply large models, contain hundreds of billions or even more parameters. The sheer number of parameters gives large models extremely strong analytical and reasoning capabilities, resulting in exceptional performance in the field of text generation.
[0005] Large language models also possess emergent capabilities, one of which is context learning. Large language models can achieve good inference results with only a small number of sample examples, and at the same time, there is no need to adjust the model weights in the process.
[0006] In summary, there is an urgent need in the fields of healthcare, elderly care, and rehabilitation to automate and precisely integrate these three domains. The emergence of large language models and their adaptation to target domains can help address this issue. However, while current large language models perform well in general domains, they perform poorly in healthcare, elderly care, and rehabilitation. Therefore, it is necessary to transform general-purpose large models into specialized large models using appropriate methods. The current research focus is on how to integrate domain knowledge and transform these models quickly, effectively, and cost-efficiently. Summary of the Invention
[0007] This invention provides a method for generating business orchestration documents in the medical, elderly care, and rehabilitation fields based on a large model, which can automatically and precisely complete the integrated orchestration of the three fields of medical care, rehabilitation, and elderly care.
[0008] A method for generating business orchestration documents in the medical, elderly care, and rehabilitation field based on a large model includes the following steps:
[0009] (1) Collect information in the fields of medical care, elderly care and rehabilitation in text form, screen the information for knowledge quality, generate retrieval vectors for knowledge fragments, and form a knowledge database for the fields of medical care, elderly care and rehabilitation.
[0010] (2) Collect business orchestration files, call prompt words to generate a large model to obtain the business orchestration file generation instructions, and store the orchestration files, generation instructions and related knowledge into the sample database.
[0011] (3) Using the content in the example database as corpus, the weight matrix of the file generation model is fine-tuned by updating the decomposition matrix of the incremental matrix;
[0012] (4) During the application process, for natural language data, the required knowledge is retrieved from the knowledge database and the required examples are retrieved from the example database to form the final prompt words. Then, the fine-tuned file is called to generate a large model and obtain the business arrangement file.
[0013] (5) Generate a score for the business orchestration file through the file scoring model, feed the score back to the file generation model until a business orchestration file that meets the requirements is generated, and give the reason for generating the business orchestration file; finally, store the generated business orchestration file in the sample database.
[0014] The specific process of step (1) is as follows:
[0015] (1-1) Using one-hot vectors constructed from natural language dictionaries as input, and judging whether they are legal natural language as output, train an artificial intelligence model. Take the model parameters of the middle fully connected layer to obtain a conversion module with natural language as input and word embedding as output.
[0016] (1-2) Collect knowledge in the fields of medical care, elderly care and rehabilitation described in natural language, including literature, expert consensus, guidance manuals, standard operating procedures and textbook materials for dealing with diseases of the elderly;
[0017] A recurrent neural network is trained and used to output the similarity between collected data and textbook text, thus filtering out high-quality knowledge.
[0018] (1-3) Divide the selected knowledge into segments according to the maximum segment size not exceeding a certain threshold. Input the segmented knowledge into the natural language-word embedding conversion module obtained in step (1-1) to obtain multidimensional vectors. Store the multidimensional vectors and knowledge in pairs into the knowledge database.
[0019] Furthermore, in steps (1-2), data with a similarity of over 80% are selected to form high-quality knowledge.
[0020] The specific process of step (2) is as follows:
[0021] (2-1) Collect business arrangement documents related to the fields of medical care, elderly care and rehabilitation, input them into the knowledge database of step (1), and retrieve the relevant knowledge;
[0022] (2-2) Using the business orchestration file from step (2-1) as input, call the prompt words to generate a large model output to generate a similar orchestration file;
[0023] (2-3) Store the formatted files, generated instructions, and related knowledge as examples in the example database.
[0024] Furthermore, in steps (2-3), the arrangement files and generation instructions correspond one-to-one, with one knowledge corresponding to one or more arrangement files.
[0025] The specific process of step (3) is as follows:
[0026] The training files are updated using the corpus in the example database to generate a large model. The model W to be updated is defined as W = W0 + ΔW, where W0 is the initial model parameter matrix or the model parameter matrix after the previous update. Updating the entire W is regarded as updating ΔW. ΔW is decomposed into A and B with smaller dimensions, ΔW = AB. Then, it is only necessary to update A and B respectively to complete the model update.
[0027] In step (4), for data in natural language form, the required knowledge is retrieved from the knowledge database, specifically as follows:
[0028] After inputting a piece of natural language text into the knowledge database, the embedding of the word to be searched is obtained. The cosine distance is calculated between the word and each vector in the knowledge database. Then, the knowledge fragments corresponding to the k vectors with the smallest cosine distance are retrieved from the database and output.
[0029] The specific process of step (5) is as follows:
[0030] (5-1) Train a reinforcement learning model with the business orchestration files in step (2) as the dataset, modifying the token as the action, the distribution of tokens in the business orchestration files as the state, and the matching degree with the business files in the dataset as the reward, to obtain a file scoring model for outputting a score of the business orchestration files.
[0031] (5-2) Input the business orchestration file generated in step (4) into the file scoring model to obtain the score, and feed it back to the file generation model. Modify the file until the score meets the set requirements.
[0032] (5-3) The document generation model is required to output the reasons for generating business orchestration documents in this way for users to refer to, and the generated documents are stored in the sample database of step (2) to update the document generation model.
[0033] The present invention also provides a business orchestration document generation system for the medical, elderly care and rehabilitation field based on a large model, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned business orchestration document generation method for the medical, elderly care and rehabilitation field.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. This invention introduces large model capabilities, which can automatically complete electronic business orchestration. Compared with purely manual orchestration methods, this invention can reduce intellectual costs, while making full use of the reuse of existing business and agile requirement iteration capabilities, thereby shortening the application construction and maintenance cycle and cost, and further improving the level of IT service business.
[0036] 2. This invention uses a low-rank decomposition weight matrix to update the model, which reduces the training time of models with a large number of parameters.
[0037] 3. This invention extracts knowledge features and example features from natural language, enhancing the expressive power of natural language. As input to the model, this enhances the model's performance.
[0038] 4. This invention automatically outputs the instructions required to generate business arrangement documents, reducing the time and effort required for manual arrangement.
[0039] 5. This invention uses knowledge from the fields of medical care, elderly care, and rehabilitation, as well as examples of business orchestration documents, as corpus to add capabilities that large models do not possess, enabling them to generate business orchestration documents and provide corresponding justifications. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating a method for generating business orchestration documents in the medical, elderly care, and rehabilitation field based on a large model, as described in this embodiment.
[0041] Figure 2 This is a schematic diagram of the overall framework of a method for generating business orchestration documents in the field of medical care, elderly care and rehabilitation based on a large model, as described in this embodiment. Detailed Implementation
[0042] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate the understanding of the present invention and do not constitute any limitation thereof.
[0043] like Figure 1 and Figure 2 As shown, a method for generating business orchestration documents in the medical, elderly care, and rehabilitation field based on a large model includes the following steps:
[0044] Step S1 involves collecting textual documents, expert consensus, guidance manuals, standard operating procedures, textbooks, and other materials in the fields of medical care, elderly care, and rehabilitation. After knowledge quality screening, retrieval vectors for knowledge fragments are generated to form a knowledge base for the fields of medical care, elderly care, and rehabilitation, which constitutes the knowledge retrieval module.
[0045] As one specific implementation method:
[0046] An artificial intelligence model can be abstractly divided into an input layer, a parameter layer, and an output layer. In this embodiment, the input layer accepts one-hot vectors. That is, assuming the number of characters in the natural language dictionary is N, and the index of a certain character is i, the one-hot vector of that character is an N-dimensional vector where the i-th position is 1 and the remaining N-1 positions are 0. Using existing Chinese and English corpora from the internet as training data, the model is trained by determining whether a word morpheme is an element in the corpus. After the model training converges, the fully connected parameter layer and the natural language one-hot vector input layer are retained. Generally, the word embedding of each text is equal to the average of the word embeddings of all characters in that text, thus obtaining the module that converts natural language into multi-dimensional floating-point vectors (referred to as word embeddings in this paper).
[0047] The intermediate parameter layer can use commonly used natural language processing models such as RNN, LSTM, TRANSFOMER, and BERT. In this embodiment, a pre-trained BERT model is used, eliminating the need for additional training costs. This embodiment only needs to retain the input layer that converts natural language into one-hot vectors and the parameter layer that converts one-hot vectors into multi-dimensional floating-point vectors, which are called word embeddings.
[0048] The intersection of knowledge in the fields of medicine, elderly care, and rehabilitation lies in the treatment and subsequent rehabilitation of geriatric diseases. Therefore, the focus should be on collecting treatment procedures and expert consensus on geriatric diseases in the medical field, rehabilitation procedures in the rehabilitation field, and care procedures for geriatric diseases in the elderly care field. In addition to knowledge in key areas, knowledge in other areas of medicine, elderly care, and rehabilitation also needs to be collected in preparation for new situations.
[0049] Train or employ a neural network model with a recurrent neural network as the backbone, taking text as input and outputting similarity scores to medical textbook text. Filter the collected knowledge based on at least 80% similarity or a similarity score adjusted by the user.
[0050] After filtering, in this embodiment, knowledge is segmented manually based on natural knowledge segmentation and a maximum limit of 4096 characters per knowledge segment. After segmentation, the knowledge is converted into fragment word embeddings through a natural language word embedding module. In this embodiment, a fragment word embedding is the average number of word embeddings for each character in a fragment. Then, the word embeddings of the knowledge fragments are stored in pairs with the knowledge fragments in a MySQL database.
[0051] This module aims to retrieve knowledge fragments related to a given text. After inputting the text into the knowledge retrieval module, the embedded terms for the search query are obtained. With vectors in the vector knowledge base Calculate cosine distance Then, the knowledge fragments corresponding to the k vectors with the minimum cosine distance are retrieved from the database and output. In this embodiment, k = 14.
[0052] Step S2: Collect business orchestration files, call the prompt words to generate the large model to generate business orchestration files, and store the orchestration files, generation instructions, and related knowledge into the example database to form the example retrieval module.
[0053] As one specific implementation method:
[0054] The business orchestration file described in this embodiment is a business orchestration file conforming to the BPMN 2.0 standard, and the fields involved are medical care, elderly care, and rehabilitation. After collecting the file, it is input into the knowledge retrieval module described in S1 to retrieve knowledge as search keywords for examples, which also serve as components of the example features.
[0055] Simultaneously, the example file will be generated via prompts from the large model API or in a dialogue format, providing instructions for generating the example file. The dialogue format is similar to, but not limited to, the following: "This is a business orchestration document [orchestration document] for the medical, elderly care, and rehabilitation fields. Now I need to use the large language model to generate such a document. Please tell me the relevant instructions." The large model has the ability f(instruction) = document, and also f... -1 (File) = Capabilities of instructions. The generated instructions are part of the example features.
[0056] Knowledge, instructions, and files are stored in an example retrieval database. In this embodiment, MySQL is used. When it is necessary to obtain example features involving knowledge, it can be retrieved based on the knowledge.
[0057] Step S3 involves using the example content as corpus and fine-tuning the weight matrix of the file generation model by updating the weight matrix decomposition matrix. Based on natural language, the required knowledge and examples are retrieved to form the final prompt words, which are then used to generate the business arrangement document using the business generation model. This constitutes the file generation module.
[0058] As one specific implementation method:
[0059] In the example database, the training files are updated to generate a large model. The model W to be updated is defined as W = W0 + ΔW, where W0 is the initial model parameter matrix or the model parameter matrix after the previous update. Updating the entire W is regarded as updating ΔW. ΔW is decomposed into A and B with smaller dimensions, ΔW = AB, where ΔW is m*n dimensional, A is m*2 dimensional, and B is 2*n dimensional. Then, only A and B need to be updated to complete the model update. This method helps to speed up the training process and reduce the computational resources consumed by training.
[0060] Using natural language, typically medical records or descriptions of symptoms, necessary knowledge features are retrieved from the knowledge base described in S1, and necessary example features are retrieved from the example library described in S2. These three are then combined as the final prompt words and input into the large model to obtain the business orchestration file.
[0061] Step S4 involves generating a score for the business orchestration file using the file quality model, feeding the score back to the larger model, and continuing this process until a business orchestration file that meets the requirements is generated. A rationale for generating the business orchestration file is also provided. Finally, the generated business orchestration file is stored in the sample database.
[0062] As one specific implementation method:
[0063] Train a reinforcement learning model that uses business orchestration files in S2 as the dataset, modifies tokens as actions, uses the distribution of business orchestration file tokens as states, and rewards the model based on the matching degree with business files in the dataset. The output is a score for the business orchestration file. The reward is defined as: T = -∑p(x)log(q(x)), where p(x) is the label of whether the hierarchy, order, attributes, and specific content of a component described in the above steps are correct (1 if correct, 0 otherwise); log(·) is the natural logarithm function; and q(x) is the probability that the model predicts the hierarchy, order, attributes, and specific content of the component.
[0064] The business orchestration file generated in S3 is input into the file scoring model to obtain a score, which is then fed back into the large file generation model. The file is modified until the score meets the set requirements.
[0065] The large model is required to output the reasons for generating the business orchestration file in this way for user reference, and the generated file is stored in the example database described in S2 to update the file to generate the large model.
[0066] The knowledge retrieval module can retrieve relevant knowledge based on natural language as required in claim 1, and this knowledge, as part of the prompt words, helps to generate higher-quality business orchestration documents.
[0067] Based on the same inventive principle, this invention also provides a business orchestration document generation system for the medical, elderly care and rehabilitation field based on a large model, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned business orchestration document generation method for the medical, elderly care and rehabilitation field.
[0068] Based on the same inventive principle, this invention also provides a business orchestration document generation system for the medical, elderly care, and rehabilitation field based on a large model, comprising:
[0069] The example retrieval module can retrieve similar and related business arrangement documents and their generation instructions based on natural language, which can then be used by the document generation module.
[0070] The example retrieval module can retrieve similar and related business arrangement documents and their generation instructions based on natural language, which can then be used by the document generation module.
[0071] The document generation module generates business arrangement documents based on prompts composed of knowledge, document generation instructions, and example documents.
[0072] The file scoring module scores the files generated by the file generation module and sends feedback to the file generation module. If the score meets the requirements, the final file is output; otherwise, the file generation module is asked to generate a new file.
[0073] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for generating business orchestration documents in the medical, elderly care, and rehabilitation field based on a large model, characterized in that, Includes the following steps: (1) Collect textual materials in the fields of medical care, elderly care, and rehabilitation; screen the materials for knowledge quality; generate retrieval vectors for knowledge fragments; and construct a knowledge database for the fields of medical care, elderly care, and rehabilitation. The specific process is as follows: (1-1) Using one-hot vectors constructed from natural language dictionaries as input, and judging whether they are legal natural language as output, train an artificial intelligence model. Take the model parameters of the middle fully connected layer to obtain a conversion module with natural language as input and word embedding as output. (1-2) Collect knowledge in the fields of medical care, elderly care and rehabilitation described in natural language, including literature, expert consensus, guidance manuals, standard operating procedures and textbook materials for dealing with diseases of the elderly; A recurrent neural network is trained and used to output the similarity between the collected data and the textbook text, and data with a similarity of more than 80% is selected to form high-quality knowledge. (1-3) Divide the selected knowledge into segments according to the maximum segment size not exceeding a certain threshold. Input the segmented knowledge into the natural language-word embedding conversion module obtained in step (1-1) to obtain multidimensional vectors. Store the multidimensional vectors and knowledge in pairs into the knowledge database. (2) Collect business orchestration files, call prompt words to generate large models to obtain the business orchestration file generation instructions, and store the orchestration files, generation instructions, and related knowledge into the sample database; (3) Using the content in the example database as corpus, the weight matrix of the file generation model is fine-tuned by updating the decomposition matrix of the increment matrix; the specific process is as follows: The training files are updated using the corpus in the example database to generate a large model, and the model to be updated is then... Defined as ,in, It is either the initial model parameter matrix or the model parameter matrix updated in the previous round, updating the entire... Consider as an update ,Will Decompose into smaller dimensions and , Then you only need to update them separately. and This means completing the model update; (4) During the application process, for natural language data, the required knowledge is retrieved from the knowledge database and the required examples are retrieved from the example database to form the final prompt words. Then, the fine-tuned file is called to generate a large model and obtain the business arrangement file. (5) Generate a score for the business orchestration file through the document scoring model, feed the score back to the document generation model until a business orchestration file with a score that meets the set requirements is generated, and give the reason for generating the business orchestration file; finally, store the generated business orchestration file in the sample database.
2. The method for generating business orchestration documents in the medical, elderly care, and rehabilitation field based on a large model according to claim 1, characterized in that, The specific process of step (2) is as follows: (2-1) Collect business arrangement documents related to the fields of medical care, elderly care and rehabilitation, input them into the knowledge database of step (1), and retrieve the relevant knowledge; (2-2) Using the business orchestration file from step (2-1) as input, call the prompt words to generate a large model output to generate a similar orchestration file; (2-3) Store the formatted files, generated instructions, and related knowledge as examples in the example database.
3. The method for generating business orchestration documents in the medical, elderly care, and rehabilitation field based on a large model according to claim 2, characterized in that, In steps (2-3), the arrangement files and generation instructions are in one-to-one correspondence, and one knowledge corresponds to one or more arrangement files.
4. The method for generating business orchestration documents in the medical, elderly care, and rehabilitation field based on a large model according to claim 1, characterized in that, In step (4), for data in natural language form, the required knowledge is retrieved from the knowledge database, specifically as follows: After inputting a piece of natural language text into a knowledge database, the embedding of the target word is obtained. The cosine distance is calculated between the target word and each vector in the knowledge database, and then the target word is retrieved from the database. Output the knowledge fragments corresponding to the vectors with the minimum cosine distance.
5. A business orchestration document generation system for the medical, elderly care, and rehabilitation field based on a large model, characterized in that, The device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the method for generating business orchestration documents in the medical and elderly care field as described in any one of claims 1-4.
Citation Information
Patent Citations
Visual application component arrangement method and system
CN111522543A
Visual task arrangement method and device and storage medium
CN113448547A
Generative large language model training method and model-based search method
CN116127020A
Application program interface scheduling system and method based on large-scale language model
CN117033030A