Method and equipment for analyzing adverse event information of clinical test

Automatically analyze subject audio information through speech recognition and knowledge-enhanced large language model, solving the problem of low efficiency in adverse event information recording in drug clinical trials, and achieving rapid and accurate adverse event information analysis and data management.

CN120496883APending Publication Date: 2025-08-15GENERAL HOSPITAL OF PLA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510567807.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In drug clinical trials, subjects' descriptions of their own symptoms are diverse and vague, resulting in researchers being inefficient in recording and understanding information about adverse events and being prone to misremembering or omissions.

Method used

Pre-trained speech recognition model and knowledge-enhanced large language model are used to automatically analyze subjects' audio information, extract key information and identify named entities to construct adverse event information.

Benefits of technology

It realizes rapid and accurate analysis of adverse event information, reduces manual interference, can efficient batch processing, strong scalability, tap potential factors, provide comprehensive information, and supports data traceability and optimization algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496883A_ABST
    Figure CN120496883A_ABST
Patent Text Reader

Abstract

The invention provides a method and equipment for analyzing adverse event information of a clinical test. The method comprises the following steps: acquiring audio information of a target object; outputting a corresponding pinyin sequence based on the audio information by using a pre-trained speech recognition model; performing text prediction on the pinyin sequence to obtain an output text; in response to the fact that the knowledge-enhanced large language model is utilized to determine that an adverse event exists in the output text, key information in the output text is extracted; performing named entity identification based on the key information to obtain name information, time information and degree information in the key information; and constructing adverse event information based on the name information, the time information and the degree information. By implementing the method and the device, the analysis processing can be automatically performed according to the description of the target object to obtain the corresponding adverse event information, compared with manual processing, the speed is higher, the method and the device are not interfered by clinical environments or other factors, and analysis tasks can be efficiently completed in batches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and device for analyzing adverse event information of clinical trials. Background Art

[0002] In drug clinical trials, subjects' descriptions of their symptoms are crucial for researchers to assess drug safety and efficacy. However, since subjects often have non-medical backgrounds, their descriptions of their symptoms are often diverse and ambiguous, often using spoken language rather than standard medical terminology. Furthermore, differences in subjects' educational backgrounds, cultural backgrounds, language skills, accents, and personal perceptions further complicate researchers' recording and understanding of adverse event (AE) information, making it easy for researchers to misremember or omit AE information.

[0003] At present, the technology for recording and analyzing subject symptoms in clinical trials mainly relies on manual operations. Specifically, subjects orally describe their symptoms to researchers during visits, and researchers record these contents in handwritten notes. The handwritten notes are then transcribed into a spreadsheet or clinical trial management system. Finally, researchers manually match the subject's description with standard terms based on the "Common Criteria for the Evaluation of Adverse Events (CTCAE)" or other medical standards, and determine the type and severity of the symptoms. However, this method has many flaws. The manual recording and organization process is inefficient, especially in projects with a large number of subjects, long trial cycles, and those requiring follow-up after discharge. The processing of adverse event data based on subject complaints often becomes a bottleneck. Summary of the Invention

[0004] In view of this, the present invention provides a method and device for analyzing adverse event information of clinical trials.

[0005] In the first aspect, an embodiment of the present invention proposes a method for analyzing adverse event information of clinical trials, including: obtaining audio information of a target object; using a pre-trained speech recognition model to output a corresponding pinyin sequence based on the audio information; performing text prediction on the pinyin sequence to obtain an output text; in response to determining the presence of an adverse event in the output text using a knowledge-enhanced large language model, extracting key information from the output text; performing named entity recognition based on the key information to obtain name information, time information and degree information in the key information; and constructing adverse event information based on the name information, time information and degree information.

[0006] In a second aspect, an embodiment of the present invention provides an analysis device for adverse event information of clinical trials, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that when the at least one processor executes, it can implement the analysis method for adverse event information of clinical trials as described in any implementation method in the first aspect.

[0007] The analysis method and device for adverse event information of clinical trials provided by the embodiments of the present invention can automatically perform analysis and processing based on the description of the target object to obtain corresponding adverse event information. Compared with manual processing, it is faster and is not affected by the clinical environment or other factors, and can complete analysis tasks efficiently and in batches. In addition, it has stronger scalability, can explore and analyze potential factors for the occurrence of adverse events, can notice more hidden clinical manifestations, and provide researchers with more comprehensive information. In addition, it can retain the output information of each stage, realize the traceability of clinical trial data, facilitate the exploration of the accumulation of system errors, and provide sufficient data support for the optimization algorithm model.

[0008] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 is an exemplary system architecture in which the present invention may be applied;

[0011] Figure 2 A flowchart of a method for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0012] Figure 3 A flowchart of another method for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0013] Figure 4 A flowchart of another method for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0014] Figure 5A flowchart of another method for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0015] Figure 6 A flowchart of another method for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0016] Figure 7 A flowchart of another method for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0017] Figure 8 A flowchart of another method for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0018] Figure 9 A structural block diagram of a device for analyzing adverse event information in clinical trials provided by an embodiment of the present invention;

[0019] Figure 10 A schematic diagram of the structure of an electronic device suitable for executing a method for analyzing adverse event information of clinical trials provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0022] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; internal connections between two components; wireless connections or wired connections. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0023] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0024] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the method, apparatus, electronic device, and computer-readable storage medium for analyzing adverse event information of clinical trials of the present invention may be applied.

[0025] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0026] Users can use terminal devices 101, 102, 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, 103 and server 105 may be installed with various applications for enabling information communication between them, such as instant messaging applications.

[0027] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here.

[0028] Server 105 can provide various services through various built-in applications. It should be noted that the data or information required to provide various services can not only be obtained from terminal devices 101, 102, and 103 via network 104, but can also be pre-stored locally on server 105 in various ways. Therefore, when server 105 detects that such data is already stored locally, it can choose to directly obtain such data locally. In this case, exemplary system architecture 100 may also not include terminal devices 101, 102, 103 and network 104.

[0029] Because analyzing and processing the aforementioned data or information may require significant computing resources and significant computing power, the clinical trial adverse event information analysis methods provided in the subsequent embodiments of the present invention are generally performed by a server 105 possessing significant computing power and resources. Accordingly, the clinical trial adverse event information analysis apparatus is generally also located within the server 105. However, it should also be noted that, when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, the terminal devices 101, 102, and 103 may also utilize the relevant applications installed thereon to perform the aforementioned operations delegated to the server 105, thereby outputting the same results as the server 105. In particular, in the presence of multiple terminal devices with varying computing power, if the relevant application determines that the terminal device in which the application is located possesses significant computing power and significant remaining computing resources, the terminal device may be enabled to perform the aforementioned operations, thereby appropriately alleviating the computing pressure on the server 105. Accordingly, the clinical trial adverse event information analysis apparatus may also be located within the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also not include the server 105 and the network 104 .

[0030] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0031] Please refer to Figure 2 , Figure 2 A flowchart of a method for analyzing adverse event information of a clinical trial provided by an embodiment of the present invention, wherein process 200 includes the following steps:

[0032] Step 201: Acquire audio information of a target object.

[0033] This step is intended to be performed by the subject of the analysis method of clinical trial adverse event information (e.g. Figure 1 The server 105 shown in FIG. 10A ) obtains audio information of a target subject. In this embodiment, the target subject may be a patient, doctor, or other relevant personnel describing a condition. The audio information may be a voice input file generated in real time by the target subject while describing the condition, or a pre-recorded audio file by the target subject, etc., although the present invention is not limited thereto. In practical applications, a portable recording device may be used to obtain the audio information.

[0034] Step 202: Use a pre-trained speech recognition model to output a corresponding pinyin sequence based on the audio information.

[0035] This step is intended to have the above-mentioned execution subject input the audio information into a pre-trained speech recognition model and output the corresponding pinyin sequence. For example, the Conformer model can be used as the pre-trained speech recognition model. The Conformer model is a deep learning architecture that combines the advantages of Transformer and Convolutional Neural Network (CNN). It is mainly used for speech recognition (Automatic Speech Recognition, ASR) and audio signal processing tasks. By fusing the capabilities of global dependency modeling (Transformer) and local feature extraction (CNN), the accuracy of speech recognition can be significantly improved.

[0036] Step 203: Perform text prediction on the pinyin sequence to obtain output text.

[0037] This step is to perform text prediction on the pinyin sequence by the execution subject to obtain the corresponding output text. In this embodiment, the output text is mainly Chinese text.

[0038] Step 204: In response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, extract key information from the output text.

[0039] This step aims to use the knowledge-enhanced large language model (such as LLM) to identify the information in the output text and determine whether there is an adverse event. If it is determined that an adverse event is present, the key information in the output text is extracted.

[0040] In medicine, an adverse event (AE) is any unfavorable or unexpected medical occurrence that occurs during or after a patient undergoes a medical intervention (such as medication, surgery, or examination). These events may or may not be related to the medical procedure, but they need to be recorded and analyzed to assess safety. This knowledge-enhanced large language model determines whether the output text contains relevant content related to adverse events and further extracts information from them.

[0041] Step 205: Perform named entity recognition based on the key information to obtain name information, time information and degree information in the key information.

[0042] This step is intended to allow the aforementioned execution subject to perform named entity recognition based on the key information, thereby obtaining the name information, time information, and severity information contained in the key information. For example, in the medical field, the recognized named entities may include: the standard medical term name corresponding to the symptom, the relevant time information, and the severity of the symptom.

[0043] Step 206: Construct adverse event information based on the name information, time information, and severity information.

[0044] This step is intended to allow the above-mentioned execution subject to construct adverse event information based on the name information, time information and degree information in the key information after obtaining this information. The constructed adverse event information can be used as a basis for subsequent analysis, display, etc.

[0045] The analysis method for clinical trial adverse event information provided by the embodiment of the present invention can automatically perform analysis and processing based on the description of the target object to obtain corresponding adverse event information. It is faster than manual processing and is not affected by the clinical environment or other factors. It can complete analysis tasks efficiently and in batches. In addition, it has stronger scalability, can explore and analyze potential factors for the occurrence of adverse events, can notice more hidden clinical manifestations, and provide researchers with more comprehensive information. In addition, it can retain the output information of each stage, realize the traceability of clinical trial data, facilitate the exploration of the accumulation of system errors, and provide sufficient data support for the optimization algorithm model.

[0046] In some optional implementations of this embodiment, the audio information of the target object may be a recording file or a voice input file, and the human voice frequency region weighting processing is performed based on the recording file or the voice file to convert it into a spectrum sequence. Figure 3 As shown, step 202, the process of using the pre-trained speech recognition model to output the corresponding pinyin sequence based on the audio information, mainly includes:

[0047] Step 301: Input the spectrum sequence into a pre-trained speech recognition model, perform feature extraction through an encoder, and obtain encoder features.

[0048] The purpose of this step is to have the above-mentioned execution entity input the spectrum sequence into the pre-trained speech recognition model, perform feature extraction through the encoder, and extract the high-dimensional representation H through the multi-head attention and convolution modules.

[0049] Step 302: Attentionally interact the current decoding state with the encoder features to generate a pinyin probability distribution for each candidate pinyin sequence.

[0050] This step aims to perform attention interaction calculation on the current decoding state and the encoder features, apply weight bias to the key vectors of symptom-related pinyin, and generate the pinyin probability distribution of each candidate pinyin sequence.

[0051] Step 303: Determine the output pinyin sequence based on the pinyin probability distribution.

[0052] This step aims to determine the output pinyin sequence based on the pinyin probability distribution.

[0053] In some optional implementations of this embodiment, such as Figure 4As shown in the figure, in step 203, the process of performing text prediction on the pinyin sequence to obtain the output text mainly includes:

[0054] Step 401: Map the pinyin of the pinyin sequence to a pinyin vector in the encoder.

[0055] In this embodiment, the process of performing text prediction on the pinyin sequence to obtain the output text can be implemented based on the encoder-decoder architecture of Transformer or Conformer. First, map the pinyin of the pinyin sequence (e.g., "zhong1") to a pinyin vector in the encoder.

[0056] Step 402: Use convolution or a sliding window to obtain the correlation relationship between adjacent pinyin vectors.

[0057] This step aims to capture the local correlation of adjacent pinyin using convolution or sliding window attention (e.g., "fa1re4" is more likely to be combined as "发热"). Further, long-distance context (such as cross-sentence reference relationships) can also be modeled through Transformer layers.

[0058] Step 403: For each pinyin vector, recall multiple candidate texts from a pre-constructed polyphone dictionary.

[0059] This step aims to recall the corresponding candidate texts for each pinyin from the pre-constructed polyphone dictionary (e.g., "zhong" → "中、种、重...").

[0060] Step 404: Determine the output text based on the multiple candidate texts and the correlation relationship.

[0061] This step aims to have the above-mentioned execution entity determine the final output text from the multiple candidate texts based on the determined correlation relationship. Further, in this step, a professional term dictionary can also be combined for text adjustment. Specifically, first determine the output candidate text from the multiple candidate texts based on the correlation relationship. Then, combine the professional term dictionary to match the text in the output candidate text to obtain the final output text. Exemplarily, for the pinyin "xiong1 tong2", although tong is in the second tone, by combining with the professional term dictionary in medicine for matching, it can be matched as "胸痛" instead of obtaining texts like "兄同" only according to the second tone.

[0062] Through the above process, the accuracy of recognizing the audio information of the target object can be further improved.

[0063] In some optional implementations of this embodiment, such as Figure 5As shown, step 204, in response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, extracts key information from the output text, mainly including:

[0064] Step 501: Generate structured prompt words based on the output text. The structured prompt words are used to guide the knowledge-enhanced large language model to perform adverse event judgment and key information extraction.

[0065] In this embodiment, the execution entity can generate structured prompt words based on the output text. The structured prompt words are used to guide the knowledge-enhanced large language model to determine whether an adverse event has occurred and extract key information. For example, the prompt words can be:

[0066] [Task] Determine whether the following description contains adverse events and extract symptoms, time, and severity:

[0067] [Example] Input: "The patient developed a mild headache that lasted for 2 hours 3 days after taking drug A." → Output: AE present, symptom = headache, time = 3 days after taking the drug, severity = mild.

[0068] [Text to be analyzed] "{user input text}".

[0069] Step 502: Using the knowledge-enhanced large language model, based on a preset professional terminology knowledge base and structured prompt words, determine whether there is a description of an adverse event in the output text.

[0070] In this embodiment, the above-mentioned execution subject can use the structured prompt word to use the knowledge-enhanced large language model to judge whether there is a description of an adverse event in the output text based on a preset professional terminology knowledge base. Among them, the preset professional terminology knowledge base can be set accordingly according to the specific scenario and field of application. The professional terminology knowledge base has strict definitions for terms in professional fields (such as medicine, law, engineering, etc.). By accessing an external knowledge base, LLM can ensure the standardization and consistency of terminology use and avoid "hallucinations" or vague expressions. When LLM lacks domain knowledge, it may generate content that seems reasonable but is actually wrong (such as confusing drug names or legal terms). The knowledge base can provide authoritative references to assist the model in correcting errors. In addition, professional terms may have multiple meanings (such as the different meanings of "cell" in biology and engineering). The knowledge base can provide contextual associations to help the model understand correctly.

[0071] Step 503: In response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, the knowledge-enhanced large language model is used to extract key information from the output text based on a preset professional terminology knowledge base and structured prompt words to obtain key information.

[0072] If the knowledge-enhanced large language model determines that there are adverse events in the output text, the execution entity can further use the knowledge-enhanced large language model to extract key information from the output text based on a preset professional terminology knowledge base and structured prompt words.

[0073] Through the above process, based on the professional terminology knowledge base, the enhanced large language model can more accurately understand the content of the output text, make accurate judgments on adverse events, and extract corresponding key information.

[0074] In some optional implementations of this embodiment, such as Figure 6 As shown, step 205, the process of performing named entity recognition based on key information, mainly includes:

[0075] Step 601: Use a ternary tagging method to tag and identify named entities in key information, and obtain candidate entities of various types that are tagged with entity start tags and entity middle tags.

[0076] In this embodiment, the ternary annotation method refers to the BIO annotation method, which is a sequence annotation method used in natural language processing (NLP) for tasks such as named entity recognition (NER). By simply marking the entity's start, inside, and outside positions, the BIO annotation method helps the model identify entity boundaries and types in text. Among them, B (Begin) indicates the beginning of a named entity. For example, "B-Person" indicates the beginning of a person's name. I (Inside) indicates the middle part of a named entity. For example, "I-Person" indicates the subsequent part of a person's name. O (Outside) indicates that the word does not belong to any named entity. For example, B-SYMPTOM (symptom start), I-SYMPTOM (symptom follow-up); B-TIME_EXACT (exact time start), I-TIME_EXACT (exact time follow-up); B-TIME_FUZZY (fuzzy time start), I-TIME_FUZZY (fuzzy time follow-up); B-GRADE (severity start), I-GRADE (grade follow-up), etc.

[0077] Step 602: Merge the entity start tag and entity middle tag of the same candidate entity to obtain name information, time information and degree information.

[0078] After labeling using the BIO labeling method, continuous B-labels and I-labels can be merged based on a unified candidate entity to obtain the corresponding named entity, that is, the name information, time information, and degree information in the key information can be identified.

[0079] During specific implementation, the time information in the above key information may be the fuzzy time obtained, such as expressions such as "a few days ago, last week". For this type of fuzzy time information, in this embodiment, the execution subject may also perform fuzzy time analysis on it to obtain a relatively certain time point. For example, the execution subject may map the fuzzy description in the time information to specific time point information based on the time parsing algorithm of Bayesian inference, combined with the visit schedule (such as the nth visit time) and / or the clinical trial timeline (such as the start date of the clinical trial). For example, "a few days ago" may correspond to the past 3-5 days; "last week" may correspond to the past 7-14 days. For the fuzzy time of "a few days ago", combined with the visit date (for example, March 20, 2025), it is parsed as March 15 to March 17, 2025, thereby obtaining specific time point information.

[0080] In some optional implementations of this embodiment, the execution entity may also combine causal reasoning to analyze the impact of confounding variables (such as age, underlying diseases, and concomitant medications) on the CTCAE grade of adverse events and make adaptive adjustments.

[0081] In this causal inference process, the variables include: Independent variables: drug (e.g., "aspirin"), dosage, patient characteristics (age, gender, underlying diseases), and concomitant medications. Dependent variable: CTCAE grade of adverse event (e.g., "grade 2"). Confounding variables: age, underlying diseases (e.g., "diabetes"), and concomitant medications (e.g., "antihypertensive drugs").

[0082] Based on the above variables, the impact of confounding variables on the adverse event classification is analyzed. Specifically, the corresponding preset standard indicators are determined based on the confounding variables; the impact relationship is determined based on the correspondence between the severity information and the preset standard indicators. For example, the relationship between symptom severity (such as "pain score 8 / 10") and abnormal laboratory indicators (such as "transaminase increased to 3 times the normal value") is examined.

[0083] After determining the impact relationship, the grading of the adverse event can be adjusted based on the impact relationship. Correspondingly, if there is a dose effect (e.g., higher doses result in more severe symptoms), the grading of the adverse event needs to be increased.

[0084] Please refer to Figure 7 , Figure 7 This is a flowchart of another method for analyzing adverse event information of a clinical trial provided by an embodiment of the present invention, wherein process 700 includes the following steps:

[0085] Step 701: Acquire audio information of a target object.

[0086] Step 702: Use the pre-trained speech recognition model to output the corresponding pinyin sequence based on the audio information.

[0087] Step 703: Perform text prediction on the pinyin sequence to obtain output text.

[0088] Step 704: In response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, extract key information from the output text.

[0089] Step 705: Perform named entity recognition based on the key information to obtain name information, time information and degree information in the key information.

[0090] Step 706: Construct adverse event information based on the name information, time information, and severity information.

[0091] The above steps 701-706 are similar to the following Figure 2 Steps 201-206 shown are consistent. For the same content, please refer to the corresponding part of the previous embodiment and will not be repeated here.

[0092] Step 707: Calculate the attention weights of the words within the same entity in the key information to obtain the entity fine-grained vector.

[0093] In this embodiment, the attention weights of the words within the entity are calculated to obtain the entity fine-grained vector, thereby strengthening the key descriptive words (such as "high fever" in "persistent high fever").

[0094] Step 708: With each entity as the center, obtain the semantic association between entities based on the entity fine-grained vector through cross attention.

[0095] In this embodiment, cross-attention is used to obtain semantic associations between entities based on fine-grained entity vectors, and attribution analysis and grading of adverse reactions are performed. This includes drug-symptom relationships, symptom-time relationships, etc. For example:

[0096] Drug-symptom: For example, "aspirin-nausea", which indicates symptoms that the drug may cause.

[0097] Symptom-time: For example, "nausea-a few days ago", which indicates the time when the symptoms occurred.

[0098] Step 709: Construct a knowledge spectrum with each entity as a node and the semantic associations between entities as edges.

[0099] After determining the semantic associations between entities, the knowledge spectrum can be constructed using each entity as a node and the semantic associations between entities as edges. For example, the nodes may include: drugs (such as "aspirin"), symptoms (such as "nausea"), organs (such as "stomach"), diseases (such as "gastritis"), genes (such as "COX-1"), and proteins. Edges are used to characterize relationship types, such as: "aspirin-inhibits-COX-1" (drug-target), "nausea-caused by gastritis" (symptom-disease), "stomach-function-digestion" (organ-function). This knowledge spectrum can be used to establish the pathological associations between drug mechanisms of action, symptom clinical manifestations, and human organ systems, forming a three-dimensional network of medical knowledge relationships that can be inferred.

[0100] Please refer to Figure 8 , Figure 8 A flowchart of another method for analyzing adverse event information of a clinical trial provided by an embodiment of the present invention, wherein process 800 includes the following steps:

[0101] Step 801: Acquire audio information of a target object.

[0102] Step 802: Use a pre-trained speech recognition model to output a corresponding pinyin sequence based on the audio information.

[0103] Step 803: Perform text prediction on the pinyin sequence to obtain output text.

[0104] Step 804: In response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, extract key information from the output text.

[0105] Step 805: Perform named entity recognition based on the key information to obtain name information, time information and degree information in the key information.

[0106] Step 806: Construct adverse event information based on the name information, time information, and severity information.

[0107] The above steps 801-806 are similar to the following Figure 2 Steps 201-206 shown are consistent. For the same content, please refer to the corresponding part of the previous embodiment and will not be repeated here.

[0108] Step 807: Construct an adverse event report based on the adverse event information.

[0109] In this embodiment, the executing entity can use a large language model to aggregate adverse event information and generate an adverse event report. The resulting formatted adverse event report has a clear structure, complete information, and is easy to analyze and archive. Furthermore, the standardized and unified adverse event report generated through the above process can eliminate human error, ensure the consistency and standardization of report content, and improve data reliability.

[0110] Furthermore, the method for analyzing clinical trial adverse event information may also include:

[0111] Step 808: Obtain correction information for the adverse event report.

[0112] The correction information may be the correction information entered by the relevant personnel after obtaining the adverse event report based on the content of the report. The correction information may include: terminology standardization adjustment, missing information filling, etc.

[0113] Step 809: Adjust the adverse event report based on the correction information.

[0114] After obtaining the correction information, the adverse event report can be adjusted in combination with the correction information.

[0115] Furthermore, to ensure the security and tamper-proof nature of adverse event reports, differential privacy can be applied to these reports. Specifically, a noise signal is obtained based on the Laplace distribution and then injected into key fields in the adverse event report using differential privacy.

[0116] Through the above process, the security and privacy of the generated report content can be improved to achieve privacy protection of relevant information of the measured object.

[0117] Furthermore, a self-supervised incremental learning method can be used to perform incremental learning on various model algorithms involved in any of the above embodiments to flexibly adapt to new modalities and tasks.

[0118] To deepen understanding, the present invention also provides a specific implementation scheme in combination with a specific application scenario. In this application scenario, the target object is a subject.

[0119] The subject's main complaint was: After taking the medicine, I felt a little uncomfortable starting around 8:30 last night. I started looking at my phone too much and felt a slight headache, so I went to bed. I probably went to bed at 10 o'clock. At first, I felt fine, but then I woke up around 3 or 4 o'clock in the middle of the night. I felt a little dizzy and went to the bathroom. When I stood up, I felt a slight headache. After going to the bathroom, the headache was still there and I couldn't sleep. It was especially uncomfortable after lying down, so I slept half-leaning, but I couldn't fall asleep. My head still hurt until morning, and I felt like it was spinning. I didn't want to eat breakfast, and then I had a blood test. It's 10 o'clock now and I still don't feel very well.

[0120] Step 1: Based on the above main complaints, adverse event information was extracted, and the adverse reaction names obtained were: headache, dizziness, insomnia, and loss of appetite.

[0121] Step 2: Perform named entity recognition on the adverse event information and obtain the following information:

[0122] Adverse reaction name:

[0123] Headache, dizziness, insomnia, loss of appetite.

[0124] Time information:

[0125] Last night at 8:30; 10 o'clock; 3 or 4 o'clock in the middle of the night; in the morning; 10 o'clock.

[0126] Degree Information:

[0127] Headache → mild to moderate;

[0128] Vertigo → mild;

[0129] Insomnia → moderate;

[0130] Loss of appetite → Mild.

[0131] Step 3: Analyze the fuzzy time to get the specific time point:

[0132] 3 or 4 o'clock in the middle of the night → 4 o'clock in the morning;

[0133] Morning → According to general custom, it is from 6:00 to 8:00 am.

[0134] Step 4: Perform semantic association on the above information to obtain the corresponding association relationship:

[0135] Drug-Time-Symptoms: Trial drug - 8:30 pm yesterday - headache;

[0136] Time-Symptoms: 3 or 4 in the middle of the night - dizziness, headache, insomnia;

[0137] Time-Symptoms: Morning-Loss of appetite.

[0138] Step 5: Perform causal reasoning:

[0139] Adverse events occurred after taking the trial drug;

[0140] Blood test results were unchanged from baseline.

[0141] Step 6: Based on the above information, a three-dimensional knowledge spectrum of drug-symptom-organ can be constructed.

[0142] The investigational drug is a highly selective inhibitor of ROCK2. ROCK is an evolutionarily conserved serine / threonine kinase that interacts with RhoAGTP. Both ROCK isoforms, ROCK1 and ROCK2, play key roles in actin cytoskeleton dynamics, controlling cell shape, cell migration, cell motility, cell proliferation, and apoptosis. Previous studies have found that inhibiting either ROCK1 or ROCK has anti-fibrotic effects. However, ROCK1 is a cardiomyocyte-specific kinase that can induce cardiac hypertrophy; cROCK1- / - mice have enhanced fibrosis, while cROCK2- / - mice have reduced cardiac hypertrophy. ROCK2 is primarily expressed in vascular and cardiac tissues, including the lungs.

[0143] Based on the knowledge graph analysis, it is believed that the adverse events are not related to the drugs.

[0144] Step 7: An adverse event report can be constructed based on the above information.

[0145] Based on the subjects' descriptions, we can analyze the following details of adverse events:

[0146] 1. Adverse reactions: headache, dizziness, insomnia, loss of appetite

[0147] 2. Time of occurrence:

[0148] -Start time: 8:30 last night

[0149] - Lasts until: This morning and until the current time (10:00)

[0150] 3. Type:

[0151] - Nervous system: headache, dizziness, insomnia

[0152] -Digestive system: loss of appetite

[0153] 4. Severity (according to CTCAE common adverse reaction classification):

[0154] - Headache: Mild to moderate (described by participants as "a little headache" or "a fair headache" and affecting sleep, but not requiring treatment or significantly interfering with daily life)

[0155] - Dizziness: Mild (Participants reported waking up in the middle of the night with a "light-headed feeling" but did not report needing treatment or experiencing significant disruption to their daily lives)

[0156] - Insomnia: Moderate (The subject mentioned, "I later leaned back to sleep, but I didn't fall asleep. My head still hurt until the morning," indicating that my sleep was disturbed to some extent)

[0157] - Loss of appetite: Mild (subject stated "I just don't want to eat breakfast" but did not mention needing medical treatment or severe malnutrition)

[0158] In summary, the subjects experienced mild to moderate headaches, dizziness and insomnia, as well as mild loss of appetite after taking the drug. These symptoms began at 8:30 pm yesterday and lasted until 10 am this morning.

[0159] Furthermore, based on the correction information input by the researcher, for example, "insomnia" can be adjusted to "mild" according to the actual situation, and the model parameters can be updated after adding noise to this adverse event information.

[0160] Further references Figure 9 As an implementation of the methods shown in the above figures, the present invention provides an embodiment of an analysis device for clinical trial adverse event information. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0161] like Figure 9 As shown, the clinical trial adverse event information analysis device 900 of this embodiment may include: an audio information acquisition module 901, a pinyin sequence generation module 902, an output text generation module 903, a key information extraction module 904, a named entity recognition module 905, and an adverse event information construction module 906. The audio information acquisition module 901 is configured to acquire audio information of a target subject; the pinyin sequence generation module 902 is configured to output a corresponding pinyin sequence based on the audio information using a pre-trained speech recognition model; the output text generation module 903 is configured to perform text prediction on the pinyin sequence to obtain an output text; the key information extraction module 904 is configured to extract key information from the output text in response to determining that an adverse event exists in the output text using a knowledge-enhanced large language model; the named entity recognition module 905 is configured to perform named entity recognition based on the key information to obtain name information, time information, and degree information in the key information; and the adverse event information construction module 906 is configured to construct adverse event information based on the name information, time information, and degree information.

[0162] In this embodiment, the specific processing of the audio information acquisition module 901, the pinyin sequence generation module 902, the output text generation module 903, the key information extraction module 904, the named entity recognition module 905 and the adverse event information construction module 906 in the clinical trial adverse event information analysis device 900 and the technical effects thereof can be referred to respectively. Figure 2 The relevant descriptions of steps 201-206 in the corresponding embodiment are not repeated here.

[0163] This embodiment exists as an embodiment of the device corresponding to the above-mentioned method embodiment. The analysis device for clinical trial adverse event information provided by this embodiment can automatically perform analysis and processing based on the description of the target object to obtain corresponding adverse event information. Compared with manual processing, it is faster and is not affected by the clinical environment or other factors. It can complete analysis tasks efficiently and in batches. In addition, it has stronger scalability, can explore and analyze potential factors for the occurrence of adverse events, can notice more hidden clinical manifestations, and provide researchers with more comprehensive information. In addition, it can retain the output information of each stage, realize the traceability of clinical trial data, facilitate the exploration of the accumulation of system errors, and provide sufficient data support for the optimization algorithm model.

[0164] Figure 10 A schematic block diagram of an example clinical trial adverse event information analysis device 1000 that can be used to implement an embodiment of the present invention is shown. Electronic equipment is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic equipment can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0165] like Figure 10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0166] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0167] The computing unit 1001 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as the analysis method of clinical trial adverse event information. For example, in some embodiments, the analysis method of clinical trial adverse event information can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the analysis method of clinical trial adverse event information described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the method for analyzing clinical trial adverse event information in any other appropriate manner (eg, by means of firmware).

[0168] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0169] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0170] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0172] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A method for analyzing adverse event information in clinical trials, characterized in that: include: Get the audio information of the target object; Outputting a corresponding pinyin sequence based on the audio information using a pre-trained speech recognition model; Performing text prediction on the pinyin sequence to obtain output text; In response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, extracting key information from the output text; Perform named entity recognition based on the key information to obtain name information, time information and degree information in the key information; Adverse event information is constructed based on the name information, time information and severity information.

2. The method according to claim 1, characterized in that The step of obtaining audio information of a target object includes: Obtaining a recording file or a voice input file of the target object; Based on the recording file or the voice input file, a weighted processing of the human voice frequency region is performed and converted into a spectrum sequence.

3. The method according to claim 2, characterized in that The method of using a pre-trained speech recognition model to output a corresponding pinyin sequence based on the audio information includes: Inputting the spectrum sequence into the pre-trained speech recognition model, performing feature extraction through an encoder, and obtaining encoder features; Performing attention interaction between the current decoding state and the encoder features to generate a pinyin probability distribution for each candidate pinyin sequence; An output pinyin sequence is determined based on the pinyin probability distribution.

4. The method according to claim 1, wherein The performing text prediction on the pinyin sequence to obtain output text includes: Mapping the pinyin of the pinyin sequence into a pinyin vector in an encoder; Use convolution or sliding window to obtain the correlation between adjacent pinyin vectors; For each of the pinyin vectors, a plurality of candidate texts are respectively recalled from a pre-built polyphonetic dictionary; The output text is determined based on the multiple candidate texts and the associated relationships.

5. The method according to claim 4, characterized in that The determining the output text based on the multiple candidate texts and the association relationship includes: Determine an output candidate text from the plurality of candidate texts based on the association relationship; In combination with a professional term dictionary, the output candidate text is matched with the text in the professional term dictionary to obtain the output text.

6. The method according to claim 1, characterized in that In response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, extracting key information from the output text includes: Generating structured prompt words based on the output text, wherein the structured prompt words are used to guide the knowledge-enhanced large language model to perform adverse event judgment and key information extraction; Using the knowledge-enhanced large language model to determine whether there is a description of an adverse event in the output text based on a preset professional terminology knowledge base and the structured prompt words; In response to determining that an adverse event exists in the output text using the knowledge-enhanced large language model, the knowledge-enhanced large language model is used to extract key information from the output text based on the preset professional terminology knowledge base and structured prompt words to obtain the key information.

7. The method according to claim 1, characterized in that The performing named entity recognition based on the key information includes: Using a ternary tagging method to tag and identify named entities in the key information, and obtaining candidate entities of various types that are tagged with entity start tags and entity middle tags; The entity start label and entity middle label of the same candidate entity are merged to obtain the name information, time information and degree information.

8. The method according to claim 7, characterized in that The performing named entity recognition based on the key information further includes: A time parsing algorithm based on Bayesian inference is used to map the fuzzy description in the time information into specific time point information in combination with the visit schedule and / or clinical trial timeline.

9. The method according to any one of claims 1 to 8, characterized in that Also includes: Obtain confounding variables related to the adverse events, including: patient characteristics, drug name, underlying disease, and concomitant medication; Analyze the influence of the confounding variables on the grading of the adverse events; The grading of the adverse event is adjusted based on the impact relationship.

10. The method according to claim 9, characterized in that The analyzing the influence of the confounding variables on the adverse events includes: Determining corresponding preset standard indicators based on the confounding variables; The influence relationship is determined based on the correspondence between the degree information and the preset standard indicator.

11. The method according to any one of claims 1 to 8, characterized in that: Also includes: Calculate the attention weights of the words within the same entity in the key information to obtain a fine-grained entity vector; Centered on each entity, the semantic association between entities is obtained based on the entity fine-grained vector through cross attention; A knowledge spectrum is constructed with each entity as a node and the semantic associations between entities as edges.

12. The method according to any one of claims 1 to 8, characterized in that: Also includes: An adverse event report is constructed based on the adverse event information.

13. The method according to claim 12, characterized in that Also includes: Obtaining correction information for the adverse event report; The adverse event report is adjusted based on the revised information.

14. The method according to claim 12 or 13, characterized in that Also includes: Obtaining a noise signal based on Laplace distribution; The noise signal is injected into the key fields of the adverse event report using differential privacy.

15. An analysis device for adverse event information of clinical trials, characterized in that: include: A processor and a memory connected to the processor; wherein the memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the processor to perform the analysis method for clinical trial adverse event information according to any one of claims 1 to 14.

Citation Information

Cited By

  • Drug adverse event grading prediction method and device based on local data

    CN121790017A