Method, system and storage medium for generating structured report based on voice data

By converting voice data into text data and combining data analysis and processing steps to generate structured reports, the problem of low writing efficiency of structured reports in the prior art is solved, and the effect of rapid entry and efficient generation of structured reports is achieved.

CN118609747BActive Publication Date: 2025-05-13ZHONGSHAN HOSPITAL AFFILIATED TO FUDAN UNIV XIAMEN HOSPITAL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410723621.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-05-13
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

The writing efficiency of structured reports in the prior art is low, especially in the diagnosis process involving multiple organs, doctors need to spend a lot of time selecting templates and filling in descriptions.

Method used

By obtaining voice data, converting it into text data, and combining data analysis and processing steps, structured report content is generated. Data analysis processing includes word segmentation processing, entity link processing, filtering processing, data matching processing and data filling processing, word segmentation processing is performed using the priority parameters of the preset classification label, and the priority of conjunctions is automatically optimized.

Benefits of technology

It realizes rapid entry of structured reports, improves the writing efficiency of structured reports, and ensures the accuracy of report content by improving the accuracy of word segmentation and the efficiency of generating structured reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118609747B_ABST
    Figure CN118609747B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and storage medium for generating a structured report based on voice data, which includes: obtaining voice data, converting it into text data and displaying it, wherein the voice data includes imaging-related content; performing data parsing processing on the text data to generate structured report content; judging whether there is newly added structured content; if so, obtaining a sub-template or chapter corresponding to the newly added structured content, and displaying it, and then automatically filling in the corresponding structured elements according to the newly added structured content to generate a structured report; if not, directly generating a structured report. The present invention realizes the rapid entry of structured reports and improves the writing efficiency of structured reports by converting voice data into text data and then generating structured report content in combination with data parsing processing steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical information technology, and in particular to a method, system and storage medium for generating a structured report based on voice data. Background Art

[0002] The imaging structured report is a productivity tool that carries the diagnostic logic. The complete imaging diagnosis process includes three stages: image recognition, differential diagnosis, and targeted answers to clinical questions. The diagnostic work form of using structured reports is that the diagnostician reads images and related case materials, continuously inputs various types of label data, and the structured report system automatically forms a diagnostic conclusion based on the built-in expert consensus rules.

[0003] Taking the upper abdominal MR scan as an example, it covers more than 20 possible diseases of multiple organs such as liver, gallbladder, pancreas, spleen, kidney, etc. In order to achieve its structured description, a dedicated structured template is usually set for each disease of each organ in the structured report system, plus normal description and anatomical variation. When a large-scale scan needs to be completed, the diagnosis range may cover 40 to 50 templates.

[0004] In many cases, although patients have important or unimportant lesions in multiple organs, the description of each lesion may only be a simple description of location, type, and size. However, if multiple organs are involved, the doctor's operation process needs to first select the template of the organ and disease, then find the structured chapters corresponding to these simple features, and then fill them out. This operation is quite time-consuming, resulting in low efficiency in writing structured reports. Summary of the invention

[0005] The main purpose of the present invention is to provide a method, system and storage medium for generating a structured report based on voice data, aiming to solve the technical problem of low writing efficiency of structured reports in the prior art.

[0006] To achieve the above-mentioned purpose, the present invention provides a method for generating a structured report based on voice data, which includes the following steps: obtaining voice data, converting it into text data and displaying it, wherein the voice data includes imaging-related content; performing data parsing processing on the text data to generate structured report content; the data parsing processing at least includes: word segmentation processing, entity linking processing, filtering processing, data matching processing and data filling processing; the word segmentation processing is performed based on the priority parameters of the preset classification label, and the priority parameters include preset default priority parameters and / or optimized priority parameters; judging whether there is newly added structured content; if so, obtaining the sub-template or chapter corresponding to the newly added structured content and displaying it, and then automatically filling in the corresponding structured elements according to the newly added structured content to generate a structured report; if not, directly generating a structured report.

[0007] Optionally, the word segmentation processing specifically includes the following steps: obtaining context data of text data; first segmenting the text data and the context data to obtain corresponding word segmentation results, and then performing entity recognition to obtain named entities in the text data and the context data; performing a first matching process on the named entities and the chapters in the structured report respectively to obtain classification labels corresponding to the named entities.

[0008] Optionally, in the first matching process, when the classification label is activated, associated words belonging to the classification label in the word segmentation vocabulary are obtained based on a preset association relationship, and priority parameters of the associated words are automatically optimized to obtain optimized priority parameters.

[0009] Optionally, automatically optimizing the priority parameter of the associated words specifically increases the priority of the associated words.

[0010] Optionally, entity linking processing is specifically as follows: based on a preset knowledge base, linking the word segmentation result with the corresponding named entity to obtain the semantic encoding corresponding to the named entity.

[0011] Optionally, the data matching process specifically includes the following steps: based on the semantic coding corresponding to the named entity, obtaining the corresponding structured report template; the chapters of the structured report template have a hierarchical relationship, and the structured report template has at least one built-in grammatical metadata tree for recording the correspondence between the structured report elements and the natural language text, and the grammatical metadata tree includes at least one grammatical metadata subtree; based on the structured report template, obtaining the grammatical metadata tree or grammatical metadata subtree corresponding to the text data; performing a second matching process on the named entities in the text data and the nodes in the grammatical metadata tree or the nodes in the grammatical metadata subtree to determine the corresponding nodes.

[0012] Optionally, the data filling process is specifically as follows: filling the semantic encoding array obtained by the entity linking process into the grammatical metadata tree or grammatical metadata subtree obtained in the data matching process to generate structured report content.

[0013] Optionally, the filtering process specifically includes: filtering the text data based on preset filtering rules, and not entering the filtered content into the structured report.

[0014] Corresponding to the method for generating a structured report based on voice data, the present invention provides a system for generating a structured report based on voice data, which includes: a voice recognition module, used to obtain voice data, convert it into text data and display it, wherein the voice data includes imaging-related content; a data analysis and processing module, used to perform data analysis on the text data to generate structured report content; the data analysis and processing at least includes: word segmentation processing, entity linking processing, filtering processing, data matching processing and data filling processing; the word segmentation processing is performed based on the priority parameters of the preset classification label, and the priority parameters include preset default priority parameters and / or optimized priority parameters; a structured content addition module, used to determine whether there is newly added structured content; if so, obtain the sub-template or chapter corresponding to the newly added structured content, display it, and then automatically fill in the corresponding structured elements according to the newly added structured content to generate a structured report; if not, directly generate a structured report.

[0015] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which is stored a program for generating a structured report based on voice data. When the program for generating a structured report based on voice data is executed by a processor, the steps of the method for generating a structured report based on voice data as described above are implemented.

[0016] The beneficial effects of the present invention are:

[0017] (1) Compared with the prior art, the present invention converts voice data into text data and generates structured report content in combination with data analysis and processing steps, thereby realizing rapid input of structured reports and improving the writing efficiency of structured reports; and performs word segmentation processing based on priority parameters of preset classification tags, which can effectively improve the accuracy of word segmentation, thereby ensuring the accuracy of the subsequently generated structured report content;

[0018] (2) Compared with the prior art, the present invention performs entity recognition by combining the context information of text data, and when the classification label is activated, it can set a higher priority for the associated words by automatically optimizing the priority parameters of the associated words, thereby further improving the accuracy of word segmentation in the process of converting text into structured content;

[0019] (3) Compared with the prior art, the present invention records the correspondence between structured report elements and natural language texts through a grammatical metadata tree, provides a basis for the conversion between natural language texts and structured content, and can improve the generation efficiency of structured reports;

[0020] (4) Compared with the prior art, the present invention can filter some text data that does not need to be additionally entered into the structured report by filtering the text data based on preset filtering rules, thereby further improving the input speed of the structured report and the generation efficiency of the structured report. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0022] Figure 1 A simplified flow chart of a method for generating a structured report based on speech data according to an embodiment of the present invention;

[0023] Figure 2 A framework diagram of a system for generating structured reports based on speech data according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0025] like Figure 1As shown, a method for generating a structured report based on voice data of the present invention includes the following steps: acquiring voice data, converting it into text data and displaying it, wherein the voice data includes imaging-related content; performing data parsing processing on the text data to generate structured report content; the data parsing processing at least includes: word segmentation processing, entity linking processing, filtering processing, data matching processing and data filling processing; the word segmentation processing is performed based on priority parameters of preset classification labels, and the priority parameters include preset default priority parameters and / or optimized priority parameters; judging whether there is newly added structured content; if so, obtaining the sub-template or chapter corresponding to the newly added structured content and displaying it, and then automatically filling in the corresponding structured elements according to the newly added structured content to generate a structured report; if not, directly generating a structured report.

[0026] It should be noted that there is a certain order relationship between the various sub-steps of data analysis processing, but it is not absolute. In this embodiment, the word segmentation processing is performed first, and then the entity linking processing is performed based on the result of the word segmentation processing, and then the data matching processing is performed based on the result of the word segmentation processing and the result of the entity linking processing, and the data filling processing is performed based on the result of the entity linking processing and the result of the data matching processing. As for the filtering processing, it is performed before or after the entity linking processing.

[0027] Preferably, the voice data is acquired by the user pressing the voice handle to acquire and recognize the voice data. The user can also turn on the monitoring mode. When a key sound is heard, such as "start recording", the voice data acquisition and recognition and subsequent processing will begin. When the voice data is converted into text data and displayed, it is specifically displayed in real time.

[0028] In this embodiment, based on ASR (automatic speech recognition) technology, unlike traditional speech recognition using acoustic models and language models, the present invention uses a Transformer transducer speech recognition model. This is an end-to-end speech recognition technology that can simplify the model training process; and compared with the traditional Transformer model, this is a streaming speech recognition technology, and users can see the recognition results in real time.

[0029] Furthermore, when the user finds that the displayed text data is wrong, the user can give up the input immediately. For example: after starting the input, the user's voice is continuously collected to obtain voice data, and the recognized text data is displayed in real time during the voice data acquisition period; according to the preset voice acquisition pause time threshold, it is judged whether the user's voice data input time interval is greater than the preset voice acquisition pause time threshold. If so, the acquired voice data is identified as a complete voice data, and a status prompt is given with different colors and marks. According to the preset waiting time, it is judged whether the user's voice input is completed. If no new voice data is obtained after exceeding the preset waiting time, it is determined that the voice data acquisition process is over. Because the user can observe the corresponding text data of his voice input in real time during the above voice data acquisition process, when it is realized that there is an error in the text data, it can enter the editing state through the preset correction method (such as shortcut keys or clicking interface buttons) before the preset waiting time ends, and correct the text data. At this time, the user can use the mouse and keyboard to complete the text correction, or directly click the corresponding button to abandon the input. If the user does not have time to correct the error or fails to discover the error in time, resulting in the structured report entering unexpected content, the user can still withdraw the operation before the next voice recording.

[0030] The present invention can help users correct erroneous content of speech recognition in a timely manner by displaying text data in real time, combined with a preset waiting time and correction method, and improve the accuracy of subsequent data analysis and processing of text data.

[0031] In this embodiment, during pre-training, the Transformer transducer model is supplemented with data of imaging-related sentences based on the open source Chinese sound data set. In addition, after pre-training, the user's voice is used for fine-tuning to achieve better use results. If the user's conditions permit, the speech recognition model can be directly installed on the user's computer to increase the response speed. However, if the user's device does not meet the requirements, the model on the LAN server can be used.

[0032] The present invention is mainly used in the preliminary report stage of imaging, targeting scenarios where one scan covers multiple organs, such as upper abdominal MR, pelvic MR, abdominal CT, head CT / MR, chest CT / MR, etc. By converting voice data into text data and combining the data analysis and processing steps to generate structured report content, the rapid entry of structured reports can be achieved, and the writing efficiency of structured reports can be improved; and word segmentation processing based on the priority parameters of the preset classification tags can effectively improve the accuracy of word segmentation to ensure the accuracy of the subsequent generation of structured report content.

[0033] In this embodiment, NLP technology is used to perform data analysis on text data to generate structured report content. The word segmentation process specifically includes the following steps: obtaining context data of the text data; first segmenting the text data and the context data to obtain corresponding word segmentation results, and then performing entity recognition to obtain named entities in the text data and the context data; performing a first matching process on the named entities and the chapters in the structured report to obtain the classification labels corresponding to the named entities.

[0034] Preferably, the BERT-BiLSTM-CRF model is used for entity recognition, such as organ name, disease type, anatomical location, etc., and the named entities are first matched with the chapters in the structured report to obtain the classification labels corresponding to the named entities, so that a group of labels related to the structured content topic can be formed, such as liver, cyst, and imaging description.

[0035] In this embodiment, during the first matching process, when a classification label is activated, associated words belonging to the classification label in the word segmentation vocabulary are obtained based on a preset association relationship, and priority parameters of the associated words are automatically optimized to obtain optimized priority parameters.

[0036] It should be noted that there is a certain correlation between associated words and named entities, but they are not exactly the same. Associated words refer to words related to a certain classification label in a certain scenario, which are used to optimize word segmentation processing to improve the accuracy and efficiency of word segmentation. Named entities are used for entity linking processing, and they can provide type and attribute information of entities in text data. Therefore, an associated word may appear in the word segmentation processing result, or it may be linked to a named entity.

[0037] In this embodiment, the activation of classification labels means that when the contextual information of the examination report, such as the examination purpose, examination items, examination site, clinical information, etc., meets the activation conditions in the preset rules (usually some preset according to the knowledge base), some classification labels are activated. For example: Preset rule 1-"If it is an upper abdominal ultrasound examination, the classification labels describing the upper abdominal organs will be activated." Of course, in order to ensure the accuracy of the activated classification labels, the preset rules may also include restricted activation conditions, for example: Preset rule 2-"If the examiner is a male, the gynecological classification label will not be activated", etc. By setting restricted activation conditions for classification labels specific to examiners of different genders, erroneous activation can be avoided. It can be understood that the preset rules can be preset according to actual needs. The above two rules are just examples for facilitating the understanding of the preset rules and do not constitute an improper limitation of the present invention.

[0038] Preferably, in the first matching process, longer keywords are matched first, and the average priority and average length of each word are considered as evaluation indicators of the matching degree.

[0039] Assuming that L represents the length of the word sequence, N represents the number of words in the word sequence, and P represents the priority attribute of all words in the word sequence, the calculation formula of the word segmentation matching degree M is: M = ΣP / N + ΣL / N. That is, the word segmentation matching degree is specifically the sum of the average priority of each word and the average length of each word.

[0040] It is understandable that the above calculation formula is only a specific embodiment, not the only possible way to calculate the word segmentation matching degree, and does not constitute an improper limitation on the present invention. The present invention performs entity recognition by combining the context information of the text data, and when the classification label is activated, it can set a higher priority for the associated words by automatically optimizing the priority parameters of the associated words, thereby further improving the accuracy of word segmentation in the process of converting text into structured content.

[0041] In this embodiment, automatically optimizing the priority parameters of the associated words specifically involves increasing the priority of the associated words.

[0042] The present invention determines which tags are available through the context of the focus section of the structured report, and automatically optimizes the priority of the associated words, so as to assign a higher priority to some words in the dictionary library. For example, when the sentence "mass with calcification, structural distortion" is segmented, usually, the part of "structural distortion" is easily divided into two words "structure" and "distortion", but if the word "structural distortion" exists in the dictionary library, and it is assigned a higher priority tag "radiological description", and this tag is currently configured to be activated during the process of voice recording imaging discovery, so "structural distortion" is divided out. For another example, if the theme of the focus section of the current structured report is "prostate" and "disease type", then the tag "lesion characteristics" associated with these two topics will be activated, so the keywords associated with "lesion characteristics" have obtained a higher priority, and the specific associated words may be some words describing location, quantity, distribution, and properties. And this association relationship between "disease type" and "lesion characteristics" is obtained from the established knowledge base.

[0043] In this embodiment, the entity linking process specifically includes: based on a preset knowledge base, linking the word segmentation result with the corresponding named entity to obtain the semantic code corresponding to the named entity.

[0044] Preferably, the knowledge base records the relationships and synonyms of medical concepts, as well as the named entities and semantic codes corresponding to the medical concepts. Since there is a mapping relationship between the named entities and the medical concepts, the semantic codes corresponding to the medical concepts are the semantic codes of the named entities corresponding to the medical concepts. The content of the structured report is also defined based on the semantic codes in the knowledge base, so the semantic codes can be matched with the content of the structured report, facilitating the use of the semantic codes in subsequent data matching and data filling processes.

[0045] In this embodiment, the data matching process specifically includes the following steps: based on the semantic coding corresponding to the named entity, obtaining the corresponding structured report template; the chapters of the structured report template have a hierarchical relationship, and the structured report template has at least one built-in grammatical metadata tree for recording the correspondence between the structured report elements and the natural language text, and the grammatical metadata tree includes at least one grammatical metadata subtree; based on the structured report template, obtaining the grammatical metadata tree or grammatical metadata subtree corresponding to the text data; performing a second matching process on the named entities in the text data and the nodes in the grammatical metadata tree or the nodes in the grammatical metadata subtree to determine the corresponding nodes, so as to facilitate the subsequent conversion of the text data into structured data (the advantage of this conversion is that there is no need to require whether all keywords in the natural language appear, and the order of appearance of the keywords in the natural language).

[0046] It should be noted that the various sections of the structured report template in this embodiment may be in a peer relationship or a superior-subordinate relationship. For example, the clinical evaluation and technical evaluation in the structured report template may be in a peer relationship, while the clinical evaluation and the inspection objectives under it are in a superior-subordinate relationship, which is specifically determined according to the settings in the structured report template.

[0047] Preferably, during the data matching process, the following preset matching rules need to be followed:

[0048] If an encoding cannot be matched with any node in the syntax metadata tree or syntax metadata subtree, the encoding is skipped.

[0049] If the average distance between the grammatical metadata tree node and other nodes matched by a certain encoding exceeds a certain threshold, or the average distance between the grammatical metadata subtree node and other nodes exceeds a certain threshold, it is considered to be ambiguous, with grammatical errors or word segmentation errors, and this encoding is skipped.

[0050] If a code can match a node a1 in the grammatical metadata tree or grammatical metadata subtree, then the minimum subtree m1 to which the current code belongs may be the minimum subtree where this node is located. Of course, there is also a small probability that it is accidental. If this code is the code of a chapter in the current structured report, it can basically be used to locate the chapter.

[0051] If a code can match multiple nodes in the tree, take a subtree t2 that can contain these nodes at the same time, then the smallest subtree m1 to which the current code belongs must belong to t2.

[0052] The distances between nodes corresponding to two adjacent codes should be close. Therefore, in a certain matching method, the smaller the sum of the distances between nodes corresponding to two adjacent codes is, the better.

[0053] Syntax metadata subtrees or structured report sections under the current structured report focus section have a higher matching priority.

[0054] Unfilled sections have higher priority than filled sections.

[0055] In this embodiment, the grammatical metadata tree is constructed together with the construction of the structured report, and is used for different purposes, such as diagnosis and imaging description often use different grammatical trees. If the voice data is used to describe imaging findings, the grammatical tree of the corresponding type under the target structured report template can be used. Each grammatical metadata carries the position of the current element in the tree, the semantic encoding, and the generation and merging rules of its natural language. Since the chapters of the structured report template have a hierarchical relationship, the grammatical metadata tree of each chapter is a subtree in the grammatical metadata tree of the structured report template as a whole. The deeper the level, the smaller the subtree. In this sense, locating the structured report chapter and locating the grammatical metadata tree subtree are equivalent. The nodes corresponding to certain keywords in the tree may appear multiple times, such as "normal", which will cause ambiguity in matching. Therefore, the expectation of the present invention is to deepen the matching level as much as possible and reduce the matching subtree. For example, there may be multiple lesion features under the same disease, and each lesion feature may have a control representing "normal". However, if the subtree can be narrowed down to a specific lesion feature, such as a morphological feature describing the disease, there will usually only be one control representing "normal", which can avoid ambiguity.

[0056] The present invention records the correspondence between structured report elements and natural language texts through a grammatical metadata tree, provides a basis for the conversion between natural language texts and structured content, and can improve the generation efficiency of structured reports.

[0057] In this embodiment, the data filling process is specifically as follows: the semantic coding array obtained by the entity linking process is filled into the grammatical metadata tree or grammatical metadata subtree obtained in the data matching process to generate structured report content.

[0058] In this embodiment, the filtering process is specifically as follows: based on the preset filtering rules, the text data is filtered, and the filtered content is not entered into the structured report. For example, a description such as "no obvious abnormality" is actually the default value of the report and does not need to be entered additionally, so it can be filtered.

[0059] The present invention filters text data based on preset filtering rules, and can filter some text data that does not need to be additionally entered into the structured report, thereby further improving the input speed of the structured report and the generation efficiency of the structured report.

[0060] In this embodiment, it is determined whether there is any newly added structured content; if so, the sub-template or chapter corresponding to the newly added structured content is obtained and displayed; then, according to the newly added structured content, the corresponding structured elements are automatically filled in to generate a structured report; if not, the structured report is directly generated.

[0061] In this embodiment, when there is new structured content, the corresponding structured sub-template or chapter can be automatically switched for display without the need for mouse or keyboard operation, which can further improve the efficiency of writing structured reports; new structured content can be added through manual input or voice input by the user, and new operations can be performed on chapters that have not yet been added to the report. The text data (if it is user voice input, it will be converted into text data first) is parsed and processed to generate structured report content.

[0062] like Figure 2 As shown, the present invention also provides a system for generating a structured report based on voice data, which includes: a voice recognition module 10, used to obtain voice data, convert it into text data and display it, and the voice data includes imaging-related content; a data analysis and processing module 20, used to perform data analysis on the text data to generate structured report content; the data analysis and processing at least includes: word segmentation processing, entity linking processing, filtering processing, data matching processing and data filling processing; the word segmentation processing is performed based on the priority parameters of the preset classification label, and the priority parameters include preset default priority parameters and / or optimized priority parameters; a structured content addition module 30, used to determine whether there is new structured content; if so, obtain the sub-template or chapter corresponding to the new structured content, display it, and then automatically fill in the corresponding structured elements according to the new structured content to generate a structured report; if not, directly generate a structured report.

[0063] The embodiment of the present invention further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the memory in the above embodiment; or a computer-readable storage medium that exists independently and is not installed in a device. The computer-readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement Figure 1 The method for generating a structured report based on speech data is shown. The computer readable storage medium may be a read-only memory, a magnetic disk or an optical disk, etc.

[0064] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, equipment embodiment and storage medium embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0065] Furthermore, in this document, the terms "comprises," "comprising," or any other variation thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.

[0066] The above description shows and describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the invention, through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not depart from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A method for generating a structured report based on speech data, characterized in that: The following steps are involved: Acquire voice data, convert it into text data and display it, wherein the voice data includes imaging-related content; Perform data analysis on text data to generate structured report content; The data analysis process at least includes: word segmentation process, entity linking process, filtering process, data matching process and data filling process; the word segmentation process is performed based on the priority parameters of the preset classification labels, and the priority parameters are optimized priority parameters; Determine whether there is newly added structured content; if so, obtain the sub-template or chapter corresponding to the newly added structured content and display it, and then automatically fill in the corresponding structured elements according to the newly added structured content to generate a structured report; if not, directly generate a structured report; The word segmentation process specifically includes the following steps: Get context data of text data; First, the text data and context data are segmented to obtain the corresponding word segmentation results, and then entity recognition is performed to obtain the named entities in the text data and context data; Performing a first matching process on the named entities and the chapters in the structured report respectively to obtain the classification labels corresponding to the named entities; In the first matching process, when the classification label is activated, the associated words belonging to the classification label in the word segmentation word library are obtained based on the preset association relationship, and the priority parameters of the associated words are automatically optimized to obtain the optimized priority parameters; Automatically optimizing the priority parameters of the associated words specifically increases the priority of the associated words.

2. The method for generating a structured report based on speech data according to claim 1, characterized in that: The entity linking process is as follows: based on a preset knowledge base, the word segmentation results are linked with the corresponding named entities to obtain the semantic encoding corresponding to the named entities.

3. The method for generating a structured report based on speech data according to claim 2, characterized in that: The data matching process specifically includes the following steps: Based on the semantic coding corresponding to the named entity, a corresponding structured report template is obtained; each chapter of the structured report template has a hierarchical relationship, and the structured report template has at least one built-in grammatical metadata tree for recording the corresponding relationship between the structured report elements and the natural language text, and the grammatical metadata tree includes at least one grammatical metadata subtree; Based on the structured report template, obtaining a grammatical metadata tree or a grammatical metadata subtree corresponding to the text data; A second matching process is performed on the named entities in the text data and the nodes in the grammatical metadata tree or the nodes in the grammatical metadata subtree to determine corresponding nodes.

4. The method for generating a structured report based on speech data according to claim 3, characterized in that: The data filling process is specifically as follows: the semantic coding array obtained by the entity linking process is filled into the grammatical metadata tree or grammatical metadata subtree obtained in the data matching process to generate structured report content.

5. The method for generating a structured report based on speech data according to claim 1, characterized in that: The filtering process is specifically as follows: based on preset filtering rules, the text data is filtered, and the filtered content is not entered into the structured report.

6. A system for generating structured reports based on speech data, characterized in that: include: A speech recognition module, used to obtain speech data, convert it into text data and display it, wherein the speech data includes imaging-related content; The data analysis and processing module is used to analyze and process text data and generate structured report content; The data analysis process at least includes: word segmentation process, entity linking process, filtering process, data matching process and data filling process; the word segmentation process is performed based on the priority parameters of the preset classification labels, and the priority parameters are optimized priority parameters; The word segmentation processing specifically includes the following steps: obtaining context data of text data; firstly performing segmentation processing on the text data and the context data to obtain corresponding word segmentation results, and then performing entity recognition to obtain named entities in the text data and the context data; performing a first matching processing on the named entities and the chapters in the structured report respectively to obtain classification labels corresponding to the named entities; in the first matching processing process, when the classification label is activated, obtaining the associated words belonging to the classification label in the word segmentation vocabulary based on the preset association relationship, and automatically optimizing the priority parameters of the associated words to obtain the optimized priority parameters; the automatic optimization of the priority parameters of the associated words specifically includes increasing the priority of the associated words; The structured content addition module is used to determine whether there is new structured content; if so, obtain the sub-template or chapter corresponding to the new structured content and display it, and then automatically fill in the corresponding structured elements according to the new structured content to generate a structured report; if not, generate a structured report directly.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program for generating a structured report based on voice data, and when the program for generating a structured report based on voice data is executed by a processor, the steps of the method for generating a structured report based on voice data as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Thyroid ultrasonic report structured scanning method based on semantic tree

    CN110399450A

  • Image report structured extraction method

    CN114328938A