Text data processing and medical data structuring method

By extracting article-level title information in medical text data and generating title paragraph pairs, the problems of redundancy and insufficient semantic understanding in medical text data processing are solved, and efficient structured processing of text data and accurate understanding of context information are achieved.

CN120126646APending Publication Date: 2025-06-10ALI HEALTH TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510081535.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art faces problems of information redundancy, insufficient semantic understanding and lack of context when processing medical text data, especially when facing longer plain text.

Method used

By obtaining the target pending text data, input it into the title acquisition model, extracting the article hierarchical title information, and generating title paragraph pairs based on this information and article paragraph information, realizing the structured processing of text data.

Benefits of technology

It significantly improves the organization and readability of text data, solves the problems of redundancy and insufficient semantic understanding, ensures the integrity and accuracy of contextual information, and thus improves the efficiency and quality of medical research and clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126646A_ABST
    Figure CN120126646A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text data processing and medical data structuring method, and the method comprises the steps: obtaining target to-be-processed text data which comprises at least one piece of article paragraph information and at least one piece of article title information; the target to-be-processed text data is input into a title acquisition model, article level title information output by the title acquisition model is acquired, and the article level title information represents the level relation between the article title information; a text data processing result is obtained based on the article level title information and the article paragraph information, and the text data processing result comprises at least one title paragraph pair. Model article level title information is obtained through titles, and text processing accuracy and efficiency are improved through subsequent pairing of article paragraphs and titles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for processing text data. Background Art

[0002] With the rapid development of information technology, opportunities for digital transformation in the medical field have emerged, especially driven by artificial intelligence and large medical models. In the face of a vast amount of text data in the medical industry, how to efficiently utilize this data to promote medical research and clinical decision-making has become a key issue that urgently needs to be solved. Existing medical text data comes from a wide range of sources, including clinical records, scientific research papers, and patient health records, etc. These data mostly exist in the form of unstructured text, such as doctor's notes, case analyses, and medical research reports, with complex content and loose structures.

[0003] Currently, although existing models can, to a certain extent, process this unstructured text data and solve the basic problems of information extraction, there are still challenges such as information redundancy, insufficient semantic understanding, and lack of context when dealing with long plain texts. Therefore, to solve the above problems, a method for processing article data is urgently needed. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide a method for processing text data, a method for structuring medical data, a method for processing text data applied to cloud devices, and a method for structuring medical data applied to cloud devices. One or more embodiments of this specification also relate to a text data processing device, a medical data structuring device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to the first aspect of the embodiments of this specification, a method for processing text data is provided, including:

[0006] Obtain target text data to be processed, where the target text data to be processed includes at least one article paragraph information and at least one article title information;

[0007] Input the target text data to be processed into a title acquisition model, and obtain the article-level title information output by the title acquisition model, where the article-level title information represents the hierarchical relationship between each article title information;

[0008] Obtain a text data processing result based on the article-level title information and each article paragraph information, where the text data processing result includes at least one title-paragraph pair.

[0009] According to the second aspect of the embodiments of this specification, a method for structuring medical data is provided, including:

[0010] Obtain target medical data to be processed, where the target medical data to be processed includes at least one article paragraph information and at least one article title information;

[0011] Input the target medical data to be processed into a title acquisition model, and obtain the article-level title information output by the title acquisition model, where the article-level title information represents the hierarchical relationship between each article title information;

[0012] Obtain structured medical data based on the article-level title information and each article paragraph information, where the structured medical data includes at least one title-paragraph pair.

[0013] According to the third aspect of the embodiments of the present specification, a text data processing method is provided, which is applied to a cloud device and includes:

[0014] Receive a text processing request sent by a terminal device, where the text processing request includes target text data to be processed, and the target text data to be processed includes at least one article paragraph information and at least one article title information;

[0015] Input the target text data to be processed into a title acquisition model, and obtain the article-level title information output by the title acquisition model, where the article-level title information represents the hierarchical relationship between each article title information;

[0016] Obtain a text data processing result based on the article-level title information and each article paragraph information, and send the text data processing result to the terminal device, where the text data processing result includes at least one title-paragraph pair.

[0017] According to the fourth aspect of the embodiments of the present specification, a medical data structuring method is provided, which is applied to a cloud device and includes:

[0018] Receive a medical data processing request sent by a terminal device, where the medical data processing request includes target medical data to be processed, and the target medical data to be processed includes at least one article paragraph information and at least one article title information;

[0019] Input the target medical data to be processed into a title acquisition model, and obtain the article-level title information output by the title acquisition model, where the article-level title information represents the hierarchical relationship between each article title information;

[0020] Obtain structured medical data based on the article-level title information and each article paragraph information, and send the structured medical data to the terminal device, where the structured medical data includes at least one title-paragraph pair.

[0021] According to a fifth aspect of the embodiments of the present specification, a computing device is provided, including:

[0022] a memory and a processor;

[0023] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned text data processing method and medical data structuring method are implemented.

[0024] According to a sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above-mentioned text data processing method and medical data structuring method are implemented.

[0025] According to a seventh aspect of the embodiments of the present specification, a computer program product is provided, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned text data processing method and medical data structuring method are implemented.

[0026] Applying the solution of the embodiments of the present specification, by obtaining the target text data to be processed and inputting it into the title acquisition model, the hierarchical title information of the article is successfully extracted, accurately representing the hierarchical relationship between each title, laying a solid foundation for subsequent structuring processing; at the same time, based on the extracted hierarchical title information and each article paragraph information, this solution generates title-paragraph pairs, significantly improving the organization and readability of the text data, and solving the problems of information redundancy and insufficient semantic understanding. In addition, this solution utilizes the advanced semantic processing ability of the large model to achieve in-depth parsing of complex medical texts, ensuring the integrity and accuracy of the context information, thereby greatly improving the efficiency and quality of medical research and clinical decision-making. By structuring a large amount of unstructured medical text data, this solution not only optimizes the data utilization method, but also promotes the digital transformation of the medical field, facilitating the in-depth development of medical research and the scientificization of clinical decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flowchart of a text data processing method provided by an embodiment of the present specification;

[0028] Figure 2 is a flowchart of a text data processing method applied to a cloud device provided by an embodiment of the present specification;

[0029] Figure 3 is an architecture diagram of a text data processing system provided by an embodiment of the present specification;

[0030] Figure 4 It is a flowchart of a medical data structuring method provided by an embodiment of this specification;

[0031] Figure 5 It is a flowchart of a medical data structuring method applied to a cloud device provided by an embodiment of this specification;

[0032] Figure 6 It is an architecture diagram of a medical data structuring system provided by an embodiment of this specification;

[0033] Figure 7 It is a processing flowchart of a scientific research paper structuring method provided by an embodiment of this specification;

[0034] Figure 8 It is a structural schematic diagram of a text data processing device provided by an embodiment of this specification;

[0035] Figure 9 It is a structural schematic diagram of a medical data structuring device provided by an embodiment of this specification;

[0036] Figure 10 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners

[0037] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.

[0038] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.

[0039] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0040] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0041] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than hundreds of millions of parameters is produced. Such a model can adapt to a wide range of downstream tasks and has good generalization ability, such as large language models (LLMs), multi-modal pre-training models, etc.

[0042] When a large model is actually applied, only a small number of samples are needed to fine-tune the pre-trained model for application to different tasks. Large models can be widely applied in fields such as natural language processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., and tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0043] First, the noun terms involved in one or more embodiments of this specification are explained.

[0044] Optical Character Recognition (OCR) is a technology that converts printed or handwritten text into a machine-readable format. By analyzing the images in the document, OCR can recognize and extract the text content, and realize digital storage and processing. This technology is widely used in document management, data input and information retrieval, etc., which helps to improve work efficiency and reduce human errors. With the continuous optimization of the algorithm, the recognition accuracy of OCR in complex backgrounds and multilingual environments is also gradually improving, which has promoted the digital transformation and intelligent development of various industries.

[0045] In this specification, a text data processing method, a medical data structuring method, a text data processing method applied to a cloud device, and a medical data structuring method applied to a cloud device are provided. One or more embodiments of this specification also relate to a text data processing device, a medical data structuring device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0046] With the rapid development of information technology and the widespread application of artificial intelligence, the medical field has gradually ushered in a major opportunity for digital transformation. Especially in the context of the development of the medical big model, how to effectively utilize massive amounts of medical text data has become a key issue that needs to be solved. Existing medical text data mostly comes from clinical records, scientific research papers, and patients' health records. These data often exist in the form of unstructured plain text, such as long doctor's notes, case analysis, medical research reports, etc., with complex content and loose structure.

[0047] In the process of data mining and model training, the unstructured nature of plain text content brings many challenges to information extraction and data processing. Existing models often face problems such as information redundancy, insufficient semantic understanding, and missing context when processing these plain text data, which directly affects the training effect of the model and the accuracy of downstream applications.

[0048] Therefore, constructing effective training data is a key step to improve the performance of large medical models. Using an accurate and automated structured technology to efficiently process and transform long plain text content can not only reduce labor costs, but also improve data availability and model learning efficiency. This process involves multiple technical links such as semantic understanding of text content, key information extraction, and information summarization.

[0049] Therefore, it is of great theoretical significance and practical application value to carry out the research on the automatic structuring technology of long medical pure text content based on large medical models. This can not only provide high-quality training data for large medical models, but also provide a more reliable data foundation for fields such as medical decision support, personalized medicine, and health management.

[0050] See Figure 1 , Figure 1 which shows a flowchart of a text data processing method provided according to an embodiment of this specification, specifically including the following steps.

[0051] Step 102: Obtain target text data to be processed, where the target text data to be processed includes at least one article paragraph information and at least one article title information.

[0052] In practical applications, the target text data to be processed refers to the text data after natural segmentation processing, which usually contains complete sentences, expressions, and semantics for subsequent analysis and processing; the article paragraph information is the specific paragraph in the text, and each paragraph contains one or more topics and carries the content of the article; the article title information refers to the title in the article, which is usually used to identify the topic of a chapter, paragraph, or specific content and is a key part of the article structure.

[0053] The target text data to be processed can be understood as the original text data after natural segmentation. This data has been transformed from the original unorganized state into a more structured form, and usually each paragraph and title have been clearly defined. This process helps to clearly distinguish each part of the article, such as: the content differences between paragraphs or the hierarchical structure of titles, making subsequent data processing more efficient and accurate. For example, after natural segmentation processing of a medical article, the target text data to be processed becomes a text containing each chapter paragraph and related titles.

[0054] The article paragraph information can be understood as the natural paragraph in the article. Each paragraph is a part of the article, usually centered around a core topic and carrying detailed content information. Obtaining the paragraph information helps to understand the structure of the article and the logical relationship of its internal content. For example, in a scientific research article, the paragraph information may include parts such as background introduction, methodology, experimental results, and discussion, and each part elaborates on an aspect of the article in detail for subsequent processing and analysis.

[0055] The title information of an article can be understood as the title used to identify the theme of a chapter or paragraph in the article. The title information not only helps to identify different parts of the article, but also helps to clarify the hierarchical relationship between each part. Titles are usually used to summarize the content of a paragraph or indicate the structural level of a chapter. In an article, the title information may include main titles, sub-titles, chapter titles, etc. For example, in a report, the titles may be "Purpose of the Experiment", "Experimental Results", and "Discussion and Conclusions" respectively. The hierarchical structure of these titles helps readers quickly locate and understand the main content of the article.

[0056] It should be noted that obtaining the target text data to be processed can be understood as extracting the content that needs to be processed from the initially obtained text data to ensure that the data can be used for subsequent processing steps; the specific method can process the obtained ultra-long text data according to the field end identifier (such as the end punctuation mark), and split it into appropriate paragraphs or sentences; it can also use natural language processing techniques, such as word segmentation and named entity recognition, to extract key paragraphs and title information from the original text; it can also identify and extract structured text blocks, such as article paragraphs, titles and their related content, etc., based on specific rules or models. This specification does not impose any restrictions on this.

[0057] In an embodiment provided in this specification, the target text data to be processed is a medical article containing multiple paragraphs, where the article paragraph information is the specific content of each paragraph, such as paragraph texts like "The vaccination plan for PID children includes...", and the article title information is the titles at all levels of the article, such as "Vaccination of PID Children" and "Types of Vaccines Used", etc.

[0058] By obtaining the target text data to be processed, the subsequent structured processing of the text data can be made more efficient.

[0059] Furthermore, obtaining the target text data to be processed includes:

[0060] Obtain the initial text data to be processed, where the text data to be processed includes at least one field end identifier;

[0061] Process the initial text data to be processed based on a preset text processing length threshold and each field end identifier to obtain the target text data to be processed.

[0062] In practical applications, the initial text data to be processed is the original, unprocessed text content, usually from documents, web pages, OCR (Optical Character Recognition) scanned images, etc., containing a large amount of unstructured data; the field end identifier is a marker indicating the end of a certain part or paragraph in the text, usually a specific punctuation mark or special character, used to identify the segmentation point of the text; the processing length threshold is a limit on the length of the input text, used to control the large length of the processed data, to optimize the model processing efficiency and avoid memory overflow.

[0063] Exemplarily, the initial text data to be processed can be understood as the raw, unprocessed text content obtained at the start of the processing process. Such data typically includes various text paragraphs, sentences, and may even contain characters with format errors or irregularities. The goal of processing these text data is usually to transform them into a more structured form for subsequent analysis and processing. The field end identifier can be understood as an identifier used to mark the end point of a field or paragraph in the text data. During the actual processing, the field end identifier helps the system identify the segmentation point of the data and ensure the correct segmentation of the text. Usually, these identifiers include punctuation marks such as full stops, question marks, exclamation marks, or some special symbols, such as line breaks. Through these identifiers, the processing system can effectively parse and reconstruct text paragraphs. For example, when processing a long document, the system can split a long paragraph into multiple smaller paragraphs according to full stops or line breaks for further analysis or structured processing.

[0064] The processing length threshold can be understood as setting a large length limit to ensure that the input text data does not exceed this limit. This threshold is often used to control the scale of the text data and avoid waste of computing resources or low model processing efficiency caused by overly long texts. In practical applications, the processing length threshold is often used to optimize the input of the model. For example, when performing natural language processing tasks, the model usually limits the length of the input text to ensure that the data can be processed within the specified time. For instance, a certain natural language processing system may set the processing length threshold to 2000 characters, and texts exceeding this length will be segmented or simplified to ensure the efficient operation of the system.

[0065] It should be noted that processing the initial text data to be processed based on the preprocessing length threshold and the end markers of each field to obtain the target text data to be processed can be understood as a text segmentation method. The specific method of this step can be to split the initial text data according to a preset length threshold, and the length of each text segment does not exceed the set maximum number of characters, so as to divide the text into multiple segments for subsequent analysis and processing; it can also be to determine the boundaries of text paragraphs according to the end markers of fields (such as punctuation marks or other specific characters) to ensure the integrity of each paragraph, and then perform natural segmentation; it can also be to combine the length threshold and the end markers of fields, while checking whether the length of the paragraph exceeds the threshold, segment the text according to the end markers of fields, so as to ensure that the length of each segment does not exceed the standard and does not damage the semantic structure. This specification makes no restrictions on this.

[0066] Using the end markers of fields in the initial text data to be processed and the preset text processing length threshold to generate the target text data to be processed can effectively ensure that the data maintains a suitable structure and length in the subsequent processing process.

[0067] Furthermore, processing the initial text data to be processed based on the preset text processing length threshold and the end markers of each field to obtain the target text data to be processed includes:

[0068] Based on the preset text processing length threshold and the end markers of each field, determine at least one reference text data to be processed in the initial text data to be processed;

[0069] Generate text repair prompt words according to the preset text repair prompt word template and each reference text data to be processed, and input the text repair prompt words into the text repair model to obtain the target text data to be processed generated by the text repair model.

[0070] In practical applications, the reference text data to be processed refers to the text segments split out during the processing; the text repair prompt word template refers to the preset input format set to guide the text repair model to generate repair content; the text repair prompt words are the specific text prompts generated according to the repair task, which are used to guide the model to optimize the data to be processed.

[0071] The reference text data to be processed can be understood as data segments separated from the initial text data to be processed through specific rules. These segments usually contain certain semantic units, such as sentences or paragraphs, for subsequent processing. For example, when processing ultra-long texts, the reference text data to be processed may divide the text into multiple smaller data blocks through the setting of end markers of fields (such as full stops) and length thresholds, and each block maintains a certain logical integrity. In this way, the original text is divided into paragraphs or segments suitable for processing, ensuring that key content will not be lost during text analysis or repair.

[0072] A text repair prompt template can be understood as a preset text format or framework that contains the rules and structures that the model needs to follow during repair. For example, during structural repair, the text repair prompt template may include certain specific formatting requirements, such as marking the start and end of each paragraph or specifying which parts need to be repaired. These templates provide clear guidance for the model, enabling it to efficiently and accurately generate output that meets the requirements during the text repair process.

[0073] Text repair prompts can be understood as the actual repair inputs generated by applying the text repair prompt template to the reference text data to be processed. These prompts are usually text fragments in a specific format that contain the specific tasks or instructions that the model needs to focus on during repair. For example, a text repair prompt may require the model to repair specific errors in a part according to the context information or supplement missing content. Guided by the text repair prompts, the model can perform a detailed repair of the input text to ensure that the generated data meets the expected standards.

[0074] Step 104: Input the target text data to be processed into the title acquisition model to obtain the article-level title information output by the title acquisition model, where the article-level title information represents the hierarchical relationship between each article title information.

[0075] In practical applications, the title acquisition model is a machine learning model used to automatically extract the article title hierarchy from text; the article-level title information refers to the data structure that describes each title in the article and its hierarchical relationship, usually presented in JSON format, indicating each title and its corresponding level and reflecting their hierarchical structure in the article.

[0076] The title acquisition model can be understood as a model that automatically identifies and extracts the title and its hierarchical structure from the input text data through machine learning algorithms. Such models usually rely on deep learning methods and use a large number of labeled datasets during training. The model determines the hierarchical relationship between different titles by identifying the title features in the text. For example, in an academic paper, the title acquisition model can help identify "Introduction" as a first-level title, followed by "Research Methods" as a second-level title, and other levels. This model plays an important role in obtaining the article structure information and understanding the article content, especially in long articles, where it can help build a clearer text structure. The article-level title information can be understood as the detailed data that describes each title in the article and its hierarchical relationship. This information is usually presented in a structured data form, such as a JSON string, indicating the order of different titles in the article and their corresponding levels.

[0077] It should be noted that the specific method for obtaining the title acquisition model can be understood as concatenating the target text data to be processed and inputting it into the title acquisition model, so that the model outputs hierarchical title information; the specific method can be to apply a pre-trained deep learning model to the target text to be processed, and automatically identify and predict the title information at all levels through the model; it can also be through a rule-driven method, according to the common identifiers of the title (such as "first", "second") and the paragraph structure information, manually extract and classify the hierarchical titles, and can also use a preset prediction template for the article hierarchical title information to make the model predict the text at the specified position to output structured article hierarchical title information, etc. This specification does not impose any restrictions on this.

[0078] In an embodiment provided in this specification, the article hierarchical title information identified by a JSON string is as follows: {"I. Vaccination for PID Patients": "First-level title", "1. Combined T / B Cell Immunodeficiency": "Second-level title", "2. Antibody-dominated Deficiency": "Second-level title", "II. Types of Vaccines Used": "First-level title"}. In this embodiment, "I. Vaccination for PID Patients" and "II. Types of Vaccines Used" are identified as first-level titles, and "1. Combined T / B Cell Immunodeficiency" and "2. Antibody-dominated Deficiency" are second-level titles.

[0079] The extraction of article hierarchical title information is particularly important for subsequent content processing, especially when it is necessary to perform structured processing, abstract generation, or automatic classification of article content. The hierarchical information provides an organizational framework for the article content, which helps to more accurately understand and analyze the article content.

[0080] Furthermore, inputting the target text data to be processed into the title acquisition model to obtain the article hierarchical title information output by the title acquisition model includes:

[0081] Obtaining a prediction template for article hierarchical title information, where the prediction template for article hierarchical title information includes at least one title information prediction position;

[0082] Inputting the target text data to be processed and the prediction template for article hierarchical title information into the title acquisition model to obtain the article hierarchical title sub-information corresponding to each title information prediction position generated by the title acquisition model;

[0083] Generating article hierarchical title information based on the article hierarchical title sub-information corresponding to each title information prediction position and the prediction template for article hierarchical title information.

[0084] In practical applications, the article-level title information prediction template is a preset input structure used to guide the title acquisition model to generate title information; the title information prediction position refers to the position in the article-level title information prediction template where the title information needs to be predicted; the article-level title sub-information is the specific title content output by the model according to the title information prediction position.

[0085] Exemplarily, the article-level title information prediction template provides a clear framework for the model, enabling it to correctly locate the title; the title information prediction position is the indication point in the template to help the model accurately predict; and the article-level title sub-information is the specific title content generated based on these indication points, usually reflecting the hierarchical structure of the article.

[0086] The article-level title information prediction template can be understood as a structured preset text format used to guide the title acquisition model to generate hierarchical title information on given text data. This template usually includes information such as title levels and title ranks, which can clearly show which parts need to be predicted and generate titles. For example, when structuring an article, the template can stipulate the main title, sub-titles, paragraph titles, etc. of the article to help the model reason and generate each title step by step according to the hierarchy.

[0087] The title information prediction position can be understood as the specific position specified in the article-level title information prediction template, used to indicate which parts need to generate or predict title information. These positions are usually preset in the template with marking symbols or formats to guide the model to extract and generate titles in the corresponding areas of the given data. For example, in some templates, specific symbols such as " <h1>”、"< / h1> <h2>", etc., to mark the positions of headings at different levels in the article, helping the model to predict headings at each level based on these marks.

[0088] The sub-information of the article-level headings can be understood as the specific heading content generated by the heading acquisition model according to the predicted positions of the heading information. They represent headings at different levels in the article, which may include the main title of the article, chapter headings, and subsection headings, etc. The heading content at each level is usually generated based on the semantics and structure of the article content, aiming to help understand the overall structure and content of each part of the article. For example, headings such as "Chapter 1: Introduction" and "1.1 Research Background" generated by the model according to the predicted positions are convenient for subsequent structured processing and content summary.

[0089] In an embodiment provided in this specification, the prediction template for the article-level heading information is: {< / h2> <h1>:"< / h1> <h2>Level heading”< / h2> <h3>, wherein,< / h3> <h1>、< / h1> <h2>and< / h2> <h3>It is the predicted position of the title information. The article-level title information generated based on the hierarchical title information of the article is {"I. Vaccination for PID patients": "First-level title", "1. Combined T / B cell immunodeficiency": "Second-level title", "2. Antibody-dominant deficiency": "Second-level title", "II. Types of vaccines used": "First-level title"}.

[0090] The sub-information of the article-level titles sequentially generated by the model are: ["I. Vaccination for PID patients", "I.", ",", "1. Combined T / B cell immunodeficiency", "II.", ",", "2. Antibody-dominant deficiency", "II.", ",", "II. Types of vaccines used", "I."}].

[0091] The title acquisition model predicts each sub-information of the article-level title through the article-level title information prediction template, enabling the title acquisition model to implement dynamic decoding restrictions, thereby improving the efficiency and accuracy of article structuring. This method accurately sets the title prediction position and title level, enabling the model to accurately identify and generate titles that conform to the article structure when processing complex texts, making the extraction of title information more accurate and avoiding the generation of redundant content. Through dynamic decoding restrictions, the model can perform directional predictions according to the preset template, reducing unnecessary consumption of computing resources and improving the response speed when processing long texts. This optimization method not only improves the model's understanding ability of the article-level structure but also effectively supports subsequent data processing and analysis, enhancing the performance of the overall text analysis system.

[0092] Furthermore, input the target text data to be processed and the article-level title information prediction template into the title acquisition model to obtain the sub-information of the article-level title corresponding to each predicted position of the title information generated by the title acquisition model, including:

[0093] Generate at least one article paragraph summary data based on the target text data to be processed;

[0094] Generate title information extraction text according to each article paragraph summary data and the target text data to be processed;

[0095] Input the title information extraction text and the article-level title information prediction template into the title acquisition model to obtain the sub-information of the article-level title corresponding to each predicted position of the title information generated by the title acquisition model.

[0096] In practical applications, the article paragraph summary data is the content after simplifying the long text paragraph, aiming to extract the core information of the paragraph; the title information extraction text is the text generated by combining the article paragraph summary data and the target text data to be processed, mainly used for further title extraction tasks.

[0097] The abstract data of article paragraphs can be understood as the refined information obtained by simplifying long paragraphs in the article. The main purpose is to remove the redundant content of the paragraphs and retain the core viewpoints that are crucial for understanding the full text. For example, when processing long papers or reports, the abstract data of article paragraphs can effectively reduce the text length, making the processing more efficient, while not losing the important semantics in the paragraphs. Through such an abstract, the accuracy and rapid response ability of the subsequent model in title extraction can be improved. During the process of content simplification, the abstract data of article paragraphs usually extracts according to the topic words, important sentences and key information points of the paragraphs, so as to ensure the integrity and relevance of the information.

[0098] The text for title information extraction can be understood as a text fragment formed by combining the abstract data of paragraphs and the target text data to be processed, which is specifically used to provide input for the title extraction model. This text is usually composed of the key information refined from the abstract data of article paragraphs and specific parts of the original text, aiming to help the title extraction model accurately identify the hierarchical titles of the article. For example, in news articles, by generating the text for title information extraction, it can help the model identify structural information such as the main title, subtitle and related chapter titles of the article. In this way, the text for title information extraction not only supports effective title extraction, but also improves the accuracy of the title level and the processing efficiency of the system.

[0099] By simplifying the content of long paragraphs and retaining key information, it is easier to extract the hierarchical titles of the article in the subsequent processing; the text for title information extraction further provides input for the title extraction model on the basis of the paragraph abstract, thereby improving the accuracy and efficiency of title extraction.

[0100] Furthermore, based on the target text data to be processed, at least one abstract data of article paragraphs is generated, including:

[0101] Determine at least one paragraph information to be compressed in each article paragraph information based on a preset abstract extraction threshold;

[0102] Input each paragraph information to be compressed into the paragraph abstract generation model, and obtain the abstract data of article paragraphs generated by the paragraph abstract generation model.

[0103] In practical applications, the paragraph information to be compressed is the paragraph fragments selected from the long article that need to be extracted for abstract; the paragraph abstract generation model is a model used to generate the simplified version of the paragraph, which reduces the length of the paragraph by extracting key information. The paragraph information to be compressed refers to the paragraphs in the article that exceed a certain length threshold, and these paragraphs are identified as the targets that need to be further compressed.

[0104] The paragraph information to be compressed can be understood as the paragraphs that need to be compressed after being screened by a preset summary extraction threshold during the text processing. To improve the processing efficiency, paragraphs with longer lengths or redundant content are often selected as the paragraph information to be compressed. For example, when processing a long report or article, if some paragraph information is verbose and contains repetitive content, then these paragraphs are selected as the paragraph information to be compressed, with the goal of removing the redundant content and retaining the more core viewpoints and information.

[0105] By processing these paragraphs to be compressed, the text content can be significantly shortened, making the subsequent processing tasks more efficient and operable. Furthermore, it can not only reduce the computational burden during data processing but also better provide concise and effective input for subsequent summary generation.

[0106] Furthermore, determining at least one piece of paragraph information to be compressed based on a preset summary extraction threshold in each article paragraph information includes:

[0107] Obtaining the paragraph text length information and paragraph position information corresponding to each article paragraph information;

[0108] Based on the preset summary extraction threshold and the paragraph text length corresponding to each article paragraph information, determining at least one piece of sub-paragraph information to be compressed in each article paragraph information, where the sub-paragraph information to be compressed is the article paragraph information with a paragraph text length less than the preset summary extraction threshold;

[0109] Based on the paragraph position information corresponding to each piece of sub-paragraph information to be compressed, obtaining the paragraph distance information between each piece of sub-paragraph information to be compressed;

[0110] According to the paragraph distance information between each piece of sub-paragraph information to be compressed, obtaining at least one piece of paragraph information to be compressed, where the paragraph information to be compressed includes at least one piece of sub-paragraph information to be compressed, and the paragraph distance information between each piece of sub-paragraph information to be compressed in the paragraph information to be compressed is less than or equal to the preset paragraph aggregation threshold.

[0111] In practical applications, the paragraph text length information is the number of characters or words contained in each paragraph; the paragraph position information is the relative position or order in which the paragraph appears in the article; the summary extraction threshold is the length limit for determining whether to perform summary processing on the paragraph; the paragraph distance information is the relative distance between different paragraphs.

[0112] The length information of a paragraph can be understood as a standard for measuring the amount of information in a paragraph, and is usually quantified by counting the number of characters or words in the paragraph. For example, if a paragraph is long and contains a lot of redundant information, then the length information of the paragraph can be used to determine whether summary processing is required. The paragraph position information can be understood as an indicator describing the position of the paragraph in the entire article, and is usually represented by the relative order of the paragraphs or the number of characters from the beginning of the article to the paragraph. By analyzing the position of the paragraph, the system can better understand the structure of the article. For example, the paragraphs at the beginning of the article may contain introductions or background information, the middle paragraphs may carry the core discussion, and the paragraphs at the end usually summarize the foregoing content. Therefore, accurately grasping the paragraph position information helps to decide which paragraphs should be processed first and which paragraphs can be summarized or deleted later.

[0113] The summary extraction threshold can be understood as a preset length limit for determining which paragraphs need to be summarized. Specifically, when the text length of a paragraph exceeds this threshold, the paragraph will be regarded as verbose content and needs to be compressed into a shorter summary. For example, when processing a document containing multiple paragraphs, it may be stipulated that any paragraph exceeding 150 words needs to be summarized. The paragraph distance information can be understood as the relative distance between different paragraphs, usually expressed as the relative order of the paragraphs in the article or the difference in the number of words between paragraphs. For example, if the distance between two paragraphs is short and the content between them is similar or highly relevant, then these two paragraphs can be combined into one summary to improve processing efficiency.

[0114] By combining this information, it is possible to effectively determine which paragraphs need to be summarized and further optimize the article structure to ensure the conciseness and readability of the processed content.

[0115] Furthermore, the title acquisition model is trained through the following steps:

[0116] Obtain the article-level title information prediction template, as well as the sample paragraph information and the sample title information corresponding to the sample paragraph information, where the article-level title information prediction template includes at least one title information prediction position;

[0117] Input the sample paragraph information and the article-level title information prediction template into the title acquisition model to obtain the article-level title information generated by the title acquisition model;

[0118] Based on each title information prediction position, determine the predicted title information in the article-level title information;

[0119] Generate a model loss value according to the sample title information and the predicted title information, and train the title acquisition model according to the model loss value until the model training stop condition is reached.

[0120] In practical applications, the sample paragraph information is the original paragraph text data provided for model training; the sample title information is the actual title data corresponding to the sample paragraph information, which is usually used for supervised learning in model training; the predicted title information is the title information obtained through model prediction, which is usually the title prediction result generated by the system based on the input sample paragraph; the model loss value is a numerical value that measures the difference between the model prediction result and the true label, and this value is used to guide model optimization; the model training stop condition is the condition set during the training process, and when these conditions are met, the training process will stop.

[0121] The sample paragraph information can be understood as the input data used in model training, and this data is usually the paragraph text extracted from actual articles. For example, in the training of a news title generation model, the sample paragraph information may be the different paragraph contents of news articles. Through these paragraphs, the model can learn how to extract key information from the original text and generate accurate titles. The quality of the sample paragraph information directly affects the training effect, so it is necessary to ensure that these paragraph contents are representative and diverse.

[0122] The sample title information can be understood as the output label corresponding to the sample paragraph information, and it is the target data for model training. Each sample paragraph has a corresponding title, and these titles are usually manually annotated or extracted from the metadata of the document. For example, in the news title generation task, the sample title information is the title corresponding to the article. By learning the relationship between the sample paragraph and the sample title, the model can generate new titles in practical applications. The sample title information provides the supervision signal required in model training and is the basis for learning.

[0123] The predicted title information can be understood as the title generated by the model after inputting new paragraph information. This is the prediction result of the model for a given paragraph after training. For example, in the test stage, when a news paragraph is input, the model will generate a predicted news title based on the existing training knowledge. This prediction result is compared with the true title to calculate the prediction accuracy of the model. The predicted title information is crucial for evaluating the performance of the model, and it directly reflects the effect and improvement space of the model.

[0124] The model loss value can be understood as an index that measures the difference between the model prediction and the actual result. Generally, the smaller the loss value, the closer the model prediction result is to the true label. In the title generation task, the loss value calculates the difference between the predicted title and the actual title. Commonly used loss functions include cross-entropy loss function, mean squared error, etc., and this specification does not impose any restrictions on this. By calculating the loss value, the model can continuously adjust the weights and parameters to reduce the prediction error and thus improve the accuracy of the model.

[0125] The model training stop condition can be understood as a termination condition set during the model training process. When the training reaches a certain accuracy or the loss value no longer decreases significantly, the training can be stopped. For example, if the loss value changes minimally after a certain number of iterations, it indicates that the model has converged and the training can be stopped to avoid overfitting. In actual operation, common stop conditions include reaching a predetermined large number of training epochs, the loss value being lower than a certain threshold, or the performance of the validation set no longer improving.

[0126] Step 106: Obtain the text data processing result based on the article hierarchical title information and each article paragraph information, where the text data processing result includes at least one title-paragraph pair.

[0127] In practical applications, the text data processing result refers to the output result obtained based on the processing and analysis of the target text data; the title-paragraph pair is the structured data that shows the pairing relationship between the article title and its corresponding content paragraph through the pairing relationship between the title and the relevant paragraph.

[0128] The text data processing result can be understood as the output after processing the original text. The processing process usually includes various technical means to generate structured data for further analysis, display, or archiving. For example, in a long article, the text is transformed into a structured format containing hierarchical titles and corresponding content paragraphs, which is convenient for further applications such as information extraction, content analysis, or generating summaries. The purpose of the text data processing result is to transform the chaotic original text into structured information that is easy to understand and utilize, which is particularly important for text analysis and automated processing.

[0129] The title-paragraph pair can be understood as the matching relationship between the article title and its relevant paragraph. Each pair of title and paragraph usually consists of a title and the text content related to it. For example, in an academic report, the section under the title "Introduction" may contain a detailed description of this part, forming a title-paragraph pair. The extraction of the title-paragraph pair can help combine the title with its corresponding content organically and clearly show which paragraph belongs to which title, which is crucial for understanding and processing the article structure. In practical applications, the title-paragraph pair is often stored in the form of structured data, and common formats are key-value pairs or lists, where each title is paired with the relevant content. For example, assume that in an article about climate change, "Influencing Factors" is used as the title, and there may be multiple paragraphs below it that detail different climate change factors, and each paragraph will be paired with the title "Influencing Factors" to form a title-paragraph pair. This structure facilitates the organization and analysis of the article.

[0130] It should be noted that obtaining the text data processing result based on the article-level title information and each article paragraph information can be understood as performing a structured analysis by integrating the title-level information and the article content, and generating an organized text processing output at the end; the specific method can be to pair the hierarchical title information with the corresponding paragraphs to form title-paragraph pairs, and perform summary extraction on them to generate a concise text processing result; it can also be to classify the paragraphs according to their themes using text classification or clustering methods based on the hierarchical structure of the title, so as to obtain the content summary under each theme; it can also be based on the semantic matching between the hierarchical title information and the paragraph content, and use a deep learning model to further semantically understand the paragraphs and automatically generate concise text content related to the title, etc. This specification does not impose any restrictions on this.

[0131] In an embodiment provided by this specification, the data processing result is title-paragraph pairs generated based on the article-level title information and each article paragraph information. For example: {"I. Vaccination of PID Patients": "The vaccination plan for PID patients includes...", "1. Combined T / B Cell Immunodeficiency": "The definition of combined T / B cell immunodeficiency is..."}. Each pair of titles is associated with its corresponding paragraph content, and it is ensured that each paragraph is appropriately paired according to the hierarchical relationship of the titles.

[0132] Applying the solution of the embodiment of this specification, the initial text data is screened and processed through a preset length threshold and field end identifier to ensure that only relevant and clearly structured content is retained, significantly improving the quality and consistency of the data; a text repair prompt word template is combined with a text repair model to optimize the structure of the target text, eliminating the noise and incomplete information in the original data, thereby enhancing the accuracy of subsequent title extraction. Using the article-level title information prediction template to accurately locate the position of the title information, effectively constructing the hierarchical relationship between the titles, ensuring the rationality and stability of the hierarchical structure; by generating paragraph summary data and extracting title information based on the summary, the text length is significantly reduced, the processing efficiency is improved, and the integrity of the context and the coherence of the semantics are maintained. In addition, during the training process, by optimizing the model loss value, the efficiency and robustness of the title acquisition model are ensured, realizing the in-depth analysis and structured processing of a large amount of unstructured medical text, greatly improving the availability of the data and the ability to support medical research and clinical decision-making.

[0133] Corresponding to the above method embodiment, this specification also provides an embodiment of a text data processing method applied to a cloud device. See Figure 2 , Figure 2 shows a flowchart of a text data processing method applied to a cloud device according to an embodiment of this specification, which specifically includes the following steps.

[0134] Step 202: Receive a text processing request sent by a device on the receiving end. The text processing request includes target text data to be processed, and the target text data to be processed includes at least one article paragraph information and at least one article title information.

[0135] Step 204: Input the target text data to be processed into a title acquisition model, and obtain the article-level title information output by the title acquisition model. The article-level title information characterizes the hierarchical relationship between each article title information.

[0136] Step 206: Obtain a text data processing result based on the article-level title information and each article paragraph information, and send the text data processing result to the device on the end side. The text data processing result includes at least one title-paragraph pair.

[0137] In practical applications, the cloud device is a server or a computing platform for processing complex computing tasks, usually having relatively high processing capabilities and storage capabilities; the device on the end side is a computing device at the user end, such as a smart phone, a personal computer or an embedded device, mainly used to receive and display processing results, and usually has relatively weak processing capabilities; the text processing request is a data request sent by the device on the end side to the cloud device, requesting the cloud device to perform a task of text data processing.

[0138] The cloud device can be understood as an efficient computing resource, in which large-scale data processing and model calculations are performed to support the needs of the device on the end side. For example, in a text data processing task, the cloud device receives a request from the device on the end side and performs complex model inferences or data processing, and returns the processing result. The advantages of the cloud device lie in its scalability and powerful computing capabilities, enabling it to process large-scale text data and perform computationally intensive tasks.

[0139] The device on the end side can be understood as a device used by end users. Its processing capabilities are usually limited by hardware conditions, so it needs to rely on the cloud device for data processing. For example, the device on the end side can be a mobile phone. The user sends a request and part of the data to the cloud, and the cloud processes it and returns the result to the device on the end side for display or further operations. The main role of the device on the end side is to serve as an interface for data input and output, initiating requests to the cloud device and displaying results.

[0140] The text processing request can be understood as a request message sent by the device on the end side to the cloud device, which contains the text data to be processed and related instructions. The cloud device analyzes and processes the text data according to the content of the request. For example, when the device on the end side processes an article, it may need the cloud device to extract the title, paragraph content of the article or perform content simplification, and the text processing request conveys these requirements to the cloud device.

[0141] The above is a schematic solution of a text data processing method applied to a cloud device in this embodiment. It should be noted that the technical solution of the text data processing method applied to the cloud device belongs to the same concept as the technical solution of the above text data processing method. For the details not described in the technical solution of the text data processing method applied to the cloud device, reference can be made to the description of the technical solution of the above text data processing method.

[0142] Applying the solution of the embodiments of this specification, through the powerful computing power of the cloud device, it can efficiently receive and process a large amount of text data requests from the end-side device, ensuring the timeliness and stability of data processing; the title acquisition model accurately extracts the hierarchical title information of the article in the cloud, clearly depicts the hierarchical relationship between the titles, and significantly improves the accuracy and consistency of the structured information; based on the extracted hierarchical title information and the article paragraph content, accurate title-paragraph pairs are generated, greatly enhancing the organization and readability of the text data, effectively reducing information redundancy and improving semantic understanding; the processing result is efficiently transmitted to the end-side device through the cloud, optimizing the data flow process, improving the overall processing efficiency and user experience. In addition, the cloud centrally manages and processes data, ensuring data security and scalability, supporting in-depth parsing and structured processing of large-scale medical texts, greatly promoting the scientific and intelligent development of medical research and clinical decision-making, and accelerating the digital transformation process in the medical field.

[0143] See Figure 3 , Figure 3 shows an architecture diagram of a text data processing system provided by an embodiment of this specification. The text data processing system may include a client 100 and a server 200;

[0144] The client 100 is used to send a text processing request to the server 200, where the text processing request includes target text data to be processed, and the target text data to be processed includes at least one article paragraph information and at least one article title information;

[0145] The server 200 is used to input the target text data to be processed into a title acquisition model, and obtain the hierarchical article title information output by the title acquisition model, where the hierarchical article title information represents the hierarchical relationship between the article title information, and the text data processing result includes at least one title-paragraph pair;

[0146] Based on the hierarchical article title information and each article paragraph information, obtain a text data processing result; send the text data processing result to the client 100;

[0147] The client 100 is further used to receive the text data processing result sent by the server 200.

[0148] Applying the solution of the embodiments of this specification, the efficient distribution and management of large-scale text processing requests are realized by adopting the architecture of the client and the server, ensuring the response speed and stability of the system when processing massive data. The hierarchical title information is extracted by the title acquisition model to accurately depict the structural relationship between the titles of the article, optimizing the organization and retrieval efficiency of the document. The process of generating title-paragraph pairs not only improves the structural degree of the text, but also enhances the readability and logic of the content, reduces information redundancy, and improves the accuracy of semantic understanding. Such a system architecture supports a flexible and scalable data processing process, enabling complex medical texts to be deeply analyzed and effectively utilized, thus significantly promoting the in-depth development of medical research and the scientificization of clinical decision-making.

[0149] The text data processing system may include multiple clients 100 and a server 200. Among them, the client 100 may be referred to as the device on the client side, and the server 200 may be referred to as the device on the cloud side. Communication connections can be established between multiple clients 100 through the server 200. In the text data processing scenario, the server 200 is used to provide text data processing services between multiple clients 100. Multiple clients 100 can be used as the sending end or the receiving end respectively to achieve communication through the server 200.

[0150] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the text data processing scenario, it can be that the user publishes a data stream to the server 200 through the client 100, and the server 200 generates a text data processing result based on the data stream and pushes the text data processing result to other clients that have established communication.

[0151] Among them, a connection is established between the client 100 and the server 200 through a network. The network provides the medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the client 100 may need to be processed such as encoding, transcoding, and compression before being published to the server 200.

[0152] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application), or a cloud application, etc. The client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0153] The server 200 can include servers that provide various services. For example, a server that provides communication services for multiple clients, or a server for background training that provides support for models used on the client, or a server that processes data sent by the client, etc. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.

[0154] It is worth noting that the text data processing method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client can also have a similar function to the server, so as to execute the text data processing method provided in the embodiments of this specification. In other embodiments, the text data processing method provided in the embodiments of this specification can also be jointly executed by the client and the server.

[0155] Corresponding to the above method embodiments, this specification also provides embodiments of a medical data structuring method. See Figure 4 , Figure 4 The figure shows a flowchart of a medical data structuring method provided according to an embodiment of this specification, which specifically includes the following steps.

[0156] Step 402: Obtain target medical data to be processed, where the target medical data to be processed includes at least one article paragraph information and at least one article title information.

[0157] Step 404: Input the target medical data to be processed into a title acquisition model, and obtain the article-level title information output by the title acquisition model, where the article-level title information represents the hierarchical relationship between each article title information.

[0158] Step 406: Obtain structured medical data based on the article-level title information and each article paragraph information, where the structured medical data includes at least one title-paragraph pair.

[0159] In practical applications, the target medical data to be processed is unprocessed medical information, which usually includes raw data such as patient medical records, diagnostic reports, laboratory test results, etc.; the structured medical data is data after being processed and organized, and these data are usually represented in the form of tables, databases or other systematic formats for easy analysis and storage.

[0160] Exemplarily, the target medical data to be processed can be understood as medical information including multiple natural paragraphs, including the patient's health status, treatment records, and other medical data, and these data are usually unstructured and may exist in the form of text, images or audio. For example, doctors' medical records, laboratory reports, and medical images in hospitals all belong to the target medical data to be processed, and these data need to be further processed for subsequent analysis. The structured medical data can be understood as medical data after being processed and transformed, and is usually presented in a form that is easy for machines to understand and analyze, such as tables, databases or specific data standard formats. The structured data can clearly define the relationships between various types of medical information, facilitating rapid access, analysis and decision-making. For example, by converting the patient's diagnosis and treatment information into a database table, the system can quickly retrieve and process this information for applications such as statistical analysis and prediction.

[0161] The above is a schematic solution of a medical data structuring method of this embodiment. It should be noted that the technical solution of this medical data structuring method and the technical solution of the above text data processing method belong to the same concept. For the details not described in detail in the technical solution of the medical data structuring method, reference can be made to the description of the technical solution of the above text data processing method.

[0162] Applying the solution of the embodiments of this specification, through the distributed architecture of the client and the server, the efficient collaboration and resource optimization of large-scale medical data processing tasks are realized, significantly improving the processing capacity and response speed of the system. The application of the title acquisition model accurately extracts the hierarchical title information of the article, clearly depicts the logical relationship between each title, and enhances the structured degree and navigability of the data. Based on the association between the hierarchical title and the paragraph content, title-paragraph pairs are generated, effectively organizing the medical text content and improving the accuracy of information retrieval and the reading experience of users. In addition, the generation of structured medical data promotes data consistency and reusability, supports the data analysis needs of medical research and the scientific basis for clinical decision-making. Overall, the optimization of the system architecture and processing flow not only improves the efficiency of medical data management, but also promotes the development of medical informatization and helps to achieve the goal of intelligent medical services.

[0163] Corresponding to the above method embodiments, this specification also provides embodiments of a medical data structuring method applied to cloud devices. Refer to Figure 5 , Figure 5 which shows a flowchart of a medical data structuring method applied to cloud devices according to an embodiment of this specification, specifically including the following steps.

[0164] Step 502: Receive a medical data processing request sent by a device on the client side, where the medical data processing request includes target medical data to be processed, and the target medical data to be processed includes at least one article paragraph information and at least one article title information.

[0165] Step 504: Input the target medical data to be processed into a title acquisition model, and obtain the hierarchical article title information output by the title acquisition model, where the hierarchical article title information represents the hierarchical relationship between each article title information.

[0166] Step 506: Obtain structured medical data based on the hierarchical article title information and each article paragraph information, and send the structured medical data to the device on the client side, where the structured medical data includes at least one title-paragraph pair.

[0167] In practical applications, a medical data processing request is a request issued for a processing task of medical data, including data acquisition, processing, and structuring; the target medical data to be processed is the original data in the medical data processing request, usually including original and unstructured data such as patient information and diagnosis results; the structured medical data is the data processed and organized in a predetermined format, which can support subsequent analysis and decision-making. A medical data processing request can be understood as a processing requirement or task request for medical data, usually proposed by a system or a user, asking to process, repair, or analyze medical data. These requests may involve batch processing of a large amount of patient data or specific data analysis for individual cases in practical applications. For example, a certain request may be to classify and summarize the diagnostic reports collected by a hospital, or to structure the medical record data of a patient for subsequent intelligent analysis.

[0168] The above is a schematic solution of a medical data structuring method applied to a cloud device in this embodiment. It should be noted that the technical solution of the medical data structuring method applied to a cloud device belongs to the same concept as the technical solution of the above medical data structuring method. For the details not described in the technical solution of the medical data structuring method applied to a cloud device, reference can be made to the description of the technical solution of the above medical data structuring method.

[0169] Applying the solution of the embodiment of this specification, through the adoption of a distributed architecture of a client and a cloud server, the efficient reception and centralized management of medical data processing requests are realized, significantly improving the response speed and stability of the system when processing a large amount of data. Deploying the title acquisition model on the cloud server ensures the accurate extraction and efficient calculation of complex hierarchical title information, optimizing the organizational structure and information retrieval ability of medical documents. The generated title-paragraph pairs not only enhance the structuring degree and readability of the text, but also effectively reduce information redundancy, improving the depth and accuracy of semantic understanding. The structured medical data is quickly transmitted to the end-side device through the cloud, ensuring the real-time availability and convenient access of the data, supporting the timeliness of clinical decision-making and the in-depth development of medical research. In addition, the high scalability and flexibility of the cloud processing architecture ensure that the system can still operate stably in the face of the increasing amount of medical data, further promoting the digital transformation and intelligent development in the medical field, and greatly improving the utilization efficiency of data and the ability to support the decision-making of medical professionals.

[0170] See Figure 6 , Figure 6 shows an architecture diagram of a medical data structuring system provided by an embodiment of this specification. The medical data structuring system may include a client 300 and a server 400;

[0171] The client 300 is used to send a medical data processing request to the server 400. The medical data processing request includes target medical data to be processed, and the target medical data to be processed includes at least one article paragraph information and at least one article title information.

[0172] The server 400 is used to input the target medical data to be processed into a title acquisition model, and acquire the article-level title information output by the title acquisition model. The article-level title information represents the hierarchical relationship between each article title information. Based on the article-level title information and each article paragraph information, structured medical data is acquired. The structured medical data includes at least one title-paragraph pair. The server 400 sends the structured medical data to the client 300.

[0173] The client 300 is further used to receive the structured medical data sent by the server 400.

[0174] Applying the solution of the embodiments of this specification, by using the distributed architecture of the client and the server, the efficient transmission and centralized management of medical data processing requests are realized, ensuring that the system still operates efficiently when processing a large amount of data. The generated title-paragraph pairs not only improve the structured level and readability of the text, but also effectively reduce information redundancy and enhance the depth of semantic understanding. The structured medical data is quickly transmitted to the client device through an efficient network transmission mechanism, ensuring the real-time availability and convenient access of the data, and supporting the timeliness of clinical decision-making and the in-depth development of medical research. In addition, the high scalability and flexibility of the system architecture ensure that it can still operate stably when facing the increasing amount of medical data, further promoting the digital transformation and intelligent development in the medical field, and significantly improving the data utilization efficiency and the ability to support the decision-making of medical professionals.

[0175] The medical data structuring system may include multiple clients 300 and a server 400. The client 300 may be referred to as an end-side device, and the server 400 may be referred to as a cloud-side device. Communication connections can be established between multiple clients 300 through the server 400. In the medical data structuring scenario, the server 400 is used to provide medical data structuring services between multiple clients 300. Multiple clients 300 can be used as senders or receivers respectively to achieve communication through the server 400.

[0176] The user can interact with the server 400 through the client 300 to receive data sent by other clients 300, or send data to other clients 300, etc. In the medical data structuring scenario, it can be that the user publishes a data stream to the server 400 through the client 300, and the server 400 generates structured medical data according to the data stream and pushes the structured medical data to other clients that have established communication.

[0177] The above is a schematic solution of a medical data structuring system according to this embodiment. It should be noted that the technical solution of this medical data structuring system and the technical solutions of the above-mentioned medical data structuring method applied to cloud devices and the text data processing system belong to the same concept. For the details not described in the technical solution of the medical data structuring system, reference can be made to the descriptions of the above-mentioned medical data structuring method applied to cloud devices and the technical solution of the text data processing system.

[0178] The following combines the attached Figure 7 , taking the application of the text data processing method provided in this specification in the scientific research paper structuring method as an example, to further illustrate the text data processing method. Among them, Figure 7 FIG. shows a flowchart of the processing process of a scientific research paper structuring method provided by an embodiment of this specification, which specifically includes the following steps.

[0179] Step 702: Perform OCR scanning on the paper-based paper A to obtain initial text data.

[0180] Step 704: Split the initial data into multiple text data to be repaired according to the end punctuation marks in the initial data and the restricted length.

[0181] Step 706: Input the multiple text data to be repaired after splitting into the field repair model one by one, obtain multiple repaired target text paragraphs, and splice the multiple target text paragraphs to generate the repaired target text data.

[0182] Step 708: Determine that each paragraph with a length greater than the threshold and a distance between the positions of the paragraphs less than the threshold in the target text data is a paragraph to be compressed.

[0183] Step 710: Input each paragraph to be compressed into the abstract extraction model to obtain the article paragraph abstract data output by the abstract extraction model.

[0184] Step 712: Replace the paragraphs in the target text data corresponding to the article paragraph abstract data with their corresponding article paragraph abstract data.

[0185] Step 714: Input the replaced text data and the article-level title information prediction template into the hierarchical title extraction model, so that the hierarchical title extraction model performs dynamic restricted decoding to obtain the article-level title information.

[0186] Step 716: Perform structuring processing on the target text data according to the article-level title information to obtain the text corresponding to the structured paper A.

[0187] Applying the solution of the embodiments of this specification, by performing OCR scanning on paper-based papers, non-digital text is successfully converted into initial data that can be processed, ensuring the integrity and accessibility of the data. The initial text is precisely segmented using end punctuation marks and length limits, effectively separating the text paragraphs to be repaired, and improving the efficiency and accuracy of subsequent processing. The application of the field repair model repairs the segmented text data paragraph by paragraph, eliminating errors and incomplete information during the scanning process, and significantly improving the quality and consistency of the text data. For the paragraphs to be compressed with lengths exceeding the threshold and close adjacent paragraph distances, paragraph summaries are generated by the summary extraction model, successfully reducing the text length, optimizing the efficiency of data processing, and maintaining the integrity of key information and semantic coherence. After replacing the original paragraphs with the summary data, dynamic limit decoding is performed in combination with the hierarchical title information prediction template to accurately extract the hierarchical title information of the article, ensuring the accurate construction of the hierarchical relationship between the titles. Structuring the target text according to the extracted hierarchical title information realizes the high structuring of scientific research papers, not only improving the readability and organization of the text, but also greatly improving the efficiency of information retrieval and data analysis, providing a solid data foundation and support for scientific research work, and promoting the in-depth development of academic research and the optimization of knowledge management.

[0188] Corresponding to the above method embodiments, this specification also provides embodiments of a text data processing device. Figure 8 The structure diagram of a text data processing device provided by an embodiment of this specification is shown. As Figure 8 shown, the device includes:

[0189] A text acquisition module 802, configured to acquire target text data to be processed, where the target text data to be processed includes at least one article paragraph information and at least one article title information;

[0190] A hierarchical relationship extraction module 804, configured to input the target text data to be processed into a title acquisition model, and acquire the hierarchical article title information output by the title acquisition model, where the hierarchical article title information represents the hierarchical relationship between each article title information;

[0191] A result acquisition module 806, configured to acquire a text data processing result based on the hierarchical article title information and each article paragraph information, where the text data processing result includes at least one title-paragraph pair.

[0192] Optionally, the text acquisition module 802 is further configured to:

[0193] Acquire initial text data to be processed, where the text data to be processed includes at least one field end identifier;

[0194] Process the initial text data to be processed based on a preset text processing length threshold and end identifiers of each field, and obtain the target text data to be processed.

[0195] Optionally, the text acquisition module 802 is further configured to:

[0196] Determine at least one reference text data to be processed in the initial text data to be processed based on a preset text processing length threshold and end identifiers of each field;

[0197] Generate a text repair prompt word according to a preset text repair prompt word template and each reference text data to be processed, and input the text repair prompt word into a text repair model to obtain the target text data to be processed generated by the text repair model.

[0198] Optionally, the hierarchical relationship extraction module 804 is further configured to:

[0199] Obtain an article hierarchical title information prediction template, where the article hierarchical title information prediction template includes at least one title information prediction position;

[0200] Input the target text data to be processed and the article hierarchical title information prediction template into a title acquisition model to obtain the article hierarchical title sub-information corresponding to each title information prediction position generated by the title acquisition model;

[0201] Generate article hierarchical title information according to the article hierarchical title sub-information corresponding to each title information prediction position and the article hierarchical title information prediction template.

[0202] Optionally, the hierarchical relationship extraction module 804 is further configured to:

[0203] Generate at least one article paragraph summary data based on the target text data to be processed;

[0204] Generate title information extraction text according to each article paragraph summary data and the target text data to be processed;

[0205] Input the title information extraction text and the article hierarchical title information prediction template into a title acquisition model to obtain the article hierarchical title sub-information corresponding to each title information prediction position generated by the title acquisition model.

[0206] Optionally, the hierarchical relationship extraction module 804 is further configured to:

[0207] Determine at least one paragraph information to be compressed in each article paragraph information based on a preset summary extraction threshold;

[0208] Input the information of each paragraph to be compressed into a paragraph summary generation model to obtain the article paragraph summary data generated by the paragraph summary generation model.

[0209] Optionally, the hierarchical relationship extraction module 804 is further configured to:

[0210] Obtain the paragraph text length information and paragraph position information corresponding to each article paragraph information;

[0211] Based on a preset summary extraction threshold and the paragraph text lengths corresponding to each article paragraph information, determine at least one paragraph sub-information to be compressed in each article paragraph information, where the paragraph sub-information to be compressed is the article paragraph information with a paragraph text length less than the preset summary extraction threshold;

[0212] Based on the paragraph position information corresponding to each paragraph sub-information to be compressed, obtain the paragraph distance information between each paragraph sub-information to be compressed;

[0213] According to the paragraph distance information between each paragraph sub-information to be compressed, obtain at least one paragraph information to be compressed, where the paragraph information to be compressed includes at least one paragraph sub-information to be compressed, and the paragraph distance information between the paragraph sub-information to be compressed in the paragraph information to be compressed is less than or equal to a preset paragraph aggregation threshold.

[0214] Optionally, the data processing device further includes a model training module, which is configured to:

[0215] Obtain an article hierarchical title information prediction template, as well as sample paragraph information and the sample title information corresponding to the sample paragraph information, where the article hierarchical title information prediction template includes at least one title information prediction position;

[0216] Input the sample paragraph information and the article hierarchical title information prediction template into a title acquisition model to obtain the article hierarchical title information generated by the title acquisition model;

[0217] Determine the predicted title information in the article hierarchical title information based on each title information prediction position;

[0218] Generate a model loss value according to the sample title information and the predicted title information, and train the title acquisition model according to the model loss value until the model training stop condition is reached.

[0219] The above is a schematic solution of a text data processing device in this embodiment. It should be noted that the technical solution of this text data processing device and the technical solution of the above text data processing method belong to the same concept. For the details not described in the technical solution of the text data processing device, reference can be made to the description of the technical solution of the above text data processing method.

[0220] Applying the solution of the embodiments of this specification, by integrating multiple processing modules, the text data processing device can effectively obtain and optimize the target text data, ensuring the accuracy and integrity of the data. The text acquisition module accurately segments the initial text by combining the length threshold and the field end identifier, effectively removing irrelevant information and retaining text content with clear structure and strong relevance. The hierarchical relationship extraction module identifies the hierarchical structure of the article through the title acquisition model and accurately extracts the title information, effectively establishing the hierarchical relationship between the article titles and providing a reliable basis for subsequent data structuring. Further optimization measures repair the errors and incomplete information in the text data paragraph by paragraph by combining the text repair model, thereby improving the quality and consistency of the processing results. In addition, the abstract generation and paragraph compression functions in the device reduce text redundancy while retaining the core information, improve the efficiency of data processing, and ensure the coherence of the context and the integrity of the information. By optimizing the loss function of the title acquisition model through the model training module, the efficiency and robustness of the system are further improved, ensuring the accurate parsing and structured processing of the text data. The collaborative work of multiple modules of the device significantly improves the automation and accuracy of text data processing, providing more efficient and accurate data support for scientific research and information management.

[0221] Corresponding to the above method embodiments, this specification also provides embodiments of a medical data structuring device. Figure 9 The structure diagram of a medical data structuring device provided by an embodiment of this specification is shown. As Figure 9 shown, the device includes:

[0222] A medical data acquisition module 902, configured to acquire target medical data to be processed, where the target medical data to be processed includes at least one article paragraph information and at least one article title information;

[0223] A hierarchical relationship extraction module 904, configured to input the target medical data to be processed into a title acquisition model, and acquire the hierarchical article title information output by the title acquisition model, where the hierarchical article title information represents the hierarchical relationship between each article title information;

[0224] A data structuring module 906, configured to acquire structured medical data based on the hierarchical article title information and each article paragraph information, where the structured medical data includes at least one title-paragraph pair.

[0225] The above is a schematic solution of a medical data structuring device according to this embodiment. It should be noted that the technical solution of this medical data structuring device and the technical solutions of the above-mentioned medical data structuring method and text data processing device belong to the same concept. For the details not described in the technical solution of the medical data structuring device, reference can be made to the descriptions of the above-mentioned medical data structuring method and text data processing device.

[0226] Applying the solution of the embodiments of this specification, by combining multiple modules for medical data processing, it is possible to efficiently obtain and optimize the data to be processed, significantly improving the degree of structuring and usability of the data. The medical data acquisition module ensures the efficient acquisition of each paragraph and title information by accurately screening the target data, laying a foundation for subsequent data processing. The hierarchical relationship extraction module uses the title acquisition model to accurately identify and extract the hierarchical relationship of titles in the article, thus effectively constructing the hierarchical structure between titles and ensuring the clarity and stability of the logical relationship of the article. The data structuring module converts the medical text into a structured format based on the title information and paragraph data, making the corresponding relationship between each paragraph and the title clear, thereby providing more accurate data support for subsequent analysis and processing. In addition, through the collaborative work of the modules, the entire device process can greatly improve the efficiency and accuracy of data processing, reduce manual intervention and potential errors, improve the availability of medical data, and provide more accurate support for medical research and clinical decision-making.

[0227] Figure 10 FIG. shows a structural block diagram of a computing device 1000 according to an embodiment of this specification. The components of the computing device 1000 include but are not limited to a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 through a bus 1030, and a database 1050 is used to store data.

[0228] The computing device 1000 further includes an access device 1040, which enables the computing device 1000 to communicate via one or more networks 1060. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1040 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).

[0229] In one embodiment of the present specification, the above components of the computing device 1000 and Figure 10 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 10 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.

[0230] The computing device 1000 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1000 can also be a mobile or stationary server.

[0231] Wherein, the processor 1020 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above text data processing method and medical data structuring method are implemented.

[0232] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above text data processing method and medical data structuring method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above text data processing method and medical data structuring method.

[0233] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above text data processing method and medical data structuring method.

[0234] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above text data processing method and medical data structuring method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above text data processing method and medical data structuring method.

[0235] An embodiment of this specification also provides a computer program product including a computer program / instructions, which, when executed by a processor, implement the steps of the above text data processing method and medical data structuring method.

[0236] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above text data processing method and medical data structuring method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the descriptions of the technical solutions of the above text data processing method and medical data structuring method.

[0237] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0238] The computer instructions include computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0239] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0240] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0241] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.< / h3>

Claims

1. A text data processing method, comprising: Acquire target text data to be processed, wherein the target text data to be processed includes at least one article paragraph information and at least one article title information; Inputting the target text data to be processed into a title acquisition model, and acquiring article-level title information output by the title acquisition model, wherein the article-level title information represents the hierarchical relationship between the title information of each article; A text data processing result is obtained based on the article-level title information and each article paragraph information, wherein the text data processing result includes at least one title-paragraph pair.

2. The method according to claim 1, obtaining the target text data to be processed, comprising: Acquire initial text data to be processed, wherein the text data to be processed includes at least one field end identifier; The initial text data to be processed is processed based on a preset text processing length threshold and end identifiers of each field to obtain target text data to be processed.

3. The method according to claim 2, processing the initial text data to be processed based on a preset text processing length threshold and each field end mark to obtain target text data to be processed, comprising: Based on a preset text processing length threshold and end identifiers of each field, determining at least one reference text data to be processed in the initial text data to be processed; A text repair prompt word is generated according to a preset text repair prompt word template and each reference text data to be processed, and the text repair prompt word is input into a text repair model to obtain target text data to be processed generated by the text repair model.

4. The method according to claim 1, inputting the target text data to be processed into a title acquisition model, and acquiring the article-level title information output by the title acquisition model, comprises: Acquire an article-level title information prediction template, wherein the article-level title information prediction template includes at least one title information prediction position; Inputting the target text data to be processed and the article-level title information prediction template into a title acquisition model, and acquiring the article-level title sub-information corresponding to each title information prediction position generated by the title acquisition model; The article-level title information is generated according to the article-level title sub-information corresponding to each title information prediction position and the article-level title information prediction template.

5. The method according to claim 4, inputting the target text data to be processed and the article-level title information prediction template into a title acquisition model, and acquiring the article-level title sub-information corresponding to each title information prediction position generated by the title acquisition model, comprising: Based on the target text data to be processed, generating at least one article paragraph summary data; Generate title information extraction text according to each article paragraph summary data and the target text data to be processed; The title information extraction text and the article-level title information prediction template are input into a title acquisition model to obtain article-level title sub-information corresponding to each title information prediction position generated by the title acquisition model.

6. The method according to claim 5, generating at least one article paragraph summary data based on the target text data to be processed, comprising: Determining at least one paragraph information to be compressed in each article paragraph information based on a preset summary extraction threshold; The information of each paragraph to be compressed is input into a paragraph summary generation model to obtain the article paragraph summary data generated by the paragraph summary generation model.

7. The method according to claim 6, wherein determining at least one paragraph information to be compressed in each article paragraph information based on a preset summary extraction threshold comprises: Obtain the paragraph text length information and paragraph position information corresponding to each article paragraph information; Based on a preset summary extraction threshold and a paragraph text length corresponding to each article paragraph information, determining at least one paragraph sub-information to be compressed in each article paragraph information, wherein the paragraph sub-information to be compressed is the article paragraph information whose paragraph text length is less than the preset summary extraction threshold; Based on the paragraph position information corresponding to each paragraph sub-information to be compressed, obtaining the paragraph distance information between each paragraph sub-information to be compressed; At least one paragraph information to be compressed is obtained according to paragraph distance information between each paragraph sub-information to be compressed, wherein the paragraph information to be compressed includes at least one paragraph sub-information to be compressed, and the paragraph distance information between each paragraph sub-information to be compressed in the paragraph information to be compressed is less than or equal to a preset paragraph aggregation threshold.

8. A method for structuring medical data, comprising: Acquire target medical data to be processed, wherein the target medical data to be processed includes at least one article paragraph information and at least one article title information; Inputting the target to-be-processed medical data into a title acquisition model, and acquiring article-level title information output by the title acquisition model, wherein the article-level title information represents a hierarchical relationship between the title information of each article; Structured medical data is acquired based on the article-level title information and each article paragraph information, wherein the structured medical data includes at least one title-paragraph pair.

9. A medical data structuring method, applied to a cloud device, comprising: A medical data processing request sent by a receiving end-side device, wherein the medical data processing request includes target medical data to be processed, and the target medical data to be processed includes at least one article paragraph information and at least one article title information; Inputting the target to-be-processed medical data into a title acquisition model, and acquiring article-level title information output by the title acquisition model, wherein the article-level title information represents a hierarchical relationship between the title information of each article; Based on the article-level title information and each article paragraph information, structured medical data is acquired, and the structured medical data is sent to a terminal-side device, wherein the structured medical data includes at least one title-paragraph pair.

10. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • File structured information extraction method and device, equipment, medium and product

    CN120849649A

  • File structured information extraction method, device, equipment, medium and product

    CN120849649B