Medical system structured data acquisition method based on web crawler technology

Through the structured data collection method of medical system based on network crawling technology, the problem of dispersed and difficult to integrate medical data is solved, and efficient and accurate data collection and integration is achieved, supporting the needs of clinical research and public health management.

CN120164593AInactive Publication Date: 2025-06-17ZHEJIANG UNIV

Patent Information

Application Number
CN202510317467.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing medical data storage and management models make data dispersed and difficult to integrate, affecting the efficiency and accuracy of clinical research, quality control and public health management.

Method used

The structured data collection method of medical system based on network crawling technology is adopted, and the automated collection and integration of multi-source medical data is realized through modules such as patient information reading, network request and redirection processing, data classification analysis, time screening and keyword matching, and structured writing.

Benefits of technology

It improves data collection efficiency, enhances data accuracy and completeness, reduces labor costs, improves the level of informatization, and provides strong data support for clinical scientific research, quality assessment and public health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164593A_ABST
    Figure CN120164593A_ABST
Patent Text Reader

Abstract

The invention discloses a medical system structured data acquisition method based on a web crawler technology, and relates to the technical field of data processing, a patient information reading module obtains a patient identifier from a pre-prepared list file, and provides the patient identifier to a network request and redirection processing module; the network request and redirection processing module constructs a network request link and sets request header information according to the patient identifier in combination with an interface specification of a hospital information system, and performs redirection processing to generate response data; the data classification analysis module performs condition matching and extraction on the response data based on a preset rule to obtain medical data; the time screening and keyword matching module extracts key data in the medical data according to a preset condition; and the structured write-in module writes the target fields in the key data into a target storage medium one by one for storage according to a preset field mapping rule. According to the invention, integration of multi-source and dispersed medical data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and more specifically, to a method for collecting structured data of a medical system based on web crawler technology. Background Art

[0002] At present, with the continuous deepening of hospital informatization construction, hospitals generally have multiple independent subsystems such as Hospital Information System (HIS), Laboratory Information Management System (LIS), Picture Archiving and Communication System (PACS), and Electronic Medical Record System (EMR). These systems respectively store various medical data generated during the patient's visit, such as inpatient records, test results, surgical documents, imaging reports, discharge summaries, etc. However, these data are often stored in different databases or information interfaces, and there is a lack of a unified data exchange and sharing mechanism, resulting in scattered data and difficult to integrate efficiently.

[0003] From the perspective of clinical research, the need for a large amount of high-quality medical data is extremely urgent. Researchers expect to explore the occurrence and development laws of diseases and evaluate the efficacy of different treatment plans by deeply analyzing multi-dimensional data such as disease diagnosis, treatment process, and treatment effect, so as to promote the innovation and development of medical research. However, the current scattered data situation makes researchers have to spend a lot of time and energy switching and searching between various systems when collecting and sorting data. Not only is the efficiency low, but also data omission or error is very likely to occur, seriously affecting the accuracy and reliability of the research.

[0004] In terms of medical quality control, timely and comprehensive access to medical data is crucial for monitoring the medical process and evaluating medical quality. For example, by analyzing data such as surgical records and test reports, potential medical risks can be discovered in time, and whether medical services meet the specifications and standards can be evaluated. However, due to the scattered data, it is difficult for quality control personnel to quickly obtain complete and accurate data, and they cannot carry out quality monitoring and improvement in a timely and effective manner, which is not conducive to the improvement of the overall medical service level of the hospital.

[0005] Public health management also relies on large-scale and high-quality medical data. In the work of infectious disease prevention and control, it is necessary to collect and analyze data such as patients' symptoms, diagnosis results, and epidemiological history in time, so as to quickly discover epidemic clues and formulate precise prevention and control strategies. However, the current scattered data state makes it difficult for public health management departments to quickly obtain comprehensive and accurate data. In key work such as epidemic prevention and control, the prevention and control opportunity may be delayed due to untimely and inaccurate information, threatening public health and safety.

[0006] With the increasing application of artificial intelligence and big data technologies in the medical field, the demand for structured and standardized medical data is growing day by day. Whether it is the training of AI-assisted diagnosis models or the optimization of clinical decision support systems, a large amount of high-quality medical data is required as support. However, the existing data storage and management models are difficult to meet the strict requirements of these emerging technologies for data, restricting the pace of intelligent development in the medical field.

[0007] Therefore, how to integrate multi-source and scattered medical data is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0008] In view of this, the present invention provides a method for collecting structured data of a medical system based on web crawler technology to solve the problems existing in the above-mentioned background technology.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] A method for collecting structured data of a medical system based on web crawler technology includes: a patient information reading module, a network request and redirection processing module, a data classification and parsing module, a time screening and keyword matching module, and a structured writing module. The patient information reading module obtains patient identifiers from a pre-prepared list file and provides the patient identifiers to the network request and redirection processing module; the network request and redirection processing module constructs a network request link and sets request header information according to the patient identifiers in combination with the interface specifications of the hospital information system, and performs redirection processing to generate response data; the data classification and parsing module performs condition matching and extraction on the response data based on preset rules to obtain medical data; the time screening and keyword matching module extracts key data from the medical data according to preset conditions; the structured writing module writes the target fields in the key data one by one into the target storage medium for storage according to the preset field mapping rules.

[0011] Preferably, the list file of the patient information reading module is an Excel table.

[0012] Preferably, the network request and redirection processing module reads the patient identifiers from the Excel table and generates the real request URL of the patient by string replacement or URL parameter splicing.

[0013] Preferably, the preset rules include: constructing retrieval conditions and parsing rules based on ultrasound reports, radiology reports, laboratory reports, pathology reports, outpatient medical records, inpatient medical records, surgical records, and discharge records.

[0014] Preferably, the preset conditions in the time screening and keyword matching module are preset time periods or keywords.

[0015] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for collecting structured data of a medical system based on web crawler technology, and the beneficial effects are as follows:

[0016] 1. Greatly improved efficiency: Through batch processing of scripts, a large amount of patient data can be sorted out in a short time.

[0017] 2. Enhanced accuracy: The automatic script matches keywords with time periods, avoiding possible omissions or copying mistakes in manual operations, and ensuring data consistency and integrity.

[0018] 3. Strong scalability: Only by making corresponding adjustments to the crawled URLs and parsing rules according to the interfaces or text formats of different hospital systems, it can be adapted to multiple hospitals or new systems within the same hospital.

[0019] 4. High degree of automation: Adopting the idea of web crawlers and scripted operations, automatically log in and access each subsystem within the authorized scope and extract data, reducing labor costs and improving the informatization level.

[0020] 5. Provide a basis for secondary analysis: The obtained structured medical data can not only directly support clinical research statistics, but also be further integrated into the hospital big data platform, providing strong support for AI-assisted diagnosis, clinical decision-making support, quality assessment and scientific research analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0022] Figure 1 It is a schematic structural diagram provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0024] The embodiments of the present invention disclose a method for collecting structured data of a medical system based on web crawler technology, as Figure 1As shown in the figure, it includes: a patient information reading module, a network request and redirection processing module, a data classification and parsing module, a time filtering and keyword matching module, and a structured writing module. The patient information reading module obtains the patient identifier from a pre-prepared list file and provides the patient identifier to the network request and redirection processing module. The network request and redirection processing module constructs a network request link and sets the request header information according to the patient identifier in combination with the interface specification of the hospital information system, and performs redirection processing to generate response data. The data classification and parsing module performs conditional matching and extraction on the response data based on preset rules to obtain medical data. The time filtering and keyword matching module extracts key data from the medical data according to preset conditions. The structured writing module writes the target fields in the key data one by one into the target storage medium for storage according to the preset field mapping rules.

[0025] In a specific embodiment, the list file of the patient information reading module is an Excel table.

[0026] In a specific embodiment, the network request and redirection processing module reads the patient identifier from the Excel table and generates the real request URL of the patient by string replacement or URL parameter splicing.

[0027] In a specific embodiment, the preset rules include: constructing corresponding retrieval conditions and parsing rules based on ultrasound reports, radiology reports, laboratory reports, pathology reports, outpatient medical records, inpatient medical records, surgical records, and discharge records.

[0028] In a specific embodiment, the preset condition in the time filtering and keyword matching module is a preset time period or keyword.

[0029] Overall idea:

[0030] Through a unified script or program entry, under the premise of legal authorization, batch requests are made to the Web interfaces or APIs of several internal hospital business systems to simulate the manual query, click, and jump processes.

[0031] According to the pre-designed time window and keywords, in-depth parsing is performed on the returned report data, test results, inpatient medical records, discharge records, etc., and appropriate tools (such as HTML parsers, JSON parsing libraries, etc.) are used to extract specific field information.

[0032] The parsed data is accurately mapped to the pre-defined structured storage locations (such as the columns of Excel or the fields in the database) to achieve batch collection and integration of multi-source medical data.

[0033] Main modules:

[0034] Patient information reading module: Prepare a list file (which can be an Excel sheet) containing patient identifiers (such as medical record numbers, PersonID, etc.) in advance for subsequent collection one by one according to the list.

[0035] Network request and redirection processing module: Responsible for initiating HTTP / HTTPS requests, carrying necessary authentication information (such as Cookies, Authorization Tokens), handling possible page redirections, and obtaining the final return.

[0036] Data classification and parsing module: Construct corresponding retrieval conditions and parsing rules according to different categories such as ultrasound reports, radiology reports, laboratory reports, pathology reports, outpatient medical records, inpatient medical records, surgical records, and discharge records; perform condition matching and extraction on the returned data.

[0037] Time filtering and keyword matching module: For the situation where there are multiple reports for each day or each patient, by setting a specific time period (such as the date range during a patient's hospitalization) or keyword filtering (such as "carotid artery", "cerebral angiography", etc.), only obtain the most relevant report content.

[0038] Structured writing module: Write the parsed target fields one by one into the target storage medium (such as the corresponding columns of an Excel sheet) according to the field mapping rules to achieve batch filling; special marks or leave blank processing can be performed on the fields for which no results are obtained.

[0039] Technical features and innovations:

[0040] Automated collection: Automatically traverse all patients in a scripted manner, eliminating the cumbersome steps of manual querying, downloading, copying, and pasting for each case, and greatly improving work efficiency.

[0041] Flexible parsing: For different types of inspection and examination data, use different crawling or parsing strategies, and perfectly compatible with multiple data formats (such as JSON, HTML, XML).

[0042] Strong scalability: When the hospital adds or adjusts information system interfaces, only need to update the crawler request address or parsing rules to continue collecting data.

[0043] High accuracy: By means of precise string matching, time filtering, document text segmentation and disassembling, etc., reduce the errors that may be brought by manual operations. Specific embodiment 1

[0045] 1) Obtaining the patient list

[0046] 1. In this embodiment, first prepare the unique identification information of several rows of patients in an Excel file, such as the medical record number or the `personID` in the hospital information system.

[0047] 2. Open the Excel file and traverse each patient record to obtain the set of target patients for which data needs to be collected in batches.

[0048] 2) Construct the network request link

[0049] 1. According to the interface specification of the hospital information system, determine the access path template or URL paradigm in advance. Taking the access to patient examination results as an example, it may be necessary to splice the patient identification in the URL, such as:

[0050] https: / / hospital domain name / ADFSAjax / ExternalCall.aspx?source=CohortDesigner&uri=InvokeSPV&PatientId=xxx

[0051] Where `xxx` is the target patient identification that needs to be replaced.

[0052] 2. For each patient, read their identification from the Excel, and dynamically generate the real request URL of the patient by string replacement or URL parameter splicing.

[0053] 3. Set the necessary request header information, including the browser proxy string, Cookies, Authorization, etc. For example:

[0054] `Cookie`: Contains session information to facilitate maintaining the logged-in state later.

[0055] `Authorization`: Carries a token or other authentication credentials to ensure legal access to data.

[0056] 3) Send the request and handle redirects

[0057] 1. Use a standard HTTP library (such as the `requests` library in Python) to send a POST or GET request to the above URL.

[0058] 2. If the server side has a jump logic for the patient request (such as giving a redirect URL in the return header), then automatically or manually follow the redirect to obtain the final real access path.

[0059] 3. Extract the internal unique identification `personID` used by the hospital system from the redirected URL or the returned response data for subsequent calls to more advanced queries.

[0060] 4) Classification-based Data Collection and Parsing

[0061] After obtaining the patient's `personID`, it is necessary to further access multiple data interfaces within the hospital to obtain different types of medical data. This can be divided into the following sub-steps:

[0062] 1. Basic Information Query

[0063] Through a dedicated interface or SQL query statement, obtain basic information such as the patient's name, gender, age, hospital admission number, outpatient number, etc.

[0064] After confirming that the corresponding fields are included in the queried JSON or result set, temporarily store them or directly write them into the output medium.

[0065] 2. Inspection and Test Report Parsing

[0066] Ultrasound Report: Construct a corresponding query statement or API request based on the `personID` to obtain a returned JSON list containing fields such as "ObservationConclusion" (Findings), "ObservationFindings" (Report Conclusion), "ConfirmDate" (Report Date), etc.

[0067] If there are a large number of ultrasound reports, filter them by setting the report date to match a certain visit period or keywords (such as "carotid artery").

[0068] Radiology Report: Similarly query radiology examination data (CT, MRI, CTA, etc.) and extract "ReportDate", "ObservationConclusion", "ObservationFindings", etc.

[0069] Test Report: Query the LIS interface to return the details of test items such as blood routine, blood lipid, blood sugar, liver and kidney function, etc. Extract fields such as "ObservationValue", "Units", "ReferencesRange", etc. and write them into the corresponding table columns such as "white blood cell count", "hemoglobin".

[0070] Pathology Report: If the patient has undergone operations such as surgery or tissue biopsy, a request can be made to the pathology system (which may also be a PACS or an independent system) to obtain "ObservationConclusion" and "ObservationFindings", and record the pathological diagnosis and pathology number.

[0071] 3. Medical Document Parsing

[0072] Inpatient document list acquisition: In the hospital document management module / system, query all inpatient-related document IDs and document types through `personID`, including admission records, discharge records, operation records, physical examination forms, etc.

[0073] Specified document details acquisition: Re-append the obtained document ID to the interface URL to download the corresponding document content (usually in HTML or XML format).

[0074] Use an HTML parser (such as in a way similar to `BeautifulSoup`) to convert the document into plain text;

[0075] Locate the required field content according to the characteristics of the document format and fixed keywords. For example, retrieve "Chief Complaint:", "History of Present Illness:", etc. in the admission record; retrieve "Brief Procedure of Operation:", "Intraoperative Diagnosis:", etc. in the operation record.

[0076] Discharge record and outpatient medical record: Similarly parse the corresponding HTML content and extract key information such as "Admission Date", "Discharge Date", "Diagnosis", "Length of Hospital Stay", "Course of Hospitalization", "Discharge Condition", etc.

[0077] 5) Time range matching and result screening

[0078] 1. Since in clinical practice, the same patient may have multiple hospitalizations and multiple reports, the present invention performs precise screening by setting a time range or keywords.

[0079] 2. Compare the hospitalization start time, operation time, or test occurrence date, and extract the report content that best matches the current hospitalization or the specified interval to avoid mixing in irrelevant data from other medical treatment cycles.

[0080] 3. When a certain piece of data does not exist within the specified interval, "not found" can be recorded or left blank to ensure the integrity of the output structure.

[0081] 6) Structured storage and writing

[0082] 1. Field comparison table: Define each column or field in advance in the target Excel or database (for example, column 3 stores the name, column 4 stores the gender, column 5 stores the age, column 40 stores the white blood cell value, etc.).

[0083] 2. Writing process:

[0084] Traverse each acquisition result in sequence and write it to the corresponding position according to the defined field order.

[0085] If multiple results are collected for a certain field, the one closest to the admission time can be selected according to the time or the latest value can be marked for writing.

[0086] If no corresponding value is found for a certain field, write a null value or a special prompt.

[0087] 3. Saving and Logging: After the collection is completed, save the final Excel or database update; record the collection time, data volume, and possible exceptions.

[0088] 7) Result Verification and Exception Handling

[0089] 1. Integrity Verification: Conduct a preliminary review of the final output for each patient, such as whether there are too many blank fields and whether the date parsing is correct.

[0090] 2. Exception Handling: If a network failure, interface unresponsiveness, or abnormal return format occurs, skip the current patient or record an exception log for subsequent re - collection or manual supplementation.

[0091] The present invention constructs an automated structured data collection process by integrating web crawler technology and multi - data - source access interfaces within a hospital. Its core lies in:

[0092] 1. First, enter the list of target patients;

[0093] 2. Dynamically splice and generate a request link and carry authentication information;

[0094] 3. Access and filter different categories of medical data one by one;

[0095] 4. Conduct text parsing on HTML / JSON results to obtain key fields;

[0096] 5. Finally, uniformly write the results into an Excel sheet or database available for analysis.

[0097] The present invention effectively solves the problem of multi - source, scattered, and difficult - to - integrate medical data, and can play an important role in various types of clinical research and hospital management processes.

[0098] In this specification, each embodiment is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0099] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for collecting structured data of a medical system based on web crawler technology, characterized in that: include: A patient information reading module, a network request and redirection processing module, a data classification and analysis module, a time screening and keyword matching module and a structured writing module. The patient information reading module obtains the patient identification from a pre-prepared list file and provides the patient identification to the network request and redirection processing module; the network request and redirection processing module constructs a network request link and sets the request header information according to the patient identification in combination with the interface specification of the hospital information system, and performs redirection processing to generate response data; the data classification and analysis module performs conditional matching and extraction on the response data based on preset rules to obtain medical data; the time screening and keyword matching module extracts key data in the medical data according to preset conditions; the structured writing module writes the target fields in the key data into the target storage medium one by one for storage according to the preset field mapping rules.

2. According to claim 1, a medical system structured data collection method based on web crawler technology is characterized in that: The list file of the patient information reading module is an Excel table.

3. The method for collecting structured data of a medical system based on web crawler technology according to claim 2, characterized in that: The network request and redirection processing module reads the patient ID from the Excel table, and generates the real request URL of the patient by string replacement or URL parameter concatenation.

4. The method for collecting structured data of a medical system based on web crawler technology according to claim 1, characterized in that: The preset rules include: constructing retrieval conditions and parsing rules based on ultrasound reports, radiology reports, test reports, pathology reports, outpatient medical records, inpatient medical records, surgical records and discharge records.

5. The method for collecting structured data of a medical system based on web crawler technology according to claim 1, characterized in that: The preset conditions in the time screening and keyword matching module are preset time periods or keywords.

Citation Information

Patent Citations

  • Web page information collection method and device based on web crawlers

    CN109657121A

  • Method for extracting data based on medical system crawler

    CN111078976A

  • Litigation case classification method and device, computer equipment and storage medium

    CN111522955A

  • Power business environment information acquisition system based on web crawler technology

    CN114443926A

  • Medical big data acquisition method and system

    CN116775973A

Cited By

  • Medical data structured extraction method based on machine learning

    CN121148571A