Text data extraction method and device, equipment, medium and product
Automatically extract target index data in text through interactive interfaces and pre-trained language models, solving the problem of software program dependence in the prior art, and achieving flexible and efficient text data extraction.
Patent Information
- Application Number
- CN202510388846.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
When extracting target index data in text, the prior art requires modifying software programs, which lacks flexibility and cannot adapt to changes in target indexes.
The target indicator combination and diagnosis and treatment description information are obtained through the interactive interface, and the prompt information is automatically determined using the pre-trained language model to generate the target indicator index data and display it in the visual interface.
It realizes the flexibility and accuracy of text data extraction without modifying software programs.
Smart Images

Figure CN120296389A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, apparatus, medium and product for extracting text data. Background Art
[0002] In the prior art, the target indicators and the data corresponding to the target indicators in the text are usually extracted based on a software program. If the target indicators to be extracted change, the software program needs to be modified, and the modification of the software program depends on professional software developers.
[0003] Therefore, it is necessary to provide a method for extracting text data that can extract the index data of any index or any combination of indexes, so as to improve the flexibility of text data extraction. Summary of the Invention
[0004] The present invention provides a method, apparatus, medium and product for extracting text data, so as to improve the flexibility of text data extraction.
[0005] According to one aspect of the present invention, there is provided a method for extracting text data, including:
[0006] Obtaining, based on an interaction interface, a target index combination and diagnostic description information of a target object, where the target index combination includes at least one target index;
[0007] Determining prompt information associated with each of the at least one target index, where the prompt information is description information for extracting the corresponding target index and the index data of the corresponding target index;
[0008] Inputting the diagnostic description information and the prompt information associated with each target index into a pre-trained language model to obtain a first data extraction result, where the first data extraction result includes the at least one target index and the index data of each of the at least one target index;
[0009] Displaying the first data extraction result in a visualization interface.
[0010] According to another aspect of the present invention, there is provided a text data extraction apparatus, including:
[0011] An obtaining module, configured to obtain a target index combination and diagnostic description information of a target object based on an interaction interface, where the target index combination includes at least one target index;
[0012] A prompt information module, configured to determine prompt information associated with each of the at least one target index, where the prompt information is description information for extracting the corresponding target index and the index data of the corresponding target index;
[0013] An extraction module, configured to input the diagnosis and treatment description information and the prompt information associated with each of the target indicators into a pre-trained language model, and obtain a first data extraction result, where the first data extraction result includes the at least one target indicator and the indicator data of each of the target indicators in the at least one target indicator;
[0014] A display module, configured to display the first data extraction result in a visualization interface.
[0015] According to another aspect of the present invention, there is provided an electronic device, where the electronic device includes:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the text data extraction method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the text data extraction method according to any embodiment of the present invention when executed.
[0020] According to another aspect of the present invention, there is provided a computer program product including a computer program, where the computer program implements the text data extraction method according to any embodiment when executed by a processor.
[0021] The technical solution of the text data extraction method provided by the embodiments of the present invention automatically determines the prompt information associated with each target indicator in the target indicator combination, and inputs the diagnosis and treatment description information and the prompt information associated with each target indicator into the pre-trained language model to obtain the first data extraction result, so that the user only needs to determine the target indicator combination and the diagnosis and treatment description information of the target object, and there are no requirements for the quantity and content of the target indicators included in the target indicator combination. Therefore, the flexibility of extracting the target indicators and the indicator data of the target indicators from the text data can be improved.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0024] Figure 1 is a flowchart of a text data extraction method provided according to an embodiment of the present invention;
[0025] Figure 2 is a display diagram of an interactive interface provided according to an embodiment of the present invention;
[0026] Figure 3 is another flowchart of a text data extraction method provided according to an embodiment of the present invention;
[0027] Figure 4 is a schematic structural diagram of a text data extraction device provided according to an embodiment of the present invention;
[0028] Figure 5 is another schematic structural diagram of a text data extraction device provided according to an embodiment of the present invention;
[0029] Figure 6 is a schematic structural diagram of an electronic device for implementing the text data extraction method of the embodiment of the present invention. Detailed implementation manners
[0030] In order to enable those skilled in the art to better understand the solution of the present invention, the following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] It should be noted that in the technical solution of the present invention, the processing of collection, use, storage, sharing, transfer, etc. of the user's personal information complies with the provisions of relevant laws and regulations, and it is necessary to inform the user and obtain the consent or authorization of the user. When applicable, technical processing such as de-identification and / or anonymization and / or encryption is performed on the user's personal information.
[0033] Figure 1 The present invention provides a flowchart of a method for extracting text data. This embodiment is applicable to the situation of extracting index data of each index in any target index combination from the diagnosis and treatment description information based on a pre-trained language model. This method can be executed by a text data extraction device, which can be implemented in the form of hardware and / or software, and the text data extraction device can be configured in the processor of an electronic device. As Figure 1 shown, the method includes:
[0034] S110. Obtain a target index combination and the diagnosis and treatment description information of a target object based on an interaction interface, where the target index combination includes at least one target index.
[0035] The interaction interface is a channel for information exchange between humans and computers. Users input information and perform operations on the computer through the interaction interface, and the computer provides information to users through the interaction interface for reading, analysis, and judgment.
[0036] The target index is the index for which index data needs to be extracted.
[0037] The diagnosis and treatment description information refers to the text information used to describe the diagnosis and treatment situation of the target object, such as the text information used to describe the outpatient diagnosis and treatment situation of the target object, or the text information used to describe the surgical situation of the target object.
[0038] In one embodiment, in response to a data extraction instruction, an interaction interface is displayed. The interaction interface includes a diagnosis and treatment description area and an index selection area including a plurality of candidate indexes; in response to an index selection operation in the index selection area and the diagnosis and treatment description information of the target object received in the diagnosis and treatment description area, the target index combination and the diagnosis and treatment description information of the target object are determined.
[0039] Specifically, the interaction interface includes a diagnosis and treatment description area and an index selection area. The diagnosis and treatment description area is used to receive the diagnosis and treatment description information about the target object. For example, add "This is a surgical record of patient A" and the actual surgical record to this diagnosis and treatment description area. The processor determines the diagnosis and treatment description information of the target object according to the information received in the diagnosis and treatment description area.
[0040] The index selection area includes multiple candidate indices. The user selects the desired target index by touching or clicking according to actual needs. The processor uses one or more candidate indices selected by the user as the target index combination. The index selection area allows the user to directly select the desired index without having to learn how to determine the prompt information for extracting the target index and the target index data, which helps improve the user experience.
[0041] In one embodiment, the interaction interface includes a diagnosis and treatment description area and an index input area. The diagnosis and treatment description area is used to receive the user's diagnosis and treatment description information about the target object. The index input area is used to receive the indices input by the user in a set format. For example, 'age' and 'weight'. The processor determines the target index combination including one or more indices according to the indices received by the index input area. The index input area enables the user to input any number of indices that conform to the predetermined format, with simple operation and good user experience.
[0042] S120. Determine the prompt information associated with each target index in the at least one target index. The prompt information is the description information for extracting the corresponding target index and the index data of the corresponding target index.
[0043] After the target index combination is determined, determine the prompt information for extracting each target index in the target index combination. For example, extract the target index A and the index data of the target index A from the diagnosis and treatment description information. This step aims to reduce the difficulty for the user to extract text data using the pre-trained language model by automatically generating the prompt information associated with each target index, and improve the user experience.
[0044] S130. Input the diagnosis and treatment description information and the prompt information associated with each target index into the pre-trained language model to obtain the first data extraction result. The first data extraction result includes the at least one target index and the index data of each target index in the at least one target index.
[0045] After the diagnosis and treatment description information and the prompt information associated with each target index are determined, input the diagnosis and treatment description information and the prompt information associated with the target index into the pre-trained language model to obtain the first data extraction result output by the model. This step uses the powerful semantic analysis ability of the pre-trained language model, the relatively accurate diagnosis and treatment description information, and the prompt information associated with each target index to extract each target index and the index data of each target index from the diagnosis and treatment description information, ensuring that the first data extraction result has relatively high accuracy.
[0046] In one embodiment, the diagnosis and treatment description information, and the prompt information associated with each target indicator are input into a default pre-trained language model to obtain a first data extraction result. This embodiment is applicable to the situation where only one pre-trained language model is configured, or when the user has no special requirements for the language model.
[0047] In one embodiment, the interaction interface is provided with a language model option. The user can select the required language model based on the language model option; the processor determines the corresponding pre-trained language model in response to the switching operation of the language model, and then inputs the diagnosis and treatment description information of the target object, and the prompt information associated with each target indicator, into the pre-trained language model to obtain a first data extraction result. This embodiment meets the needs of different users for different language models by providing multiple language models.
[0048] In one embodiment, the diagnosis and treatment description information, and the prompt information associated with each target indicator are input into at least three pre-trained language models to obtain at least three first data extraction results; according to the comparison result of the at least three first data extraction results, the first data extraction result with the highest accuracy among the at least three first data set extraction results is determined, and this first data extraction result is used as the final first data extraction result. This embodiment can ensure that the final first data extraction result has high accuracy.
[0049] S140. Display the first data extraction result in the visualization interface.
[0050] After the first data extraction result is determined, display the first data extraction result in the visualization interface so that the user can intuitively view the at least one target indicator and the indicator data of each target indicator.
[0051] In one embodiment, the visualization interface is the aforementioned interaction interface. Specifically, the interaction interface includes a result display area, and the first data extraction result is displayed in the result display area.
[0052] The technical solution of the text data extraction method provided by the embodiments of the present invention automatically determines the prompt information associated with each target indicator in the target indicator combination, and inputs the diagnosis and treatment description information and the prompt information associated with each target indicator into the pre-trained language model to obtain the first data extraction result, so that the user only needs to determine the target indicator combination and the diagnosis and treatment description information of the target object, and there are no quantity and content requirements for the target indicators included in the target indicator combination. Therefore, the flexibility of extracting target indicators and the indicator data of target indicators from text data can be improved.
[0053] Figure 2The flowchart of the text data extraction method provided by the embodiment of the present invention. In this embodiment, the step of displaying the first data extraction result is refined on the basis of the foregoing embodiment. As Figure 2 shown, the method includes:
[0054] S210. Obtain the target index combination and the diagnosis and treatment description information of the target object based on the interaction interface, where the target index combination includes at least one target index.
[0055] S220. Determine the prompt information associated with each target index in the at least one target index, where the prompt information is the description information used to extract the corresponding target index and the index data of the corresponding target index.
[0056] S230. Input the diagnosis and treatment description information and the prompt information associated with each target index into a pre-trained language model to obtain a first data extraction result, where the first data extraction result includes the at least one target index and the index data of each target index in the at least one target index.
[0057] S2401. When the interaction interface includes a code result display area, display the first data extraction result in the form of structured code in the code result display area.
[0058] The interaction interface of this embodiment further includes at least one of a code result display area and a table result display area. If the interaction interface includes a code display area, the first data extraction result is displayed in the form of structured code in the code display area. The first data extraction result displayed in the form of structured code can be exported, which helps to simplify the subsequent data analysis process.
[0059] S2402. When the interaction interface includes a table result display area, display the first data extraction result in the form of a table in the table result display area.
[0060] Figure 3 The schematic diagram of the interaction interface provided by the embodiment of the present invention. The interaction interface includes a diagnosis and treatment description area in the upper left corner, an index selection area in the upper right corner, a table result display area in the lower left corner, and a code result display area in the lower right corner. The interaction interface includes both a table result display area and a code result display area, which can meet different usage requirements of users for the first data extraction result.
[0061] In one embodiment, when the interaction interface includes a table result display area, display the first data extraction result in the form of a read-only table in the table result display area. The first data extraction result in this embodiment can only be read but cannot be edited, and is used to provide the user with the original first data extraction result.
[0062] In one embodiment, when the interaction interface includes a table result display area, the first data extraction result is displayed in the table result display area in the form of an editable table. The user can edit the target indicator or indicator data in any target cell in the editable table based on the diagnosis and treatment description information in the diagnosis and treatment description area; the processor responds to the modification operation on any target cell in the table, completes the modification of the corresponding target indicator or corresponding indicator data in the editable table, and obtains a second data extraction result, where the target cell is a cell including the target indicator or the indicator data of the target indicator. This second data extraction result is the data extraction result after the user's correction. By presenting the first data extraction result in the form of an editable table, the user is provided with a modifiable first data extraction result, enabling the user to modify the required target indicator or target indicator data as needed, ensuring that the user can obtain the original first data result and a second data extraction result with higher accuracy after self-modification.
[0063] In one embodiment, determine the combination of indicator data for which the second data extraction result has changed compared to the first data extraction result; determine the accuracy score of the first data extraction result based on the number of indicator data included in the combination of indicator data and the number of indicators included in the first data extraction result; and display the accuracy score in the interaction interface.
[0064] The combination of indicator data that has changed is the combination of indicator data modified by the user, which may include zero, one, or more indicator data. Determine the ratio of the number of all indicator data modified by the user to the number of indicator data included in the first data extraction result, and determine the accuracy score of the first data extraction result based on this ratio. It can be understood that the smaller this ratio, the smaller the number of modified indicator data and the larger the number of unmodified indicator data, that is, the larger the number of accurate indicator data, so the higher the accuracy score of the first data extraction result; the larger this ratio, the larger the number of modified indicator data and the smaller the number of unmodified indicator data, that is, the smaller the number of accurate indicator data, so the lower the accuracy score of the first data extraction result. This accuracy score is preferably set in the table result display area to facilitate the user to view the first data extraction result and / or the second data extraction result, as well as this accuracy score at the same time.
[0065] The technical solution provided by the embodiments of the present invention meets various usage requirements of users for the first data extraction result through different display methods of the first data extraction result, and the user experience is better.
[0066] Figure 4 It is a schematic structural diagram of the text data extraction device provided by the embodiments of the present invention. As Figure 4 shown, the device includes:
[0067] An acquisition module 31, configured to acquire a target index combination and diagnostic and treatment description information of a target object based on an interaction interface, where the target index combination includes at least one target index;
[0068] A prompt information module 32, configured to determine prompt information associated with each of the at least one target index, where the prompt information is description information for extracting a corresponding target index and index data of the corresponding target index;
[0069] An extraction module 33, configured to input the diagnostic and treatment description information and the prompt information associated with each target index into a pre-trained language model to obtain a first data extraction result, where the first data extraction result includes the at least one target index and index data of each of the at least one target index;
[0070] A display module 34, configured to display the first data extraction result in a visualization interface.
[0071] In one embodiment, the interaction interface includes a diagnostic and treatment description area and an index selection area including a plurality of candidate indexes, and the acquisition module 31 is specifically configured to:
[0072] In response to an index selection operation in the index selection area and the diagnostic and treatment description information of the target object received in the diagnostic and treatment description area, determine the target index combination and the diagnostic and treatment description information of the target object.
[0073] In one embodiment, the interaction interface further includes at least one of a code result display area and a table result display area, and the display module 34 is specifically configured to:
[0074] A code display unit, configured to, when the interaction interface includes a code result display area, display the first data extraction result in the form of structured code in the code result display area;
[0075] A table display unit, configured to, when the interaction interface includes a table result display area, display the first data extraction result in the form of a table in the table result display area.
[0076] In one embodiment, when the interaction interface includes a table result display area, the table displayed in the table result display area is an editable table, and the table display unit is further configured to:
[0077] In response to a modification operation on any target cell in the table, complete the modification of the corresponding target index or the corresponding index data in the editable table to obtain a second data extraction result, where the target cell is a cell including the target index or the index data of the target index.
[0078] In one embodiment, the table display unit is further configured to:
[0079] Determine an index data combination in which the second data extraction result changes compared to the first data extraction result;
[0080] Determine the accuracy score of the first data extraction result according to the number of index data included in the index data combination and the number of indexes included in the first data extraction result;
[0081] Display the accuracy score in the interaction interface.
[0082] In one embodiment, the interaction interface includes a language model option. As Figure 5 shown, the device further includes a model selection module 35, and the model selection module 35 is configured to:
[0083] In response to a language model switching operation, determine the corresponding pre-trained language model.
[0084] The technical solution of the text data extraction method provided by the embodiments of the present invention automatically determines the hint information associated with each target index in the target index combination, and inputs the diagnosis and treatment description information and the hint information associated with each target index into the pre-trained language model, so as to obtain the first data extraction result, enabling the user to only need to determine the target index combination and the diagnosis and treatment description information of the target object, and there are no quantity and content requirements for the target indexes included in the target index combination. Therefore, the flexibility of extracting target indexes and index data of target indexes from text data can be improved.
[0085] The text data extraction device provided by the embodiments of the present invention can execute the text data extraction method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0086] Figure 6 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0087] As Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0088] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0089] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the text data extraction method.
[0090] In some embodiments, the text data extraction method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the text data extraction method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the text data extraction method by any other appropriate means (e.g., by means of firmware).
[0091] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0092] The computer program for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0093] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0095] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0096] The computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of high management difficulty and weak business scalability existing in traditional physical hosts and VPS services.
[0097] An embodiment of the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the text data extraction method provided in any embodiment of the present application.
[0098] In the process of implementing the computer program product, computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0099] It should be understood that the various forms of the process shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0100] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for extracting text data, characterized in that, including: obtaining, based on an interaction interface, a target indicator combination and medical treatment description information of a target object, where the target indicator combination includes at least one target indicator; determining prompt information associated with each of the at least one target indicator, where the prompt information is description information for extracting the corresponding target indicator and the indicator data of the corresponding target indicator; inputting the medical treatment description information and the prompt information associated with each target indicator into a pre-trained language model to obtain a first data extraction result, where the first data extraction result includes the at least one target indicator and the indicator data of each of the at least one target indicator; displaying the first data extraction result in a visualization interface.
2. The method according to claim 1, characterized in that, The interaction interface includes a medical treatment description area and an indicator selection area including a plurality of candidate indicators. The obtaining, based on the interaction interface, of the target indicator combination and the medical treatment description information of the target object includes: responding to an indicator selection operation in the indicator selection area and the medical treatment description information of the target object received in the medical treatment description area to determine the target indicator combination and the medical treatment description information of the target object.
3. The method according to claim 1, wherein The interaction interface further includes at least one of a code result display area and a table result display area. After obtaining the first data extraction result, it further includes: in the case where the interaction interface includes a code result display area, displaying the first data extraction result in the form of structured code in the code result display area; in the case where the interaction interface includes a table result display area, displaying the first data extraction result in the form of a table in the table result display area.
4. The method according to claim 3, wherein In the case where the interaction interface includes a table result display area, the table displayed in the table result display area is an editable table. After displaying the first data extraction result in the form of a table in the table result display area, it further includes: responding to a modification operation on any target cell in the table to complete the modification of the corresponding target indicator or the corresponding indicator data in the editable table to obtain a second data extraction result, where the target cell is a cell including the target indicator or the indicator data of the target indicator.
5. The method according to claim 4, wherein After obtaining the second data extraction result, it further includes: determining an indicator data combination in which the second data extraction result changes compared with the first data extraction result; determining an accuracy score of the first data extraction result according to the number of indicator data included in the indicator data combination and the number of indicators included in the first data extraction result; displaying the accuracy score in the interaction interface.
6. The method according to claim 1, characterized in that, The interaction interface includes a language model option. Before inputting the medical treatment description information and the prompt information associated with each target indicator into a pre-trained language model to obtain a first data extraction result, it further includes: responding to a language model switching operation to determine the corresponding pre-trained language model.
7. A text data extraction device, characterized in that, including: an obtaining module for obtaining, based on an interaction interface, a target indicator combination and medical treatment description information of a target object, where the target indicator combination includes at least one target indicator; A prompt information module, configured to determine prompt information associated with each of the at least one target metric, where the prompt information is descriptive information for extracting the corresponding target metric and the metric data of the corresponding target metric; An extraction module, configured to input the diagnosis and treatment description information and the prompt information associated with each of the target metrics into a pre-trained language model to obtain a first data extraction result, where the first data extraction result includes the at least one target metric and the metric data of each of the at least one target metric; A display module, configured to display the first data extraction result in a visualization interface.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the text data extraction method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the text data extraction method according to any one of claims 1-6 is implemented.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, the text data extraction method according to any one of claims 1-6 is implemented.