Text extraction method and device, medium and product

By using the target text extraction model based on the logistics industry and regional corpus in the logistics industry, the problem of lack of domain knowledge and generalization capabilities in general technology is solved, and more efficient and accurate text key information extraction is achieved.

CN120146042APending Publication Date: 2025-06-13CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510230896.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

General technology lacks domain knowledge and generalization capabilities in the logistics industry, resulting in low accuracy in extracting text key information.

Method used

The target text extraction model trained by the logistics industry-based training text data set and regional corpus data set is used to extract the required target text from the pending text.

Benefits of technology

It improves the accuracy and efficiency of text extraction, and ensures that the model has domain knowledge and good generalization capabilities in the logistics industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146042A_ABST
    Figure CN120146042A_ABST
Patent Text Reader

Abstract

The invention provides a text extraction method and device, a medium and a product, relates to the technical field of information extraction, and aims at solving the problem that the accuracy of key information extraction is relatively low due to the fact that a general technology lacks logistics industry field knowledge and is relatively poor in generalization ability. The text extraction method comprises the steps of obtaining a to-be-processed text and a text extraction task cue word; wherein the to-be-processed text comprises the text content of the logistics industry and the regional feature expression text of the region to which the to-be-processed text belongs. And based on the text extraction task cue word and the target text extraction model, extracting a target text required by the text extraction task from the to-be-processed text. The target text extraction model is obtained by training based on a training text data set and a regional corpus data set of the logistics industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information extraction technology, and in particular, to a text extraction method, device, medium, and product. Background Art

[0002] In the logistics industry, to promote the healthy development of the industry, institutions in different regions will release various relevant text information. However, in these text information, there are different expression names for the same text content. This affects the recognition and extraction of text information, and further affects the effective utilization of text information.

[0003] In general technology, mainly based on the named entity recognition technology in the field of natural language understanding, key information in the text is extracted. On the one hand, through training with a large amount of labeled data, key information in the text is recognized based on traditional deep learning algorithms. On the other hand, with the rapid development of artificial intelligence technology, large models can be used for natural language processing to recognize key information in text information.

[0004] However, both traditional deep learning algorithms and large models lack knowledge in the logistics industry field and have relatively poor generalization ability. Therefore, the accuracy of extracting key information by general technology is relatively low. Summary of the Invention

[0005] Embodiments of the present disclosure provide a text extraction method, device, medium, and product, aiming to solve the technical problem that general technology lacks knowledge in the logistics industry field and has relatively poor generalization ability, resulting in relatively low accuracy of extracting key information.

[0006] To achieve the above object, this application adopts the following technical solutions:

[0007] In the first aspect, a text extraction method is provided, including: obtaining a text to be processed and a text extraction task prompt; the text to be processed includes text content in the logistics industry and text expressing regional characteristics of the region where the text to be processed belongs; based on the text extraction task prompt and a target text extraction model, extracting target text required for the text extraction task from the text to be processed; the target text extraction model is trained based on a training text data set in the logistics industry and a regional corpus data set.

[0008] Optionally, extracting target text required for the text extraction task from the text to be processed based on the text extraction task prompt and the target text extraction model includes: determining the text extraction task prompt as the text prefix of the text to be processed, and inputting the text to be processed with the added text prefix into the target text extraction model to obtain the target text.

[0009] Optionally, the text extraction method further includes: obtaining an initial text extraction model; training the initial text extraction model based on a training text data set to obtain a text extraction model for the logistics industry; the training text data set includes text data representing the text content of the logistics industry; training the text extraction model for the logistics industry based on a regional corpus data set to obtain a target text extraction model; the regional corpus data set includes corpus data representing the language expression forms of the same text in different regions.

[0010] Optionally, the text extraction method further includes: obtaining an initial training text and preprocessing the initial training text to obtain a target training text; the preprocessing includes at least one of removing noise data and unifying the text format; extracting key text from the target training text based on a preset structured rule template; the key text includes professional vocabulary and high-frequency vocabulary in the target training text; the high-frequency vocabulary is a vocabulary with an occurrence frequency greater than a preset frequency; performing annotation processing on the key text based on an annotation rule to obtain a preliminary annotation result of the key text; correcting the preliminary annotation result based on a pre-constructed annotation database to obtain a target annotation result of the key text; constructing a training text data set based on the initial training text, the target training text, the key text, and the target annotation result of the key text.

[0011] Optionally, the text extraction method further includes: updating the annotation rule according to the preliminary annotation result and the target annotation result.

[0012] Optionally, the text extraction method further includes: obtaining the region to which the target training text belongs; extracting the language features of the target training text; constructing a regional corpus data set according to the correspondence between the region and the language features.

[0013] In a second aspect, a text extraction device is provided, and the text extraction device includes: a communication unit and a processing unit; the communication unit is configured to obtain a text to be processed and a text extraction task prompt word; the text to be processed includes the text content of the logistics industry and the regional feature expression text of the region to which the text to be processed belongs; the processing unit is configured to extract the target text required for the text extraction task from the text to be processed based on the text extraction task prompt word and the target text extraction model; the target text extraction model is trained based on a training text data set of the logistics industry and a regional corpus data set.

[0014] Optionally, the processing unit is specifically configured to: determine the text extraction task prompt word as the text prefix of the text to be processed, and input the text to be processed with the added text prefix into the target text extraction model to obtain the target text.

[0015] Optionally, the communication unit is further configured to obtain an initial text extraction model; the processing unit is further configured to train the initial text extraction model based on a training text dataset to obtain a logistics industry text extraction model; the training text dataset includes text data representing text content in the logistics industry; the processing unit is further configured to train the logistics industry text extraction model based on a regional corpus dataset to obtain a target text extraction model; the regional corpus dataset includes corpus data representing language expression forms of the same text in different regions.

[0016] Optionally, the communication unit is further configured to obtain an initial training text and preprocess the initial training text to obtain a target training text; the preprocessing includes at least one of removing noise data and unifying the text format; the processing unit is further configured to extract key text from the target training text based on a preset structured rule template; the key text includes professional vocabulary and high-frequency vocabulary in the target training text; the high-frequency vocabulary is a vocabulary with an appearance frequency greater than a preset frequency; the processing unit is further configured to perform annotation processing on the key text based on an annotation rule to obtain a preliminary annotation result of the key text; the processing unit is further configured to correct the preliminary annotation result based on a pre-constructed annotation database to obtain a target annotation result of the key text; the processing unit is further configured to construct a training text dataset based on the initial training text, the target training text, the key text, and the target annotation result of the key text.

[0017] Optionally, the processing unit is further configured to update the annotation rule according to the preliminary annotation result and the target annotation result.

[0018] Optionally, the communication unit is further configured to obtain the region to which the target training text belongs; the processing unit is further configured to extract the language features of the target training text; the processing unit is further configured to construct a regional corpus dataset according to the correspondence between the region and the language features.

[0019] In a third aspect, a text extraction device is provided, including a memory and a processor; the memory is used to store computer execution instructions, and the processor is connected to the memory through a bus; when the text extraction device runs, the processor executes the computer execution instructions stored in the memory to enable the text extraction device to execute the text extraction method in the first aspect.

[0020] The text extraction device may be an electronic device or a part of a device in the electronic device, such as a chip system in the electronic device. The chip system is used to support the electronic device to implement the functions involved in the first aspect and any of its possible implementation manners, for example, to obtain and determine the data and / or information involved in the above text extraction method. The chip system includes a chip and may also include other discrete devices or circuit structures.

[0021] Fourthly, a computer-readable storage medium is provided. The computer-readable storage medium includes computer-executable instructions. When the computer-executable instructions run on a computer, the computer is caused to execute the text extraction method described in the first aspect.

[0022] Fifthly, a computer program product is further provided. The computer program product includes a computer program or instructions. When the computer instructions run on a text extraction device, the text extraction device is caused to execute the text extraction method described in the first aspect above.

[0023] It should be noted that the above computer instructions can be stored in whole or in part on a computer-readable storage medium. Among them, the computer-readable storage medium can be packaged together with the processor of the text extraction device, or can be separately packaged from the processor of the text extraction device. The embodiments of the present application do not limit this.

[0024] For the descriptions of the second aspect, the third aspect, the fourth aspect, and the fifth aspect in the present application, reference can be made to the detailed description of the first aspect.

[0025] In the embodiments of the present application, the name of the above text extraction device does not limit the device or functional module itself. In actual implementation, these devices or functional modules may appear under other names. For example, the receiving unit can also be called a receiving module, a receiver, etc. As long as the functions of each device or functional module are similar to those of the present application and fall within the scope of the claims of the present application and their equivalent technologies.

[0026] The technical solutions provided by the present application at least bring the following beneficial effects:

[0027] Based on any of the above aspects, embodiments of the present application provide a text extraction method. First, a text to be processed and a text extraction task prompt word are obtained. Among them, the text to be processed includes text content in the logistics industry and text expressing regional characteristics of the region to which the text to be processed belongs. Then, based on the text extraction task prompt word and a target text extraction model, the target text required for the text extraction task is extracted from the text to be processed. Among them, the target text extraction model is trained based on a training text data set in the logistics industry and a regional corpus data set.

[0028] As can be seen from the above, the present application uses a target text extraction model trained based on a training text data set in the logistics industry and a regional corpus data set to extract the target text required for the text extraction task from the text to be processed. Since the target text extraction model is trained based on a training text data set in the logistics industry and a regional corpus data set, the target text extraction model has domain knowledge in the logistics industry and has good generalization ability in the logistics industry. In this way, based on the text extraction task prompt word and the target text extraction model, the present application can quickly and accurately extract the target text.

[0029] The beneficial effects of the first aspect, the second aspect, the third aspect, the fourth aspect, and the fifth aspect in this application can all refer to the analysis of the above beneficial effects, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic structural diagram of a text extraction system provided by an embodiment of this application;

[0031] Figure 2 It is a schematic hardware structure diagram of a communication device provided by an embodiment of this application;

[0032] Figure 3 It is a schematic flowchart of a text extraction method provided by an embodiment of this application;

[0033] Figure 4 It is a schematic flowchart of another text extraction method provided by an embodiment of this application;

[0034] Figure 5 It is a schematic flowchart of another text extraction method provided by an embodiment of this application;

[0035] Figure 6 It is a schematic flowchart of text extraction model training provided by an embodiment of this application;

[0036] Figure 7 It is a schematic flowchart of another text extraction method provided by an embodiment of this application;

[0037] Figure 8 It is a schematic flowchart of another text extraction method provided by an embodiment of this application;

[0038] Figure 9 It is a schematic structural diagram of a text extraction device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0040] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0041] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the words such as "first" and "second" do not limit the quantity and execution order.

[0042] Before introducing the text extraction method provided by the present application in detail, the application scenarios and implementation environments involved in the present application will be briefly introduced.

[0043] First, the application scenarios involved in the present application will be briefly introduced.

[0044] As described in the background art, in the logistics industry, in order to promote the healthy development of the industry, institutions in different regions will release various relevant text information. Identify and extract the key information in the text information, and establish an information database for the logistics industry to achieve the normalization processing of the text information. In this way, information sharing and cooperation can be realized, enabling each unit to obtain the required information and promoting the healthy development of the logistics industry. However, in these text information, there are different expression names for the same text content (for example, there are abbreviations, aliases for the document-issuing institutions in the text, and even different expressions for the same document-issuing institution in different regions). This affects the identification and extraction of text information and further affects the effective utilization of text information.

[0045] In the general technology, the named entity recognition technology in the field of natural language understanding is mainly used to extract the key information in the text. On the one hand, through training with a large amount of labeled data, the key information in the text is recognized based on traditional deep learning algorithms (for example, the Bidirectional Encoder Representations from Transformers (BERT) algorithm). On the other hand, with the rapid development of artificial intelligence technology, large models can be used for natural language processing to identify the key information in the text information.

[0046] However, traditional deep learning algorithms rely extremely on a large amount of labeled data during the training process, and both the quality and quantity of the labeled data will affect the accuracy of the deep learning algorithm in extracting key information. During the training process, it is necessary to repeatedly optimize the model and manually label data for a long time. Moreover, when facing Named Entity Recognition (NER) problems (such as polysemy and synonymy), it is impossible to accurately extract key information. Therefore, traditional deep learning algorithms are time-consuming, costly, and have low accuracy in extracting key information.

[0047] Although large models have powerful language expression capabilities, there are significant differences in language characteristics, terms, and named entities in different industries or fields. Therefore, when a large model processes NER problems in the logistics industry, it is impossible to quickly and accurately identify key information in text information.

[0048] As can be seen from the above, general technologies lack knowledge in the logistics industry field and have relatively poor generalization ability, resulting in low accuracy in extracting key information.

[0049] In view of the above problems, the embodiments of the present application provide a text extraction method. First, obtain the text to be processed and the text extraction task prompt word. Among them, the text to be processed includes the text content of the logistics industry and the text expressing regional characteristics of the region to which the text to be processed belongs. Then, based on the text extraction task prompt word and the target text extraction model, extract the target text required for the text extraction task from the text to be processed. The target text extraction model is trained based on the training text data set of the logistics industry and the regional corpus data set.

[0050] As can be seen from the above, the present application uses the target text extraction model trained based on the training text data set of the logistics industry and the regional corpus data set to extract the target text required for the text extraction task from the text to be processed. Since the target text extraction model is trained based on the training text data set of the logistics industry and the regional corpus data set, the target text extraction model has knowledge in the logistics industry field and has good generalization ability in the logistics industry. In this way, based on the text extraction task prompt word and the target text extraction model, the present application can quickly and accurately extract the target text.

[0051] The implementation environment of the above text extraction method can be the text extraction system provided by the embodiments of the present application.

[0052] Figure 1 Shows the structural schematic diagram of the text extraction system provided by the embodiments of the present application. As Figure 1 shown, the text extraction system includes: a text extraction device 101 and a data storage device 102.

[0053] Among them, the text extraction device 101 and the data storage device 102 are communicatively connected.

[0054] In practical applications, the text extraction device 101 can be connected to any number of data storage devices 102. For ease of understanding, Figure 1 a case where one text extraction device 101 is connected to one data storage device 102 is taken as an example for illustration.

[0055] In the embodiments of the present application, the data storage device 102 is used to provide data for text extraction (such as the text to be processed, text extraction task prompt words, initial training text, initial text extraction model, etc.) to the text extraction device 101, so that the text extraction device 101 can implement text extraction according to the data sent by the data storage device 102.

[0056] Optionally, the physical devices of the text extraction device 101 and the data storage device 102 can be servers, terminals, or other types of electronic devices, and the embodiments of the present application do not limit this.

[0057] Optionally, the above terminal can be a device that provides voice and / or data connectivity to users, a handheld device with a wireless connection function, or other processing devices connected to a wireless modem. The wireless terminal can communicate with one or more core networks via a radio access network (RAN). The wireless terminal can be a mobile terminal, such as a mobile phone (or a "cellular" phone) and a computer with a mobile terminal, or a portable, pocket-sized, handheld, computer-integrated, or vehicle-mounted mobile device that exchanges language and / or data with the wireless access network. For example, mobile phones, tablets, laptop computers, netbooks, personal digital assistants (PDAs).

[0058] Optionally, the above server can be a server in a server cluster (composed of multiple servers), a chip in the server, a system-on-chip in the server, or can also be implemented by a virtual machine (VM) deployed on a physical machine. The embodiments of the present application do not limit this.

[0059] Optionally, the text extraction device 101 and the data storage device 102 can be two independently set devices, or can be integrated in the same device. When the text extraction device 101 and the data storage device 102 are integrated in the same device, the data storage device 102 can be a storage module (such as a database, etc.) of the text extraction device 101.

[0060] It is easy to understand that when the text extraction device 101 and the data storage device 102 are integrated in the same device, the communication method between the text extraction device 101 and the data storage device 102 is the communication between internal modules of the device. In this case, the communication process between the two is the same as that between the text extraction device 101 and the data storage device 102 when they are independent of each other.

[0061] For the sake of easy understanding, this application takes the text extraction device 101 and the data storage device 102 being independent of each other as an example for illustration.

[0062] The text extraction device in the text extraction system includes components such as Figure 2 those included. Taking the Figure 2 shown communication device as an example, the hardware structure of the text extraction device will be introduced.

[0063] Figure 2 As shown, it is a schematic diagram of a hardware structure of the communication device provided by an embodiment of this application. The communication device includes a processor 21, a memory 22, a communication interface 23, and a bus 24. The processor 21, the memory 22, and the communication interface 23 can be connected through the bus 24.

[0064] The processor 21 is the control center of the communication device, which can be a single processor or a collective term for multiple processing elements. For example, the processor 21 can be a general-purpose central processing unit (CPU), or other general-purpose processors, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0065] As an embodiment, the processor 21 can include one or more CPUs, such as Figure 2 the CPU0 and CPU1 shown in

[0066] The memory 22 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0067] In one possible implementation, the memory 22 can exist independently of the processor 21. The memory 22 can be connected to the processor 21 via the bus 24 and is used to store instructions or program codes. When the processor 21 invokes and executes the instructions or program codes stored in the memory 22, the text extraction method provided in the following embodiments of the present application can be implemented.

[0068] In the embodiments of the present application, for a communication device, different software programs are stored in the memory 22, so the functions implemented by the communication device are different. The functions performed by each device will be described in conjunction with the following flowcharts.

[0069] In another possible implementation, the memory 22 can also be integrated with the processor 21.

[0070] The communication interface 23 is used for the communication device to connect to other devices via a communication network. The communication network can be an Ethernet, a radio access network, a wireless local area network (WLAN), etc. The communication interface 23 can include a receiving unit for receiving data and a transmitting unit for transmitting data.

[0071] The bus 24 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 2 only a thick line is shown herein, but it does not mean that there is only one bus or one type of bus.

[0072] It should be noted that Figure 2 the structure shown in Figure 2 does not constitute a limitation on the communication device. Except for

[0073] the components shown, the communication device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0074] The text extraction method provided in the embodiments of the present application is applied to Figure 1 the text extraction device 101 in the text extraction system shown in Figure 3 As shown in

[0075] S301. Obtain the text to be processed and the text extraction task prompt.

[0076] The text to be processed includes the text content of the logistics industry and the text expressing the regional characteristics of the region to which the text to be processed belongs.

[0077] In the embodiments of the present application, the text extraction device can obtain the text to be processed and the text extraction task prompt. Among them, the text extraction task prompt is used to represent the field name of the target text extracted from the text to be processed. It can be understood that the text to be processed usually includes multiple fields. Therefore, the text extraction device can obtain the text to be processed and the text extraction task prompt. In this way, the text extraction device can extract the target text required for the text extraction task from the text to be processed according to the text extraction task prompt.

[0078] The text to be processed includes: the issuing agency, the document number, the issuing date, the article title, and the article body.

[0079] The text extraction tasks include: extracting the issuing agency, extracting the document number, extracting the issuing date, and summarizing the article summary.

[0080] The text extraction task prompt (prompt) is used to guide the target text extraction model to generate text that conforms to the text extraction task.

[0081] Optionally, the text extraction task prompt includes: the format of the target text (for example, a dictionary in JavaScript Object Notation (JSON) format).

[0082] The keys of the above dictionary are of string type, and the values of the dictionary may be of string type, numerical type, or array type.

[0083] In some embodiments, the value of the issuing agency is usually a string. In the case where the same text to be processed is jointly issued by multiple issuing agencies, the value of the issuing agency is an array of strings.

[0084] It can be understood that for the same text, the text expressing the characteristics of different regions is different.

[0085] S302. Based on the text extraction task prompt and the target text extraction model, extract the target text required for the text extraction task from the text to be processed.

[0086] The target text extraction model is trained based on the training text data set of the logistics industry and the regional corpus data set.

[0087] In the embodiments of the present application, the target text extraction model is trained based on a training text data set in the logistics industry and a regional corpus data set. Therefore, the target text extraction model has the ability to understand texts in the logistics industry and the ability to recognize texts with high accuracy. That is, when the target text extraction model faces texts with regional feature expressions, it can determine the text content expressed by them. Therefore, the text extraction device can extract the target text required for the text extraction task from the text to be processed based on the text extraction task prompt and the target text extraction model.

[0088] The training text data set in the logistics industry includes: text data representing the text content in the logistics industry.

[0089] The regional corpus data set includes: corpus data representing the language expression forms of the same text in different regions.

[0090] It can be understood that the model depends on the construction of the algorithm.

[0091] The algorithms adopted by the training target text extraction model in the embodiments of the present application include: key information extraction algorithms, traditional deep learning methods, or rule-based extraction methods.

[0092] In some embodiments, combined Figure 3 , such as Figure 4 shown, in S302, based on the text extraction task prompt and the target text extraction model, extracting the target text required for the text extraction task from the text to be processed specifically includes:

[0093] S401. Determine the text extraction task prompt as the text prefix of the text to be processed, and input the text to be processed with the added text prefix into the target text extraction model to obtain the target text.

[0094] In the embodiments of the present application, the text extraction device can determine the text extraction task prompt as the text prefix of the text to be processed, and input the text to be processed with the added text prefix into the target text extraction model. In this way, the text extraction device can extract the target text required for the text extraction task from the text to be processed.

[0095] In one implementable manner, the text extraction device inputs the text to be processed with the added text prefix into the target text extraction model, obtains a structured output result, and extracts structured information from the structured output result, that is, the target text.

[0096] The target text includes: the issuing agency of the text to be processed, the document number of the text to be processed, the issuing date of the text to be processed, and / or, the article summary of the text to be processed.

[0097] Exemplarily, the text content of the text to be processed is as follows: "The total logistics volume in Region A in the first three quarters increased by 3.6% year-on-year. Institution B, 2024-12-31-12:30. In the first three quarters of this year, the total logistics volume in Region A was 3 million yuan, a year-on-year increase of 3.6%. Looking at each quarter, the growth rate was above 5%. In the first three quarters, the total logistics volume in more than 90% of the industrial fields showed a year-on-year positive growth, and the total logistics volume of industrial products in more than 90% of the regions in Region A increased, driving an increase in the physical volume of industrial product logistics by more than 60%".

[0098] In the case where the text content of the text to be processed includes the above content, the text extraction device can determine the issuing institution: Institution B, the issuing date: December 31, 2024, and the article summary: The total logistics volume in Region A in the first three quarters increased by 3.6% year-on-year.

[0099] In some embodiments, if the text format of the text to be processed is different, the text extraction task prompt words are different, that is, the input content of the target text extraction model is different.

[0100] Optionally, the input content of the target text extraction model is the text extraction task prompt words + the text to be processed (the text to be processed is serialized raw data).

[0101] Exemplarily, when the text format of the text to be processed is in HyperText Markup Language (HTML) format, the input content of the target text extraction model is "Please extract the key text information in the logistics industry according to the following HTML content and output it in JSON format <p class = \"bgbg\" style = \"width: 100%;\"> <span style="font-weight:bold;">Document Title: Notice of the First Region, Institution A in the First Region on Issuing the Implementation Plan for the Development of Cold Chain Logistics in the First Region <p class="bgbgs"> <span style="font-weight:bold;">Document Number: No. 21 〔2022〕 <p class="bgbgs"><spanstyle="font-weight:bold;">Issuing Agency: First Region's Agency A <p class="bgbg"> <span style="font-weight:bold;">Date of Adoption: 2022-01-30 <p class="bgbg"> <span style="font-weight:bold;">Date of Issuance: 2022-01-30 <p class = \"bgbgs\"> <span style="font-weight:bold;">Subject Classification: Other <p class = \"bgbg\"> <span style="font-weight:bold;">Effectiveness Status: Effective ".

[0102] Correspondingly, the output content of the target text extraction model is ""File Title": "Notice of the First Region, Institution A in the First Region on Issuing the Implementation Plan for the Development of Cold Chain Logistics in the First Region", "Document Number": "〔2022〕21", "Issuing Date": "2022-01-30", "Subject Classification": "Other", "Effectiveness Status": "Effective"".

[0103] Another exemplary case is when the text format of the text to be processed is HTML format. The input content of the target text extraction model is "Please extract the key information of the logistics industry text from the following HTML content and output it in JSON format <div data-v-05ae83ca=""class="law-typebox">Effectiveness Level: Administrative Regulatory Document <div data-v-05ae83ca=""class="law-typebox">Document Number:

[2020] No. 1 Implementation Date: 2020-02-12Timeliness: Currently effective <div data-v-05ae83ca=""class="change-history"> <span data-v-05ae83ca="">Change History: <!----> ”.

[0104] Correspondingly, the output content of the target text extraction model is ""Effective Level": "Administrative Regulatory Document", "Document Number": "〔2020〕1", "Implementation Date": "2020-02-12", "Timeliness": "Currently effective", "Change History": """.

[0105] Another exemplary case is when the text format of the text to be processed is text (Text, TXT) format. The input content of the target text extraction model is "Please extract the key information of the logistics industry text from the following text content and output it in JSON format Index Number: 000019713O09 / 2024-00055 Document Number:

[0106] 〔2024〕135 Publication Date: November 27, 2024 Subject Terms: Transportation Logistics; Cost Reduction, Quality Improvement, and Efficiency Enhancement Industry Classification: Road Freight Transport; Waterborne Freight Transport; Railway Engineering Construction and Railway Transport Industry; Others".

[0107] Correspondingly, the output content of the target text extraction model is ""Index Number":

[0108] "000019713O09 / 2024-00055", "Document Number": "〔2024〕135", "Publication Date":

[0109] "2024-11-27", "Subject Terms": ["Transportation Logistics", "Cost Reduction, Quality Improvement, and Efficiency Enhancement"], "Agency Classification": "Transportation Service Department", "Subject Classification": "Policy Document", "Industry Classification": ["Road Freight Transport", "Waterborne Freight Transport", "Railway Engineering Construction and Railway Transport Industry", "Others"]". In some embodiments, in combination with Figure 4 , such as Figure 5 shown, the text extraction method further includes:

[0110] S501. Obtain an initial text extraction model.

[0111] In the embodiments of the present application, the text extraction device may select an initial text extraction model, and train the initial text extraction model through a training text data set and a regional corpus data set to obtain a target text extraction model with the ability to understand logistics industry texts and a text recognition ability with high accuracy.

[0112] It can be understood that the text extraction device may select a general large model that meets the test regulations as the initial text extraction model. That is, the general large model has the ability to extract texts.

[0113] Combined with Figure 1 , the text extraction device 101 can obtain the initial text extraction model from the data storage device 102.

[0114] S502. Train the initial text extraction model based on the training text data set to obtain a logistics industry text extraction model.

[0115] The training text data set includes text data representing the text content of the logistics industry.

[0116] In the embodiments of the present application, the training text data set includes text data representing the text content of the logistics industry. Therefore, the text extraction device can train the initial text extraction model based on the training text data set. In this way, the initial text extraction model can obtain the general texts in the logistics industry, the statistical rules of general text expressions, and the semantic information of general texts. The text extraction device can obtain a logistics industry text extraction model with the ability to understand logistics industry texts.

[0117] Optionally, the training text data set includes: text data in the logistics industry and text data in the general field.

[0118] It can be understood that the text data in the general field is used to increase the data volume of the training text data set, so that the initial text extraction model can be fully trained and a logistics industry text extraction model with the ability to understand logistics industry texts can be obtained.

[0119] S503. Train the logistics industry text extraction model based on the regional corpus data set to obtain the target text extraction model.

[0120] The regional corpus data set includes corpus data representing the language expression forms of the same text in different regions.

[0121] In the embodiments of the present application, the regional corpus dataset includes corpus data for representing the language expression forms of the same text in different regions. Therefore, the text extraction device can train a text extraction model for the logistics industry based on the regional corpus dataset. In this way, the text extraction model for the logistics industry obtains text data of the same text in different regional language expression forms. The text extraction device can obtain a target text extraction model with the ability to understand text in the logistics industry and a high-accuracy text recognition ability.

[0122] In the process of the text extraction device training the text extraction model for the logistics industry based on the regional corpus dataset, a regional adaptation training strategy is adopted to strengthen the training of the regional corpus dataset and improve the accuracy of the target text extraction model in recognizing text. When the target text extraction model processes various to-be-processed texts with complex structures involving regional information, the target text extraction model can better handle text extraction tasks in specific regions.

[0123] As can be seen from the above, the target text extraction model obtained in the present application has a unique ability, that is, it can extract the target text, and supplement the target text by integrating the text content of the to-be-processed text and the region to which the to-be-processed text belongs to generate a more accurate target text. In addition, the target text extraction model can also integrate the text content of the to-be-processed text and a specific dictionary in the logistics industry to supplement and explain detailed information for the case of multiple words with the same meaning, and output it in a structured manner. In this way, the target text extraction model realizes the resolution of texts / fields with abbreviations or ambiguities, and solves the NER problem. Therefore, the present application uses deep semantic analysis and a specific dictionary in the logistics industry to optimize the abbreviation resolution problem, ensuring the accuracy and integrity of target text extraction, and improving the target text extraction model's understanding and processing ability of complex knowledge in the logistics industry.

[0124] Exemplarily, Figure 6 A schematic flowchart of training a target text extraction model is shown.

[0125] First, the text extraction device can obtain an initial text extraction model. Among them, the initial text extraction model is a basic large model.

[0126] Then, the text extraction device can perform unsupervised pre-training on the initial text extraction model based on the training text dataset (that is, input a large amount of text files related to the logistics industry, web news, etc. for unsupervised pre-training) to obtain a text extraction model for the logistics industry. In this way, the text extraction model for the logistics industry obtained by the text extraction device has the ability to understand text in the logistics industry.

[0127] Next, the text extraction device can fine-tune the text extraction model for the logistics industry based on the regional corpus dataset to obtain the target text extraction model. In this way, the target text extraction model obtained by the text extraction device has a high-accuracy text recognition ability.

[0128] As can be seen from the above, the text extraction device performs unsupervised pre-training on the initial text extraction model based on the training text dataset, and further fine-tunes it based on the regional corpus dataset. In this way, the target text extraction model obtained in this application has the ability to understand text in the logistics industry and a high-accuracy text recognition ability. Therefore, the target text extraction model can deeply understand the text semantics, context relationship of the text to be processed, and the language characteristics of the region to which the text to be processed belongs. Further, the target text extraction model can extract the target text required for the text extraction task from the text to be processed. Since the target text extraction model can deeply understand the language characteristics of the region to which the text to be processed belongs, the target text extraction model can output accurate and unambiguous target text.

[0129] In some embodiments, in combination with Figure 5 , such as Figure 7 shown, the text extraction method further includes:

[0130] S701. Obtain the initial training text, and preprocess the initial training text to obtain the target training text.

[0131] The preprocessing includes at least one of removing noise data and unifying the text format.

[0132] It can be understood that the initial training text is obtained from different sources (for example, logistics industry reports, logistics industry-related documents, logistics industry web links). Therefore, the initial training text has noise data, and the initial training texts from different sources have different format problems. The text extraction device can obtain the initial training text and preprocess the initial training text to obtain the target training text.

[0133] Optionally, the text extraction device can obtain the initial training text from different sources.

[0134] Optionally, in combination with Figure 1 , the text extraction device 101 can obtain the initial training text from the data storage device 102.

[0135] The formats of the initial training text include: Portable Document Format (PDF), TXT, HTML format, and word file format.

[0136] The text extraction device preprocesses the initial training text, including: removing noise data (e.g., removing irrelevant characters and garbled codes) and unifying the text format to obtain a target training text with correct format and complete quality.

[0137] S702. Extract the key text of the target training text based on a preset structured rule template.

[0138] The key text includes professional terms and high-frequency terms in the target training text.

[0139] The high-frequency terms are terms whose occurrence frequency is greater than a preset frequency.

[0140] In the embodiments of the present application, the text extraction device can extract the key text of the target training text based on a preset structured rule template. For example, professional terms, high-frequency terms, the issuing agency of the target training text, the document number of the target training text, the issuing date of the target training text, and / or the article summary of the target training text.

[0141] It can be understood that the data storage device includes multiple preset structured rule templates. The structured rule template includes: text format, field name, and the structured rule for extracting the target text corresponding to the field name in this text format.

[0142] Exemplarily, a preset structured rule template is shown in Table 1.

[0143]

[0144]

[0145] Table 1

[0146] As can be seen from Table 1, when the text format is "TXT" and the field name is "issuing date", the structured rule is "match the regular expression "issuing date: \d{4} / \d{2} / \d{2}"".

[0147] S703. Perform annotation processing on the key text based on the annotation rules to obtain a preliminary annotation result of the key text.

[0148] In the embodiments of the present application, the text extraction device can perform annotation processing on the key text based on the annotation rules to annotate the field name corresponding to the key text. In this way, the text extraction device can obtain a training text data set with annotation results and train the initial text extraction model based on the training text data set.

[0149] S704. Correct the preliminary annotation result based on a pre-constructed annotation database to obtain the target annotation result of the key text.

[0150] In the embodiments of the present application, the text extraction device may correct the preliminary annotation result based on a pre-constructed annotation database to obtain the target annotation result of the key text. In this way, the target annotation result obtained by the text extraction device is more accurate, and the text extraction model trained for the logistics industry has a more accurate understanding of the text in the logistics industry.

[0151] In a feasible manner, the text extraction device may correct the preliminary annotation result based on a pre-constructed annotation database to obtain the target annotation result of the key text.

[0152] In some embodiments, the text extraction device extracts the key text of the target training text based on a preset structured rule template, and performs annotation processing on the key text based on the annotation rule to obtain the preliminary annotation result of the key text. However, due to the problems of missing or incorrect key fields in the text to be processed, the preliminary annotation result determined by the text extraction device is incorrect (for example, the issue date and the release date are confused).

[0153] Exemplarily, Table 2 shows a key text and the preliminary annotation result of the key text.

[0154]

[0155] Table 2

[0156] It can be seen from Table 2 that the above information includes an issue date and a release date, and the issue date and the release date may be incorrect.

[0157] Therefore, in another feasible manner, the text extraction device may obtain the result of the staff's review and correction of the preliminary annotation result, and use the result of the staff's correction of the preliminary annotation result as the target annotation result.

[0158] S705. Construct a training text dataset based on the initial training text, the target training text, the key text, and the target annotation result of the key text.

[0159] In the embodiments of the present application, the text extraction device may construct a training text dataset based on the initial training text, the target training text, the key text, and the target annotation result of the key text. In this way, the text extraction device can train the initial text extraction model based on the training text dataset.

[0160] In some embodiments, the text extraction method further includes:

[0161] Update the annotation rule according to the preliminary annotation result and the target annotation result.

[0162] The annotation rules include: the processing method for annotation symbols, the processing method for special situations, and the correspondence between regions and the abbreviations of document-issuing agencies.

[0163] In an implementable manner, the text extraction device can also update the annotation rules according to the preliminary annotation result and the target annotation result.

[0164] Specifically, first, the text extraction device performs annotation processing on the key text based on the annotation rules to obtain the preliminary annotation result of the key text. Then, the text extraction device can obtain the result of the staff's review and correction of the preliminary annotation result, and use the result of the staff's correction of the preliminary annotation result as the target annotation result. Next, the text extraction device compares the preliminary annotation result and the target annotation result of some text to be processed, and updates the annotation rules. Then, repeat the above steps until all initial training texts are processed.

[0165] The result of the staff's review and correction of the preliminary annotation result includes: the field name corresponding to the key text, and the result of the adjustment of the field boundary of the key text.

[0166] In some embodiments, combined with Figure 7 , such as Figure 8 shown, the text extraction method further includes:

[0167] S801. Obtain the region to which the target training text belongs.

[0168] Optionally, combined with Figure 1 , the text extraction device 101 can directly obtain the region to which the target training text belongs from the data storage device.

[0169] Optionally, the text extraction device can determine the region to which the target training text belongs according to the text content of the target training text.

[0170] S802. Extract the language features of the target training text.

[0171] The text extraction device can extract the language features of the target training text.

[0172] S803. Construct a regional corpus dataset according to the correspondence between the region and the language features.

[0173] In the embodiments of the present application, for the same text, the language expression forms in different regions are different. In order to obtain the correspondence between the region to which the target training text belongs and the language features of the target training text, as well as the language expression forms of the same text in different regions. The text extraction device can obtain the region to which the target training text belongs and extract the language features of the target training text, so that the text extraction device constructs a regional corpus dataset according to the correspondence between the region and the language features.

[0174] The regional corpus dataset includes: corpus datasets from different regions.

[0175] For the same text, the language expression forms in different regions are different.

[0176] Exemplarily, the document-issuing agencies in the first region and the second region are both xxx document-issuing agencies, but the document numbers in the first region and the second region are different.

[0177] The text extraction device can construct a regional corpus dataset (which can also be called a regional Q&A dataset) according to the language features of the target training text in different regions.

[0178] The regional corpus dataset includes document-issuing agencies and document numbers from different regions.

[0179] In this way, the text extraction device can train the logistics industry text extraction model based on the regional corpus dataset, and the target text extraction model can also deeply learn the language features of different regions based on the language features of the target training text. In this way, the target text extraction model obtained by the text extraction device has a high accuracy in text recognition.

[0180] The above mainly introduces the solution provided in the embodiments of the present application from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0181] The embodiments of the present application can divide the text extraction device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. Optionally, the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. There can be other division methods in actual implementation.

[0182] Figure 9 Shows a schematic structural diagram of a text extraction device provided in an embodiment of the present application. As Figure 9 shown, the text extraction device includes: a communication unit 901 and a processing unit 902;

[0183] A communication unit 901 for obtaining a text to be processed and a text extraction task prompt word; the text to be processed includes text content in the logistics industry and a regional feature expression text of the region to which the text to be processed belongs; a processing unit 902 for extracting a target text required for the text extraction task from the text to be processed based on the text extraction task prompt word and a target text extraction model; the target text extraction model is trained based on a training text data set in the logistics industry and a regional corpus data set.

[0184] Optionally, the processing unit 902 is specifically configured to: determine the text extraction task prompt word as a text prefix of the text to be processed, and input the text to be processed with the added text prefix into the target text extraction model to obtain the target text.

[0185] Optionally, the communication unit 901 is further configured to obtain an initial text extraction model; the processing unit 902 is further configured to train the initial text extraction model based on the training text data set to obtain a text extraction model in the logistics industry; the training text data set includes text data for representing text content in the logistics industry; the processing unit 902 is further configured to train the text extraction model in the logistics industry based on the regional corpus data set to obtain the target text extraction model; the regional corpus data set includes corpus data for representing language expression forms of the same text in different regions.

[0186] Optionally, the communication unit 901 is further configured to obtain an initial training text and preprocess the initial training text to obtain a target training text; the preprocessing includes at least one of removing noise data and unifying the text format; the processing unit 902 is further configured to extract key texts of the target training text based on a preset structured rule template; the key texts include professional words and high-frequency words in the target training text; the high-frequency words are words with an occurrence frequency greater than a preset frequency; the processing unit 902 is further configured to perform annotation processing on the key texts based on an annotation rule to obtain a preliminary annotation result of the key texts; the processing unit 902 is further configured to correct the preliminary annotation result based on a pre-constructed annotation database to obtain a target annotation result of the key texts; the processing unit 902 is further configured to construct a training text data set based on the initial training text, the target training text, the key texts, and the target annotation result of the key texts.

[0187] Optionally, the processing unit 902 is further configured to update the annotation rule according to the preliminary annotation result and the target annotation result.

[0188] Optionally, the communication unit 901 is further configured to obtain the region to which the target training text belongs; the processing unit 902 is further configured to extract the language features of the target training text; the processing unit 902 is further configured to construct a regional corpus data set according to the correspondence between the region and the language features.

[0189] The embodiments of the present application also provide a computer-readable storage medium, which includes computer-executable instructions. When the computer-executable instructions run on a computer, the computer is enabled to execute the text extraction method provided in the above embodiments.

[0190] The embodiments of the present application also provide a computer program, which can be directly loaded into a memory and contains software code. After being loaded and executed by a computer, the computer program can implement the text extraction method provided in the above embodiments.

[0191] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes a computer-readable storage medium and a communication medium, where the communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0192] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0193] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form. The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0194] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the general technology, or all or part of this technical solution, may be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0195] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A text extraction method, characterized in that: The method comprises: Obtaining a text to be processed and a prompt word for a text extraction task; the text to be processed includes text content of the logistics industry and a text expressing regional characteristics of the region to which the text to be processed belongs; Based on the text extraction task prompt words and the target text extraction model, the target text required for the text extraction task is extracted from the text to be processed; the target text extraction model is trained based on the training text data set and the regional corpus data set of the logistics industry.

2. The method according to claim 1, characterized in that The step of extracting the target text required for the text extraction task from the text to be processed based on the text extraction task prompt word and the target text extraction model includes: The text extraction task prompt word is determined as the text prefix of the text to be processed, and the text to be processed with the text prefix added is input into the target text extraction model to obtain the target text.

3. The method according to claim 1, characterized in that The method further comprises: Get the initial text extraction model; Based on the training text data set, the initial text extraction model is trained to obtain a logistics industry text extraction model; the training text data set includes text data for representing text content of the logistics industry; Based on the regional corpus data set, the logistics industry text extraction model is trained to obtain a target text extraction model; the regional corpus data set includes corpus data for representing the language expression forms of the same text in different regions.

4. The method according to claim 3, characterized in that The method further comprises: Acquire an initial training text, and preprocess the initial training text to obtain a target training text; the preprocessing includes at least one of removing noise data and unifying the text format; Based on a preset structured rule template, extract the key text of the target training text; the key text includes professional vocabulary and high-frequency vocabulary in the target training text; the high-frequency vocabulary is a vocabulary with an appearance frequency greater than a preset frequency; Annotating the key text based on the annotation rules to obtain a preliminary annotation result of the key text; Based on a pre-built annotation database, the preliminary annotation result is modified to obtain a target annotation result of the key text; The training text data set is constructed based on the initial training text, the target training text, the key text, and the target annotation result of the key text.

5. The method according to claim 4, characterized in that The method further comprises: The labeling rule is updated according to the preliminary labeling result and the target labeling result.

6. The method according to claim 3, characterized in that The method further comprises: Obtaining the region to which the target training text belongs; Extracting language features of the target training text; The regional corpus dataset is constructed according to the correspondence between the region and the language feature.

7. A text extraction device, characterized in that: The device comprises: a communication unit and a processing unit; The communication unit is used to obtain the text to be processed and the prompt words of the text extraction task; the text to be processed includes the text content of the logistics industry and the regional feature expression text of the region to which the text to be processed belongs; The processing unit is used to extract the target text required for the text extraction task from the text to be processed based on the text extraction task prompt words and the target text extraction model; the target text extraction model is trained based on the training text data set and the regional corpus data set of the logistics industry.

8. A text extraction device, characterized in that: include: A processor and a memory; wherein the memory is used to store one or more programs, and the one or more programs include computer-executable instructions. When the device is running, the processor executes the computer-executable instructions stored in the memory to enable the device to perform the method described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: When the computer-executable instructions stored in the computer-readable storage medium are executed by a processor of a text extraction device, the text extraction device can perform the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product comprises: a computer program or instructions, and when the computer program or instructions are run on a computer, the computer is caused to perform the method according to any one of claims 1 to 6.