Information intelligent extraction method and device, electronic equipment and storage medium

By acquiring target information indicators, parsing them, and inputting them into a large language model, the problem of low efficiency in manual information acquisition in existing technologies is solved, and intelligent and efficient information extraction is achieved.

CN117332037BActive Publication Date: 2026-07-24SHENZHEN VALUE ONLINE INFORMATION POLYTRON TECH INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN VALUE ONLINE INFORMATION POLYTRON TECH INC
Filing Date
2023-09-19
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, companies rely on manual methods to obtain information from the capital market, which is inefficient, error-prone, and makes it difficult to intelligently extract useful information from massive amounts of data.

Method used

By acquiring target information indicators, identifying target files, parsing target files to obtain relevant text content, and inputting the information indicators into a pre-trained large language model, target information is intelligently extracted.

Benefits of technology

It improves the accuracy and efficiency of information extraction, reduces human intervention, and enables the efficient extraction of useful information from massive amounts of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117332037B_ABST
    Figure CN117332037B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of information processing, and provides an information intelligent extraction method and device, an electronic equipment and a storage medium. The method comprises the following steps: obtaining a target information index; determining a target file, wherein the target file is an information file to be extracted with the target information index; analyzing the target file to obtain target text content related to the target information index; inputting the target information index and the target text content into a target large language model to extract target information corresponding to the target information index. By using the method, useful information can be intelligently extracted from a large amount of information, and the accuracy and efficiency of information extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, and in particular to an intelligent information extraction method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the sustained and stable development of the national economy, the capital market has also developed rapidly. The development and popularization of the internet have led to an explosive increase in various types of information and data. The capital market generates a large amount of business information daily, including capital market-related regulations and financial data. For enterprises, obtaining relevant information from the capital market is extremely important.

[0003] In current technologies, companies typically rely on manual methods to obtain information from capital markets, which is inefficient and prone to errors. How to intelligently extract useful information from massive amounts of data, minimize human intervention, and improve the accuracy and efficiency of information extraction is a problem that needs to be addressed. Summary of the Invention

[0004] This application provides an intelligent information extraction method, device, electronic device, and storage medium, which can solve the problem in the prior art of lacking effective monitoring of abnormal behavior in the vehicle driving environment, thus failing to effectively ensure the safety of drivers and passengers during driving.

[0005] In a first aspect, embodiments of this application provide an intelligent information extraction method, including:

[0006] Obtain target information indicators;

[0007] The target file is determined, wherein the target file is an information file from which the target information indicators are to be extracted;

[0008] The target file is parsed to obtain the target text content related to the target information indicators;

[0009] The target information indicators and the target text content are input into the target large language model to extract the target information corresponding to the target information indicators.

[0010] In one possible implementation of the first aspect, prior to the step of determining the target file, the method further includes:

[0011] Monitor the designated information platform and capture information files from the designated information platform according to preset capture rules;

[0012] The captured information file is stored in a designated database;

[0013] The step of determining the target file includes:

[0014] Obtain user instructions;

[0015] Based on the user instructions, an information file is selected from the designated information database to determine the target file for extracting the information indicators.

[0016] In one possible implementation of the first aspect, the step of parsing the target file to obtain target text content related to the target information indicator includes:

[0017] Determine the word count and / or page count of the target file;

[0018] If the number of words and / or the number of pages falls within the first preset word and page range, then the first parsing method is selected to parse the target file and obtain the target text content related to the target information indicator;

[0019] If the number of words and / or the number of pages falls within the second preset word and page range, then the second parsing method is selected to parse the target file and obtain the target text content related to the target information indicator.

[0020] In one possible implementation of the first aspect, the step of selecting a first parsing method to parse the target file and obtain target text content related to the target information indicator includes:

[0021] The target file is traversed page by page based on the target information indicators;

[0022] Based on the page traversal results, extract the target text content related to the target information indicators from the target file.

[0023] In one possible implementation of the first aspect, the target text content includes a title and paragraphs. After the step of selecting a second parsing method to parse the target file and obtain the target text content related to the target information indicator, the method includes:

[0024] The target file is disassembled to obtain the title and paragraphs of the disassembled target file;

[0025] Calculate the similarity between the target information index and the title;

[0026] Based on the similarity, titles and paragraphs related to the target information indicators are determined.

[0027] In one possible implementation of the first aspect, after the step of inputting the target information indicator and the target text content into the target large language model and extracting the target information corresponding to the target information indicator, the method further includes:

[0028] Retrieve the matching rules configured by the user;

[0029] The target information is output after being matched using regular expressions according to the matching rules.

[0030] In one possible implementation of the first aspect, after the step of inputting the target information indicator and the target text content into the target large language model and extracting the target information corresponding to the target information indicator, the method further includes:

[0031] The target large language model is trained again based on the target text content and the extracted target information.

[0032] Secondly, embodiments of this application provide an intelligent information extraction device, comprising:

[0033] The target indicator acquisition unit is used to acquire target information indicators;

[0034] The target file determination unit is used to determine the target file, which is an information file from which the target information indicators are to be extracted.

[0035] The target content acquisition unit is used to parse the target file and acquire target text content related to the target information indicators;

[0036] The target information extraction unit is used to input the target information indicators and the target text content into the target large language model and extract the target information corresponding to the target information indicators.

[0037] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the intelligent information extraction method as described in the first aspect above.

[0038] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the intelligent information extraction method as described in the first aspect above.

[0039] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the information intelligent extraction method described in the first aspect above.

[0040] In this embodiment, by acquiring target information indicators, a target file is determined. The target file is an information file from which the target information indicators are to be extracted. Then, the target file is parsed to obtain target text content related to the target information indicators. The target information indicators and the target text content are then input into a target large language model to intelligently extract the target information corresponding to the target information indicators. This eliminates the need for manual searching of useful information from massive amounts of data, greatly improving the accuracy and efficiency of information extraction. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating the implementation of the intelligent information extraction method provided in this application embodiment;

[0043] Figure 2 This is a flowchart illustrating the specific implementation of the intelligent information extraction method for capturing and storing information files provided in this application embodiment;

[0044] Figure 3 This is a flowchart illustrating a specific implementation of the intelligent information extraction method for determining a target file provided in this application embodiment;

[0045] Figure 4 This is a flowchart illustrating a specific implementation of step S103 in the intelligent information extraction method provided in this application embodiment;

[0046] Figure 4.1 This is a flowchart illustrating a specific implementation of the intelligent information extraction method provided in this application, in which a first parsing method is selected to parse the target file;

[0047] Figure 4.2 This is a flowchart illustrating a specific implementation of the intelligent information extraction method provided in this application, in which the target file is parsed using a second parsing method.

[0048] Figure 5 This is a flowchart illustrating a specific implementation of regular expression matching in the intelligent information extraction method provided in this application embodiment;

[0049] Figure 6 This is a structural block diagram of the intelligent information extraction device provided in the embodiments of this application;

[0050] Figure 7 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0051] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0052] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0053] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0054] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0055] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0056] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0057] The information intelligent extraction method provided in this application embodiment can be applied to various types of terminal devices or servers that need to perform information intelligent extraction, specifically including electronic devices such as mobile phones, tablets, wearable devices, laptops, and desktop computers.

[0058] Figure 1 The implementation flow of the intelligent information extraction method provided in this application embodiment is illustrated. The method flow includes steps S101 to S104. The specific implementation principle of each step is as follows:

[0059] S101: Obtain target information indicators.

[0060] The target information metrics are the information metrics required for this intelligent information extraction, such as repurchase data and equity incentive data.

[0061] In this embodiment, the target information indicator is specified by the user. There can be one or more target information indicators. In some implementations, the types of the multiple target information indicators can be different.

[0062] In some implementations, the intelligent information extraction method provided in this embodiment is applied to an intelligent information extraction system, where the user inputs target information indicators on a designated page of the intelligent information extraction system.

[0063] S102: Determine the target file, which is an information file from which the target information indicators are to be extracted.

[0064] In this embodiment, the target file can be specified by the user.

[0065] As one possible implementation of this application, such as Figure 2 As shown, prior to the steps described above for determining the target file, the following steps are also included:

[0066] A1: Monitor the specified information platform and capture information files from the specified information platform according to the preset capture rules.

[0067] In this embodiment, information files are captured on the designated information platform according to preset capture rules. The information files can be unstructured data, which is data with irregular or incomplete structures, lacking a predefined data model, and inconvenient to represent using a two-dimensional logical table in a database. Unstructured data includes all formats of office documents, text, images, XML, HTML, various reports, etc.

[0068] A2: Store the captured information file in the specified database. The specified database can be an Elasticsearch database.

[0069] In this embodiment of the application, information monitoring is performed on designated information platforms, and web crawlers are set up to crawl information files published on the designated information platforms. The designated information platforms include, but are not limited to, online platforms (such as financial forums, stock market forums, technology forums, regulatory agency websites, financial associations and other financial professional websites) and instant messaging platforms (such as QQ and WeChat). For example, for information platforms such as Weibo, designated regulatory agency websites, financial associations and other financial professional websites, and stock market forums, web crawlers are set up to automatically capture a large number of information files on the information platforms and store the captured information files in a designated database.

[0070] In this embodiment, by monitoring the designated information platform and automatically capturing information files, the latest information files can be obtained. The captured information files are stored in the designated database, which is conducive to the retrieval of information at any time.

[0071] In one possible implementation, the captured information files may be of various file types, such as Word, PDF, Excel, etc. In this embodiment, the captured information files are normalized according to their file types and then stored in the specified database.

[0072] Figure 3 This application illustrates a specific implementation of the intelligent information extraction method provided in its embodiments, which involves determining a target file, where the target file is an information file from which the target information indicators are to be extracted. Details are as follows:

[0073] B1: Obtain user instructions. The user instructions are used to instruct the determination of target files. The user instructions may indicate the determination of multiple target files.

[0074] B2: Based on the user instruction, select an information file from the designated information database to determine the target file for extracting the information indicators.

[0075] In this embodiment, an information file is selected from a designated information database as the target file for extracting the information indicators according to user instructions.

[0076] In one possible implementation, the information files in the designated information database are stored according to file type. For example, information files are stored according to categories such as announcements, web page text, and organizational documents. In this embodiment, the user instruction is used to indicate the file type, and based on the user instruction, the information file corresponding to the file type is selected from the designated information database as the target file for extracting the information indicators.

[0077] In this embodiment, by having the user specify the target file for extracting target information indicators, the scope of information extraction can be effectively narrowed, thereby improving the accuracy and efficiency of information extraction.

[0078] S103: Parse the target file to obtain target text content related to the target information indicators.

[0079] In this embodiment of the application, target text content related to target information indicators is obtained by parsing the target file.

[0080] The target text content related to the target information indicator refers to the text content that includes the target information corresponding to the target information indicator. That is, the target text content includes the target information corresponding to the target information indicator. For example, if the target information indicator is repurchase data, the target text content includes the repurchase data indicator and the related numerical values; if the target information indicator is equity incentive data, the target text content includes equity incentive data and the related numerical values.

[0081] In some implementations, target text content related to a target information indicator refers to text content that includes target information corresponding to the target information indicator, and / or text content that includes information indicators associated with the target information indicator.

[0082] As one possible implementation of this application Figure 4 The specific implementation flow of step S103 in the intelligent information extraction method provided in this application embodiment is shown below:

[0083] C1: Determine the number of words and / or pages in the target file.

[0084] C2: If the number of words and / or the number of pages falls within the first preset word and page range, then the first parsing method is selected to parse the target file and obtain the target text content related to the target information indicator.

[0085] C3: If the number of words and / or pages falls within a second preset word / page range, then the second parsing method is selected to parse the target file and obtain the target text content related to the target information indicator. The minimum value of the second preset word / page range is greater than the maximum value of the first preset word / page range.

[0086] In this embodiment of the application, different parsing methods are selected according to the number of words and / or pages of the target file, which helps to ensure the accuracy and effectiveness of file parsing.

[0087] As one possible implementation of this application, the first parsing method includes page-by-page traversal, such as... Figure 4.1 The specific implementation process of the intelligent information extraction method provided in this application embodiment, which selects a first parsing method to parse the target file and obtain target text content related to the target information indicator, includes:

[0088] C21: If the number of words and / or the number of pages belongs to the first preset number of words and pages range, then the target file is traversed page by page based on the target information index; the maximum value of the first preset number of words and pages range is less than the minimum value of the second preset number of words and pages range.

[0089] In this embodiment, if the number of words is less than or equal to a preset word count threshold, and / or the number of pages is less than or equal to a preset page count threshold, then the target file is traversed page by page based on the target information index, that is, the target information index is searched for on each page.

[0090] C22: Based on the page traversal results, extract the target text content related to the target information indicator from the target file. The content related to the target information indicator is the page content in the target file that contains the target information indicator.

[0091] In this embodiment, for target files with a small number of words and / or pages, the pages containing the target information indicators are locked by traversing each page, which can avoid missed detections.

[0092] As one possible implementation of this application, the target text content includes a title and paragraphs, such as... Figure 4.2 The specific implementation process of the intelligent information extraction method provided in this application embodiment, which selects a second parsing method to parse the target file and obtain target text content related to the target information indicator, includes:

[0093] C31: Disassemble the target file to obtain the title and paragraphs of the disassembled target file. The file disassembly can be performed using Python, and specific details can be found in existing technologies, which will not be elaborated here.

[0094] C32: Calculate the similarity between the target information index and the title.

[0095] In one possible implementation, a cosine similarity algorithm can be used to calculate the similarity between the target information index and the title.

[0096] C33: Based on the similarity, determine the titles and paragraphs related to the target information indicators.

[0097] Content related to the target information indicator refers to the titles and paragraphs related to the target information indicator. The titles and paragraphs related to the target information indicator are those whose similarity to the target information indicator is greater than or equal to a preset similarity threshold.

[0098] In some implementations, if the similarity between the extracted title and the target information indicator is less than the preset similarity threshold, a user information extraction anomaly is indicated, suggesting that the target file may be incorrect.

[0099] In this embodiment, for target files with a large number of words and / or pages, the title and paragraphs are obtained by decomposing the files. Then, based on the similarity between the decomposed title and the target information indicator, the paragraphs that may contain the target information indicator are located, thereby determining the target file content related to the information indicator.

[0100] S104: Input the target information indicators and the target text content into the target large language model, and extract the target information corresponding to the target information indicators.

[0101] Large Language Model (LLM) refers to a deep learning model trained on a large amount of text data that can generate natural language text or understand the meaning of language text. In the embodiments of this application, the target large language model is a pre-trained large language model.

[0102] In this embodiment, the target large language model is configured with rules based on the target information indicators. The rule configuration is used to instruct the target large language model to extract the target information indicators and match the rules corresponding to the target information indicators.

[0103] In one possible implementation, before inputting the target information indicators and the target text content into the target large language model, the target text content undergoes data cleaning and standardization. Data cleaning includes removing noise, non-text characters, HTML tags, etc.

[0104] In one possible implementation, a large language model is pre-trained and configured as needed to obtain the target large language model. The configuration of the large language model includes output rule configuration.

[0105] In one possible implementation, indicator samples, text samples, and answer samples are collected, and the text samples undergo data cleaning and standardization. A large language model is trained using the indicator samples, the cleaned and standardized text samples, and the answer samples. The large language model is considered complete when the answer output by the large language model is the same as the answer sample, or when the error between the two is within a preset error range. The various model parameters of the current large language model are saved to obtain the target large language model.

[0106] In one possible implementation, the target large language model is retrained based on the target text content and the extracted target information. Continuing to train the target large language model using the target information extracted from it can further improve its performance.

[0107] In one possible implementation, the process of extracting target information from the target large language model is monitored in real time. If extraction fails, it is retried. If extraction still fails after a specified number of retries, for example, after 3 retries, the extraction process is logged.

[0108] As one possible implementation of this application, such as Figure 5 As shown, in the information intelligent extraction method provided in this application embodiment, after the step of inputting the target information indicator and the target text content into the target large language model and extracting the target information corresponding to the target information indicator, the method further includes:

[0109] D1: Retrieves the matching rules configured by the user. Matching rules include information output rules, etc.

[0110] D2: Output the target information after performing regular expression matching according to the matching rules. Specifically, output the target information after matching with isomorphic regular expressions. A regular expression is an object that describes a character pattern. Regular expressions can be used for powerful pattern matching, text retrieval, and replacement functions.

[0111] In this embodiment, users are allowed to configure matching rules, and the target information is output after being matched with regular expressions according to the matching rules. This helps to improve the efficiency of information extraction, so that the extracted information meets the user's needs, thereby improving the user experience.

[0112] As can be seen from the above, in this embodiment of the application, by obtaining target information indicators, a target file is determined. The target file is an information file from which the target information indicators are to be extracted. Then, the target file is parsed to obtain target text content related to the target information indicators. The target information indicators and the target text content are then input into the target large language model to intelligently extract the target information corresponding to the target information indicators. This eliminates the need for manual searching of useful information from massive amounts of information, and can greatly improve the accuracy and efficiency of information extraction.

[0113] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0114] Corresponding to the information intelligent extraction method described in the above embodiments, Figure 6 A structural block diagram of the intelligent information extraction device provided in the embodiments of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0115] Reference Figure 6The intelligent information extraction device includes: a target indicator acquisition unit 61, a target file determination unit 62, a target content acquisition unit 63, and a target information extraction unit 64, wherein:

[0116] The target indicator acquisition unit 61 is used to acquire target information indicators;

[0117] The target file determination unit 62 is used to determine the target file, which is an information file from which the target information indicators are to be extracted.

[0118] The target content acquisition unit 63 is used to parse the target file and acquire target text content related to the target information indicators;

[0119] The target information extraction unit 64 is used to input the target information indicators and the target text content into the target large language model and extract the target information corresponding to the target information indicators.

[0120] As one possible implementation of this application, the above-mentioned intelligent information extraction device further includes:

[0121] The file capture unit is used to monitor a specified information platform and capture information files from the specified information platform according to preset capture rules;

[0122] The file storage unit is used to store the captured information file to a designated database.

[0123] The aforementioned target file determination unit 62 includes:

[0124] The instruction acquisition module is used to acquire user instructions;

[0125] The file determination module is used to select an information file from the specified information database based on the user instruction and determine it as the target file for extracting the information indicators.

[0126] As one possible implementation of this application, the target content acquisition unit 63 includes:

[0127] The content information determination module is used to determine the number of words and / or pages of the target file;

[0128] The first parsing module is used to select a first parsing method to parse the target file if the number of words and / or the number of pages belongs to a first preset number of words and pages range, so as to obtain the target text content related to the target information indicator;

[0129] The second parsing module is used to select the second parsing method to parse the target file if the number of words and / or the number of pages falls within a second preset range of word and page counts, thereby obtaining target text content related to the target information indicators.

[0130] As one possible implementation of this application, the first parsing module includes:

[0131] The page traversal submodule is used to traverse the target file page by page based on the target information indicators;

[0132] The first content extraction submodule is used to extract target text content related to the target information indicators from the target file based on the page traversal results.

[0133] As one possible implementation of this application, the target text content includes a title and paragraphs, and the second parsing module includes:

[0134] The content decomposition submodule is used to decompose the target file to obtain the title and paragraphs of the decomposed target file;

[0135] A similarity calculation submodule is used to calculate the similarity between the target information index and the title;

[0136] The second content extraction submodule is used to determine the titles and paragraphs related to the target information indicators based on the similarity.

[0137] As one possible implementation of this application, the above-mentioned intelligent information extraction device further includes:

[0138] The rule retrieval unit is used to retrieve the matching rules configured by the user.

[0139] The matching output unit is used to output the target information after performing regular expression matching according to the matching rules.

[0140] As one possible implementation of this application, the above-mentioned intelligent information extraction device further includes:

[0141] The training unit is used to retrain the target large language model based on the target text content and the extracted target information.

[0142] As can be seen from the above, in this embodiment of the application, by obtaining target information indicators, a target file is determined. The target file is an information file from which the target information indicators are to be extracted. Then, the target file is parsed to obtain target text content related to the target information indicators. The target information indicators and the target text content are then input into the target large language model to intelligently extract the target information corresponding to the target information indicators. This eliminates the need for manual searching of useful information from massive amounts of information, and can greatly improve the accuracy and efficiency of information extraction.

[0143] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0144] This application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements... Figures 1 to 5 The steps of any intelligent information extraction method are represented.

[0145] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements... Figures 1 to 5 The steps of any intelligent information extraction method are represented.

[0146] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform the following: Figures 1 to 5 The steps of any intelligent information extraction method are represented.

[0147] Figure 7 This is a schematic diagram of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 7 of this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70. When the processor 70 executes the computer program 72, it implements the steps in the various information intelligent extraction method embodiments described above, for example... Figure 1 Steps S101 to S104 are shown. Alternatively, when the processor 70 executes the computer program 72, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 6 The functions of units 61 to 64 are shown.

[0148] For example, the computer program 72 may be divided into one or more modules / units, which are stored in the memory 71 and executed by the processor 70 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program 72 in the electronic device 7.

[0149] The electronic device 7 can be an in-vehicle intelligent terminal. The electronic device 7 may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that... Figure 7This is merely an example of electronic device 7 and does not constitute a limitation on electronic device 7. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 7 may also include input / output devices, network access devices, buses, etc.

[0150] The processor 70 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0151] The memory 71 can be an internal storage unit of the electronic device 7, such as a hard disk or memory. The memory 71 can also be an external storage device of the electronic device 7, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 71 can include both internal and external storage units of the electronic device 7. The memory 71 is used to store the computer program and other programs and data required by the electronic device. The memory 71 can also be used to temporarily store data that has been output or will be output.

[0152] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0153] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0154] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0155] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0156] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An intelligent information extraction method, characterized in that, include: Obtain target information indicators; The target file is determined, wherein the target file is an information file from which the target information indicators are to be extracted; The target file is parsed to obtain the target text content related to the target information indicators; The target information indicators and the target text content are input into the target large language model to extract the target information corresponding to the target information indicators; The step of parsing the target file to obtain target text content related to the target information indicators includes: Determine the word count and / or page count of the target file; If the number of words and / or the number of pages falls within the first preset word and page range, then the first parsing method is selected to parse the target file and obtain the target text content related to the target information indicator; If the number of words and / or the number of pages falls within the second preset word and page range, then the second parsing method is selected to parse the target file and obtain the target text content related to the target information indicator. The minimum value of the second preset word and page range is greater than the maximum value of the first preset word and page range. The step of selecting a first parsing method to parse the target file and obtain target text content related to the target information indicator includes: The target file is traversed page by page based on the target information indicators; Based on the page traversal results, extract the target text content related to the target information indicators from the target file; The target text content includes a title and paragraphs. After the step of selecting the second parsing method to parse the target file and obtain the target text content related to the target information indicator, the following steps are included: The target file is disassembled to obtain the title and paragraphs of the disassembled target file; Calculate the similarity between the target information index and the title; Based on the similarity, titles and paragraphs related to the target information indicators are determined.

2. The method according to claim 1, characterized in that, Prior to the step of determining the target file, the following is also included: Monitor the designated information platform and capture information files from the designated information platform according to preset capture rules; The captured information file is stored in a designated database; The step of determining the target file includes: Obtain user instructions; Based on the user instructions, an information file is selected from the specified database to be determined as the target file for extracting the information indicators.

3. The method according to claim 1, characterized in that, After the step of inputting the target information indicator and the target text content into the target large language model and extracting the target information corresponding to the target information indicator, the method further includes: Retrieve the matching rules configured by the user; The target information is output after being matched using regular expressions according to the matching rules.

4. The method according to any one of claims 1 to 3, characterized in that, After the step of inputting the target information indicator and the target text content into the target large language model and extracting the target information corresponding to the target information indicator, the method further includes: The target large language model is trained again based on the target text content and the extracted target information.

5. An intelligent information extraction device, characterized in that, include: The target indicator acquisition unit is used to acquire target information indicators; The target file determination unit is used to determine the target file, which is an information file from which the target information indicators are to be extracted. The target content acquisition unit is used to parse the target file and acquire target text content related to the target information indicators; The target information extraction unit is used to input the target information indicators and the target text content into the target large language model and extract the target information corresponding to the target information indicators. The content information determination module is used to determine the number of words and / or pages of the target file; The first parsing module is used to select a first parsing method to parse the target file if the number of words and / or the number of pages belongs to a first preset number of words and pages range, so as to obtain the target text content related to the target information indicator; The second parsing module is used to select a second parsing method to parse the target file if the number of words and / or the number of pages belongs to a second preset number of words and pages range, and to obtain target text content related to the target information indicator. The minimum value of the second preset number of words and pages range is greater than the maximum value of the first preset number of words and pages range. The first parsing module includes: The page traversal submodule is used to traverse the target file page by page based on the target information indicators; The first content extraction submodule is used to extract target text content related to the target information index from the target file based on the page traversal results. The target text content includes a title and paragraphs, and the second parsing module includes: The content decomposition submodule is used to decompose the target file to obtain the title and paragraphs of the decomposed target file; A similarity calculation submodule is used to calculate the similarity between the target information index and the title; The second content extraction submodule is used to determine the titles and paragraphs related to the target information indicators based on the similarity.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the intelligent information extraction method as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the intelligent information extraction method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Key information extraction method, device and equipment and readable storage medium

    CN112182141A