Information labeling method and device, electronic equipment and storage medium
By using an automated method for matching and labeling structured and unstructured data, this approach solves the problems of time-consuming and costly generation of event graph training datasets in existing technologies, achieving fast and low-cost labeling results.
Patent Information
- Application Number
- CN202111285525.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-11-01
AI Technical Summary
In existing technologies, the training dataset for generating event graphs requires the full participation of technical personnel, resulting in long processing times and high costs, while relying on manual annotation methods is inefficient.
By matching structured data with unstructured information, matching and correspondence relationships are established, and labels are used to automatically annotate unstructured information, reducing manual intervention.
It enables the rapid and low-cost construction of labeled datasets for events, and can accurately label different data formats and synonym fragments in unstructured data.
Smart Images

Figure CN113961672B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to data processing technology and data generation technology. More specifically, the present disclosure provides an information labeling method and device, an electronic device and a storage medium. BACKGROUND
[0002] The event graph takes events as the core and accurately describes event information and the correlation between events. At present, the construction of the event graph mainly uses pre-training language models such as BERT (Bidirectional Encoder Representation From Transformer) and ERINE (Enhanced Language Representation With Informative Entities) that have better performance. The performance of these models often depends on the size and distribution of the training data set. In related technologies, the training data set is generally generated by manual labeling, which requires professional technical personnel to participate throughout the process and undergo multiple acceptance. SUMMARY
[0003] Based on this, the present disclosure provides an information labeling method, device, equipment and storage medium.
[0004] According to a first aspect, an information labeling method is provided, which includes: matching a plurality of structured data related to a target event and unstructured information related to the target event; in response to at least one structured data matching the unstructured information successfully, taking the at least one structured data matching the unstructured information successfully as target data, obtaining a matching relationship between each target data and at least one unstructured information segment; establishing a corresponding relationship between at least one unstructured information segment and a label according to the matching relationship, the label corresponding to the target data; and labeling the unstructured information using the label according to the corresponding relationship.
[0005] According to a second aspect, an information labeling apparatus is provided, which comprises: a matching module configured to match a plurality of structured data related to a target event and unstructured information related to the target event; an obtaining module configured to, in response to successful matching of at least one structured data and the unstructured information, obtain each target data matched successfully with the unstructured information as target data, and obtain a matching relationship between each target data and at least one unstructured information segment; a establishing module configured to establish a corresponding relationship between the at least one unstructured information segment and a label corresponding to the target data according to the matching relationship; and a labeling module configured to label the unstructured information with the label according to the corresponding relationship.
[0006] According to a third aspect, an electronic device is provided, which comprises: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the present disclosure.
[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to enable a computer to perform the method provided by the present disclosure.
[0008] According to a fifth aspect, a computer program product is provided, which comprises a computer program, and the computer program, when executed by a processor, implements the method provided by the present disclosure.
[0009] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:
[0011] Figure 1 is an exemplary system architecture schematic diagram of an information labeling method and apparatus according to an embodiment of the present disclosure;
[0012] Figure 2 is a flowchart of an information labeling method according to an embodiment of the present disclosure;
[0013] Figure 3A is a schematic diagram of unstructured information according to an embodiment of the present disclosure;
[0014] Figure 3B is a schematic diagram of a plurality of structured data according to an embodiment of the present disclosure;
[0015] Figure 3C is a schematic diagram of a plurality of structured data according to another embodiment of the present disclosure;
[0016] Figure 3D is a schematic diagram of a matching relationship according to an embodiment of the present disclosure;
[0017] Figure 3E is a schematic diagram of a matching relationship according to another embodiment of the present disclosure;
[0018] Figure 3F is a schematic diagram of a corresponding relationship according to an embodiment of the present disclosure;
[0019] Figure 4 is a block diagram of an information labeling apparatus according to an embodiment of the present disclosure; and
[0020] Figure 5 is a block diagram of an electronic device to which an information labeling method can be applied according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.
[0022] In the related art, in order to obtain a high-quality training data set, an explicit labeling requirement and a labeling scheme can be specified first. Related professional technicians clean, evaluate, extract, and analyze the data according to the labeling scheme to form a training data that preliminarily meets the labeling requirement. Then, a third-party professional technician can be hired to review the training data that preliminarily meets the labeling requirement to further improve the quality of the training data.
[0023] The training data set is generated by using the artificial labeling method, which is time-consuming and costly.
[0024] Figure 1 is an exemplary system architecture schematic diagram of an information labeling method and apparatus according to an embodiment of the present disclosure. It should be noted that, Figure 1 The system architecture shown is only an example of a system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.
[0025] As Figure 1As shown, the system architecture 100 according to this embodiment can include a plurality of terminal devices 101, a network 102, and a server 103. The network 102 is a medium to provide a communication link between the terminal devices 101 and the server 103. The network 102 can include various connection types, such as wired and / or wireless communication links, and the like.
[0026] A user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, and the like. The terminal device 101 can be various electronic devices, including but not limited to a smartphone, a tablet computer, a laptop computer, and the like.
[0027] The information labeling method provided by the embodiments of the present disclosure can generally be executed by the server 103. Accordingly, the information labeling apparatus provided by the embodiments of the present disclosure can generally be disposed in the server 103. The information labeling method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103. Accordingly, the information labeling apparatus provided by the embodiments of the present disclosure can also be disposed in a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103.
[0028] Figure 2 is a flowchart of an information labeling method according to one embodiment of the present disclosure.
[0029] As shown, the method 200 can include operation S210 to operation S240. Figure 2
[0030] In operation S210, a plurality of structured data related to a target event and unstructured information related to the target event are matched.
[0031] In the embodiments of the present disclosure, the target event can correspond to a target object.
[0032] For example, the target event can be an event related to the target object. In one example, the target event can be a stock price rising event of a certain company.
[0033] In the embodiments of the present disclosure, the structured data can be obtained from a third-party database.
[0034] For example, a data table containing data of the target object can be obtained from the third database. In one example, the table header fields of the data table include date, subject, stock price, and the like. The data corresponding to the subject table header field can be the name of the target object, such as “a certain company”.
[0035] In the embodiments of the present disclosure, the unstructured information can be a text.
[0036] For example, the unstructured information can be "Today, XX company rose sharply, and closed at 68.12 yuan".
[0037] In the embodiments of the present disclosure, the unstructured information can be an image with text.
[0038] For example, the unstructured information can be an image with text "Today, XX company rose sharply, and closed at 68.12 yuan". In an example, the text can be recognized from the image by using a text recognition technology such as an OCR (Optical Character Recognition) technology, as unstructured information required for subsequent operations.
[0039] In the embodiments of the present disclosure, the plurality of structured data and the unstructured information can be subjected to string matching.
[0040] For example, for each structured data, exact matching can be performed according to the structured data and the unstructured information.
[0041] In an example, the exact matching can refer to a matching mode of judging whether the structured data is completely identical to a segment in the unstructured information. For example, the structured data is "XX company", and it can be judged whether a segment in the unstructured information is completely identical to the structured data.
[0042] In the embodiments of the present disclosure, in response to a failure of the exact matching, a longest common substring of the structured data and the unstructured information can be determined for matching.
[0043] For example, there is no segment in the unstructured information (such as a text) that is completely identical to the structured data (such as "XX company"). In this case, a longest common substring of the structured data and the unstructured information can be determined for matching. In an example, the unstructured information can be "Today, XX technology company rose sharply, and closed at 68.12 yuan". The longest common substring of the unstructured information and the structured data ("XX company") can be "XX company". The longest common substring can be used for matching.
[0044] In some examples, in response to that the structured data and the unstructured information have a longest common substring, it can be directly considered that the matching is successful.
[0045] In the embodiments of the present disclosure, for the plurality of structured data, an extended data set of each structured data can be obtained, to obtain extended data sets of the plurality of structured data.
[0046] For example, for each structured data, the type of the structured data can be determined.
[0047] For example, the structured data type can include a plurality of predetermined types. For example, a first predetermined type can be a date type. For example, a second predetermined type can be a ratio type. For example, a third predetermined type can be an asset type. For example, a fourth predetermined type can be a string type.
[0048] In one example, in response to the type of the structured data being the first predetermined type, an extension data set including a plurality of first type extension data of the structured data can be obtained. In one example, the time information represented by the plurality of first type extension data is consistent with the structured data, and each of the first type extension data can be obtained by performing a format conversion on the structured data. For example, the structured data is “202X-1X-1X, 00:00:00”, and the first type extension data can be “202X year 1X month 1X day”, “today”, “1X month 1X day of this year”, “202X.1X.1X”, and the like.
[0049] In one example, in response to the type of the structured data being the second predetermined type, an extension data set including a plurality of second type extension data of the structured data can be obtained. In one example, the numerical value represented by the plurality of second type extension data is consistent with the structured data, and each of the second type extension data can be obtained by performing a numerical value conversion on the structured data. For example, the structured data is “10%”, and the second type extension data can be “0.1”, “10.00%”, “0.10”, “1x10 -1 ” and the like.
[0050] In one example, in response to the type of the structured data being the third predetermined type, an extension data set including a third type extension data of the structured data can be obtained. In one example, each of the third type extension data can be obtained by performing an asset conversion on the structured data, and the asset information represented by the plurality of third type extension data is consistent with the structured data. For example, the structured data is “¥68.12”, and the third type extension data can be “¥1208” and “$10.66” “HK$82.89” and the like.
[0051] In one example, in response to the type of the structured data being the fourth predetermined type, an extension data set including a fourth type extension data of the structured data can be obtained. In one example, each of the fourth type extension data can be a synonym of the structured data. For example, the structured data is “certain company”, and the fourth type extension data can be “certain technology company”, “certain limited company” and the like. For example, the founder of the “certain company” can be obtained by querying a third-party external knowledge base, and the fourth type extension data can also be “Li’s company”.
[0052] In one example, the fourth predetermined type of structured data can also include a string obtained according to other structured data. For example, a data table including stock prices of yesterday and today is obtained from a third database, and a comparison between the two can obtain the structured data of "stock price rising". The fourth type of extended data corresponding to the structured data can be "rising", "turning red", "greatly rising", and the like.
[0053] In operation S220, in response to successful matching of at least one structured data and unstructured information, the at least one structured data matched successfully with the unstructured information can be taken as target data, and a matching relationship between each target data and at least one unstructured information segment can be obtained.
[0054] For example, the unstructured information can be "Today, XX company greatly rises, and closes at 68.12 yuan". The structured data can be a data table, and the table header fields of the data table include date, subject, stock price, and the like. The structured data corresponding to the date is "202X-1X-1X, 00:00:00", the structured data corresponding to the subject is "XX company", and the structured data corresponding to the stock price is "¥68.12". The matching relationship can be represented as: " '202X-1X-1X, 00:00:00' - 'today' ", " 'XX company' - 'XX company' ", and " '¥68.12' - '68.12 yuan' ".
[0055] In operation S230, according to the matching relationship, a corresponding relationship between at least one unstructured information segment and a label can be established.
[0056] In the embodiments of the present disclosure, the label can correspond to the target data.
[0057] For example, the label can be a table header field corresponding to the target data. Further, according to the matching relationship in the embodiments of operation S220 described above, the corresponding relationship can be: " 'date' - 'today' ", "'subject' - 'XX company' ", and "'stock price' - '68.12 yuan' ".
[0058] For example, the label can also be the structured data itself. In one example, the corresponding relationship described above can also include "'stock price rising' - 'greatly rising' ".
[0059] In the embodiments of the present disclosure, a position of each unstructured information segment in the unstructured information can be obtained.
[0060] For example, the unstructured information can be placed in a plane coordinate system, and the coordinates of each unstructured information segment can be obtained.
[0061] In this embodiment of the disclosure, a correspondence is established between the location of each tag and at least one unstructured information fragment based on the matching relationship.
[0062] For example, a correspondence can be established between each label and the coordinates of at least one unstructured information fragment.
[0063] In operation S240, unstructured information is labeled using tags based on the correspondence.
[0064] For example, an annotation set of unstructured information can be generated directly based on multiple labels and the above correspondences related to the multiple labels, so as to perform annotation.
[0065] The embodiments disclosed herein enable the rapid and low-cost construction of labeled datasets for events. Accurate labeling of fragments or synonym fragments in different data formats within unstructured data is possible.
[0066] Figure 3A This is a schematic diagram of unstructured information according to an embodiment of the present disclosure.
[0067] like Figure 3A As shown, the unstructured information 301 can be a piece of text extracted from a webpage.
[0068] Figure 3B This is a schematic diagram of multiple structured data according to an embodiment of the present disclosure.
[0069] like Figure 3B As shown, multiple structured data can be obtained from third-party databases, such as "202X-1X-1X, 00:00:00" 3021, "Company X" 3022, and "68.12" 3033.
[0070] Figure 3C This is a schematic diagram of multiple structured data according to another embodiment of the present disclosure.
[0071] like Figure 3C As shown, further processing can be performed on multiple structured data obtained from third-party databases to obtain new structured data, such as "Stock Price Increase 3024". The structured data "Stock Price Increase 3024" is obtained by further processing stock prices on different dates.
[0072] Figure 3D This is a schematic diagram of a matching relationship according to an embodiment of the present disclosure.
[0073] like Figure 3D As shown, according to, for example Figure 3A The unstructured information 301 shown and, for example Figure 3BBy matching the multiple structured data shown, we can obtain, for example... Figure 3D The matching relationship shown is 303.
[0074] Figure 3E This is a schematic diagram of the matching relationship according to another embodiment of the present disclosure.
[0075] like Figure 3E As shown, according to, for example Figure 3A The unstructured information 301 shown and, for example Figure 3C By matching the multiple structured data shown, we can obtain, for example... Figure 3E The matching relationship shown is 304.
[0076] Figure 3F This is a schematic diagram illustrating the correspondence according to an embodiment of the present disclosure.
[0077] like Figure 3F As shown, according to, for example Figure 3E The matching relationship 304 shown and the labels corresponding to multiple structured data can yield, for example... Figure 3F The correspondence shown is 305.
[0078] It should be noted that, Figure 3D to Figure 3E The matching relationship shown Figure 3F The correspondences shown are all presented in tabular form. However, the matching or correspondence relationships obtained through the methods of this disclosure can also be represented in other forms.
[0079] It should be noted that, according to, for example Figure 3F The correspondence shown in 305 allows for the labeling of unstructured information. During the labeling process, a single label can be used to label the entire unstructured information, such as using the label "stock price increase" to label unstructured information 301.
[0080] Figure 4 This is a block diagram of an information labeling device according to an embodiment of the present disclosure.
[0081] like Figure 4 As shown, the device 400 may include a matching module 410, an acquisition module 420, an establishment module 430, and an annotation module 440.
[0082] The matching module 410 is used to match multiple structured data related to the target event and unstructured information related to the target event.
[0083] The obtaining module 420 is configured to, in response to successful matching of at least one piece of structured data to the unstructured information, take the at least one piece of structured data successfully matched to the unstructured information as target data, and obtain a matching relationship between each piece of target data and at least one piece of unstructured information segment.
[0084] The establishing module 430 is configured to establish a corresponding relationship between at least one piece of unstructured information segment and a label according to the matching relationship, the label corresponding to the target data.
[0085] The labeling module 440 is configured to label the unstructured information by using the label according to the corresponding relationship.
[0086] In some embodiments, the matching module comprises a first matching submodule configured to perform string matching on the plurality of pieces of structured data and the unstructured information.
[0087] In some embodiments, the first matching submodule comprises a first matching unit configured to, for each piece of structured data, perform exact matching on the structured data and the unstructured information; and a second matching unit configured to, in response to failure of the exact matching, determine a longest common substring of the structured data and the unstructured information for matching.
[0088] In some embodiments, the matching module comprises a first obtaining submodule configured to, for the plurality of pieces of structured data, obtain an extended data set of each piece of structured data to obtain a plurality of extended data sets of the plurality of pieces of structured data; and a second matching submodule configured to perform matching according to the plurality of extended data sets of the plurality of pieces of structured data and the unstructured information.
[0089] In some embodiments, the first obtaining module comprises: a determining unit, configured to determine, for each structured data, a type of the structured data; a first obtaining unit, configured to, in response to the type of the structured data being a first predetermined type, obtain an extension data set comprising a plurality of first type extension data of the structured data, wherein the plurality of first type extension data represent time information consistent with the structured data, and each first type extension data is obtained by performing a format conversion on the structured data once; a second obtaining unit, configured to, in response to the type of the structured data being a second predetermined type, obtain an extension data set comprising a plurality of second type extension data of the structured data, wherein the plurality of second type extension data represent numerical values consistent with the structured data, and each second type extension data is obtained by performing a numerical value conversion on the structured data once; a third obtaining unit, configured to, in response to the type of the structured data being a third predetermined type, obtain an extension data set comprising third type extension data of the structured data, wherein each third type extension data is obtained by performing an asset conversion on the structured data once; and a fourth obtaining unit, configured to, in response to the type of the structured data being a fourth predetermined type, obtain an extension data set comprising fourth type extension data of the structured data, wherein each fourth type extension data is a synonym of the structured data.
[0090] In some embodiments, the establishing module comprises: a second obtaining sub-module, configured to obtain a position of each unstructured information segment in the unstructured information; and an establishing sub-module, configured to establish a corresponding relationship between each label and the position of at least one unstructured information segment according to the matching relationship.
[0091] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0092] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0093] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0094] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0095] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0096] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as information annotation methods. For example, in some embodiments, the information annotation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the information annotation method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the information annotation method by any other suitable means (e.g., by means of firmware).
[0097] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0098] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.
[0099] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0100] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0101] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0102] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0103] It should be understood that various forms of flow shown above can be used, re-ordered, added to, or deleted from without departing from the spirit of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.
[0104] The specific embodiments described above have been shown by way of example, and anyone skilled in the art should understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Accordingly, the scope of the present disclosure is not intended to be limited to the particular embodiments described above.
Claims
1. A method for information labeling, comprising: matching a plurality of structured data related to a target event and unstructured information related to the target event, wherein the plurality of structured data are a plurality of event information in a data table; in response to at least one structured data matching the unstructured information successfully, taking the at least one structured data matching the unstructured information successfully as target data to obtain a matching relationship between each target data and at least one unstructured information segment; establishing a corresponding relationship between at least one unstructured information segment and a label corresponding to the target data according to the matching relationship; and labeling the unstructured information using the label according to the corresponding relationship to obtain a labeled data set of the target event; wherein the structured data comprises date type structured data; the method further comprises: processing structured data of different dates to obtain new structured data of the target event as new event information of the target event.
2. The method of claim 1, wherein, The matching the plurality of structured data related to a target event and unstructured information related to the target event comprises: performing string matching on the plurality of structured data and the unstructured information.
3. The method of claim 2, wherein, The string matching comprises: for each structured data, performing exact matching on the structured data and the unstructured information; in response to the exact matching failing, determining a maximum common substring of the structured data and the unstructured information to perform matching.
4. The method according to any one of claims 1 to 3, wherein, The matching the plurality of structured data related to a target event and unstructured information related to the target event comprises: for the plurality of structured data, obtaining an extended data set of each structured data to obtain a plurality of extended data sets of the structured data; performing matching according to the plurality of extended data sets of the structured data and the unstructured information.
5. The method of claim 4, wherein, The obtaining an extended data set of each structured data comprises: for each structured data, determining a type of the structured data; in response to the type of the structured data being a first predetermined type, obtaining an extended data set comprising a plurality of first type extended data containing the structured data, wherein the plurality of first type extended data represent time information consistent with the structured data, and each first type extended data is obtained by performing a format conversion on the structured data once; in response to the type of the structured data being a second predetermined type, obtaining an extended data set comprising a plurality of second type extended data containing the structured data, wherein the plurality of second type extended data represent numerical values consistent with the structured data, and each second type extended data is obtained by performing a numerical value conversion on the structured data once; in response to the type of the structured data being a third predetermined type, obtaining an extended data set comprising a third type extended data containing the structured data, wherein each third type extended data is obtained by performing an asset conversion on the structured data once; In response to the type of the structured data being a fourth predetermined type, an extension data set containing fourth type extension data of the structured data is obtained, wherein each fourth type extension data is a synonym of the structured data.
6. The method of claim 1, wherein, The establishing, according to the matching relationship, of the corresponding relationship between at least one unstructured information segment and a label comprises: obtaining a position of each unstructured information segment in the unstructured information; establishing, according to the matching relationship, a corresponding relationship between each label and the position of at least one unstructured information segment.
7. An information labeling apparatus, comprising: a matching module configured to match a plurality of structured data related to a target event and unstructured information related to the target event, wherein the plurality of structured data is a plurality of event information in a data table; an obtaining module configured to, in response to successful matching of at least one structured data and the unstructured information, obtain a matching relationship between each target data and at least one unstructured information segment, wherein the target data is at least one structured data successfully matched with the unstructured information; an establishing module configured to establish, according to the matching relationship, a corresponding relationship between at least one unstructured information segment and a label corresponding to the target data; and a labeling module configured to label the unstructured information using the label according to the corresponding relationship to obtain a labeled data set of the target event. The structured data comprises date type structured data; and the apparatus further comprises: a processing module configured to process structured data of different dates to obtain new structured data of the target event as new event information of the target event.
8. The apparatus of claim 7, wherein, The matching module comprises: a first matching sub-module configured to perform string matching on the plurality of structured data and the unstructured information.
9. The apparatus of claim 8, wherein, The first matching sub-module comprises: a first matching unit configured to, for each structured data, perform exact matching on the structured data and the unstructured information; and a second matching unit configured to, in response to failure of the exact matching, determine a maximum common substring of the structured data and the unstructured information to perform matching.
10. The apparatus of any one of claims 7 to 9, wherein, The matching module comprises: a first obtaining sub-module configured to, for the plurality of structured data, obtain an extension data set of each structured data to obtain a plurality of extension data sets of the plurality of structured data; and a second matching sub-module configured to perform matching according to the plurality of extension data sets of the plurality of structured data and the unstructured information.
11. The apparatus of claim 10, wherein, The first obtaining sub-module comprises: a determining unit configured to, for each structured data, determine a type of the structured data; a first obtaining unit configured to, in response to the type of the structured data being a first predetermined type, obtain an extension data set containing a plurality of first type extension data of the structured data, wherein time information represented by the plurality of first type extension data is consistent with the structured data, and each first type extension data is obtained by performing a format conversion on the structured data once. a second obtaining unit, configured to, in response to the type of the structured data being a second predetermined type, obtain an extension data set containing a plurality of second-type extension data of the structured data, wherein the numerical values represented by the plurality of second-type extension data are consistent with the structured data, and each second-type extension data is obtained by performing a numerical conversion on the structured data; a third obtaining unit, configured to, in response to the type of the structured data being a third predetermined type, obtain an extension data set containing third-type extension data of the structured data, wherein each third-type extension data is obtained by performing an asset conversion on the structured data; a fourth obtaining unit, configured to, in response to the type of the structured data being a fourth predetermined type, obtain an extension data set containing fourth-type extension data of the structured data, wherein each fourth-type extension data is a synonym of the structured data.
12. The apparatus of claim 7, wherein, The establishing module comprises: a second obtaining sub-module, configured to obtain the position of each unstructured information segment in the unstructured information; an establishing sub-module, configured to establish a corresponding relationship between each label and the position of at least one unstructured information segment according to the matching relationship. 13.An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.
14. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to make the computer execute the method according to any one of claims 1 to 6. 15.A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Knowledge data providing method and device, electronic equipment and storage medium
CN109739964A
Automated generation of structured training data from unstructured documents
US20210248420A1