Training method, data processing method, device, equipment, medium and program product

By combining the code tags and location information of the web page source code, using pre-trained models to construct input vectors and train the data processing model, the problems of low compatibility and high training cost in the existing technology are solved, and efficient text element extraction is achieved.

CN114443931BActive Publication Date: 2025-07-25CHINA CONSTRUCTION BANK +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210189215.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-07-25
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

In the prior art, handwriting rules have low compatibility and low efficiency, while plain text training deep learning models rely on large-scale manual data annotations, which are highly trained.

Method used

By obtaining the code tags and text content included in the source code of the web page, combining location information, using a pre-trained natural language processing model, constructing input vectors, training a bidirectional long and short-term memory network and a fully connected layer data processing model, and automatically extracting text elements.

Benefits of technology

It reduces training costs, improves the compatibility and efficiency of data processing, and avoids the need to manually label large amounts of plain text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114443931B_ABST
    Figure CN114443931B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method, apparatus, device, storage medium, and program product for a data processing model. The method includes: obtaining a first web page, wherein a source code of the first web page includes a first code tag and first text content to be processed; combining at least one of a second code tag associated with a first text segment and first position information with the text of the first text segment to obtain a first input vector; using the first input vector and an element tag of the first text segment as training samples to train the data processing model. Embodiments of the present disclosure can reduce training costs and are no longer limited to the method of formulating a data extraction rule for one type of web page in the prior art, improving compatibility and processing efficiency. The present disclosure also provides a data processing method, apparatus, device, storage medium, and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more particularly to a training method, a data processing method, an apparatus, a device, a medium, and a program product. Background Art

[0002] When obtaining text information on the Internet, not only the text information itself is obtained, but also the structured metadata of the text itself is required.

[0003] In related technologies, information crawling can be performed by writing rules manually. For example, for different web pages, different regular expressions are used to extract data. It is also possible to pre-process the text into plain text, and then manually label the structured metadata in the plain text to train a deep learning model.

[0004] The above-mentioned method of writing rules manually has low compatibility, and different regular expressions need to be written for different web pages, resulting in low efficiency. And the method of training a deep learning model using plain text depends on a large amount of manual data annotation, and the training cost is relatively high. Therefore, it is an urgent problem to be solved to propose an automated data processing method with both low cost and high efficiency. Summary of the Invention

[0005] In view of the above problems, the present disclosure provides a training method, a data processing method, an apparatus, a device, a medium, and a program product for improving data processing efficiency.

[0006] In one aspect of the embodiments of the present disclosure, a method for training a data processing model is provided, including: obtaining a first web page, where the source code of the first web page includes a first code tag and a first text content to be processed, the first text content includes M text segments, and M is an integer greater than or equal to 1; combining at least one of the second code tag associated with the first text segment and the first position information with the text of the first text segment to obtain a first input vector, where the first text segment is any one of the M text segments, the first code tag includes the second code tag, and the first position information is the position information of the first text segment in the M text segments; using the first input vector and the element tag of the first text segment as training samples to train the data processing model.

[0007] According to an embodiment of the present disclosure, the method further includes obtaining the element tag of the first text segment, specifically including: determining the element category of the first text segment; determining the second position information of the first text segment in the element category; and annotating the element tag based on the element category and the second position information.

[0008] According to an embodiment of the present disclosure, obtaining the first input vector includes obtaining a text vector, specifically including: inputting the text of the first text segment into a pre-trained model, where the pre-trained model includes a pre-trained natural language processing model; obtaining the text vector output by the pre-trained model.

[0009] According to an embodiment of the present disclosure, obtaining the first input vector includes obtaining a code label vector, specifically including: determining S first code labels associated with the first text content, where the S first code labels include the second code label; performing vector encoding on each of the S first code labels; obtaining the code label vector of the second code label according to the result of the vector encoding.

[0010] According to an embodiment of the present disclosure, obtaining the first input vector includes obtaining a position vector, specifically including: determining the first order of the first text segment among the M text segments, where the first position information includes the first order; obtaining the position vector based on the first order.

[0011] According to an embodiment of the present disclosure, the data processing model includes a bidirectional long short-term memory network layer and a fully connected layer. Training the data processing model includes: using the first input vector as the input of the bidirectional long short-term memory network layer; using the output of the bidirectional long short-term memory network layer as the input of the fully connected layer; calculating a loss function based on the output of the fully connected layer and the element label, and updating the parameters of the data processing model according to the loss function, where the output of the fully connected layer includes the predicted element label of the first text segment.

[0012] According to an embodiment of the present disclosure, the data processing model further includes a normalization layer. Before using the first input vector as the input of the bidirectional long short-term memory network layer, it further includes: performing normalization processing on the first input vector through the normalization layer.

[0013] According to an embodiment of the present disclosure, the data processing model further includes a dropout layer. Before using the output of the bidirectional long short-term memory network layer as the input of the fully connected layer, it further includes: processing the output of the bidirectional long short-term memory network layer through the dropout layer.

[0014] According to an embodiment of the present disclosure, the source code of the first web page is obtained using hypertext markup language, and the first code label includes a hypertext markup language label.

[0015] Another aspect of the embodiments of the present disclosure provides a data processing method, including: obtaining a second web page, where the source code of the second web page includes a third code tag and second text content to be processed; inputting the second text content and the third code tag into a data processing model, where the data processing model is obtained by training through the method described above; processing the text of the second text segment according to the predicted element tag of the second text segment output by the data processing model, where the second text content includes at least one text segment, and the second text segment is any one of the at least one text segment.

[0016] Another aspect of the embodiments of the present disclosure provides a training device for a data processing model, including: a first obtaining module, configured to obtain a first web page, where the source code of the first web page includes a first code tag and first text content to be processed, the first text content includes M text segments, and M is an integer greater than or equal to 1; an input vector module, configured to combine at least one of the second code tag associated with the first text segment and the first position information with the text of the first text segment to obtain a first input vector, where the first text segment is any one of the M text segments, the first code tag includes the second code tag, and the first position information is the position information of the first text segment among the M text segments; a model training module, configured to use the first input vector and the element tag of the first text segment as training samples to train the data processing model.

[0017] Another aspect of the embodiments of the present disclosure provides a data processing device, including: a second obtaining module, configured to obtain a second web page, where the source code of the second web page includes a third code tag and second text content to be processed; a data input module, configured to input the second text content and the third code tag into a data processing model, where the data processing model is obtained by training through a training device; a data processing module, configured to process the text of the second text segment according to the predicted element tag of the second text segment output by the data processing model, where the second text content includes at least one text segment, and the second text segment is any one of the at least one text segment.

[0018] Another aspect of the embodiments of the present disclosure provides an electronic device, including: one or more processors; a storage device, configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method described above.

[0019] Another aspect of the embodiments of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method described above.

[0020] On the other hand, an embodiment of the present disclosure further provides a computer program product, including a computer program which, when executed by a processor, implements the method described above.

[0021] The above-mentioned one or more embodiments have the following beneficial effects:

[0022] Through the training method of the embodiments of the present disclosure, by combining code labels, position information, and integrating the semantics of natural language, a data processing model with strong compatibility can be trained with a small amount of labeled data. Specifically, the method of using the first position information of the first text segment in the first text content and / or the second code label to combine the text of the first text segment to construct the first input vector and train to obtain the data processing model can avoid manual element label annotation of a large amount of pure text, reduce the training cost, and is no longer limited to the method of formulating a data extraction rule for one type of web page in the prior art, improving the compatibility and processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0024] Figure 1 Schematically shows an application scenario diagram of the training method or data processing method according to an embodiment of the present disclosure;

[0025] Figure 2 Schematically shows a flowchart of the training method of the data processing model according to an embodiment of the present disclosure;

[0026] Figure 3 Schematically shows a flowchart of obtaining a text vector according to an embodiment of the present disclosure;

[0027] Figure 4 Schematically shows a flowchart of obtaining a code label vector according to an embodiment of the present disclosure;

[0028] Figure 5 Schematically shows a flowchart of obtaining a position vector according to an embodiment of the present disclosure;

[0029] Figure 6 Schematically shows a flowchart of obtaining an element label according to an embodiment of the present disclosure;

[0030] Figure 7 Schematically shows a flowchart of training a data processing model according to an embodiment of the present disclosure;

[0031] Figure 8 Schematically shows a flowchart of the data processing method according to an embodiment of the present disclosure;

[0032] Figure 9 Schematically shows a structural block diagram of a training device according to an embodiment of the present disclosure;

[0033] Figure 10 Schematically shows a structural block diagram of a data processing device according to an embodiment of the present disclosure;

[0034] Figure 11 Schematically shows a block diagram of an electronic device suitable for implementing a training method or a data processing method according to an embodiment of the present disclosure. Detailed implementation manners

[0035] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0036] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0037] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0038] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0039] For the text information disclosed on a web page, such as news, policy documents, court judgments, financial reports, etc., different styles have different structured metadata. The structured metadata can be embodied in the form of text elements. Taking news as an example, its text elements can include a title, a release time, a release source, etc. Taking a policy document as an example, its text elements can include - a title, a document number, a release date, a release agency, an accepting unit, etc.

[0040] When obtaining information from a web page, the text information itself and the text elements to which different texts belong can be obtained simultaneously.

[0041] If the method of using handwriting rules is adopted, rules may be set specifically for different styles of writing, and may also be set specifically for different web pages of the same style. Even if the languages used for programming the web page source code are the same, rules may need to be reset because of different tags used in that language.

[0042] If a deep learning model is trained using plain text, its disadvantage is that the plain text loses information such as the original source code structure, font size, and code tags of the web page. Therefore, a large amount of training data is required to obtain a model with better robustness, and the training cost is too high.

[0043] The embodiments of the present disclosure use a data processing model trained based on a neural network by integrating the structure, tags, position information of the original web page source code and the semantic information of the text itself to automatically extract text elements. The trained model can be used for extracting structured information from network texts and can also be used for information entry in the text management terminal.

[0044] Figure 1 An application scenario diagram of the training method or data processing method according to the embodiments of the present disclosure is schematically shown.

[0045] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0046] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0047] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0048] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0049] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0050] are merely illustrative. According to implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 - 8 scenario described below, the training method and data processing method of the data processing model of the embodiments of the present disclosure will be described in detail through

[0051] Figure 2 FIG. schematically shows a flowchart of a training method of a data processing model according to an embodiment of the present disclosure.

[0052] As Figure 2 shown, the training method of the data processing model of this embodiment includes operation S210 to operation S230.

[0053] In operation S210, a first web page is obtained, where the source code of the first web page includes a first code tag and first text content to be processed, and the first text content includes M text segments, and M is an integer greater than or equal to 1;

[0054] According to an embodiment of the present disclosure, the source code of the first web page can be obtained using HyperText Markup Language (HTML), and the first code tag includes HyperText Markup Language tags. HyperText Markup Language is a standard markup language for creating web pages. Specifically, HTML is a basic technology that is often used with CSS and JavaScript by many websites to design the user interfaces of web pages, web applications, and mobile applications. A web browser can read an HTML file and render it into a visual web page. HTML describes the structural semantics of a website as it is presented along with the thread, making it a markup language.

[0055] In some embodiments, the source code of the first web page can also be obtained using languages such as XHTML, Jade, Haml, Slim, Xml, etc.

[0056] Taking an HTML web page as an example, a web page may include body information (i.e., the first text content), and may also include information such as headers, footers, images, script codes, CSS styles, advertisement links, etc. These information can be represented by HTML. For example, "a paragraph of text" can be represented as " A paragraph ". " ” is a first code label. The text segment of the first text content can be determined according to the code label. For example, both ends of the above "a paragraph" have " ”, it can be considered as a text segment.

[0057] In operation S220, at least one of the second code label associated with the first text segment and the first position information is combined with the text of the first text segment to obtain a first input vector, where the first text segment is any one of the M text segments, the first code label includes the second code label, and the first position information is the position information of the first text segment among the M text segments;

[0058] Exemplarily, the first text content is associated with multiple first code labels, and each text segment is associated with some of these code labels. For example, at the beginning and end of the first text segment, there is " ”, then " ” is the second code label of the first text segment.

[0059] Exemplarily, the first input vector can be obtained based on the second code label and text of the first text segment, can also be obtained based on the first position information and text of the first text segment, or can be obtained based on the second code label, the first position information and text of the first text segment.

[0060] In operation S230, the first input vector and the element label of the first text segment are used as training samples to train the data processing model.

[0061] Exemplarily, the first input vector can be input into the data processing model, and the predicted element label is output. The purpose of training is to make the predicted element label the same as the element label, so that the data processing model can accurately determine which text element each text segment belongs to.

[0062] According to the embodiments of the present disclosure, the method of constructing the first input vector by using the first position information of the first text segment in the first text content and / or the second code label representing the first text segment, and training to obtain the data processing model can avoid manual element label annotation for a large amount of pure text, reduce the training cost, and enable the data processing model to learn the relationship between the text segment and the code label, the text semantics between text segments, and the position relationship during the training process, so as to no longer be limited to the method of formulating a data extraction rule for a single type of web page in the prior art, improving the compatibility and processing efficiency.

[0063] Figure 3 Schematically shows a flowchart of obtaining a text vector in operation S220 according to an embodiment of the present disclosure.

[0064] As Figure 3 shown, obtaining the first input vector in operation S220 includes obtaining a text vector, specifically including operations S310 to S320.

[0065] In operation S310, the text of the first text segment is input into a pre-trained model, where the pre-trained model includes a pre-trained natural language processing model;

[0066] Exemplarily, the pre-trained model can include a BERT model, a RoBerta-wwm-ext-large model, an ERNIE model, a NEZHA model, or an XLNet model, etc.

[0067] In operation S320, the text vector output by the pre-trained model is obtained.

[0068] Exemplarily, the text of the first text segment can be encoded using the text encoding ID of the Bert model. After encoding the text data, the Bert model can extract text semantic information to obtain a text vector.

[0069] According to an embodiment of the present disclosure, the process of obtaining a text vector using a pre-trained model can incorporate the understanding of the text context to extract text semantic information, thereby improving the overall recognition effect.

[0070] Figure 4 Schematically shows a flowchart of obtaining a code label vector in operation S220 according to an embodiment of the present disclosure.

[0071] As Figure 4 shown, obtaining the first input vector in operation S220 includes obtaining a code label vector, specifically including operations S410 to S430.

[0072] In operation S410, determine S first code labels associated with the first text content, where the S first code labels include a second code label; specifically, the source code includes the S first code labels and may also include other first code labels.

[0073] In some embodiments, since there are some other descriptive contents in the HTML web page, such as css styles and scripts, etc. They can be removed to retain the first text content and the associated code labels. The removal result is as follows (only for example).

[0074] <h4>Notice of the Office of XX Bureau on Strengthening XX Work< / h4> <h4>

[0075] XX Zi

[2021] No. 30

[0076] Bureaus of XX, Centers of XX in all provinces, autonomous regions, and municipalities directly under the Central Government

[0077] Strengthening patent navigation by industrial field is an important part of the work of promoting the application of intellectual property rights, and is of great significance for improving innovation efficiency, saving innovation costs, and strengthening patent protection.

[0078] Notice is hereby given.

[0079] Office of XX Bureau

[0080] X month X day, 20XX

[0081] Referring to the above example, after elimination, it can be seen that the title uses the h4 tag in html, the document number uses the p tag, and the paragraphs use the p tag. For the last paragraph, the span tag is used. The h4, p, and span tags are three first code tags associated with the first text content. Taking the first text segment as "Notice of XX Bureau Office on Strengthening XX Work" as an example, the h4 tag is the second code tag associated with the first text segment. Each text segment can be determined based on whether there are HTML tags at both ends of a piece of text.

[0082] In some embodiments, if the title uses the h5 tag, then the regular expression for h4 may fail, resulting in the reconstruction of the regular expression. However, using the data processing model of the embodiments of the present disclosure can avoid the problem of low compatibility.

[0083] According to an embodiment of the present disclosure, S types of first code tags can be as shown in Table 1 (for example only).

[0084] Table 1

[0085]

[0086] In operation S420, vector encoding is performed on each of the S types of first code tags;

[0087] In some embodiments, the one-hot encoding method can be used to implement operation S420. One-hot encoding uses an S-bit status register to encode the S types of first code tags. Each first code tag has its own independent register bit, that is, only one bit is 1 and the rest are zero values. Referring to Table 1, there are 11 types of first code tags in total, so there will be corresponding encoding results for each first code tag. For example, the encoding result of h1 is (1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0). In other embodiments, the word embedding method can be used to implement operation S420.

[0088] In operation S430, according to the result after vector encoding, the code tag vector of the second code tag is obtained.

[0089] Taking the "Notice of XX Bureau Office on Strengthening XX Work" mentioned above as an example, the code label vector of its associated h4 label can use the result of one-hot encoding, that is, (0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0). Or the result of one-hot encoding can be input into a pre-trained model to further extract semantic information, and the output of the pre-trained model is used as the code label vector. 11 categories of labels are shown in Table 1. Therefore, for each text segment, we have a 1*11 dimensional vector.

[0090] According to an embodiment of the present disclosure, by obtaining a code label vector as a component of the first input vector, the data processing model can learn the connection between the code label and the text content. Compared with the pure text training method that removes the code label, the training cost and manual annotation cost are reduced, and the model robustness is improved.

[0091] Figure 5 Schematically shows a flowchart of obtaining a position vector in operation S220 according to an embodiment of the present disclosure.

[0092] Such as Figure 5 shown, obtaining the first input vector in operation S220 may include a position vector, specifically including operations S510 to S520.

[0093] In operation S510, determine the first order of the first text segment among M text segments, where the first position information includes the first order;

[0094] Exemplarily, the first order may be the sorting of the first text segment among M text segments, that is, which text segment it is. The above-mentioned "Notice of XX Bureau Office on Strengthening XX Work" is the first text segment, and the order is 1.

[0095] In operation S520, obtain a position vector based on the first order.

[0096] Exemplarily, the calculation method of the relative position information may be: text segment order / all text segment lengths. For example, if M is 10 in total, then the relative position encoding of the first text segment is 0.1, the relative position encoding of the second text segment is 0.2, and the relative position encoding of the last text segment is 1. The relative position encoding can be directly used as a 1*1 dimensional position vector, or the relative position encoding can be input into a pre-trained model for encoding to obtain a 1*1 dimensional position vector.

[0097] According to an embodiment of the present disclosure, the position vector of each text segment is used as a component of the first input vector, so that the data processing model can learn the connection between the position information and the text content. In some embodiments, the connection between the code label, the position information, and the text content can also be learned. This reduces the training cost and the manual annotation cost, and improves the model robustness.

[0098] According to the embodiments of the present disclosure, referring to Figures 3 - 5 , the following is an example of obtaining the first input vector by combining HTML tags, position information and text content.

[0099] First, the HTML tags, location information and text content are vectorized. For the above-mentioned "Notice of the Office of the XX Bureau on Strengthening the Work of XX", the following three input vectors can be obtained (only as an example).

[0100] Html tag vector input: (0, 0, 0, 1, 0,0,0,0,0, 0, 0)

[0101] Relative position information input: (0.1)

[0102] Bert dictionary encoding input: (101, 1744, 2157, 4761, ..., 4761, 102)

[0103] Taking Bert dictionary encoding input as an example, Bert's dictionary encoding input is a simple vocabulary query. For example, the code corresponding to "知" is 4761, so "知" intellectual property rights and "通" are the same ID; 101 represents the beginning of a sentence, and 102 represents the end of a sentence (just an example).

[0104] Secondly, extract text semantic information. For example, directly load the pre-trained language model weights of Bert to extract semantic information from all input texts. Bert can extract an abstract sentence text vector representation as a 768-dimensional vector.

[0105] At this point, the above three input vectors become:

[0106] Html tag vector: 1*11 dimension vector

[0107] Relative position vector: 1*1 dimension vector

[0108] Bert semantic vector: 1*768 dimensional vector

[0109] Finally, the three input vectors are concatenated together horizontally to construct a 1*780 dimensional vector, which is the first input vector.

[0110] Figure 6 The flowchart of obtaining element labels according to an embodiment of the present disclosure is schematically shown.

[0111] like Figure 6 As shown, the process of obtaining the element tag of the first text segment in this embodiment includes operations S610 to S630.

[0112] In operation S610, determining the element category of the first text segment;

[0113] Exemplarily, the element category may refer to a type of text element. For example, the element category of the above-mentioned "Notice of the Office of the XX Bureau on Strengthening XX Work" is the title element. The above-mentioned "XX Bureau, XX Center of each province, autonomous region, municipality directly under the Central Government" is the receiving unit element.

[0114] In operation S620, second position information of the first text segment in the element category is determined;

[0115] Exemplarily, the first text content may include text elements such as title, document number, publication date, publishing agency, and receiving unit. Each element category may include one or more text segments. In the case of including multiple text segments, there will be a sequence between the text segments. The second position information is the position information of the first text segment in the element category to which it belongs. The second position information may include the beginning, the middle, the end, or the whole. Among them, the whole means that the element category only includes one text segment.

[0116] In operation S630 , the feature tags are annotated based on the feature category and the second location information.

[0117] Exemplarily, the marking methods of B, I, E, O and S can be used, that is, if the text segment is the beginning of an element, it is a B-element, the middle part of an element is an I-element, the end of an element is an E-element, a single sentence is the entire element (i.e., the whole), it is an S-element; if it does not belong to any element, it is O. The element label of the above-mentioned "Notice of the Office of the XX Bureau on Strengthening XX Work" can be S-title. The element label of the above-mentioned "Hereby Notice" can be O. In some embodiments, if the title part includes three text segments, the element labels of the three text segments include B-title, I-title and E-title.

[0118] For example, if the five text elements are title, document number, release date, issuing agency, and receiving unit, there may be 5*4+1 combinations (each text element may be combined with B, I, E, and S. If it does not belong to a text element, it is marked with O alone), a total of 21 element label forms. Therefore, each text segment will have a 1*21 dimension, a vector output that adds up to 1, and each position in the vector represents the probability of the output being the label.

[0119] According to an embodiment of the present disclosure, accurate annotation of element labels can improve the training effect, enabling the data processing model to have high prediction accuracy.

[0120] Figure 7 A flowchart of training a data processing model according to an embodiment of the present disclosure is schematically shown.

[0121] As Figure 7 shown, the data processing model includes a bidirectional long short-term memory network layer (i.e., BI-LSTM layer) and a fully connected layer (i.e., Dense layer). The training data processing model of this embodiment includes operation S720, operation S740, operation S750, and operation S760.

[0122] In operation S720, the first input vector is used as the input of the bidirectional long short-term memory network layer;

[0123] Exemplarily, a bidirectional LSTM recurrent neural network structure (BI-LSTM) is used in operation S720. As a recurrent neural network that can be used to process ultra-long sequences, LSTM can remember information before the sequence input, while BI-LSTM can record information after the input and is a bidirectional recurrent neural network. The input of the BI-LSTM layer is a 1*780-dimensional vector, and the output is a 1*512-dimensional vector.

[0124] In operation S740, the output of the bidirectional long short-term memory network layer is used as the input of the fully connected layer. Exemplarily, the input of the Dense layer is a 1*512-dimensional vector, and the output is a 1*21-dimensional vector.

[0125] In operation S750, a loss function is calculated based on the output of the fully connected layer and the element label to update the parameters of the data processing model according to the loss function. Among them, the output of the fully connected layer includes the predicted element label of the first text segment.

[0126] Exemplarily, the Dense layer can use the softmax activation function and will output a 21-dimensional vector whose values add up to 1. The value at each position in the 21-dimensional vector represents the probability of a certain element label, and the predicted element label of the first text segment can be determined from the maximum probability and output.

[0127] In operation S760, the parameters of the data processing model are updated. For example, the weight coefficients in the BI-LSTM layer and the Dense layer are updated. Among them, the loss function can adopt cross-entropy, and then training can be performed according to the convergence degree of the cross-entropy. In some embodiments, the loss function can also adopt functions such as mean square error or exponential loss function.

[0128] As Figure 7 As shown, the data processing model further includes a normalization layer, and the training data processing model of this embodiment may include operation S710.

[0129] Exemplarily, operation S710 is executed before operation S720. After the first input vector is normalized by the normalization layer, it is input into the bidirectional long short-term memory network layer. The role of the normalization layer is to normalize the input features to facilitate the training convergence of the model.

[0130] As Figure 7 shown, the data processing model further includes a dropout layer (i.e., Dropout layer), and the training data processing model of this embodiment further includes operation S730. Operation S730 is executed before operation S740 to process the output of the bidirectional long short-term memory network layer so as to be input into the fully connected layer. The Dropout layer can prevent the model from overfitting.

[0131] It should be noted that Figure 7 the structure of the data processing model in

[0132] Figure 8 is only an example. Without departing from the concept of the present invention, part or all of the network model structure can be replaced. For example, the BI-LSTM layer can be replaced by other common neural network layers such as GRU and RNN; the normalization layer is not a necessary network layer, but it is very helpful for network convergence.

[0133] As Figure 8 shown, the data processing method of this embodiment includes operations S810 to S830.

[0134] In operation S810, at least one second web page is obtained, where the source code of the second web page includes a third code tag and the second text content to be processed;

[0135] Exemplarily, the second web page can obtain the source code in the same language as the first web page, such as using hypertext markup language, and the third code tag is an HTML tag. The second text content and the first text content can be of the same type of style and have the same or similar text elements.

[0136] In operation S820, the second text content and the third code tag are input into the data processing model, where the data processing model can be trained by the method described above Figures 2 - 7 and obtained.

[0137] Exemplarily, in the same way as obtaining the first input vector, at least one of the third code tag associated with each text segment in the second text content and the corresponding position information is combined with the text data of the corresponding text segment to obtain a second input vector. Refer to Figure 7 Input the second input vector into the trained data processing model to obtain an output result. When using the trained data processing model to process data, the dropout layer can be cancelled, and there is no longer a need to obtain the loss function.

[0138] In operation S830, process the text of the second text segment according to the predicted element labels of the second text segment output by the data processing model, where the second text content includes at least one text segment, and the second text segment is any one of the at least one text segment.

[0139] Exemplarily, the predicted element labels of the second text segment may include the element category to which it belongs and location information. For example, when the output is (0.1, 0.8, 0.05, …, 0.002), it means that the probability of the second position is the highest, and the second position corresponds to I - title, indicating that this passage should be the middle part of the title element. Finally, the B - elements, I - elements, and E - elements of the same element category in the second text content can be concatenated. Or directly use S as an element category, and the respective text elements and corresponding text contents in the second web page can be obtained. In some embodiments, the extracted information can also be used for data entry in the text management terminal.

[0140] According to an embodiment of the present disclosure, a pre - trained data processing model can be utilized. The input of the model is the source code of a web page, and the output can be structured text elements. It can cover and process a relatively large range of second web pages. Even for second web pages with different code tags or different styles, efficient processing can be performed.

[0141] Based on the above - mentioned training method and data processing method, the present disclosure also provides a training device and a data processing device. The following will be described in detail in conjunction with Figure 9 and Figure 10 for a detailed description.

[0142] Figure 9 Schematically shows a structural block diagram of a training device 900 according to an embodiment of the present disclosure.

[0143] As Figure 9 shown, the training device 900 for the data processing model in this embodiment includes a first acquisition module 910, an input vector module 920, and a model training module 930.

[0144] The first acquisition module 910 can perform operation S210 to obtain a first web page, where the source code of the first web page includes a first code tag and first text content to be processed, and the first text content includes M text segments, where M is an integer greater than or equal to 1;

[0145] The input vector module 920 may perform operation S220 to combine at least one of the second code label associated with the first text segment and the first position information with the text of the first text segment to obtain a first input vector, where the first text segment is any one of the M text segments, the first code label includes the second code label, and the first position information is the position information of the first text segment among the M text segments;

[0146] According to an embodiment of the present disclosure, the input vector module 920 may also perform operations S310 to S320, operations S410 to S430, and operations S510 to S520, which will not be elaborated herein.

[0147] The model training module 930 may perform operation S230 to use the first input vector and the feature label of the first text segment as training samples to train the data processing model. The model training module 930 may also perform operations S710 to S760, which will not be elaborated herein.

[0148] According to an embodiment of the present disclosure, the training device 900 may further include a feature label module, and the feature label module may perform operations S610 to S630, which will not be elaborated herein.

[0149] According to an embodiment of the present disclosure, the training device 900 can combine the input of code labels, the input of relative positions, and text semantics to construct and train a deep learning model, greatly reducing the training cost and manual annotation cost, and improving the model robustness.

[0150] Figure 10 The structural block diagram of a data processing device 1000 according to an embodiment of the present disclosure is schematically shown.

[0151] As Figure 10 shown, the data processing device 1000 of this embodiment includes a second acquisition module 1010, a data input module 1020, and a data processing module 1030.

[0152] The second acquisition module 1010 may perform operation S810 to acquire a second web page, where the source code of the second web page includes a third code label and a second text content to be processed;

[0153] The data input module 1020 may perform operation S820 to input the second text content and the third code label into the data processing model, where the data processing model is obtained by training with the training device 900;

[0154] The data processing module 1030 may perform operation S830 to process the text of the second text segment according to the predicted element tags of the second text segment output by the data processing model, where the second text content includes at least one text segment, and the second text segment is any one of the at least one text segment.

[0155] It should be noted that the implementation manners, the technical problems solved, the functions achieved, and the technical effects achieved by each module / unit / sub-unit, etc. in some embodiments of the device are the same as or similar to those of the corresponding steps in some embodiments of the method, and will not be elaborated herein.

[0156] According to an embodiment of the present disclosure, any plurality of modules in the training device 900 or the data processing device 1000 may be combined and implemented in one module, or any one of them may be split into multiple modules. Or, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module.

[0157] According to an embodiment of the present disclosure, at least one module in the training device 900 or the data processing device 1000 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or may be implemented in any one of the three implementation manners of software, hardware, and firmware or in any appropriate combination of several of them. Or, at least one module in the training device 900 or the data processing device 1000 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function may be executed.

[0158] Figure 11 A block diagram of an electronic device suitable for implementing the training method or the data processing method according to an embodiment of the present disclosure is schematically shown.

[0159] As Figure 11 As shown, the electronic device 1100 according to an embodiment of the present disclosure includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage section 1108 into a random access memory (RAM) 1103. The processor 1101 may include, for example, a general-purpose microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1101 may also include on-board memory for caching purposes. The processor 1101 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0160] In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are stored. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. The processor 1101 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 1102 and / or the RAM 1103. It should be noted that the program may also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0161] According to an embodiment of the present disclosure, the electronic device 1100 may further include an input / output (I / O) interface 1105, and the input / output (I / O) interface 1105 is also connected to the bus 1104. The electronic device 1100 may further include one or more of the following components connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from it can be installed into the storage section 1108 as needed.

[0162] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present disclosure are implemented.

[0163] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 1102 and / or RAM 1103 and / or one or more memories other than ROM 1102 and RAM 1103.

[0164] Embodiments of the present disclosure also include a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method provided by the embodiments of the present disclosure.

[0165] When the computer program is executed by the processor 1101, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0166] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 1109, and / or installed from the removable medium 1111. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0167] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 1109, and / or installed from the removable medium 1111. When the computer program is executed by the processor 1101, the above-described functions defined in the system of the embodiments of the present disclosure are executed. According to the embodiments of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0168] According to the embodiments of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by connecting through the Internet using an Internet service provider).

[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0170] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0171] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure. < / h4>

Claims

1. A training method for a data processing model, comprising: Obtaining a first web page, wherein the source code of the first web page includes a first code tag and first text content to be processed, the first text content includes M text segments, and M is an integer greater than or equal to 1; wherein, the first code tag is a code tag associated with the first text content; Combining at least one of the second code tag associated with the first text segment and the first position information with the text of the first text segment to obtain a first input vector, wherein the first text segment is any one of the M text segments, the first code tag includes the second code tag, and the first position information is the position information of the first text segment among the M text segments; Obtaining an element tag for the first text segment, including: Determining the element category of the first text segment; determining the second position information of the first text segment in the element category; annotating the element tag based on the element category and the second position information; Using the first input vector and the element tag of the first text segment as training samples to train the data processing model to obtain a predicted element tag for the first text segment.

2. The method according to claim 1, wherein The obtaining of the first input vector includes obtaining a text vector, specifically including: Inputting the text of the first text segment into a pre-trained model, wherein the pre-trained model includes a pre-trained natural language processing model; Obtaining the text vector output by the pre-trained model.

3. The method according to claim 2, wherein, The obtaining of the first input vector includes obtaining a code tag vector, specifically including: Determining S types of first code tags associated with the first text content, wherein the S types of first code tags include the second code tag; Performing vector encoding on each of the S types of first code tags; Obtaining the code tag vector of the second code tag according to the result of the vector encoding.

4. The method according to any one of claims 2 or 3, wherein The obtaining of the first input vector includes obtaining a position vector, specifically including: Determining the first order of the first text segment among the M text segments, wherein the first position information includes the first order; Obtaining the position vector based on the first order.

5. The method according to claim 1, wherein, The data processing model includes a bidirectional long short-term memory network layer and a fully connected layer, and the training of the data processing model includes: Using the first input vector as the input of the bidirectional long short-term memory network layer; Using the output of the bidirectional long short-term memory network layer as the input of the fully connected layer; Calculating a loss function based on the output of the fully connected layer and the element tag, and updating the parameters of the data processing model according to the loss function, wherein the output of the fully connected layer includes the predicted element tag of the first text segment.

6. The method according to claim 5, wherein The data processing model further includes a normalization layer, and before using the first input vector as the input of the bidirectional long short-term memory network layer, it further includes: Performing normalization processing on the first input vector through the normalization layer.

7. The method according to claim 5, wherein The data processing model further includes a dropout layer, and before using the output of the bidirectional long short-term memory network layer as the input of the fully connected layer, it further includes: Process the output of the bidirectional long short-term memory network layer through the waiver layer.

8. The method according to claim 1, wherein The source code of the first web page is obtained using Hypertext Markup Language, and the first code tag includes Hypertext Markup Language tags.

9. A data processing method, comprising: Obtain a second web page, wherein the source code of the second web page includes a third code tag and second text content to be processed; Input the second text content and the third code tag into a data processing model, wherein the data processing model is obtained by training using the method according to any one of claims 1 to 8; Process the text of the second text segment according to the predicted element tags of the second text segment output by the data processing model, wherein the second text content includes at least one text segment, and the second text segment is any one of the at least one text segment.

10. A training device for a data processing model, comprising: A first acquisition module for acquiring a first web page, wherein the source code of the first web page includes a first code tag and first text content to be processed, the first text content includes M text segments, and M is an integer greater than or equal to 1; wherein the first code tag is a code tag associated with the first text content; An input vector module for combining at least one of the second code tag associated with the first text segment and the first position information with the text of the first text segment to obtain a first input vector, wherein the first text segment is any one of the M text segments, the first code tag includes the second code tag, and the first position information is the position information of the first text segment among the M text segments; An element tag acquisition module for obtaining the element tags of the first text segment, comprising: Determine the element category of the first text segment; determine the second position information of the first text segment in the element category; label the element tags based on the element category and the second position information; A model training module for using the first input vector and the element tags of the first text segment as training samples to train the data processing model and obtain the predicted element tags of the first text segment.

11. A data processing device, comprising: A second acquisition module for acquiring a second web page, wherein the source code of the second web page includes a third code tag and second text content to be processed; A data input module for inputting the second text content and the third code tag into a data processing model, wherein the data processing model is obtained by training using the device according to claim 10; A data processing module for processing the text of the second text segment according to the predicted element tags of the second text segment output by the data processing model, wherein the second text content includes at least one text segment, and the second text segment is any one of the at least one text segment.

12. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 9.

13. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to execute the method according to any one of claims 1 to 9.

14. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Source code file multi-service label automatic classification method

    CN110069252A

  • Text analysis method and device, equipment, medium and program product

    CN115081450A