Page element identification method and device, equipment and storage medium
Through the method of calculating conditional probability by word segmentation processing and word segmentation statistical model, page element types are identified, which solves the problem of low accuracy of machine crawling in the prior art and improves recognition accuracy.
Patent Information
- Application Number
- CN202311424149.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-02
AI Technical Summary
In the prior art, when determining page element types by machine crawling, the accuracy rate is low.
By obtaining the text information of the target page, performing word segmentation processing, the conditional probability of each word segmentation is calculated using the word segmentation statistical model to identify the element type of the page element.
Improve the accuracy of page element recognition, and enhance the certainty of element types by identifying element types from a text semantic perspective.
Smart Images

Figure CN119918537A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of page search technology, and in particular, to a page element recognition method, device, equipment and storage medium. Background Art
[0002] In order to facilitate users to find page elements of a certain element type in a timely manner, it is necessary to first determine the element types corresponding to the multiple page elements included in each page. The user can enter the search term corresponding to the element type (for example, the indicator name) to find the page element corresponding to the element type.
[0003] In the related art, a machine crawling method can be used to determine the element types corresponding to the multiple page elements included in each page. The specific steps are: according to the element type, multiple position information corresponding to the element type is determined, and for each position information, the text corresponding to the position information is crawled through a script, and the text is used as the page element corresponding to the element type.
[0004] However, the inventors have discovered that the related art has at least the following technical problems: the accuracy of the element types obtained by machine crawling is low. Summary of the invention
[0005] The embodiments of the present disclosure provide a method, apparatus, device and storage medium for identifying page elements to improve the recognition accuracy of page elements.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for identifying page elements, including:
[0007] Obtaining first text information corresponding to a page element to be identified in a target page;
[0008] Performing word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes a plurality of first word segmentations, and a first arrangement order of the plurality of first word segmentations in the first word segmentation sequence is consistent with a second arrangement order of the plurality of first word segmentations in the first text information;
[0009] For each first participle, determine a conditional probability corresponding to the first participle by using a participle statistical model, wherein the conditional probability is used to represent a probability that the first participle is located at a current position in the first arrangement sequence;
[0010] The element type corresponding to the page element is identified according to the conditional probabilities respectively corresponding to the multiple first participles.
[0011] In a second aspect, an embodiment of the present disclosure provides a device for identifying page elements, including:
[0012] An acquisition module, used to acquire first text information corresponding to a page element to be identified in a target page;
[0013] a word segmentation module, configured to perform word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes a plurality of first word segmentations, and a first arrangement order of the plurality of first word segmentations in the first word segmentation sequence is consistent with a second arrangement order of the plurality of first word segmentations in the first text information;
[0014] a determination module, configured to determine, for each first participle, a conditional probability corresponding to the first participle by using a participle statistical model, wherein the conditional probability is used to represent a probability that the first participle is located at a current position in the first arrangement sequence;
[0015] The identification module is used to identify the element type corresponding to the page element according to the conditional probabilities respectively corresponding to the multiple first word segmentations.
[0016] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0017] The memory stores computer-executable instructions;
[0018] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the page element identification method described in the first aspect and various possible designs of the first aspect.
[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the page element recognition method described in the first aspect and various possible designs of the first aspect is implemented.
[0020] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the page element recognition method as described in the first aspect and various possible designs of the first aspect.
[0021] The present embodiment provides a method, device, equipment and storage medium for identifying page elements. The method includes obtaining first text information corresponding to a page element to be identified in a target page; performing word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes multiple first word segmentations, and the first arrangement order of the multiple first word segmentations in the first word segmentation sequence is consistent with the second arrangement order of the multiple first word segmentations in the first text information; for each first word segmentation, determining the conditional probability corresponding to the first word segmentation through a word segmentation statistical model, wherein the conditional probability is used to represent the probability that the first word segmentation is located at the current position in the first arrangement order; identifying the element type corresponding to the page element according to the conditional probabilities corresponding to each of the multiple first word segmentations. In the disclosed embodiment, since the conditional probabilities corresponding to each of the multiple first word segmentations in the first word segmentation sequence can be determined through the word segmentation statistical model, the element type is obtained through the conditional probabilities corresponding to each of the multiple first word segmentations; wherein the word segmentation statistical model can identify the element type from the perspective of text semantics, thereby improving the accuracy of the determined element type. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0023] Figure 1 A schematic diagram of an application scenario of a page element recognition method provided by an embodiment of the present disclosure;
[0024] Figure 2 A flow chart of a page element recognition method provided by an embodiment of the present disclosure;
[0025] Figure 3 A flowchart of another page element identification method provided by an embodiment of the present disclosure;
[0026] Figure 4 A schematic diagram of a page element recognition method provided by an embodiment of the present disclosure;
[0027] Figure 5 A structural block diagram of a page element recognition device provided by an embodiment of the present disclosure;
[0028] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0031] On the page of a data product, there are many business-related page elements, such as the indicator names of data indicators (transaction amount, number of transaction persons, transaction order volume, sales transaction amount, etc.). When the business volume increases, there are thousands of indicator names involved in the entire product, and these indicator names are scattered across multiple pages. When a user uses a data product and needs to find a data indicator, it is impossible to quickly locate the specific page containing the indicator name of the data indicator.
[0032] In order to facilitate users to find page elements of a certain element type in a timely manner, it is necessary to first determine the element types corresponding to the multiple page elements included in each page. The user can enter the search term corresponding to the element type (for example, the indicator name) to find the page element corresponding to the element type.
[0033] In the related art, a machine crawling method can be used to determine the element types corresponding to each of the multiple page elements included in each page. The specific steps are: according to the element type, multiple position information corresponding to the element type is determined, and for each position information, the text corresponding to the position information is crawled through a script, and the text is used as the page element corresponding to the element type. However, the accuracy of the element type obtained by machine crawling is low.
[0034] It can be seen that how to improve the accuracy of identifying the element types corresponding to each of the multiple page elements included in the page is a problem that needs to be solved urgently.
[0035] In order to solve the above problem, the present embodiment provides the following technical concept: first, obtain the first text information corresponding to the page element to be identified in the target page; perform word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes multiple first word segmentations, and the first arrangement order of the multiple first word segmentations in the first word segmentation sequence is consistent with the second arrangement order of the multiple first word segmentations in the first text information; then, for each first word segmentation, determine the conditional probability corresponding to the first word segmentation through a word segmentation statistical model, wherein the conditional probability is used to represent the probability that the first word segmentation is located at the current position in the first arrangement order; finally, identify the element type corresponding to the page element according to the conditional probabilities corresponding to each of the multiple first word segmentations.
[0036] In this case, since the conditional probabilities corresponding to each of the multiple first participles in the first participle sequence can be determined by the word segmentation statistical model, and then the element type is obtained by the conditional probabilities corresponding to each of the multiple first participles; wherein the word segmentation statistical model can identify the element type from the perspective of text semantics, thereby improving the accuracy of the determined element type.
[0037] The application scenarios of the embodiments of the present disclosure are explained below:
[0038] The page element identification method provided by the embodiment of the present disclosure can be applied in multiple scenarios, for example, in the scenario of maintaining the "indicator name" in the page. Figure 1 A schematic diagram of an application scenario of a page element recognition method provided by an embodiment of the present disclosure. Figure 1 As shown, the page includes multiple page elements. The type of each page element is different. Among them, through the page element identification method, the page element corresponding to the "indicator name" can be identified from multiple page elements. Among them, the page elements corresponding to the "indicator name" are: "transaction amount", "transaction order volume", "transaction number", "refund amount". Among them, other page elements such as "home page" and "real-time overview" are not "indicator names". The data processing method provided by the embodiment of the present disclosure is described in detail using a detailed embodiment below.
[0039] Figure 2 The following is a flow chart of a method for identifying page elements provided in an embodiment of the present disclosure. The execution subject of the method for identifying page elements may be a terminal or a server. In the embodiment of the present disclosure, the execution subject is taken as an example to illustrate. Figure 2 , the page element identification method includes:
[0040] S201: Obtain first text information corresponding to a page element to be identified in a target page.
[0041] In the embodiments of the present disclosure, the target page usually includes multiple page elements. The page elements can be presented in various forms. Optionally, the page elements can include images, questions, trend charts, etc. Exemplarily, the page elements can include XX trend charts, indicator names, data corresponding to the indicator names, function point names, product information, serial numbers, avatars, function controls, etc. The acquisition method can be crawling by using crawlers or scripts.
[0042] In one embodiment of the present disclosure, Figure 4 As shown, in order to improve the accuracy and efficiency of recognition, the acquired original data may be preprocessed first. Accordingly, obtaining the first text information corresponding to the page element to be recognized in the target page may include: obtaining the original text information corresponding to the page element to be recognized in the target page; filtering the text information in the original text information that does not conform to the preset format to obtain the first text information corresponding to the page element.
[0043] Specifically, the text information that does not conform to the preset format may be text information such as null values and illegal values. Since the above null values and illegal values are filtered out by performing data cleaning on the original data, the effectiveness of the obtained page elements is improved.
[0044] S202: Perform word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes a plurality of first word segmentations, and a first arrangement order of the plurality of first word segmentations in the first word segmentation sequence is consistent with a second arrangement order of the plurality of first word segmentations in the first text information.
[0045] In the embodiment of the present disclosure, performing word segmentation processing on the first text information is to divide the characters (eg, Chinese) in the text information into multiple character combinations (eg, words). Optionally, the first text information can be subjected to word segmentation processing by a word segmentation tool.
[0046] Exemplarily, the first word segmentation sequence obtained by performing word segmentation processing on "transaction amount" is: "transaction", "amount"; wherein the first arrangement order of multiple first word segmentations in the first word segmentation sequence is consistent with the second arrangement order of the multiple first word segmentations in the first text information, that is, "transaction" is located at the first position in the first arrangement order, and "amount" is located after "transaction".
[0047] For example, the first segmentation sequence obtained by segmenting "product exposure number" is: "product", "exposure", "number of people". The first arrangement order of the multiple first segmentations in the first segmentation sequence is consistent with the second arrangement order of the multiple first segmentations in the first text information, that is, "product" is located at the first position in the first arrangement order, "exposure" is located after "product", and "number of people" is located after "exposure".
[0048] S203. For each first participle, determine the conditional probability corresponding to the first participle through the participle statistical model, wherein the conditional probability is used to represent the probability that the first participle is located at the current position in the first arrangement sequence.
[0049] In the disclosed embodiments, the element type can be identified from the perspective of text semantics. Optionally, when a page element of the "indicator name" type is selected from a page, it is necessary to determine whether the element type of the page element to be identified belongs to the "indicator name". From the perspective of text semantics, first determine the conditional probability of each first participle, which represents the conditional probability that the first participle belongs to the "indicator name" and is located at the current position in the first arrangement order. It should be noted that the present application can assume that the element type of the page element to be identified belongs to the "indicator name", then the conditional probability can be expressed as: the probability that the first participle is located at the current position in the first arrangement order.
[0050] In some embodiments, for each first participle, determining the conditional probability corresponding to the first participle through a participle statistical model may include: if the first participle is located at the first position in the first arrangement sequence, inputting the first participle into the participle statistical model, outputting a first conditional probability corresponding to the first participle, and the first conditional probability represents the probability that the first participle is located at the first position in the first participle sequence; if the first participle is not located at the first position in the first arrangement sequence, inputting the first participle and at least one second participle located before the first participle into the participle statistical model, and outputting a second conditional probability corresponding to the first participle, and the second conditional probability represents the probability that the first participle is located after at least one second participle.
[0051] Assuming that the first word segmentation sequence corresponding to the first text information I is: W1, W2, W3...Wn, then W1 is input into the word segmentation statistical model, and the conditional probability corresponding to the output of the first word segmentation W1 can be expressed as P(W1), which represents the probability of W1 appearing in the "indicator name".
[0052] W1 and W2 are input into the word segmentation statistical model, and the conditional probability of outputting the first word segmentation W2 can be expressed as P(W2|W1), which represents the probability of W2 appearing after W1 in the "indicator name".
[0053] Input W1, W2, and W3 into the word segmentation statistical model, and output the conditional probability of the first word W3, which can be expressed as P(W3|W1, W2), indicating the probability of the first word W3 appearing in the "indicator name" given that the preceding words are W1 and W2.
[0054] Input W1, W2, W3...Wn-1, and Wn into the word segmentation statistical model, and the output of the first word segmentation Wn can be expressed as: P(Wn|W1, W2, W3...Wn-1), which means the probability of the first word segmentation Wn appearing in the "indicator name" under the premise that the preceding words are W1, W2, W3...Wn-1.
[0055] In other embodiments, in order to simplify the calculation process and improve the calculation efficiency, it can be assumed that the probability of each word segmentation is only related to the previous word segmentation. Accordingly, for each first word segmentation, the conditional probability corresponding to the first word segmentation is determined by the word segmentation statistical model, which may include: if the first word segmentation is located in the first position in the first arrangement sequence, the first word segmentation statistical model is input, and the first conditional probability corresponding to the first word segmentation is output, and the first conditional probability represents the probability that the first word segmentation is located in the first position in the first word segmentation sequence; if the first word segmentation is not located in the first position in the first arrangement sequence, the first word segmentation and the second word segmentation before the first word segmentation are input into the word segmentation statistical model, and the second conditional probability corresponding to the first word segmentation is output, and the second conditional probability represents the probability that the first word segmentation is located after the second word segmentation.
[0056] In this case, assuming that the first word segmentation sequence corresponding to the first text information I is: W1, W2, W3...Wn, then W1 is input into the word segmentation statistical model, and the conditional probability corresponding to the output of the first word segmentation W1 can be expressed as P(W1), which represents the probability of W1 appearing in the "indicator name".
[0057] W1 and W2 are input into the word segmentation statistical model, and the conditional probability of outputting the first word segmentation W2 can be expressed as P(W2|W1), which represents the probability of W2 appearing after W1 in the "indicator name".
[0058] Input W2 and W3 into the word segmentation statistical model, and output the conditional probability of the first word segmentation W3, which can be expressed as P(W3|W2), indicating the probability of W3 appearing after W2 in the "Indicator Name".
[0059] Input Wn-1 and Wn into the word segmentation statistical model, and the output of the first word segmentation Wn can be expressed as: P(Wn|Wn-1) represents the probability of Wn appearing after Wn-1 in the "indicator name".
[0060] S204: Identify the element type corresponding to the page element according to the conditional probabilities corresponding to each of the plurality of first participles.
[0061] In the embodiment of the present disclosure, the element type of the page element can be determined by the product of multiple first participles. Accordingly, this step is: determine the product of the conditional probabilities corresponding to each of the multiple first participles; if the product is greater than a preset classification threshold, it is determined that the element type corresponding to the page element is the preset element type; if the product is not greater than the preset classification threshold, it is determined that the element type corresponding to the page element is not the preset element type.
[0062] The preset classification threshold may be represented by T. In the embodiment of the present disclosure, the value of the preset classification threshold is not specifically limited and may be set and modified as needed.
[0063] Exemplarily, assuming that the first text information is "number of people exposed to the product", the corresponding first word segmentation sequence is "product", "exposure", "number of people". The conditional probability of each word segmentation is determined to be: P(product), P(exposure|product), P(number of people|exposure). The product of the conditional probabilities corresponding to each of the multiple first word segmentations is determined to be: P(product)*P(exposure|product)*P(number of people|exposure). If P(product)*P(exposure|product)*P(number of people|exposure) is greater than T, it is determined that the element type corresponding to the page element is the preset element type (indicator name); if P(product)*P(exposure|product)*P(number of people|exposure) is not greater than T, it is determined that the element type corresponding to the page element is not the preset element type (indicator name).
[0064] The disclosed embodiment provides a method for identifying page elements, including: obtaining first text information corresponding to a page element to be identified in a target page; performing word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes multiple first word segmentations, and the first arrangement order of the multiple first word segmentations in the first word segmentation sequence is consistent with the second arrangement order of the multiple first word segmentations in the first text information; for each first word segmentation, determining the conditional probability corresponding to the first word segmentation through a word segmentation statistical model, wherein the conditional probability is used to represent the probability that the first word segmentation is located at the current position in the first arrangement order; identifying the element type corresponding to the page element according to the conditional probabilities corresponding to each of the multiple first word segmentations. In the disclosed embodiment, since the conditional probabilities corresponding to each of the multiple first word segmentations in the first word segmentation sequence can be determined through the word segmentation statistical model, the element type is obtained through the conditional probabilities corresponding to each of the multiple first word segmentations; wherein the word segmentation statistical model can identify the element type from the perspective of text semantics, thereby improving the accuracy of the determined element type.
[0065] One point that needs to be explained is that, in one embodiment of the present disclosure, after step S204 identifies the element type corresponding to the page element based on the conditional probabilities corresponding to each of the multiple first participles, it may also include: S205, if it is determined that the element type corresponding to the page element is a preset element type (indicator name), then the page element is associated with the target page, thereby obtaining the association data between the page and the indicator name.
[0066] Figure 3 A flowchart of another method for identifying page elements provided by an embodiment of the present disclosure. In an embodiment of the present disclosure, a method for determining the first conditional probability or the second conditional probability corresponding to the first word segmentation by using a word segmentation statistical model in step S203 is described in detail.
[0067] Optionally, the word segmentation statistical model includes first conditional probabilities corresponding to each of the multiple sample word segments. Figure 3 , the specific method for determining the first conditional probability corresponding to the first participle includes:
[0068] S301. According to a first participle, search for a target participle that is the same as the first participle from multiple sample participles.
[0069] S302. Determine the first conditional probability corresponding to the target word segmentation as the first conditional probability corresponding to the first word segmentation, wherein the first conditional probability corresponding to the target word segmentation is determined based on the first number of occurrences of the target word segmentation in the target corpus and the number of multiple second text information included in the target corpus, wherein the target corpus includes second text information corresponding to each of multiple page elements of a preset element type.
[0070] Alternatively, if Figure 4 As shown, before determining the first conditional probability, the target corpus can be determined first, and the word segmentation statistical model can be obtained by training the target corpus. Accordingly, before inputting the first word segmentation into the word segmentation statistical model and outputting the first conditional probability corresponding to the first word segmentation, it also includes: performing word segmentation processing on each second text information to obtain multiple sample word segments; for each sample word segmentation, counting the first occurrence number of the sample word segmentation in the target corpus, and determining the ratio of the first occurrence number to the number of multiple second text information included in the target corpus as the first conditional probability corresponding to the sample word segmentation; and obtaining the word segmentation statistical model through the first conditional probabilities corresponding to each of the multiple sample word segments.
[0071] Optionally, the word segmentation statistical model includes a plurality of second conditional probabilities corresponding to each of the word segmentation groups, wherein each word segmentation group includes two sorted sample word segments. Figure 3 , the specific method for determining the second conditional probability corresponding to the first participle includes:
[0072] S303: According to the first participle and the second participle located before the first participle, search from multiple participle groups for a target participle group including the first participle and the second participle, wherein the second participle is located before the first participle.
[0073] S304. Determine the second conditional probability corresponding to the target word segmentation group as the second conditional probability corresponding to the first word segmentation, wherein the second conditional probability corresponding to the target word segmentation group is determined based on the second number of occurrences of the target sample word segmentation that is ranked higher among the two ranked target sample word segmentations corresponding to the target word segmentation group in the target corpus and the third number of occurrences of the target word segmentation group in the target corpus, wherein the target corpus includes second text information corresponding to each of a plurality of page elements of a preset element type.
[0074] Alternatively, if Figure 4 As shown, before determining the second conditional probability, the target corpus can be determined first, and the word segmentation statistical model can be obtained by training the target corpus. Accordingly, the first word segmentation and the second word segmentation located before the first word segmentation are input into the word segmentation statistical model, and before outputting the second conditional probability corresponding to the first word segmentation, it also includes: performing word segmentation processing on each second text information to obtain multiple word segmentation groups, wherein each word segmentation group includes two sorted sample word segmentations; for each word segmentation group, counting the second number of occurrences of the sample word segmentation with the front sorting in the two sorted sample word segmentations corresponding to the word segmentation group in the target corpus and the third number of occurrences of the word segmentation group in the target corpus, and determining the ratio of the third number to the second number as the second conditional probability corresponding to the word segmentation group; and obtaining the word segmentation statistical model through the second conditional probabilities corresponding to each of the multiple word segmentation groups.
[0075] It should be noted that the word segmentation statistical model may include the first conditional probabilities corresponding to each of the multiple sample word segments and the second conditional probabilities corresponding to each of the multiple word segmentation groups, so that only one word segmentation statistical model is needed to determine the first conditional probability or the second conditional probability corresponding to the first word segmentation. In addition, the word segmentation statistical model only needs to be trained once, which also saves the time of model training.
[0076] In one embodiment of the present disclosure, in order to ensure smoothness of probability estimation during the application of the target model, a preset probability may be set for the newly appearing word segmentation, which is lower than the probability of the appearance of the word segmentation of the target text information appearing in the corpus. Accordingly, the method may also include: if the first number of times the target word segmentation appears in the target corpus is 0, the first conditional probability of the target word segmentation is set to the preset probability. Similarly, if the second number of times the sample word segmentation that is ranked higher in the two sorted sample word segmentations corresponding to the word segmentation group appears in the target corpus or the third number of times the word segmentation group appears in the target corpus is 0, the second conditional probability corresponding to the first word segmentation is set to the preset probability.
[0077] Specifically, considering that the number of multiple target text information in the corpus pre - used for training is limited and may not fully cover the first text information extracted from the target page. If there are new word segments in the first text information, assuming that the word segment "commodity" does not appear in the corpus, then there is no relevant probability for this "commodity" in the target model. At this time, in order to ensure the smoothness of probability estimation during the application of the target model, a preset probability can be set for this newly - appeared word segment, and this preset probability is lower than the appearance probability of the word segments of the target text information that appear in the corpus.
[0078] In an embodiment of the present disclosure, in order to ensure a high accuracy of the word - segment statistical model, new index names need to be added to the target corpus so as to retrain to obtain a new word - segment statistical model. Correspondingly, the method may further include: determining the product of the conditional probabilities respectively corresponding to multiple first word segments; if the product is not greater than the preset classification threshold and is greater than the training classification threshold, then determine the text information corresponding to the multiple first word segments as new index names (that is, the training samples of the word - segment statistical model).
[0079] In an embodiment of the present disclosure, the trained word - segment statistical model can also be evaluated. A variable T (0 < T < 1) can be defined as the judgment threshold of the word - segment statistical model. When P > T, it is considered that this text is an index name, and vice versa. The magnitude of the T value determines the accuracy and recall rate of the model, and can be adjusted as appropriate in different application scenarios.
[0080] Exemplarily, during the evaluation process, 500 test samples can be extracted, and the positive - negative sample ratio is 1:4 (the index text is the positive sample, and the non - index text is the negative sample), and the recall rate and accuracy of the model are evaluated by fixing the T value:
[0081] T 0.7 0.75 0.8 0.85 0.9 0.95 Accuracy 0.943 0.955 0.975 0.982 0.990 0.998 Recall 0.979 0.973 0.971 0.968 0.958 0.939
[0082] Figure 5 The block diagram of a page element recognition device provided by an embodiment of the present disclosure. As Figure 5 shown, the page element recognition device includes: an acquisition module 501, a word - segment module 502, a determination module 503, and a recognition module 504.
[0083] Among them, the acquisition module 501 is used to acquire the first text information corresponding to the page element to be recognized in the target page;
[0084] The word - segment module 502 is used to perform word - segment processing on the first text information to obtain a first word - segment sequence, where the first word - segment sequence includes multiple first word segments, and the first arrangement order of the multiple first word segments in the first word - segment sequence is the same as the second arrangement order of the multiple first word segments in the first text information;
[0085] A determination module 503 is used to determine, for each first participle, a conditional probability corresponding to the first participle by using a participle statistical model, wherein the conditional probability is used to represent the probability that the first participle is located at a current position in the first arrangement sequence;
[0086] The identification module 504 is used to identify the element type corresponding to the page element according to the conditional probabilities respectively corresponding to the multiple first word segmentations.
[0087] According to one or more embodiments of the present disclosure, the determination module 503 determines, for each first participle, the conditional probability corresponding to the first participle through a participle statistical model, specifically including: if the first participle is located at the first position in the first arrangement sequence, the first participle is input into the participle statistical model, and a first conditional probability corresponding to the first participle is output, and the first conditional probability represents the probability that the first participle is located at the first position in the first participle sequence; if the first participle is not located at the first position in the first arrangement sequence, the first participle and the second participle located before the first participle are input into the participle statistical model, and a second conditional probability corresponding to the first participle is output, and the second conditional probability represents the probability that the first participle is located after the second participle.
[0088] According to one or more embodiments of the present disclosure, the word segmentation statistical model includes first conditional probabilities corresponding to each of a plurality of sample word segmentations; accordingly, the determination module 503 inputs the first word segmentation into the word segmentation statistical model and outputs the first conditional probability corresponding to the first word segmentation, specifically including: according to the first word segmentation, querying a target word segmentation identical to the first word segmentation from the plurality of sample word segmentations; determining the first conditional probability corresponding to the target word segmentation as the first conditional probability corresponding to the first word segmentation, wherein the first conditional probability corresponding to the target word segmentation is determined based on the number of first occurrences of the target word segmentation in a target corpus and the number of a plurality of second text information included in the target corpus, wherein the target corpus includes second text information corresponding to each of a plurality of page elements of a preset element type.
[0089] According to one or more embodiments of the present disclosure, the device also includes: a training module; the training module is used to perform word segmentation processing on each of the second text information to obtain multiple sample word segmentations; for each sample word segmentation, the number of first appearances of the sample word segmentation in the target corpus is counted, and the ratio of the first number of times to the number of multiple second text information included in the target corpus is determined as the first conditional probability corresponding to the sample word segmentation; the word segmentation statistical model is obtained through the first conditional probabilities corresponding to each of the multiple sample word segmentations.
[0090] According to one or more embodiments of the present disclosure, the word segmentation statistical model includes second conditional probabilities corresponding to multiple word segmentation groups, each of which includes two sorted sample word segmentations; accordingly, the determination module 503 inputs the first word segmentation and the second word segmentation located before the first word segmentation into the word segmentation statistical model, and outputs the second conditional probability corresponding to the first word segmentation, specifically including: according to the first word segmentation and the second word segmentation located before the first word segmentation, querying from the multiple word segmentation groups a target word segmentation group including the first word segmentation and the second word segmentation, and the second word segmentation is located before the first word segmentation; determining the second conditional probability corresponding to the target word segmentation group as the second conditional probability corresponding to the first word segmentation, the second conditional probability corresponding to the target word segmentation group is based on the second number of occurrences of the target sample word segmentation with the front ranking in the two sorted target sample word segmentations corresponding to the target word segmentation group in the target corpus and the third number of occurrences of the target word segmentation group in the target corpus, wherein the target corpus includes second text information corresponding to multiple page elements of a preset element type, respectively.
[0091] According to one or more embodiments of the present disclosure, the device also includes: a training module; the training module is used to perform word segmentation processing on each of the second text information to obtain multiple word segmentation groups, wherein each word segmentation group includes two sorted sample word segmentations; for each word segmentation group, the second number of occurrences of the sample word segmentation with a higher sorting among the two sorted sample word segmentations corresponding to the word segmentation group and the third number of occurrences of the word segmentation group in the target corpus are counted, and the ratio of the third number to the second number is determined as the second conditional probability corresponding to the word segmentation group; the word segmentation statistical model is obtained through the second conditional probabilities corresponding to each of the multiple word segmentation groups.
[0092] According to one or more embodiments of the present disclosure, the identification module 504 identifies the element type corresponding to the page element according to the conditional probabilities respectively corresponding to the multiple first participles, specifically including: determining the product of the conditional probabilities respectively corresponding to the multiple first participles; if the product is greater than a preset classification threshold, determining that the element type corresponding to the page element is the preset element type; if the product is not greater than the preset classification threshold, determining that the element type corresponding to the page element is not the preset element type.
[0093] According to one or more embodiments of the present disclosure, the acquisition module 501 acquires the first text information corresponding to the page element to be identified in the target page, specifically including: acquiring the original text information corresponding to the page element to be identified in the target page; filtering the text information in the original text information that does not conform to a preset format to obtain the first text information corresponding to the page element.
[0094] The acquisition module 501, the word segmentation module 502, the determination module 503 and the recognition module 504 are connected in sequence. The data processing device provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.
[0095] In order to implement the above embodiment, the embodiment of the present disclosure also provides an electronic device.
[0096] Figure 6 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present disclosure. Figure 6 The electronic device 600 may be a terminal device or a server. The terminal device may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0097] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0098] Typically, the following devices may be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 608 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0099] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0100] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0101] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0102] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0103] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0104] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0105] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of a unit does not limit the unit itself in some cases. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses".
[0106] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0107] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0108] In a first aspect, according to one or more embodiments of the present disclosure, a method for identifying page elements is provided, comprising:
[0109] Obtaining first text information corresponding to a page element to be identified in a target page;
[0110] Performing word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes a plurality of first word segmentations, and a first arrangement order of the plurality of first word segmentations in the first word segmentation sequence is consistent with a second arrangement order of the plurality of first word segmentations in the first text information;
[0111] For each first participle, determine a conditional probability corresponding to the first participle by using a participle statistical model, wherein the conditional probability is used to represent a probability that the first participle is located at a current position in the first arrangement sequence;
[0112] The element type corresponding to the page element is identified according to the conditional probabilities respectively corresponding to the multiple first participles.
[0113] According to one or more embodiments of the present disclosure, for each first participle, determining the conditional probability corresponding to the first participle through a participle statistical model includes: if the first participle is located at the first position in the first arrangement sequence, inputting the first participle into the participle statistical model, and outputting a first conditional probability corresponding to the first participle, the first conditional probability representing the probability that the first participle is located at the first position in the first participle sequence; if the first participle is not located at the first position in the first arrangement sequence, inputting the first participle and a second participle located before the first participle into the participle statistical model, and outputting a second conditional probability corresponding to the first participle, the second conditional probability representing the probability that the first participle is located after the second participle.
[0114] According to one or more embodiments of the present disclosure, the word segmentation statistical model includes first conditional probabilities corresponding to each of a plurality of sample word segmentations; accordingly, the first word segmentation is input into the word segmentation statistical model, and the first conditional probability corresponding to the first word segmentation is output, including: according to the first word segmentation, querying a target word segmentation that is the same as the first word segmentation from the plurality of sample word segmentations; determining the first conditional probability corresponding to the target word segmentation as the first conditional probability corresponding to the first word segmentation, wherein the first conditional probability corresponding to the target word segmentation is determined based on the first number of occurrences of the target word segmentation in a target corpus and the number of a plurality of second text information included in the target corpus, wherein the target corpus includes second text information corresponding to each of a plurality of page elements of a preset element type.
[0115] According to one or more embodiments of the present disclosure, before inputting the first word segmentation into the word segmentation statistical model and outputting the first conditional probability corresponding to the first word segmentation, it also includes: performing word segmentation processing on each of the second text information to obtain multiple sample word segmentations; for each sample word segmentation, counting the number of first occurrences of the sample word segmentation in the target corpus, and determining the ratio of the first number of times to the number of multiple second text information included in the target corpus as the first conditional probability corresponding to the sample word segmentation; and obtaining the word segmentation statistical model through the first conditional probabilities corresponding to each of the multiple sample word segmentations.
[0116] According to one or more embodiments of the present disclosure, the word segmentation statistical model includes second conditional probabilities corresponding to a plurality of word segmentation groups, wherein each word segmentation group includes two sorted sample word segments;
[0117] Correspondingly, the first participle and the second participle located before the first participle are input into the participle statistical model, and the second conditional probability corresponding to the first participle is output, including: according to the first participle and the second participle located before the first participle, querying from the multiple participle groups a target participle group including the first participle and the second participle, and the second participle is located before the first participle; determining the second conditional probability corresponding to the target participle group as the second conditional probability corresponding to the first participle, the second conditional probability corresponding to the target participle group is determined based on the second number of occurrences of the target sample participle that is ranked higher among the two ranked target sample participles corresponding to the target participle group in the target corpus and the third number of occurrences of the target participle group in the target corpus, wherein the target corpus includes second text information corresponding to each of a plurality of page elements of a preset element type.
[0118] According to one or more embodiments of the present disclosure, before the first participle and the second participle located before the first participle are input into the participle statistical model, and the second conditional probability corresponding to the first participle is output, the method further includes: performing participle processing on each of the second text information to obtain multiple participle groups, wherein each participle group includes two sorted sample participles; for each participle group, counting the second number of occurrences of the sample participle that is ranked higher in the two sorted sample participles corresponding to the participle group in the target corpus and the third number of occurrences of the participle group in the target corpus, and determining the ratio of the third number to the second number as the second conditional probability corresponding to the participle group; and obtaining the participle statistical model through the second conditional probabilities corresponding to each of the multiple participle groups.
[0119] According to one or more embodiments of the present disclosure, identifying the element type corresponding to the page element based on the conditional probabilities respectively corresponding to the multiple first participles includes: determining the product of the conditional probabilities respectively corresponding to the multiple first participles; if the product is greater than a preset classification threshold, determining that the element type corresponding to the page element is the preset element type; if the product is not greater than the preset classification threshold, determining that the element type corresponding to the page element is not the preset element type.
[0120] According to one or more embodiments of the present disclosure, obtaining the first text information corresponding to the page element to be identified in the target page includes: obtaining the original text information corresponding to the page element to be identified in the target page; filtering the text information in the original text information that does not conform to a preset format to obtain the first text information corresponding to the page element.
[0121] In a second aspect, according to one or more embodiments of the present disclosure, a device for identifying page elements is provided, the device comprising:
[0122] An acquisition module, used to acquire first text information corresponding to a page element to be identified in a target page;
[0123] a word segmentation module, configured to perform word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes a plurality of first word segmentations, and a first arrangement order of the plurality of first word segmentations in the first word segmentation sequence is consistent with a second arrangement order of the plurality of first word segmentations in the first text information;
[0124] a determination module, configured to determine, for each first participle, a conditional probability corresponding to the first participle by using a participle statistical model, wherein the conditional probability is used to represent a probability that the first participle is located at a current position in the first arrangement sequence;
[0125] The identification module is used to identify the element type corresponding to the page element according to the conditional probabilities respectively corresponding to the multiple first word segmentations.
[0126] According to one or more embodiments of the present disclosure, the determination module determines, for each first participle, the conditional probability corresponding to the first participle through a participle statistical model, specifically including: if the first participle is located at the first position in the first arrangement sequence, the first participle is input into the participle statistical model, and a first conditional probability corresponding to the first participle is output, and the first conditional probability represents the probability that the first participle is located at the first position in the first participle sequence; if the first participle is not located at the first position in the first arrangement sequence, the first participle and a second participle located before the first participle are input into the participle statistical model, and a second conditional probability corresponding to the first participle is output, and the second conditional probability represents the probability that the first participle is located after the second participle.
[0127] According to one or more embodiments of the present disclosure, the word segmentation statistical model includes first conditional probabilities corresponding to each of a plurality of sample word segmentations; accordingly, the determination module inputs the first word segmentation into the word segmentation statistical model and outputs a first conditional probability corresponding to the first word segmentation, specifically including: according to the first word segmentation, querying a target word segmentation that is the same as the first word segmentation from the plurality of sample word segmentations; determining the first conditional probability corresponding to the target word segmentation as the first conditional probability corresponding to the first word segmentation, wherein the first conditional probability corresponding to the target word segmentation is determined based on the number of first occurrences of the target word segmentation in a target corpus and the number of a plurality of second text information included in the target corpus, wherein the target corpus includes second text information corresponding to each of a plurality of page elements of a preset element type.
[0128] According to one or more embodiments of the present disclosure, the device also includes: a training module; the training module is used to perform word segmentation processing on each of the second text information to obtain multiple sample word segmentations; for each sample word segmentation, the number of first appearances of the sample word segmentation in the target corpus is counted, and the ratio of the first number of times to the number of multiple second text information included in the target corpus is determined as the first conditional probability corresponding to the sample word segmentation; the word segmentation statistical model is obtained through the first conditional probabilities corresponding to each of the multiple sample word segmentations.
[0129] According to one or more embodiments of the present disclosure, the word segmentation statistical model includes second conditional probabilities corresponding to multiple word segmentation groups, each of which includes two sorted sample word segmentations; accordingly, the determination module inputs the first word segmentation and the second word segmentation located before the first word segmentation into the word segmentation statistical model, and outputs the second conditional probability corresponding to the first word segmentation, specifically including: according to the first word segmentation and the second word segmentation located before the first word segmentation, querying from the multiple word segmentation groups a target word segmentation group including the first word segmentation and the second word segmentation, and the second word segmentation is located before the first word segmentation; determining the second conditional probability corresponding to the target word segmentation group as the second conditional probability corresponding to the first word segmentation, the second conditional probability corresponding to the target word segmentation group is based on the second number of occurrences of the target sample word segmentation ranked higher in the two sorted target sample word segmentations corresponding to the target word segmentation group in the target corpus and the third number of occurrences of the target word segmentation group in the target corpus, wherein the target corpus includes second text information corresponding to multiple page elements of a preset element type, respectively.
[0130] According to one or more embodiments of the present disclosure, the device also includes: a training module; the training module is used to perform word segmentation processing on each of the second text information to obtain multiple word segmentation groups, wherein each word segmentation group includes two sorted sample word segmentations; for each word segmentation group, the second number of occurrences of the sample word segmentation with a higher sorting among the two sorted sample word segmentations corresponding to the word segmentation group and the third number of occurrences of the word segmentation group in the target corpus are counted, and the ratio of the third number to the second number is determined as the second conditional probability corresponding to the word segmentation group; the word segmentation statistical model is obtained through the second conditional probabilities corresponding to each of the multiple word segmentation groups.
[0131] According to one or more embodiments of the present disclosure, the identification module identifies the element type corresponding to the page element based on the conditional probabilities respectively corresponding to the multiple first participles, specifically including: determining the product of the conditional probabilities respectively corresponding to the multiple first participles; if the product is greater than a preset classification threshold, determining that the element type corresponding to the page element is the preset element type; if the product is not greater than the preset classification threshold, determining that the element type corresponding to the page element is not the preset element type.
[0132] According to one or more embodiments of the present disclosure, the acquisition module acquires the first text information corresponding to the page element to be identified in the target page, specifically including: acquiring the original text information corresponding to the page element to be identified in the target page; filtering the text information in the original text information that does not conform to a preset format to obtain the first text information corresponding to the page element.
[0133] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0134] The memory stores computer-executable instructions;
[0135] The processor executes the computer-executable instructions stored in the memory to implement the data processing method described in the first aspect and various possible designs of the first aspect.
[0136] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer execution instructions. When a processor executes the computer execution instructions, the data processing method described in the first aspect and various possible designs of the first aspect is implemented.
[0137] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the data processing method described in the first aspect and various possible designs of the first aspect.
[0138] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0139] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0140] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A page element recognition method, characterized in that: include: Obtaining first text information corresponding to a page element to be identified in a target page; Performing word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes a plurality of first word segmentations, and a first arrangement order of the plurality of first word segmentations in the first word segmentation sequence is consistent with a second arrangement order of the plurality of first word segmentations in the first text information; For each first participle, determine a conditional probability corresponding to the first participle by using a participle statistical model, wherein the conditional probability is used to represent a probability that the first participle is located at a current position in the first arrangement sequence; The element type corresponding to the page element is identified according to the conditional probabilities respectively corresponding to the multiple first participles.
2. The method according to claim 1, characterized in that The determining, for each first participle, a conditional probability corresponding to the first participle by using a participle statistical model includes: If the first participle is located at the first position in the first arrangement sequence, inputting the first participle into the participle statistical model, and outputting a first conditional probability corresponding to the first participle, wherein the first conditional probability represents the probability that the first participle is located at the first position in the first participle sequence; If the first participle is not located at the first position in the first arrangement order, the first participle and the second participle located before the first participle are input into the participle statistical model, and a second conditional probability corresponding to the first participle is output, where the second conditional probability represents the probability that the first participle is located after the second participle.
3. The method according to claim 2, characterized in that The word segmentation statistical model includes first conditional probabilities corresponding to each of a plurality of sample word segments; Accordingly, inputting the first word segmentation into the word segmentation statistical model and outputting a first conditional probability corresponding to the first word segmentation includes: According to the first participle, searching for a target participle that is the same as the first participle from the multiple sample participles; The first conditional probability corresponding to the target participle is determined as the first conditional probability corresponding to the first participle, wherein the first conditional probability corresponding to the target participle is determined based on the first number of occurrences of the target participle in a target corpus and the number of multiple second text information included in the target corpus, wherein the target corpus includes second text information corresponding to each of multiple page elements of a preset element type.
4. The method according to claim 3, characterized in that Before inputting the first participle into the participle statistical model and outputting the first conditional probability corresponding to the first participle, the method further includes: Performing word segmentation processing on each of the second text information to obtain a plurality of sample word segments; For each sample segmentation, counting the number of first occurrences of the sample segmentation in the target corpus, and determining a ratio of the first number of occurrences to the number of the plurality of second text information included in the target corpus as a first conditional probability corresponding to the sample segmentation; The word segmentation statistical model is obtained through the first conditional probabilities respectively corresponding to the multiple sample word segments.
5. The method according to claim 2, characterized in that: The word segmentation statistical model includes second conditional probabilities corresponding to a plurality of word segmentation groups, wherein each word segmentation group includes two sorted sample word segments; Accordingly, the first participle and the second participle located before the first participle are input into the participle statistical model, and the second conditional probability corresponding to the first participle is output, including: According to the first participle and the second participle located before the first participle, searching from the plurality of participle groups for a target participle group including the first participle and the second participle, wherein the second participle is located before the first participle; The second conditional probability corresponding to the target word segmentation group is determined as the second conditional probability corresponding to the first word segmentation, and the second conditional probability corresponding to the target word segmentation group is determined based on the second number of times a target sample segmentation with a front ranking among two sorted target sample segmentations corresponding to the target word segmentation group appears in a target corpus and the third number of times the target word segmentation group appears in the target corpus, wherein the target corpus includes second text information corresponding to each of a plurality of page elements of a preset element type.
6. The method according to claim 5, characterized in that The method further includes: inputting the first participle and the second participle located before the first participle into the participle statistical model, and outputting the second conditional probability corresponding to the first participle. Performing word segmentation processing on each of the second text information to obtain a plurality of word segmentation groups, wherein each word segmentation group includes two sorted sample word segments; For each word group, count the second number of occurrences of the sample word that is ranked higher among the two ranked sample word groups corresponding to the word group in the target corpus and the third number of occurrences of the word group in the target corpus, and determine the ratio of the third number to the second number as the second conditional probability corresponding to the word group; The word segmentation statistical model is obtained through the second conditional probabilities respectively corresponding to the multiple word segmentation groups.
7. The method according to claim 1, characterized in that The step of identifying the element type corresponding to the page element according to the conditional probabilities respectively corresponding to the plurality of first participles includes: Determine the product of the conditional probabilities corresponding to each of the plurality of first participles; If the product is greater than a preset classification threshold, it is determined that the element type corresponding to the page element is a preset element type; if the product is not greater than the preset classification threshold, it is determined that the element type corresponding to the page element is not the preset element type.
8. The method according to any one of claims 1 to 7, characterized in that: The step of obtaining first text information corresponding to a page element to be identified in a target page includes: Obtaining original text information corresponding to the page element to be identified in the target page; The text information that does not conform to the preset format in the original text information is filtered to obtain the first text information corresponding to the page element.
9. A page element recognition device, characterized in that: include: An acquisition module, used to acquire first text information corresponding to a page element to be identified in a target page; a word segmentation module, configured to perform word segmentation processing on the first text information to obtain a first word segmentation sequence, wherein the first word segmentation sequence includes a plurality of first word segmentations, and a first arrangement order of the plurality of first word segmentations in the first word segmentation sequence is consistent with a second arrangement order of the plurality of first word segmentations in the first text information; a determination module, configured to determine, for each first participle, a conditional probability corresponding to the first participle by using a participle statistical model, wherein the conditional probability is used to represent a probability that the first participle is located at a current position in the first arrangement sequence; The identification module is used to identify the element type corresponding to the page element according to the conditional probabilities respectively corresponding to the multiple first word segmentations.
10. An electronic device, characterized in that: include: Processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the page element identification method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the page element recognition method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying page elements according to any one of claims 1 to 8 is implemented.