Data processing system for determining final application text

By designing a data processing system, obtaining intermediate text and list of text to be selected based on the target text, combining natural language processing technology and multi-source data integration, the problem of low efficiency and insufficient accuracy of determining the final application text in the prior art is solved, and efficient and accurate application text determination is achieved.

CN120144748AActive Publication Date: 2025-06-13ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510241493.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

In the prior art, determining the final application text is inefficient and failing to intelligently analyze the target text, resulting in low accuracy.

Method used

A data processing system is designed to execute a computer program through a processor and memory to realize the following steps: obtain several first intermediate texts based on the target text, obtain a list of to-select text corresponding to the candidate text based on these intermediate texts, and finally obtain a list of target application text corresponding to the to-select text list to determine the final application text. The system deeply analyzes the target text and integrates multi-source heterogeneous data through matching the target title list, target priority list and specified conditional text list.

Benefits of technology

The efficiency of obtaining the final application text is improved, the target text is deeply analyzed through natural language processing technology, and multi-source heterogeneous data is integrated, making the final application documents obtained more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144748A_ABST
    Figure CN120144748A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing system for determining a final application text, the system comprises a processor and a memory in which a computer program is stored, and when the computer program is executed by the processor, the following steps are realized: based on a target text, obtaining a plurality of first intermediate texts, obtaining a to-be-selected text list corresponding to a candidate text, and obtaining a target application text list corresponding to the to-be-selected text list to obtain a final application text, and based on the to-be-selected text, obtaining a target title list set, a target priority list set, a specified condition text list set and a target application text corresponding to the to-be-selected text. According to the method, the policy text is classified, the classified policy text is graded, screening is carried out based on the requirements of the candidate objects to determine the final application text, the efficiency of obtaining the final application text is improved, deep analysis is carried out on the target text through the natural language processing technology, and the accuracy of the obtained final application file is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and particularly to a data processing system for determining a final application text. Background Art

[0002] In the context of the current global economic integration and the rapid change of the policy environment, there are more and more development opportunities and challenges. In order to promote the development of specific industries, encourage technological innovation, optimize the industrial structure and other goals, governments and regional organizations around the world frequently introduce various policy measures. These policies often include a series of specific standards and requirements, aiming to screen eligible objects to enjoy policy dividends. However, due to the large amount of content included in the policies and the large number of corresponding achievements of enterprises or relevant personnel, how to accurately match and determine the final application text has become an urgent problem to be solved.

[0003] In the prior art, the method for determining the final application text is as follows: an artificial method is adopted, and by reading policy documents one by one, the standards therein are manually compared with one's own situation, and through manual interpretation and analysis, the specific content and requirements of the policy are understood so as to determine the final application text. Due to the huge amount of data, text screening based on the needs of candidate objects is not considered, resulting in low efficiency in obtaining the final application text, and the intelligent analysis of the target text is not carried out, so that the accuracy of the obtained final application document is relatively low. Summary of the Invention

[0004] In view of the above technical problems, the technical solution adopted by the present invention is as follows: a data processing system for determining a final application text, the system includes: a processor and a memory storing a computer program, and when the computer program is executed by the processor, the following steps are implemented:

[0005] S100, based on the target text, obtain a plurality of first intermediate texts, where the first intermediate texts are texts obtained by splitting the target text based on a preset category.

[0006] S200, based on the plurality of first intermediate texts, obtain a list of candidate texts corresponding to the candidate text, where the list of candidate texts includes a plurality of candidate texts, and the candidate texts are first intermediate texts that meet the preset conditions obtained based on a plurality of second intermediate texts corresponding to the candidate text, and the second intermediate texts are texts obtained by splitting the candidate text based on a preset category.

[0007] S300. Obtain the target application text list corresponding to the list of candidate texts to obtain the final application text. Among them, the target application text list includes several target application texts, each candidate text corresponds to one target application text, and the final application text includes all the target application texts in the target application text list. Among them, in S300, the target application text is obtained through the following steps:

[0008] S1. Based on the candidate text, obtain the target title list set E = {E 1 , E 2 , E 3}, E 2 = {E 2 1 , ……, E 2 i , ……, E 2 n}, E 3 = {E 3 1 , ……, E 3 i , ……, E 3 n}, E 3 i = {E 3 i1 , ……, E 3 ij , ……, E 3 im(i)}, E 1 is the first target title, E 2 i is the i-th second target title in the second target title list E 1 corresponding to E 2 , E 3 ij is the j-th third target title in the third target title list corresponding to E 2 i .

[0009] S2. According to E and the second intermediate text corresponding to the candidate text, obtain the target priority list set F = {F 1 , ……, F i , ……, F n}, F i = {F i1 , ……, F ij , ……, F im(i)}, F ij is the target priority corresponding to the candidate path corresponding to E 3 ij .

[0010] S3. Obtain E according to the text to be selected 3 The corresponding specified condition text list set Q 3 ={Q 3 1 , ……, Q 3 i , ……, Q 3 n}, where Q 3 i ={Q 3 i1 , ……, Q 3 ij , ……, Q 3 im(i)}, where Q 3 ij is the specified condition text list corresponding to E 3 ij Among them, the specified condition text list includes several specified condition texts, and the specified condition text is a conditional statement with a judgment nature obtained based on the text between two adjacent third target headings in the text to be selected

[0011] S4. According to the specific target requirements and Q 3 , obtain the target application text corresponding to the text to be selected. Among them, the second intermediate text corresponding to the text to be selected is compared with the specified condition text lists in Q 3 in sequence to obtain intermediate scores until the sum of the obtained intermediate scores is not less than the target score to determine the target application text. Among them, the target order is the order in which the target priorities in F are sorted from largest to smallest, and the intermediate score is the total score corresponding to the third target heading obtained by matching the second intermediate text corresponding to the text to be selected with the specified condition texts included in Q 3 ij

[0012] Compared with the prior art, the present invention has obvious beneficial effects. By means of the above technical solutions, a data processing system for determining the final application text provided by the present invention can achieve considerable technological progressiveness and practicality, and has wide utilization value in the industry. It has at least the following beneficial effects

[0013] ​The present invention can obtain a number of first intermediate texts based on a target text, obtain a list of candidate texts corresponding to a candidate text based on the number of first intermediate texts, and obtain a list of target application texts corresponding to the list of candidate texts to obtain a final application text. Among them, based on the candidate text, a set of target title lists is obtained, and based on the set of target title lists and the second intermediate text corresponding to the candidate text, a set of target priority lists is obtained. A set of specified condition text lists is obtained, and based on specific target requirements and the set of specified condition text lists, a target application text corresponding to the candidate text is obtained. It can be seen that by classifying policy texts, grading the classified policy texts, and screening based on the needs of candidate objects to determine the final application text, the efficiency of obtaining the final application text is improved. By deeply analyzing the target text through natural language processing technology and integrating multi-source heterogeneous data, the accuracy of the obtained final application document is relatively high.

[0014] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented in accordance with the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart of a computer program executed by a data processing system for determining a final application text according to Embodiment 1 of the present invention;

[0016] Figure 2 It is a flowchart of step S300 according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0018] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0019] Embodiment 1

[0020] This embodiment provides a data processing system for determining a final application text. The system includes: a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented, as Figure 1 shown:

[0021] S100, based on the target text, obtain a number of first intermediate texts, where the first intermediate texts are texts obtained by splitting the target text based on a preset category.

[0022] Specifically, the preset category is a preset subject category involved in the application, such as preset categories like patents, software copyrights, and qualification certificates.

[0023] S200, based on a number of first intermediate texts, obtain a list of candidate texts corresponding to the candidate text, where the list of candidate texts includes a number of candidate texts, and the candidate texts are first intermediate texts that meet the preset conditions obtained based on a number of second intermediate texts corresponding to the candidate text. The second intermediate texts are texts obtained by splitting the candidate text based on the preset category.

[0024] Specifically, the candidate text is the text corresponding to a series of achievements that the candidate object has in the historical time period. Among them, the candidate text can be understood as the achievements corresponding to the candidate object, such as candidate texts like patents, software copyrights, and qualification certificates obtained.

[0025] Furthermore, the candidate object is an object to be matched with the target text to confirm whether it meets the requirements corresponding to the target text, such as: companies, enterprises, and scientific research personnel.

[0026] Specifically, the time span corresponding to the historical time period is the same as the time span required in the target text.

[0027] Specifically, the preset condition is that the similarity between the category corresponding to the second intermediate text and the category corresponding to the first intermediate text is 1, where the category corresponding to the second intermediate text is one of the preset categories, and the category corresponding to the first intermediate text is one of the preset categories.

[0028] S300. Obtain the target application text list corresponding to the candidate text list to obtain the final application text, where the target application text list includes several target application texts, each candidate text corresponds to one target application text, and the final application text includes all the target application texts in the target application text list. Among them, in S300, the target application text is obtained through the following steps, as Figure 2 shown:

[0029] S1. Based on the candidate text, obtain the target title list set E = {E 1 , E 2 , E 3}, E 2 = {E 2 1 , ……, E 2 i , ……, E 2 n}, E 3 = {E 3 1 , ……, E 3 i , ……, E 3 n}, E 3 i = {E 3 i1 , ……, E 3 ij , ……, E 3 im(i)}, E 1 is the first target title, E 2 i is the i-th second target title in the second target title list E 1 corresponding to E 2 , E 3 ij is the j-th third target title in the third target title list corresponding to E 2 i .

[0030] Specifically, the first target title is the first-level title in the candidate text.

[0031] Specifically, the second target title is the second-level title under the first-level title in the candidate text.

[0032] Specifically, the third target title is a third-level title under a second-level title in the text to be selected.

[0033] Furthermore, those skilled in the art are aware that any method for determining a title from a text in the prior art falls within the protection scope of the present invention and will not be elaborated herein.

[0034] S2. According to E and the second intermediate text corresponding to the text to be selected, obtain the target priority list set F = {F 1 , ……, F i , ……, F n}, where F i = {F i1 , ……, F ij , ……, F im(i)}, and F ij is the target priority corresponding to the candidate path corresponding to E 3 ij .

[0035] Specifically, the candidate path corresponding to E 3 ij is E 1 -E 2 i -E 3 ij , where E 1 is connected to E 2 i , and E 2 i is connected to E 3 ij . It can be understood that a tree-like structure model including N candidate paths is formed based on E. Among them, each second target title in E 1 is connected to E 2 , and each third target title in E 2 i is connected to E 3 i .

[0036] Specifically, N is the number of candidate paths in the target key model, and N meets the following conditions:

[0037] N = ∑ n i=1 m(i).

[0038] Specifically, in S2, F ij is obtained through the following steps:

[0039] S21. Obtain a first key statement list corresponding to a second intermediate text corresponding to a text to be selected, where the first key statement list includes a plurality of first key statements, and the first key statements are key statements obtained from the second intermediate text corresponding to the text to be selected.

[0040] Specifically, those skilled in the art know that any method of obtaining key statements from a text in the prior art falls within the protection scope of the present invention and will not be elaborated herein.

[0041] Specifically, the first key statements are statements obtained by dividing the second intermediate text corresponding to the text to be selected according to punctuation marks.

[0042] S22. Obtain a second key statement list corresponding to E 3 ij where the second key statement list includes a plurality of second key statements, and the second key statements are key statements obtained from the key text corresponding to E 3 ij The key text corresponding to E 3 ij is the text between the third target title corresponding to E in the text to be selected and the third target title corresponding to E 3 ij The E 3 i(j+1) is the (j + 1)-th third target title in the third target title list corresponding to E. 3 i(j+1) is E 2 i The third target title corresponding to E

[0043] Specifically, the obtaining method of the second key statements is the same as that of the first key statements.

[0044] S23. According to the first key statement list and the second key statement list, obtain the first candidate score η 3 ij corresponding to E 1 ij where η 1 ij meets the following conditions:

[0045] η 1 ij = ∑ s(ij) r=1 λ r ij where λ r ij is E 3 ijThe number of the first key statements in the first key statement list that the corresponding r-th second key statement matches, where r = 1... s(ij), and s(ij) is E 3 ij The number of the second key statements in the corresponding second key statement list.

[0046] Specifically, the method for the first key statement to match the second key statement is to judge by obtaining the similarity between the first key statement and the second key statement. Among them, those skilled in the art know that any method for calculating the similarity between statements in the prior art falls within the protection scope of the present invention and will not be elaborated herein.

[0047] S24, obtain E 2 i The corresponding third key statement list, where the second key statement list includes several second key statements, and the third key statement is any key statement other than the second key statement among the key statements obtained from the E 2 i corresponding key text, and the E 2 i corresponding key text is the text between the E 2 i corresponding second target title and the E 2 i+1 corresponding second target title, and the E 2 i+1 is E 1 corresponding second target title list E 2 The (i + 1)-th second target title in.

[0048] S25, according to the first key statement list and the third key statement list, obtain E 2 i The corresponding second candidate score η 2 i , where η 2 i meets the following conditions:

[0049] η 2 i = ∑ f(i) e=1 β e i , β e i is E 2 i The number of the first key statements in the first key statement list that the corresponding e-th third key statement matches, where e = 1... f(i), and f(i) is E 2 iThe number of the third key statements in the corresponding third key statement list.

[0050] Specifically, the method for the third key statement to match the first key statement is the same as that for the second key statement to match the first key statement.

[0051] S26. According to η 1 ij and η 2 i , obtain F ij , where F ij meets the following conditions:

[0052] F ij = η 1 ij + η 2 i .

[0053] S3. According to the candidate text, obtain the specified condition text list set Q 3 corresponding to E 3 = {Q 3 1 , ……, Q 3 i , ……, Q 3 n}, Q 3 i = {Q 3 i1 , ……, Q 3 ij , ……, Q 3 im(i)}, Q 3 ij is the specified condition text list corresponding to E 3 ij , where the specified condition text list includes several specified condition texts, and the specified condition text is a conditional statement with a judgment nature obtained based on the text between two adjacent third target headings in the candidate text.

[0054] Specifically, the specified condition text can be understood as: in a candidate text, there will be text corresponding to the previous heading between two adjacent third headings, and there will be some conditional statements in these texts. For example, if the number of patent applications in the mechanical field is 3 or more, it is recorded as 3 points, and if the number of patent applications in the mechanical field is less than 3, it is recorded as 1 point, etc.

[0055] S4. According to the specific target requirements and Q 3 , obtain the target application text corresponding to the candidate text, where the second intermediate text corresponding to the candidate text is sequentially arranged in the target order and compared with Q 3Compare with the specified condition text list in it, and sequentially obtain the intermediate scores until the sum of the obtained intermediate scores is not less than the target score to determine the target application text. Among them, the target order is the order in which the target priorities in F are sorted from largest to smallest, and the intermediate score is the second intermediate text corresponding to the text to be selected and Q 3 ij The total score corresponding to the third target title obtained by matching the specified condition text included.

[0056] Specifically, the specific target requirement is the score that a candidate object obtained from the target requirements wants to confirm by matching the first intermediate text under a specific preset category.

[0057] Specifically, those skilled in the art know that any method of condition matching based on text description in the prior art falls within the protection scope of the present invention and will not be elaborated here.

[0058] Specifically, the target application text is the text screened from the second intermediate text corresponding to the text to be selected that meets the requirements of the text under the determined third target title. It can be understood that: after sorting the target priorities in F from largest to smallest, the third target titles corresponding to the target priorities will also obtain a new sorting order. Starting from the third target title corresponding to the largest target priority, match the second intermediate text corresponding to the text to be selected with several specified condition texts corresponding to the third target title corresponding to the largest target priority. This third target title will correspond to a score. Using the same method, sequentially obtain a score corresponding to the third target title corresponding to the second sorted target priority, and several scores will be obtained. When the sum of the obtained scores is not less than the target score, obtain the content under the third target title corresponding to the several obtained scores, and screen out the text that meets the requirements of these contents from the second intermediate text corresponding to the text to be selected as the target application text.

[0059] As described above, classifying the policy text, grading the classified policy text, and screening based on the needs of the candidate object to determine the final application text improve the efficiency of obtaining the final application text. By deeply analyzing the target text through natural language processing technology and integrating multi-source heterogeneous data, the accuracy of the obtained final application document is relatively high.

[0060] A data processing system for determining a final application text provided in this embodiment, the system includes: a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented: Based on the target text, obtain a number of first intermediate texts. Based on the number of first intermediate texts, obtain a list of candidate texts corresponding to the candidate text. Obtain a list of target application texts corresponding to the list of candidate texts to obtain the final application text. Among them, based on the candidate text, obtain a set of target title lists. According to the set of target title lists and the second intermediate text corresponding to the candidate text, obtain a set of target priority lists. Obtain a set of specified condition text lists. According to specific target requirements and the set of specified condition text lists, obtain the target application text corresponding to the candidate text. It can be seen that by classifying the policy text, grading the classified policy text, and screening based on the needs of the candidate object to determine the final application text, the efficiency of obtaining the final application text is improved. By deeply analyzing the target text through natural language processing technology and integrating multi-source heterogeneous data, the accuracy of the obtained final application document is relatively high.

[0061] Embodiment 2

[0062] This embodiment provides a data processing system for determining a final application document. The system includes: a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented:

[0063] S100, obtain the target text, where the target text is a criterion text including several rules formulated to evaluate whether the target object meets a certain condition.

[0064] Specifically, the target text is, for example: target texts such as policies. Among them, the target object can be an enterprise, a team, or an individual.

[0065] S200, obtain the candidate text corresponding to the candidate object, where the candidate object is an object to be matched with the target text to confirm whether it meets the requirements corresponding to the target text, and the candidate text is the text corresponding to a series of achievements of the candidate object within the historical time period.

[0066] Specifically, the candidate object is an object to be matched with the target text to confirm whether it meets the requirements corresponding to the target text, such as: companies, enterprises, and scientific research personnel.

[0067] Specifically, the candidate text can be understood as the achievements corresponding to the candidate object, such as candidate texts such as patents, software copyrights, and qualification certificates.

[0068] Specifically, the time span corresponding to the historical time period is the same as the time span required in the target text.

[0069] S300. When the candidate object has a target requirement, the target text and the candidate text are compared using the first processing method to obtain the final application text corresponding to the candidate object.

[0070] Specifically, the target requirement is the target score corresponding to the candidate object, where the target score is the score that the candidate object wants to confirm for matching the target text.

[0071] Specifically, in S300, the final application text is obtained through the following steps:

[0072] S301. Based on the target text, obtain a number of first intermediate texts, where the first intermediate texts are texts obtained by splitting the target text based on a preset category.

[0073] Specifically, the preset category is the theme category involved in the preset application, such as preset categories like patents, software copyrights, and qualification certificates.

[0074] S302. Based on a number of first intermediate texts, obtain a list of candidate texts corresponding to the candidate text, where the list of candidate texts includes a number of candidate texts, and the candidate texts are first intermediate texts that meet the preset conditions obtained based on a number of second intermediate texts corresponding to the candidate text. The second intermediate texts are texts obtained by splitting the candidate text based on a preset category.

[0075] Specifically, the preset condition is that the similarity between the category corresponding to the second intermediate text and the category corresponding to the first intermediate text is 1, where the category corresponding to the second intermediate text is one of the preset categories, and the category corresponding to the first intermediate text is one of the preset categories.

[0076] S303. Obtain a list of target application texts corresponding to the list of candidate texts to obtain the final application text, where the list of target application texts includes a number of target application texts, each candidate text corresponds to one target application text, and the final application text includes all the target application texts in the list of target application texts.

[0077] Specifically, in S303, the following steps are also included:

[0078] S1. Based on the candidate text, obtain a set of target title lists E = {E 1 ,E 2 ,E 3}, E 2 = {E 2 1 ,……,E 2 i ,……,E 2 n}, E3 = {E 3 1 , ……, E 3 i , ……, E 3 n}, E 3 i = {E 3 i1 , ……, E 3 ij , ……, E 3 im(i)}, E 1 is the first target title, E 2 i is E 1 the corresponding list of the second target titles E 2 the i-th second target title in, E 3 ij is E 2 i the j-th third target title in the corresponding list of the third target titles corresponding to E.

[0079] Specifically, the first target title is the first-level title in the text to be selected.

[0080] Specifically, the second target title is the second-level title under the first-level title in the text to be selected.

[0081] Specifically, the third target title is the third-level title under the second-level title in the text to be selected.

[0082] Furthermore, those skilled in the art know that any method for determining titles from text in the prior art falls within the protection scope of the present invention, which will not be elaborated here.

[0083] S2. According to E and the second intermediate text corresponding to the text to be selected, obtain the target priority list set F = {F 1 , ……, F i , ……, F n}, F i = {F i1 , ……, F ij , ……, F im(i)}, F ij is the target priority corresponding to the candidate path corresponding to E 3 ij Specifically, the candidate path corresponding to E

[0084] is E 3 ij - E 1 - E 2 i - E3 ij , where E 1 is connected to E 2 i connected, E 2 i is connected to E 3 ij connected, where it can be understood that a tree - like structure model including N candidate paths is formed based on E. Among them, in the tree - like structure model, E 1 is connected to each second target title in E 2 connected, E 2 i is connected to E 3 i connected to each third target title in E

[0085] Specifically, N is the number of candidate paths in the target key model, and N meets the following conditions:

[0086] N = ∑ n i=1 m(i).

[0087] Specifically, in S2, F is obtained through the following steps ij :

[0088] S21, obtain the first key statement list corresponding to the second intermediate text of the candidate text. Among them, the first key statement list includes several first key statements, and the first key statements are the key statements obtained from the second intermediate text corresponding to the candidate text.

[0089] Specifically, those skilled in the art know that any method of obtaining key statements from text in the prior art falls within the protection scope of the present invention and will not be elaborated here.

[0090] S22, obtain the second key statement list corresponding to E 3 ij Among them, the second key statement list includes several second key statements, and the second key statements are the key statements obtained from the key text corresponding to E 3 ij The key text corresponding to E 3 ij is the text between the third target title corresponding to E in the candidate text and the third target title corresponding to E 3 ij The E 3 i(j+1) is the text between the third target title corresponding to E 3 i(j+1) is E 2 iThe (j + 1)-th third target title in the corresponding third target title list.

[0091] S23. Obtain E according to the first key statement list and the second key statement list 3 ij The corresponding first candidate score η 1 ij , where η 1 ij Meets the following conditions:

[0092] η 1 ij = ∑ s(ij) r=1 λ r ij λ r ij Is the number of the first key statements in the first key statement list that the r-th second key statement in E 3 ij Matches, r = 1... s(ij), and s(ij) is the number of second key statements in the second key statement list corresponding to E 3 ij Corresponding to the second key statement list.

[0093] Specifically, the method for the first key statement to match the second key statement is to judge by obtaining the similarity between the first key statement and the second key statement. Among them, those skilled in the art know that any method for calculating the similarity between statements in the prior art falls within the protection scope of the present invention and will not be elaborated herein.

[0094] S24. Obtain the third key statement list corresponding to E 2 i , where the second key statement list includes several second key statements, and the third key statement is any key statement other than the second key statement among the key statements obtained from the key text corresponding to E 2 i The key text corresponding to E 2 i Is the text between the second target title corresponding to E in the candidate text and the second target title corresponding to E 2 i The second target title corresponding to E 2 i+1 The second target title corresponding to E 2 i+1 Is E 1 The (i + 1)-th second target title in the second target title list E corresponding to E 2 in.

[0095] S25. Obtain E according to the first key statement list and the third key statement list2 i The corresponding second candidate score η 2 i , where η 2 i meets the following conditions:

[0096] η 2 i = ∑ f(i) e=1 β e i , β e i is the number of the first key sentences in the first key sentence list that the e-th third key sentence corresponding to E 2 i matches, e = 1... f(i), and f(i) is the number of the third key sentences in the third key sentence list corresponding to E 2 i corresponding to.

[0097] Specifically, the method for the third key sentence to match the first key sentence is the same as the method for the second key sentence to match the first key sentence.

[0098] S26, obtain F 1 ij according to η 2 i and η ij , where F ij meets the following conditions:

[0099] F ij = η 1 ij + η 2 i .

[0100] S3, obtain the specified condition text list set Q 3 corresponding to E 3 = {Q 3 1 , ……, Q 3 i , ……, Q 3 n}, Q 3 i = {Q 3 i1 , ……, Q 3 ij , ……, Q 3 im(i)}, Q 3 ij is the 3 ijThe corresponding list of specified condition texts, where the list of specified condition texts includes several specified condition texts, and the specified condition texts are conditional statements with judgment properties obtained based on the text between two adjacent third target headings in the text to be selected.

[0101] Specifically, the specified condition text can be understood as follows: in a text to be selected, there will be text corresponding to the previous heading between two adjacent third headings, and there will be some conditional statements in these texts. For example, if the number of patent applications in the mechanical field is 3 or more, it is recorded as 3 points, and if the number of patent applications in the mechanical field is less than 3, it is recorded as 1 point, and so on.

[0102] S4. According to the specific target requirement and Q 3 , obtain the target application text corresponding to the text to be selected. Among them, the second intermediate text corresponding to the text to be selected is compared with the list of specified condition texts in Q 3 in turn, and the intermediate scores are obtained in turn until the sum of the obtained intermediate scores is not less than the target score to determine the target application text. Among them, the target order is the order in which the target priorities in F are sorted from large to small, and the intermediate score is the total score corresponding to the third target heading obtained by matching the second intermediate text corresponding to the text to be selected with the specified condition texts included in Q 3 ij Specifically, the specific target requirement is the score that a candidate object obtained from the target requirement wants to confirm for matching the first intermediate text under a specific preset category.

[0103] Specifically, those skilled in the art know that any method for condition matching based on text description in the prior art falls within the protection scope of the present invention and will not be elaborated here.

[0104]

[0105] ​Specifically, the target application text is the text that meets the requirements of the text under the determined third target title selected from the second intermediate text corresponding to the candidate text. It can be understood that: after sorting the target priorities in F from largest to smallest, the third target titles corresponding to the target priorities will also obtain a new sorting order. Starting from the third target title corresponding to the largest target priority, the second intermediate text corresponding to the candidate text is matched with several specified condition texts corresponding to the third target title corresponding to the largest target priority. This third target title corresponds to a score. Using the same method, in sequence, obtain a score corresponding to the third target title corresponding to the second sorted target priority, and several scores will be obtained. When the sum of the obtained scores is not lower than the target score, obtain the content under the third target title corresponding to the obtained scores, and screen out the text that meets the requirements of these contents from the second intermediate text corresponding to the candidate text as the target application text.

[0106] As described above, classifying the policy text, grading the classified policy text, and screening based on the needs of the candidate object to determine the final application text improve the efficiency of obtaining the final application text. By deeply analyzing the target text through natural language processing technology and integrating multi-source heterogeneous data, the accuracy of the obtained final application document is relatively high.

[0107] S400. When the candidate object has no target needs, the second processing method is used to compare the target text with the candidate text to obtain the final application text corresponding to the candidate object.

[0108] Specifically, the target application text is obtained through the following steps in S400:

[0109] S401. Based on the target text, obtain several first intermediate texts, where the first intermediate text is the text obtained by splitting the target text based on a preset category.

[0110] Specifically, the preset category is the preset theme category involved in the application, such as preset categories like patents, software copyrights, and qualification certificates.

[0111] S402. Based on several first intermediate texts, obtain the list of candidate texts corresponding to the candidate text, where the list of candidate texts includes several candidate texts, and the candidate text is the first intermediate text that meets the preset conditions obtained based on several second intermediate texts corresponding to the candidate text. The second intermediate text is the text obtained by splitting the candidate text based on a preset category.

[0112] Specifically, the preset condition is that the similarity between the category corresponding to the second intermediate text and the category corresponding to the first intermediate text is 1, where the category corresponding to the second intermediate text is one of the preset categories, and the category corresponding to the first intermediate text is one of the preset categories.

[0113] S403. Obtain the target application text list corresponding to the list of candidate texts to obtain the final application text, where the target application text list includes several target application texts, each candidate text corresponds to one target application text, and the final application text includes all the target application texts in the target application text list.

[0114] Specifically, S403 further includes the following steps:

[0115] S4031. Based on the candidate text, obtain the target title list set E = {E 1 , E 2 , E 3}, E 2 = {E 2 1 , ……, E 2 i , ……, E 2 n}, E 3 = {E 3 1 , ……, E 3 i , ……, E 3 n}, E 3 i = {E 3 i1 , ……, E 3 ij , ……, E 3 im(i)}, E 1 is the first target title, E 2 i is the i-th second target title in the second target title list E 1 corresponding to E 2 , and E 3 ij is the j-th third target title in the third target title list corresponding to E 2 i .

[0116] Specifically, the first target title is the first-level title in the candidate text.

[0117] Specifically, the second target title is the second-level title under the first-level title in the candidate text.

[0118] Specifically, the third target title is a third-level title under a second-level title in the text to be selected.

[0119] Furthermore, those skilled in the art are aware that any method for determining a title from a text in the prior art falls within the protection scope of the present invention and will not be elaborated herein.

[0120] S4032. Obtain E according to the text to be selected 3 The corresponding specified condition text list set Q 3 ={Q 3 1 , ……, Q 3 i , ……, Q 3 n}, where Q 3 i ={Q 3 i1 , ……, Q 3 ij , ……, Q 3 im(i)}, and Q 3 ij is the specified condition text list corresponding to E 3 ij . Among them, the specified condition text list includes several specified condition texts, and the specified condition text is a conditional statement with a judgment nature obtained based on the text between two adjacent third target titles in the text to be selected.

[0121] Specifically, the specified condition text can be understood as follows: in a text to be selected, there will be text corresponding to the previous title between two adjacent third titles, and there will be some conditional statements in these texts. For example, if the number of patent applications in the mechanical field is 3 or more, it is recorded as 3 points, and if the number of patent applications in the mechanical field is less than 3, it is recorded as 1 point, and so on for conditional statements.

[0122] S4033. Obtain the corresponding specified priority list set J according to Q 3 , and obtain J 3 corresponding to Q 3 ={J 3 1 , ……, J 3 i , ……, J 3 n}, where J 3 i ={J 3 i1 , ……, J 3 ij , ……, J 3 im(i)}, and J3 ij For Q 3 ij The corresponding specified priority.

[0123] Specifically, the specified priority is the intermediate score obtained by comparing the first intermediate text corresponding to the text to be selected with the specified condition text in the specified condition text list. Among them, the intermediate score is the total score corresponding to the third target title obtained by matching the text included in the first intermediate text corresponding to the text to be selected with the specified condition text of Q 3 ij Including the specified condition text.

[0124] S4034, according to J 3 , obtain the target application document corresponding to the text to be selected. Among them, when J 3 ij ≠0, the text in the first intermediate text corresponding to the text to be selected that matches the text content under the third target title corresponding to J 3 ij Is used as the target application text corresponding to the text to be selected.

[0125] As described above, according to whether the candidate object has a target requirement, the first processing method and the second processing method are respectively used to match the target text and the candidate text to obtain the final application text. Based on the different requirements of the candidate object, different matching methods are used for its result text to obtain the final application text, so that the accuracy of the obtained final application text is relatively high.

[0126] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for the purpose of illustration and not for the purpose of limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A data processing system for determining a final application text, characterized in that: The system comprises: a processor and a memory storing a computer program, and when the computer program is executed by the processor, the following steps are implemented: S100, based on the target text, obtaining a plurality of first intermediate texts, wherein the first intermediate texts are texts obtained by splitting the target text based on preset categories; S200, based on a plurality of first intermediate texts, obtaining a list of to-be-selected texts corresponding to the candidate texts, wherein the list of to-be-selected texts includes a plurality of to-be-selected texts, the to-be-selected texts being first intermediate texts that meet preset conditions and are obtained based on a plurality of second intermediate texts corresponding to the candidate texts, and the second intermediate texts being texts obtained by splitting the candidate texts based on preset categories; S300, obtaining a target application text list corresponding to a list of to-be-selected texts to obtain a final application text, wherein the target application text list includes a plurality of target application texts, each to-be-selected text corresponds to a target application text, and the final application text includes all target application texts in the target application text list, wherein the target application text is obtained in S300 by the following steps: S1, based on the selected text, obtain the target title list set E = {E 1 , E 2 , E 3 }, E 2 ={E 2 1, ..., E 2 i ,……,E 2 n }, E 3 ={E 3 1, ..., E 3 i ,……,E 3 n }, E 3 i ={E 3 i1 ,……,E 3 ij ,……,E 3 im(i) }, E 1 For the first target title, E 2 i For E 1 Corresponding second target title list E 2 The i-th second target title in E 3 ij For E 2 i The jth third target title in the corresponding third target title list; S2, according to E and the second intermediate text corresponding to the selected text, obtain the target priority list set F = {F1, ..., F i , ..., F n }, F i ={F i1 , ..., F ij , ..., F im(i) }, F ij For E 3 ij The target priority of the corresponding candidate path; S3, according to the text to be selected, obtain E 3 The corresponding specified condition text list set Q 3 = {Q 3 1, ..., Q 3 i , ..., Q 3 n }, Q 3 i = {Q 3 i1 , ..., Q 3 ij , ..., Q 3 im(i) }, Q 3 ij For E 3 ij A corresponding specified condition text list, wherein the specified condition text list includes a plurality of specified condition texts, and the specified condition texts are conditional statements with judgment properties obtained based on the text between two adjacent third target titles in the selected text; S4, according to specific target needs and Q 3 , obtain the target application text corresponding to the selected text, wherein the second intermediate text corresponding to the selected text is sequentially compared with Q in the target order 3 The target application text is determined by comparing the specified condition text list in the target text list with the target text list, obtaining the intermediate scores in turn until the sum of the intermediate scores obtained is not less than the target score, wherein the target order is the order in which the target priorities in F are sorted from large to small, and the intermediate score is the difference between the second intermediate text corresponding to the candidate text and Q. 3 ij The total score corresponding to the third target title is obtained by matching the specified conditional text.

2. The data processing system for determining the final application text according to claim 1, characterized in that: The target text is a criterion text including several rules formulated to evaluate whether the target object meets certain conditions.

3. The data processing system for determining the final application text according to claim 1, characterized in that: The candidate text is a text corresponding to a series of achievements possessed by the candidate object in a historical period of time, wherein the candidate object is an object to be matched with the target text to confirm whether it meets the corresponding requirements of the target text.

4. The data processing system for determining the final application text according to claim 3, characterized in that: The time span corresponding to the historical time period is consistent with the time span required in the target text.

5. The data processing system for determining the final application text according to claim 1, characterized in that: In S2, F is obtained by the following steps: ij : S21, obtaining a first key sentence list corresponding to a second intermediate text corresponding to the text to be selected, wherein the first key sentence list includes a plurality of first key sentences, and the first key sentences are key sentences obtained from the second intermediate text corresponding to the text to be selected; S22, get E 3 ij The corresponding second key sentence list, wherein the second key sentence list includes a plurality of second key sentences, and the second key sentence is from E 3 ij The key sentences obtained from the corresponding key text, the E 3 ij The corresponding key text is E in the selected text 3 ij The corresponding third target title is E 3 i(j+1) The text between the corresponding third target titles, E 3 i(j+1) For E 2 i The j+1th third target title in the corresponding third target title list; S23, obtaining E according to the first key sentence list and the second key sentence list 3 ij The corresponding first candidate score η 1 ij , where η 1 ij Meet the following conditions: η 1 ij =∑ s(ij) r=1 λ r ij ,λ r ij For E 3 ij The number of first key sentences in the first key sentence list that the corresponding r-th second key sentence matches, r=1...s(ij), s(ij) is E 3 ij The number of second key sentences in the corresponding second key sentence list; S24, Get E 2 i The corresponding third key sentence list, wherein the second key sentence list includes a plurality of second key sentences, and the third key sentence is from E 2 i Any key sentence other than the second key sentence in the key sentence obtained from the corresponding key text, the E 2 i The corresponding key text is E in the selected text 2 i The corresponding second target title is E 2 i+1 The text between the corresponding second target titles, E 2 i+1 For E 1 Corresponding second target title list E 2 The i+1th second target title in ; S25, obtaining E according to the first key sentence list and the third key sentence list 2 i The corresponding second candidate score η 2 i , where η 2 i Meet the following conditions: η 2 i =∑ f(i) e=1 β e i , β e i For E 2 i The number of corresponding e-th third key sentence matched to the first key sentence in the first key sentence list, e=1...f(i), f(i) is E 2 i The number of third key sentences in the corresponding third key sentence list. S26, according to η 1 ij and η 2 i , get F ij , where F ij Meet the following conditions: F ij =the 1 ij +n 2 i 。 6. The data processing system for determining the final application text according to claim 5, characterized in that: The method for matching the first key sentence with the second key sentence is to determine the similarity between the first key sentence and the second key sentence.

7. The data processing system for determining the final application text according to claim 1, characterized in that: The target application text is a text that meets the requirements of the text under the determined third target title and is screened out from the second intermediate text corresponding to the selected text.

Citation Information

Patent Citations

  • Text importance calculation method and device and equipment and storage medium

    CN109670183A

  • Method, equipment and medium for mining PDF (Portable Document Format) file

    CN114116616A

  • Government affair service field multi-strategy fusion dialogue method based on knowledge graph

    CN116628172A

  • Policy knowledge query method and device, electronic equipment and storage medium

    CN118093852A

  • Policy matching system based on big data

    CN118394992A