A data processing system for determining a final application text
By constructing a data processing system for the target text and utilizing natural language processing technology and text classification, the problems of low efficiency and insufficient accuracy in existing technologies have been solved, achieving efficient and accurate final application text determination.
Patent Information
- Application Number
- CN202510241493.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-03-03
AI Technical Summary
In the existing technology, the methods for determining the final application text are inefficient and have low accuracy, failing to effectively utilize the needs of candidate objects for intelligent parsing.
A data processing system is used to acquire intermediate text of the target text, construct a list of target titles, priorities, and specified condition texts, classify and rank the texts based on the needs of the candidate objects, perform deep analysis using natural language processing technology, and integrate multi-source heterogeneous data.
It improves the efficiency and accuracy of determining the final application text, meets the specific needs of candidate candidates, and achieves efficient text screening and matching.
Smart Images

Figure CN120144748B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of text processing technology, and in particular to a data processing system for determining the final application text. Background Technology
[0002] In the context of today's global economic integration and rapidly changing policy environment, we face more and more development opportunities and challenges. In order to promote the development of specific industries, encourage technological innovation, and optimize industrial structure, governments and regional organizations frequently introduce various policies and measures. These policies often contain a series of specific standards and requirements, aiming to select eligible entities to enjoy policy benefits. However, due to the large amount of content included in the policies and the numerous achievements corresponding to enterprises or related personnel, how to accurately match and determine the final application text has become an urgent problem to be solved.
[0003] In existing technologies, the method for determining the final application text is as follows: it is done manually by reading policy documents one by one, comparing the standards in them with one's own situation, and interpreting and analyzing them manually to understand the specific content and requirements of the policies in order to determine the final application text. Due to the huge amount of data, the text screening based on the needs of the candidate objects is not considered, resulting in low efficiency in obtaining the final application text. Furthermore, the lack of intelligent parsing of the target text leads to low accuracy of the obtained final application documents. Summary of the Invention
[0004] To address the aforementioned technical problems, the present invention adopts the following technical solution: a data processing system for determining the final application text, the system comprising: a processor and a memory storing a computer program, wherein when the computer program is executed by the processor, the following steps are implemented:
[0005] S100, based on the target text, obtain several first intermediate texts, wherein the first intermediate texts are texts obtained by splitting the target text based on a preset category.
[0006] S200, based on several first intermediate texts, obtain a list of candidate texts corresponding to candidate texts, wherein the list of candidate texts includes several candidate texts, and the candidate texts are first intermediate texts that meet preset conditions obtained based on several second intermediate texts corresponding to candidate texts, and the second intermediate texts are texts obtained by splitting candidate texts based on preset categories.
[0007] S300, obtain the target application text list corresponding to the candidate text list to obtain the final application text, wherein the target application text list includes several target application texts, each candidate text corresponds to one target application text, and the final application text includes all target application texts in the target application text list. In S300, the target application texts are obtained through the following steps:
[0008] S1, Based on the candidate text, obtain the target title list set E = {E 1 E 2 E 3}, E 2 ={E 2 1, ..., E 2 i , ..., E 2 n}, E 3 ={E 3 1, ..., E 3 i , ..., E 3 n}, E 3 i ={E 3 i1 , ..., E 3 ij , ..., E 3 im(i)}, E 1 For the first target title, E 2 i For E 1 The corresponding second target title list E 2 The i-th second target title in E 3 ij For E 2 i The j-th third target title in the corresponding third target title list.
[0009] S2, based on E and the second intermediate text corresponding to the candidate text, obtain the target priority list set F = {F1, ..., F2}. i , ..., F n}, F i ={F i1 , ..., F ij , ..., F im(i)}, F ij For E 3 ij The target priority of the corresponding candidate path.
[0010] S3, based on the candidate text, obtain E 3The corresponding specified conditional text list set Q 3 ={Q 3 1, ..., Q 3 i Q 3 n}, Q 3 i ={Q 3 i1 Q 3 ij Q 3 im(i)}, Q 3 ij For E 3 ij The corresponding list of specified condition texts includes several specified condition texts, which are conditional statements with judgment properties obtained based on the text between two adjacent third target titles in the candidate text.
[0011] S4, based on specific target requirements and Q 3 Obtain the target application text corresponding to the candidate text, wherein the second intermediate text corresponding to the candidate text is sequentially compared with Q according to the target order. 3 The target application text is determined by comparing the specified conditional text list in F with the target text list in Q. Intermediate scores are obtained sequentially until the sum of the intermediate scores is not less than the target score. The target order is determined by sorting the target priorities in F from highest to lowest. The intermediate score is the score between the second intermediate text corresponding to the candidate text and the score in Q. 3 ij The total score is obtained by matching the specified conditional text to the third target title.
[0012] Compared with the prior art, the present invention has significant advantages. Through the above technical solution, the data processing system for determining the final application text provided by the present invention achieves considerable technical progress and practicality, and has broad industrial application value. It has at least the following advantages:
[0013] This invention can obtain a number of first intermediate texts based on a target text, obtain a list of candidate texts corresponding to the candidate texts based on the number of first intermediate texts, obtain a list of target application texts corresponding to the list of candidate texts, and obtain the final application text. Specifically, based on the candidate texts, a target title list is obtained; based on the target title list and the second intermediate texts corresponding to the candidate texts, a target priority list is obtained; a list of texts with specified conditions is obtained; and based on specific target requirements and the list of texts with specified conditions, the target application texts corresponding to the candidate texts are obtained. It can be seen that classifying policy texts, ranking the classified policy texts, and filtering based on the requirements of candidate objects to determine the final application text improves the efficiency of obtaining the final application text. Furthermore, by using natural language processing technology to deeply analyze the target text and integrating multi-source heterogeneous data, the accuracy of the obtained final application document is high.
[0014] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0015] Figure 1 A flowchart of the execution computer program for a data processing system for determining the final application text, provided in Embodiment 1 of the present invention;
[0016] Figure 2 This is a flowchart of step S300 provided in Embodiment 1 of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0019] Example 1
[0020] This embodiment provides a data processing system for determining the final application text. The system includes a processor and a memory storing a computer program. When the computer program is executed by the processor, it performs the following steps, such as... Figure 1 As shown:
[0021] S100, based on the target text, obtain several first intermediate texts, wherein the first intermediate texts are texts obtained by splitting the target text based on a preset category.
[0022] Specifically, the preset category is a preset subject category involved in the application, such as preset categories like patents, software copyrights, and qualification certificates.
[0023] S200, based on several first intermediate texts, obtain a list of candidate texts corresponding to candidate texts, wherein the list of candidate texts includes several candidate texts, and the candidate texts are first intermediate texts that meet preset conditions obtained based on several second intermediate texts corresponding to candidate texts, and the second intermediate texts are texts obtained by splitting candidate texts based on preset categories.
[0024] Specifically, the candidate text refers to the text corresponding to a series of achievements of the candidate object within a historical period. The candidate text can be understood as the achievements corresponding to the candidate object, such as the candidate texts of patents, software copyrights, and qualification certificates obtained.
[0025] Furthermore, the candidate object is an object that needs to be matched with the target text to confirm whether it meets the requirements corresponding to the target text, such as: companies, enterprises, and researchers.
[0026] Specifically, the time span corresponding to the historical time period is consistent with the time span required in the target text.
[0027] Specifically, the preset condition is: the similarity between the category corresponding to the second intermediate text and the category corresponding to the first intermediate text is 1, wherein the category corresponding to the second intermediate text is one of the preset categories, and the category corresponding to the first intermediate text is one of the preset categories.
[0028] S300: Obtain the target application text list corresponding to the candidate text list to obtain the final application text. The target application text list includes several target application texts, with each candidate text corresponding to one target application text. The final application text includes all target application texts in the target application text list. In S300, the target application texts are obtained through the following steps: Figure 2 As shown:
[0029] S1, Based on the candidate text, obtain the target title list set E = {E 1 E 2 E 3}, E 2 ={E 2 1, ..., E 2 i , ..., E 2 n}, E 3 ={E 3 1, ..., E 3 i , ..., E 3 n}, E 3 i ={E 3 i1 , ..., E 3 ij , ..., E 3 im(i)}, E 1 For the first target title, E 2 i For E 1 The corresponding second target title list E 2 The i-th second target title in E 3 ij For E 2 i The j-th third target title in the corresponding third target title list.
[0030] Specifically, the first target title is a first-level heading in the text to be selected.
[0031] Specifically, the second target title is the second-level title under the first-level title in the text to be selected.
[0032] Specifically, the third target title is a third-level title under a second-level title in the candidate text.
[0033] Furthermore, those skilled in the art will recognize that any method for determining the title from text in the prior art falls within the protection scope of this invention, and will not be elaborated further here.
[0034] S2, based on E and the second intermediate text corresponding to the candidate text, obtain the target priority list set F = {F1, ..., F2}. i , ..., F n}, F i ={F i1 , ..., F ij , ..., F im(i)}, F ij For E 3 ij The target priority of the corresponding candidate path.
[0035] Concrete, E 3 ij The corresponding candidate path is E 1 -E 2 i -E 3 ij , of which E 1 With E 2 i Connection, E 2 i With E 3 ij The connection, which can be understood as constructing a tree structure model based on E, including N candidate paths, wherein E in the tree structure model 1 With E 2 Each second target title link in E 2 i With E 3 i Each third target header in the link.
[0036] Specifically, N represents the number of candidate paths in the target critical model, and N satisfies the following condition:
[0037] N = ∑ n i=1 m(i).
[0038] Specifically, F is obtained in S2 through the following steps. ij :
[0039] S21, obtain a list of first key statements corresponding to the second intermediate text corresponding to the candidate text, wherein the list of first key statements includes several first key statements, and the first key statements are key statements obtained from the second intermediate text corresponding to the candidate text.
[0040] Specifically, as those skilled in the art will know, any existing method for extracting key sentences from text falls within the protection scope of this invention, and will not be elaborated further here.
[0041] Specifically, the first key statement is the statement obtained by dividing the selected text into second intermediate texts according to punctuation marks.
[0042] S22, obtain E 3 ij The corresponding second key statement list, wherein the second key statement list includes several second key statements, the second key statement being from E 3 ij The key statements obtained from the corresponding key text, the E 3 ij The corresponding key text is E in the candidate text. 3 ij The corresponding third target title and E 3 i(j+1) The text between the corresponding third target headings, E 3 i(j+1) For E 2 i The (j+1)th third target title in the corresponding third target title list.
[0043] Specifically, the method for obtaining the second key statement is the same as the method for obtaining the first key statement.
[0044] S23, Based on the first key statement list and the second key statement list, obtain E 3 ij The corresponding first candidate score η 1 ij , where η 1 ij The following conditions must be met:
[0045] η 1 ij =∑ s(ij) r=1 λ r ij , λ r ij For E 3 ijThe number of times the r-th second key statement matches the first key statement in the list of first key statements, r = 1...s(ij), where s(ij) is E. 3 ij The number of second key statements in the corresponding list of second key statements.
[0046] Specifically, the method for matching the first key statement with the second key statement is to determine the similarity between the first key statement and the second key statement. As those skilled in the art know, any method for calculating the similarity between statements in the prior art falls within the protection scope of this invention, and will not be elaborated here.
[0047] S24, obtain E 2 i The corresponding third key statement list, wherein the second key statement list includes several second key statements, and the third key statement is from E 2 i The E is any key statement obtained from the corresponding key text, excluding the second key statement. 2 i The corresponding key text is E in the candidate text. 2 i The corresponding second target title and E 2 i+1 The text between the corresponding second target headings, E 2 i+1 For E 1 The corresponding second target title list E 2 The (i+1)th second target title in the [theory].
[0048] S25, based on the first key statement list and the third key statement list, obtain E 2 i The corresponding second candidate score η 2 i , where η 2 i The following conditions must be met:
[0049] η 2 i =∑ f(i) e=1 β e i ,β e i For E 2 i The number of times the e-th third key statement matches the first key statement in the list of first key statements, e = 1...f(i), where f(i) is E. 2 iThe number of third key statements in the corresponding third key statement list.
[0050] Specifically, the method for matching the third key statement with the first key statement is the same as the method for matching the second key statement with the first key statement.
[0051] S26, according to η 1 ij and η 2 i , obtain F ij , of which F ij The following conditions must be met:
[0052] F ij =η 1 ij +η 2 i .
[0053] S3, based on the candidate text, obtain E 3 The corresponding specified conditional text list set Q 3 ={Q 3 1, ..., Q 3 i Q 3 n}, Q 3 i ={Q 3 i1 Q 3 ij Q 3 im(i)}, Q 3 ij For E 3 ij The corresponding list of specified condition texts includes several specified condition texts, which are conditional statements with judgment properties obtained based on the text between two adjacent third target titles in the candidate text.
[0054] Specifically, the specified condition text can be understood as follows: in a candidate text, there will be text corresponding to the previous title between two adjacent third titles. These texts will contain certain conditional statements, such as: if the number of patent applications in the mechanical field is 3 or more, it is scored as 3 points; if the number of patent applications in the mechanical field is less than 3, it is scored as 1 point, etc.
[0055] S4, based on specific target requirements and Q 3 Obtain the target application text corresponding to the candidate text, wherein the second intermediate text corresponding to the candidate text is sequentially compared with Q according to the target order. 3The target application text is determined by comparing the specified conditional text list in F with the target text list in Q. Intermediate scores are obtained sequentially until the sum of the intermediate scores is not less than the target score. The target order is determined by sorting the target priorities in F from highest to lowest. The intermediate score is the score between the second intermediate text corresponding to the candidate text and the score in Q. 3 ij The total score is obtained by matching the specified conditional text to the third target title.
[0056] Specifically, the specific target requirement refers to the score that the candidate object obtained from the target requirement wants to confirm when matching the first intermediate text under a certain preset category.
[0057] Specifically, as those skilled in the art will know, any existing method for condition matching based on text description falls within the protection scope of this invention, and will not be elaborated further here.
[0058] Specifically, the target application text is the text selected from the second intermediate text corresponding to the candidate text that meets the requirements of the text under the determined third target title. This can be understood as follows: after sorting the target priorities in F from largest to smallest, the third target titles corresponding to the target priorities will also get a new sorting order. Starting from the third target title corresponding to the highest target priority, the second intermediate text corresponding to the candidate text is matched with several specified condition texts corresponding to the third target title corresponding to the highest target priority. This third target title will correspond to a score. In the same way, the score corresponding to the third target title corresponding to the second target priority is obtained in order. Several scores will be obtained. When the sum of the obtained scores is not lower than the target score, the content under the third target title corresponding to the obtained scores is obtained. The text that meets these content requirements is selected from the second intermediate text corresponding to the candidate text as the target application text.
[0059] The above-mentioned classification of policy texts, the grading of classified policy texts, and the screening based on the needs of candidate objects to determine the final application text improve the efficiency of obtaining the final application text. By using natural language processing technology to perform in-depth analysis of the target text and integrating multi-source heterogeneous data, the accuracy of the obtained final application documents is high.
[0060] This embodiment provides a data processing system for determining the final application text. The system includes a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented: based on the target text, a number of first intermediate texts are obtained; based on the number of first intermediate texts, a list of candidate texts corresponding to the candidate texts is obtained; a list of target application texts corresponding to the list of candidate texts is obtained to obtain the final application text. Specifically, based on the candidate texts, a list of target titles is obtained; based on the list of target titles and the second intermediate texts corresponding to the candidate texts, a list of target priorities is obtained; a list of texts with specified conditions is obtained; and based on specific target requirements and the list of texts with specified conditions, the target application texts corresponding to the candidate texts are obtained. It can be seen that classifying policy texts, ranking the classified policy texts, and filtering based on the requirements of candidate objects to determine the final application text improves the efficiency of obtaining the final application text. Furthermore, by using natural language processing technology to deeply analyze the target text and integrating multi-source heterogeneous data, the accuracy of the obtained final application document is high.
[0061] Example 2
[0062] This embodiment provides a data processing system for determining the final application document. The system includes a processor and a memory storing a computer program. When the computer program is executed by the processor, the following steps are implemented:
[0063] S100, Obtain target text, wherein the target text is a criterion text comprising several rules for evaluating whether a target object meets certain conditions.
[0064] Specifically, the target text may be, for example, policy text, and the target object may be an enterprise, team, or individual.
[0065] S200, obtain the candidate text corresponding to the candidate object, wherein the candidate object is an object to be matched with the target text to confirm whether it meets the requirements corresponding to the target text, and the candidate text is the text corresponding to a series of achievements of the candidate object in the historical time period.
[0066] Specifically, the candidate objects are objects that need to be matched with the target text to confirm whether they meet the requirements corresponding to the target text, such as companies, enterprises, and researchers.
[0067] Specifically, the candidate text can be understood as the achievement corresponding to the candidate object, such as candidate texts for patents, software copyrights, and qualification certificates.
[0068] Specifically, the time span corresponding to the historical time period is consistent with the time span required in the target text.
[0069] S300: When a candidate object has a target requirement, the first processing method is used to compare the target text with the candidate text to obtain the final application text corresponding to the candidate object.
[0070] Specifically, the target requirement is the target score corresponding to the candidate object, wherein the target score is the score that the candidate object wants to confirm for matching the target text.
[0071] Specifically, the final application text is obtained through the following steps in S300:
[0072] S301, based on the target text, obtain several first intermediate texts, wherein the first intermediate texts are texts obtained by splitting the target text based on a preset category.
[0073] Specifically, the preset category is a preset subject category involved in the application, such as preset categories like patents, software copyrights, and qualification certificates.
[0074] S302, based on several first intermediate texts, obtain a list of candidate texts corresponding to candidate texts, wherein the list of candidate texts includes several candidate texts, and the candidate texts are first intermediate texts that meet preset conditions obtained based on several second intermediate texts corresponding to candidate texts, and the second intermediate texts are texts obtained by splitting candidate texts based on preset categories.
[0075] Specifically, the preset condition is: the similarity between the category corresponding to the second intermediate text and the category corresponding to the first intermediate text is 1, wherein the category corresponding to the second intermediate text is one of the preset categories, and the category corresponding to the first intermediate text is one of the preset categories.
[0076] S303, obtain the target application text list corresponding to the candidate text list to obtain the final application text, wherein the target application text list includes a number of target application texts, each candidate text corresponds to one target application text, and the final application text includes all target application texts in the target application text list.
[0077] Specifically, S303 also includes the following steps:
[0078] S1, Based on the candidate text, obtain the target title list set E = {E 1 E 2 E 3}, E 2 ={E 2 1, ..., E 2 i , ..., E 2 n}, E 3 ={E3 1, ..., E 3 i , ..., E 3 n}, E 3 i ={E 3 i1 , ..., E 3 ij , ..., E 3 im(i)}, E 1 For the first target title, E 2 i For E 1 The corresponding second target title list E 2 The i-th second target title in E 3 ij For E 2 i The j-th third target title in the corresponding third target title list.
[0079] Specifically, the first target title is a first-level heading in the text to be selected.
[0080] Specifically, the second target title is the second-level title under the first-level title in the text to be selected.
[0081] Specifically, the third target title is a third-level title under a second-level title in the candidate text.
[0082] Furthermore, those skilled in the art will recognize that any method for determining the title from text in the prior art falls within the protection scope of this invention, and will not be elaborated further here.
[0083] S2, based on E and the second intermediate text corresponding to the candidate text, obtain the target priority list set F = {F1, ..., F2}. i , ..., F n}, F i ={F i1 , ..., F ij , ..., F im(i)}, F ij For E 3 ij The target priority of the corresponding candidate path.
[0084] Concrete, E 3 ij The corresponding candidate path is E 1 -E 2 i -E 3 ij , of which E 1 With E2 i Connection, E 2 i With E 3 ij The connection, which can be understood as constructing a tree structure model based on E, including N candidate paths, wherein E in the tree structure model 1 With E 2 Each second target title link in E 2 i With E 3 i Each third target header in the link.
[0085] Specifically, N represents the number of candidate paths in the target critical model, and N satisfies the following condition:
[0086] N = ∑ n i=1 m(i).
[0087] Specifically, F is obtained in S2 through the following steps. ij :
[0088] S21, obtain a list of first key statements corresponding to the second intermediate text corresponding to the candidate text, wherein the list of first key statements includes several first key statements, and the first key statements are key statements obtained from the second intermediate text corresponding to the candidate text.
[0089] Specifically, as those skilled in the art will know, any existing method for extracting key sentences from text falls within the protection scope of this invention, and will not be elaborated further here.
[0090] S22, obtain E 3 ij The corresponding second key statement list, wherein the second key statement list includes several second key statements, the second key statement being from E 3 ij The key statements obtained from the corresponding key text, the E 3 ij The corresponding key text is E in the candidate text. 3 ij The corresponding third target title and E 3 i(j+1) The text between the corresponding third target headings, E 3 i(j+1) For E 2 i The (j+1)th third target title in the corresponding third target title list.
[0091] S23, Based on the first key statement list and the second key statement list, obtain E 3 ij The corresponding first candidate score η 1 ij , where η 1 ij The following conditions must be met:
[0092] η 1 ij =∑ s(ij) r=1 λ r ij , λ r ij For E 3 ij The number of times the r-th second key statement matches the first key statement in the list of first key statements, r = 1...s(ij), where s(ij) is E. 3 ij The number of second key statements in the corresponding list of second key statements.
[0093] Specifically, the method for matching the first key statement with the second key statement is to determine the similarity between the first key statement and the second key statement. As those skilled in the art know, any method for calculating the similarity between statements in the prior art falls within the protection scope of this invention, and will not be elaborated here.
[0094] S24, obtain E 2 i The corresponding third key statement list, wherein the second key statement list includes several second key statements, and the third key statement is from E 2 i The E is any key statement obtained from the corresponding key text, excluding the second key statement. 2 i The corresponding key text is E in the candidate text. 2 i The corresponding second target title and E 2 i+1 The text between the corresponding second target headings, E 2 i+1 For E 1 The corresponding second target title list E 2 The (i+1)th second target title in the [theory].
[0095] S25, based on the first key statement list and the third key statement list, obtain E 2 i The corresponding second candidate score η 2 i, where η 2 i The following conditions must be met:
[0096] η 2 i =∑ f(i) e=1 β e i ,β e i For E 2 i The number of times the e-th third key statement matches the first key statement in the list of first key statements, e = 1...f(i), where f(i) is E. 2 i The number of third key statements in the corresponding third key statement list.
[0097] Specifically, the method for matching the third key statement with the first key statement is the same as the method for matching the second key statement with the first key statement.
[0098] S26, according to η 1 ij and η 2 i , obtain F ij , of which F ij The following conditions must be met:
[0099] F ij =η 1 ij +η 2 i .
[0100] S3, based on the candidate text, obtain E 3 The corresponding specified conditional text list set Q 3 ={Q 3 1, ..., Q 3 i Q 3 n}, Q 3 i ={Q 3 i1 Q 3 ij Q 3 im(i)}, Q 3 ij For E 3 ijThe corresponding list of specified condition texts includes several specified condition texts, which are conditional statements with judgment properties obtained based on the text between two adjacent third target titles in the candidate text.
[0101] Specifically, the specified condition text can be understood as follows: in a candidate text, there will be text corresponding to the previous title between two adjacent third titles. These texts will contain certain conditional statements, such as: if the number of patent applications in the mechanical field is 3 or more, it is scored as 3 points; if the number of patent applications in the mechanical field is less than 3, it is scored as 1 point, etc.
[0102] S4, based on specific target requirements and Q 3 Obtain the target application text corresponding to the candidate text, wherein the second intermediate text corresponding to the candidate text is sequentially compared with Q according to the target order. 3 The target application text is determined by comparing the specified conditional text list in F with the target text list in Q. Intermediate scores are obtained sequentially until the sum of the intermediate scores is not less than the target score. The target order is determined by sorting the target priorities in F from highest to lowest. The intermediate score is the score between the second intermediate text corresponding to the candidate text and the score in Q. 3 ij The total score is obtained by matching the specified conditional text to the third target title.
[0103] Specifically, the specific target requirement refers to the score that the candidate object obtained from the target requirement wants to confirm when matching the first intermediate text under a certain preset category.
[0104] Specifically, as those skilled in the art will know, any existing method for condition matching based on text description falls within the protection scope of this invention, and will not be elaborated further here.
[0105] Specifically, the target application text is the text selected from the second intermediate text corresponding to the candidate text that meets the requirements of the text under the determined third target title. This can be understood as follows: after sorting the target priorities in F from largest to smallest, the third target titles corresponding to the target priorities will also get a new sorting order. Starting from the third target title corresponding to the highest target priority, the second intermediate text corresponding to the candidate text is matched with several specified condition texts corresponding to the third target title corresponding to the highest target priority. This third target title will correspond to a score. In the same way, the score corresponding to the third target title corresponding to the second target priority is obtained in order. Several scores will be obtained. When the sum of the obtained scores is not lower than the target score, the content under the third target title corresponding to the obtained scores is obtained. The text that meets these content requirements is selected from the second intermediate text corresponding to the candidate text as the target application text.
[0106] The above-mentioned classification of policy texts, the grading of classified policy texts, and the screening based on the needs of candidate objects to determine the final application text improve the efficiency of obtaining the final application text. By using natural language processing technology to perform in-depth analysis of the target text and integrating multi-source heterogeneous data, the accuracy of the obtained final application documents is high.
[0107] S400, when the candidate does not have a target requirement, the second processing method is used to compare the target text with the candidate text to obtain the final application text corresponding to the candidate.
[0108] Specifically, in S400, the target application text is obtained through the following steps:
[0109] S401, Based on the target text, obtain a number of first intermediate texts, wherein the first intermediate texts are texts obtained by splitting the target text based on a preset category.
[0110] Specifically, the preset category is a preset subject category involved in the application, such as preset categories like patents, software copyrights, and qualification certificates.
[0111] S402, based on several first intermediate texts, obtain a list of candidate texts corresponding to candidate texts, wherein the list of candidate texts includes several candidate texts, and the candidate texts are first intermediate texts that meet preset conditions obtained based on several second intermediate texts corresponding to candidate texts, and the second intermediate texts are texts obtained by splitting candidate texts based on preset categories.
[0112] Specifically, the preset condition is: the similarity between the category corresponding to the second intermediate text and the category corresponding to the first intermediate text is 1, wherein the category corresponding to the second intermediate text is one of the preset categories, and the category corresponding to the first intermediate text is one of the preset categories.
[0113] S403, obtain the target application text list corresponding to the candidate text list to obtain the final application text, wherein the target application text list includes a number of target application texts, each candidate text corresponds to one target application text, and the final application text includes all target application texts in the target application text list.
[0114] Specifically, S403 also includes the following steps:
[0115] S4031, Based on the candidate text, obtain the target title list set E = {E 1 E 2 E 3}, E 2 ={E 2 1, ..., E 2 i , ..., E 2 n}, E 3 ={E 3 1, ..., E 3 i , ..., E 3 n}, E 3 i ={E 3 i1 , ..., E 3 ij , ..., E 3 im(i)}, E 1 For the first target title, E 2 i For E 1 The corresponding second target title list E 2 The i-th second target title in E 3 ij For E 2 i The j-th third target title in the corresponding third target title list.
[0116] Specifically, the first target title is a first-level heading in the text to be selected.
[0117] Specifically, the second target title is the second-level title under the first-level title in the text to be selected.
[0118] Specifically, the third target title is a third-level title under a second-level title in the candidate text.
[0119] Furthermore, those skilled in the art will recognize that any method for determining the title from text in the prior art falls within the protection scope of this invention, and will not be elaborated further here.
[0120] S4032, Based on the candidate text, obtain E 3 The corresponding specified conditional text list set Q 3 ={Q 3 1, ..., Q 3 i Q 3 n}, Q 3 i ={Q 3 i1 Q 3 ij Q 3 im(i)}, Q 3 ij For E 3 ij The corresponding list of specified condition texts includes several specified condition texts, which are conditional statements with judgment properties obtained based on the text between two adjacent third target titles in the candidate text.
[0121] Specifically, the specified condition text can be understood as follows: in a candidate text, there will be text corresponding to the previous title between two adjacent third titles. These texts will contain certain conditional statements, such as: if the number of patent applications in the mechanical field is 3 or more, it is scored as 3 points; if the number of patent applications in the mechanical field is less than 3, it is scored as 1 point, etc.
[0122] S4033, according to Q 3 Get Q 3 The corresponding priority list set J 3 ={J 3 1, ..., J 3 i , ..., J 3 n}, J 3 i ={J 3 i1 , ..., J 3 ij , ..., J 3 im(i)}, J 3 ij For Q 3ij The corresponding specified priority.
[0123] Specifically, the specified priority is the intermediate score obtained by comparing the first intermediate text corresponding to the candidate text with the specified condition text in the specified condition text list, wherein the intermediate score is the score between the first intermediate text corresponding to the candidate text and the Q... 3 ij The total score is obtained by matching the specified conditional text to the third target title.
[0124] S4034, according to J 3 Obtain the target application file corresponding to the candidate text, where J 3 ij When ≠0, the first intermediate text corresponding to the candidate text will be compared with J. 3 ij The text matching the text content under the corresponding third target title is used as the target application text corresponding to the candidate text.
[0125] As described above, depending on whether the candidate object has a target requirement, the target text and candidate text are matched using the first processing method and the second processing method respectively to obtain the final application text. Based on the different requirements of the candidate object, different matching methods are used for its result text to obtain the final application text, so that the accuracy of obtaining the final application text is high.
[0126] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.
Claims
1. A data processing system for determining a final application text, characterized in that The system comprises a processor and a memory storing a computer program, when the computer program is executed by the processor, the following steps are implemented: S100, based on the target text, a plurality of first intermediate texts are obtained, wherein the first intermediate text is a text obtained by splitting the target text based on a preset category; S200, based on the plurality of first intermediate texts, a candidate text corresponding to a candidate text list is obtained, wherein the candidate text list comprises a plurality of candidate texts, the candidate text is a first intermediate text satisfying a preset condition obtained based on a plurality of second intermediate texts corresponding to the candidate text, the second intermediate text is a text obtained by splitting the candidate text based on a preset category; S300, a target application text list corresponding to the candidate text list is obtained to obtain a final application text, wherein the target application text list comprises a plurality of target application texts, each candidate text corresponds to a target application text, the final application text comprises all target application texts in the target application text list, wherein in S300, the target application text is obtained by the following steps: S1, based on the candidate text, obtaining a target title list set E={E 1 , E 2 , E 3}, E 2 ={E 2 1, …, E 2 i , …, E 2 n}, E 3 ={E 3 1, …, E 3 i , …, E 3 n}, E 3 i ={E 3 i1 , …, E 3 ij , …, E 3 im(i)}, E 1 is the first target title, E 2 i is the i-th second target title in the second target title list E 1 corresponding to E 2 , E 3 ij is the j-th third target title in the third target title list corresponding to E 2 i , the value range of i is 1 to n, n is the number of second target titles, the value range of j is 1 to m(i), m(i) is the number of third target titles in the third target title list corresponding to E 2 i . S2, obtaining a target priority list set F={F1,..., F i} n} from the second intermediate text corresponding to E and the candidate text to be selected i F={F i1 ,..., F ij ,..., F im(i)} ij , where F 3 ij The target priority corresponding to the candidate path corresponding to E S3, according to the candidate text, obtaining E 3 a corresponding specified condition text list set Q 3 ={Q 3 1, …, Q 3 i , …, Q 3 n}, Q 3 i ={Q 3 i1 , …, Q 3 ij , …, Q 3 im(i)}, Q 3 ij for E 3 ij a corresponding specified condition text list, wherein the specified condition text list comprises a plurality of specified condition texts, and the specified condition texts are condition sentences with judgment properties obtained based on texts between two adjacent third target titles in the candidate text; S4, according to the specific target demand and Q 3 , obtain the target application text corresponding to the selected text, wherein the second intermediate text corresponding to the selected text is sequentially compared with the specified condition text list in Q 3 , and the intermediate score is sequentially obtained until the sum of the obtained intermediate scores is not less than the target score to determine the target application text, wherein the target order is the order of the target priority in F sorted from large to small, and the intermediate score is the second intermediate text corresponding to the selected text and Q 3 ij The total score corresponding to the third target title obtained by matching the specified condition text included in Q 2. The data processing system for determining a final application text according to claim 1, characterized in that, The target text is a criterion text comprising a plurality of rules formulated to evaluate whether a target object meets a certain condition.
3. The data processing system for determining a final application text of claim 1, wherein, The candidate text is a text corresponding to a series of achievements possessed by a candidate object in a historical time period, wherein the candidate object is an object to be matched with the target text to confirm whether it meets the requirements corresponding to the target text.
4. The data processing system for determining a final application text according to claim 3, characterized in that, The time span corresponding to the historical time period is consistent with the time span required in the target text.
5. The data processing system for determining a final application text of claim 1, wherein, In S2 F is obtained by the following steps ij : S21, a first key sentence list corresponding to the second intermediate text of the candidate text is obtained, wherein the first key sentence list comprises a plurality of first key sentences, the first key sentence is a key sentence obtained from the second intermediate text corresponding to the candidate text; S22, obtaining E 3 ij a corresponding second key sentence list, wherein the second key sentence list includes a plurality of second key sentences, the second key sentences being obtained from E 3 ij a key sentence obtained from a corresponding key text, the E 3 ij a corresponding key text being a candidate text, the E 3 ij a corresponding third target title and E 3 i(j+1) text between corresponding third target titles, the E 3 i(j+1) being E 2 i a (j+1)th third target title in a corresponding third target title list; S23, obtaining E according to the first key sentence list and the second key sentence list 3 ij a corresponding first candidate score η 1 ij wherein η 1 ij satisfies the following condition: η 1 ij =∑ s(ij) r=1 λ r ij ,λ r ij for E 3 ij the number of first key sentences in the first key sentence list to which the corresponding rth second key sentence matches, r = 1 … s(ij), s(ij) is the number of second key sentences in the corresponding second key sentence list 3 ij the number of second key sentences in the corresponding second key sentence list; S24, obtaining E 2 i a corresponding third key sentence list, wherein the second key sentence list comprises a plurality of second key sentences, the third key sentence is any key sentence in the key text obtained from E 2 i corresponding to any key sentence in the key text obtained in addition to the second key sentence, the E 2 i corresponding to the key text is the candidate text E 2 i corresponding to the second target title E 2 i+1 corresponding to the text between the second target titles E 2 i+1 is E 1 corresponding to the second target title list E 2 the i+1th second target title in S25, according to the first key sentence list and the third key sentence list, obtaining E 2 i a corresponding second candidate score η 2 i wherein η 2 i satisfies the following conditions: η 2 i =∑ f(i) e=1 β e i ,β e i for E 2 i the number of first key sentences in the first key sentence list to which the corresponding e-th third key sentence matches, e = 1... f(i), f(i) is E 2 i the number of third key sentences in the corresponding third key sentence list; S26, according to η 1 ij and η 2 i , obtaining F ij where F ij complies with the following conditions: F ij =η 1 ij +η 2 i .
6. The data processing system for determining a final application text according to claim 5, characterized in that, The method for matching the first key sentence with the second key sentence is to judge by obtaining the similarity between the first key sentence and the second key sentence.
7. The data processing system for determining a final application text of claim 1, wherein, The target application text is a text meeting the requirements of the determined third target title text selected from the second intermediate text corresponding to the candidate text.
Citation Information
Patent Citations
Text importance calculation method and device and equipment and storage medium
CN109670183A
Policy matching system and method, computer equipment and storage medium
CN119357377A