A project material text analysis method, device and system

By segmenting and mining textual knowledge of project materials, combining it with the method of adjusting eccentric variables and optimizing natural language processing technology, we solved the problems of manual time consumption and inconsistent results in the review of project application materials, and achieved more efficient and accurate text review.

CN120429435BActive Publication Date: 2025-10-10BEIJING HUIQI YIDIANTONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410093443.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-10-10
Estimated Expiration
2044-01-23

AI Technical Summary

Technical Problem

The existing method for reviewing the text of project application materials relies on manual reading, which is time-consuming and labor-intensive and results in inconsistent review results. Natural language processing technology lacks accuracy in reviewing complex texts.

Method used

Segment the project material text, conduct text knowledge mining and eccentric variable adjustment, make project audit estimates through knowledge representation of segmented text, and optimize text audit algorithms to improve accuracy.

Benefits of technology

It improves the efficiency and accuracy of project material text review and ensures the reliability and consistency of the review results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429435B_ABST
    Figure CN120429435B_ABST
Patent Text Reader

Abstract

The application provides a project material text analysis method, device and system. When a target project material text is subjected to project review estimation, the text is segmented to obtain segmented texts, each segmented text is subjected to text knowledge mining to obtain a segmented text knowledge representation, and first project review estimation is performed to obtain a review estimation result of each segmented text. A bias variable of each segmented text is determined, the bias variable is used to adjust each segmented text knowledge representation, a segmented text bias knowledge representation is obtained, and second project review estimation is performed on the target project material text based on the segmented text bias knowledge representation to obtain a text review estimation result. The bias variable used to adjust each segmented text knowledge representation represents the influence of the segmented text on the text review estimation result of the target project material text, the obtained segmented text bias knowledge representation can better represent the target project material text, and the evaluation of the target project material text is more reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of natural language processing technology, and in particular to a method, device, and system for analyzing project material text. Background Art

[0002] During the various project application, approval, and evaluation processes, detailed review and analysis of project application materials are required. These application materials typically include a large amount of text. Due to the length of project application materials and the numerous key points that require review, reviewers often need to spend a considerable amount of time and effort to read and understand them. Traditional methods for reviewing project application materials rely primarily on manual reading and analysis, which has several limitations. First, reviewers need to spend a considerable amount of time reading and understanding the text, which can easily lead to omissions and errors. Second, due to the different professional backgrounds and experience of different reviewers, their understanding and analysis of the text may also vary, resulting in inconsistent review results.

[0003] To address these issues, natural language processing technology is widely used in the analysis of project application materials. Natural language processing can automatically analyze and process text, extracting key information and helping reviewers quickly understand the content. However, existing natural language processing technology still lacks accuracy in project application materials due to the high complexity of the text and the large number of review categories. Summary of the Invention

[0004] In view of this, the embodiments of the present disclosure at least provide a method, device and system for analyzing project material text.

[0005] The technical solution of the embodiment of the present disclosure is implemented as follows:

[0006] In a first aspect, an embodiment of the present disclosure provides a project material text analysis method, which is applied to a project material text review system. The method includes:

[0007] Segmenting the target project material text to obtain a plurality of segmented texts included in the target project material text;

[0008] Performing text knowledge mining on each of the segmented texts to obtain a segmented text knowledge representation of each of the segmented texts;

[0009] Performing a first item review and estimation on each of the segmented texts using the segmented text knowledge representation of each segmented text to obtain a segmented text review and estimation result for each segmented text;

[0010] For each of the segmented texts, determining an eccentricity variable of the segmented text based on the segmented text review estimation result of the segmented text, and determining a segmented text eccentricity knowledge representation of the segmented text based on the segmented text knowledge representation of the segmented text and the eccentricity variable, wherein the eccentricity variable represents the influence of the segmented text on the text review estimation result of the target project material text;

[0011] The second project review estimation is performed on the target project material text through the segmented text eccentricity knowledge representation of the multiple segmented texts to obtain a text review estimation result of the target project material text.

[0012] In some embodiments, performing text knowledge mining on each of the segmented texts to obtain a segmented text knowledge representation of each of the segmented texts includes:

[0013] Performing text knowledge mining in the first feature space on each of the segmented texts to obtain a transitional segmented text knowledge representation in the first feature space of each of the segmented texts;

[0014] splicing the transition segment text knowledge representations of the first feature spaces of the plurality of segment texts to obtain a transition segment text knowledge matrix of the target project material text;

[0015] Performing text knowledge mining in a second feature space on the transition segmented text knowledge matrix to obtain a segmented text knowledge matrix of the target project material text, wherein the segmented text knowledge matrix includes the segmented text knowledge representation of the second feature space of each segmented text, wherein the dimension of the second feature space is lower than the dimension of the first feature space;

[0016] The determining of the eccentricity variable of the segmented text by using the segmented text review estimation result of the segmented text includes:

[0017] Obtaining the number of project audit categories estimated in the first project audit, and determining an average distribution variable of the number of project audit categories;

[0018] Determining a deviation value between a segmented text review estimate result of the segmented text and the average distribution variable;

[0019] Using the deviation value as the eccentricity variable of the segmented text;

[0020] The determining of the segmented text eccentricity knowledge representation of the segmented text by using the segmented text knowledge representation of the segmented text and the eccentricity variable includes:

[0021] Multiplying the segmented text knowledge representation of the segmented text by the eccentricity variable of the segmented text to obtain the segmented text eccentricity knowledge representation of the segmented text;

[0022] The step of performing a second project review and estimation on the target project material text by using the segmented text eccentricity knowledge representation of the plurality of segmented texts to obtain a text review and estimation result of the target project material text includes:

[0023] Performing knowledge integration processing on the segmented text eccentricity knowledge representations of the plurality of segmented texts to obtain an integrated segmented text knowledge representation;

[0024] Through the integrated segmented text knowledge representation, a second project review estimate is performed on the target project material text to obtain a text review estimate result of the target project material text.

[0025] In some embodiments, the first project review estimate is completed by a segment dimension classification operator of a text review algorithm, and the second project review estimate is completed by a text dimension classification operator of the text review algorithm;

[0026] The method further comprises:

[0027] Obtaining a text learning sample set for optimizing the text audit algorithm, the text learning sample set comprising a plurality of text learning samples, the text learning samples matching text supervision labels, the text learning samples comprising a plurality of segmented text learning samples, the segmented text learning samples matching segmented text supervision labels;

[0028] By using the segmented dimension classification operator, a first item review and estimation is performed on each segmented text learning example in each of the text learning examples to obtain a segmented review and estimation result of each segmented text learning example;

[0029] For each of the text learning examples, complete the following processing:

[0030] Based on the segmented text supervision labels of each segmented text learning example in the text learning example, performing a second project review estimation on the text learning example through the text dimension classification operator to obtain a sample review estimation result of the text learning example;

[0031] Determine the algorithm cost of the text audit algorithm by using the deviation value between the segmented audit prediction result and the segmented text supervision mark of each segmented text learning example, and the deviation value between the sample audit prediction result and the text supervision mark of the text learning example;

[0032] The algorithm parameters of the text audit algorithm are modified according to the algorithm cost value to optimize the text audit algorithm.

[0033] In some embodiments, determining the algorithm cost of the text audit algorithm by using the deviation value between the segmented audit prediction result and the segmented text supervision mark of each segmented text learning example, and the deviation value between the sample audit prediction result and the text supervision mark of the text learning example, includes:

[0034] For each of the segmented text learning examples, determining the transition segmentation dimension cost of the text review algorithm through the segmented review estimation result of the segmented text learning example and the deviation value of the segmented text supervision label;

[0035] Accumulating the transition segment dimension costs of the text audit algorithm to obtain the segment dimension cost of the text audit algorithm;

[0036] Determining the text dimension cost of the text audit algorithm by using the sample audit estimation result of the text learning sample and the deviation value of the text supervision mark;

[0037] By using the first cost eccentricity variable of the segment dimension cost and the second cost eccentricity variable of the text dimension cost, the segment dimension cost and the text dimension cost are subjected to eccentric addition processing to obtain the algorithm cost value of the text audit algorithm;

[0038] The first item review and estimation is performed on each of the segmented text learning examples in each of the text learning examples by using the segmented dimension classification operator to obtain a segmented review and estimation result of each of the segmented text learning examples, including:

[0039] Acquiring segmented text example knowledge of each segmented text learning example in each of the text learning examples;

[0040] By using the segmented dimension classification operator and the segmented text example knowledge of each segmented text learning example, a first project review estimation is performed on each segmented text learning example to obtain a segmented review estimation result of each segmented text learning example;

[0041] The segmented text supervision labeling of each segmented text learning example in the text learning example, performing a second project review and estimation on the text learning example through the text dimension classification operator, and obtaining a sample review and estimation result of the text learning example, includes:

[0042] Determining the segmented text eccentricity variable of each segmented text learning example through the segmented text supervision label of each segmented text learning example in the text learning example, and determining the segmented text sample eccentricity knowledge of each segmented text learning example through the segmented text sample knowledge and the segmented text eccentricity variable of each segmented text learning example;

[0043] By using the segmented text sample eccentricity knowledge of multiple segmented text learning samples in the text learning sample, the second project review estimation is performed on the text learning sample through the text dimension classification operator to obtain the sample review estimation result of the text learning sample.

[0044] In some embodiments, determining the segmented text eccentricity variable of each segmented text learning example by using the segmented text supervision label of each segmented text learning example in the text learning example includes:

[0045] For each of the segmented text learning examples in the text learning examples, the following processing is completed:

[0046] Obtaining the number of project audit categories estimated in the first project audit, and determining an average distribution variable of the number of project audit categories;

[0047] Determining a deviation value between the segmented text supervision label of the segmented text learning example and the average distribution variable;

[0048] The deviation value is used as the segmented text eccentricity variable of the segmented text learning example.

[0049] In some embodiments, before completing the following processing for each of the text learning examples, the method further includes:

[0050] Determine the core segmented texts of each project review category obtained by the first project review estimate from the multiple segmented text learning examples in the text learning example set based on the segmented review estimate results of each segmented text learning example in the text learning example set;

[0051] For each of the project review categories, a basic knowledge representation and a corresponding basic supervision mark of the project review category are obtained, and based on the core segmented text of the project review category, the basic knowledge representation is momentum adjusted to obtain an adjusted basic knowledge representation of the project review category;

[0052] determining a similarity measurement result between the segmented text knowledge representation of the segmented text learning sample and the adjusted basic knowledge representation of each of the project review categories, and determining the basic supervision label corresponding to the target adjusted basic knowledge representation with the largest similarity measurement result as the segmented text basic supervision label of the segmented text learning sample;

[0053] momentum-adjusting the segmented text supervision label of each of the segmented text learning samples based on the segmented text basic supervision label of the segmented text learning sample, to obtain an adjusted segmented text supervision label of the segmented text learning sample;

[0054] performing second project review estimation on the text learning sample based on the segmented text supervision label of each of the segmented text learning samples in the text learning sample, by using the text dimension classification operator, to obtain a sample review estimation result of the text learning sample, including:

[0055] performing second project review estimation on the text learning sample based on the adjusted segmented text supervision label of each of the segmented text learning samples in the text learning sample, by using the text dimension classification operator, to obtain a sample review estimation result of the text learning sample;

[0056] determining the algorithm generation value of the text review algorithm based on the deviation value of the segmented review estimation result and the segmented text supervision label of each of the segmented text learning samples, and the deviation value of the sample review estimation result and the text supervision label of the text learning sample, including:

[0057] determining the algorithm generation value of the text review algorithm based on the deviation value of the segmented review estimation result and the adjusted segmented text supervision label of each of the segmented text learning samples, and the deviation value of the sample review estimation result and the text supervision label of the text learning sample.

[0058] In some embodiments, the segmented review estimation result includes a segmented estimation support coefficient of the segmented text learning sample belonging to each of the project review categories;

[0059] determining the core segmented text of each of the project review categories obtained by the first project review estimation from the plurality of segmented text learning samples in the text learning sample set based on the segmented review estimation result of each of the segmented text learning samples in the text learning sample set, including:

[0060] for each of the project review categories, the following processing process is completed:

[0061] Determining, from the plurality of segmented text learning examples in the text learning example set, a preset number of target segmented text learning examples that belong to the project review category and have the largest segmented estimated support coefficients;

[0062] The preset number of target segmented text learning examples are determined as the preset number of core segmented texts of the project review category.

[0063] In some embodiments, the number of the core segmented texts is S, where S is greater than 1. The step of momentum-adjusting the basic knowledge representation based on the core segmented texts of the project review category to obtain the adjusted basic knowledge representation of the project review category includes:

[0064] Momentum-adjust the basic knowledge representation through the first core segmented text of the project review category to obtain a first transition basic knowledge representation of the project review category;

[0065] Momentum-adjust the jth transition basic knowledge representation of the project review category through the kth core segmented text of the project review category to obtain the kth transition basic knowledge representation of the project review category, where j=k-1, and j is a positive integer less than or equal to S;

[0066] Wandering through the core segmented text, obtaining the Sth transition basic knowledge representation of the project review category, and using the Sth transition basic knowledge representation of the project review category as the adjusted basic knowledge representation of the project review category;

[0067] Alternatively, there are multiple core segmented texts, and the basic knowledge representation is adjusted by momentum based on the core segmented texts of the project review category to obtain the adjusted basic knowledge representation of the project review category, including:

[0068] Obtaining a core segmented text knowledge representation of each core segmented text of the project review category, and determining an average segmented text knowledge representation of a plurality of the core segmented text knowledge representations;

[0069] Obtaining a first knowledge representation eccentricity variable of the average segmented text knowledge representation and a second knowledge representation eccentricity variable of the basic knowledge representation;

[0070] Performing eccentric addition processing on the average segmented text knowledge representation and the basic knowledge representation through the first knowledge representation eccentricity variable and the second knowledge representation eccentricity variable to obtain an adjusted basic knowledge representation of the project review category;

[0071] The step of adjusting the segmented text supervision label of the segmented text learning example by momentum based on the segmented text basic supervision label of the segmented text learning example to obtain the adjusted segmented text supervision label of the segmented text learning example includes:

[0072] Obtaining a first label eccentricity variable of the segmented text basic supervision label and a second label eccentricity variable of the segmented text supervision label;

[0073] The segmented text basic supervision label and the segmented text supervision label are subjected to eccentric addition processing by using the first label eccentricity variable and the second label eccentricity variable to obtain the adjusted segmented text supervision label of the segmented text learning example.

[0074] In a second aspect, the present disclosure provides a project material text analysis device, comprising:

[0075] A text segmentation module, configured to segment the target project material text to obtain a plurality of segmented texts included in the target project material text;

[0076] A knowledge mining module is used to perform text knowledge mining on each of the segmented texts to obtain a segmented text knowledge representation of each of the segmented texts;

[0077] A first estimation module is configured to perform a first project review estimation on each of the segmented texts based on the segmented text knowledge representation of each segmented text, and obtain a segmented text review estimation result for each segmented text;

[0078] an eccentricity calculation module for determining, for each segmented text, an eccentricity variable of the segmented text based on a segmented text review estimation result of the segmented text, and determining a segmented text eccentricity knowledge representation of the segmented text based on the segmented text knowledge representation of the segmented text and the eccentricity variable, wherein the eccentricity variable represents an influence of the segmented text on the text review estimation result of the target project material text;

[0079] The second estimation module is used to perform a second project review estimation on the target project material text through the segmented text eccentricity knowledge representation of the multiple segmented texts, and obtain a text review estimation result of the target project material text.

[0080] In a third aspect, the present disclosure provides a project material text review system, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above-mentioned method when executing the program.

[0081] The present disclosure at least includes the following beneficial effects:

[0082] In the project material text analysis method, device and system provided by the embodiments of the present disclosure, when performing project review and estimation on the target project material text, the target project material text is first segmented to obtain a plurality of segmented texts included in the target project material text, and then text knowledge mining is performed on each segmented text to obtain a segmented text knowledge representation of each segmented text, and a first project review and estimation is performed on each segmented text through the segmented text knowledge representation of each segmented text to obtain a segmented text review and estimation result of each segmented text; thereby, the eccentricity variable of each segmented text is determined through the segmented text review and estimation result of each segmented text, and the eccentricity adjustment is performed on each segmented text knowledge representation through the segmented text knowledge representation and eccentricity variable of each segmented text to obtain a segmented text eccentricity knowledge representation of each segmented text, and finally, the second project review and estimation is performed on the target project material text through the segmented text eccentricity knowledge representation of each segmented text to obtain a text review and estimation result of the target project material text. Since the eccentricity variable of the eccentricity adjustment of the knowledge representation of each segmented text represents the influence of the segmented text on the text review estimation result of the target project material text, the eccentric knowledge representation of each segmented text obtained by eccentric adjustment can better represent the target project material text, making the text review estimation result obtained through the project review estimation of the eccentric knowledge representation of each segmented text more accurate and the evaluation of the target project material text more reliable.

[0083] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0085] Figure 1 A schematic diagram of the implementation flow of a project material text analysis method provided in an embodiment of the present disclosure.

[0086] Figure 2 A schematic diagram of the composition structure of a project material text analysis device provided in an embodiment of the present disclosure.

[0087] Figure 3 A schematic diagram of the hardware entity of a project material text review system provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0088] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the technical solutions of the present disclosure are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting the present disclosure. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0089] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or sequence of "first / second / third" may be interchanged where permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0090] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are for the purpose of describing the present disclosure only and are not intended to limit the present disclosure.

[0091] The present disclosure provides a method for analyzing project material text, which can be executed by a processor of a project material text review system. The project material text review system can be a device with data processing capabilities, such as a server, a laptop, a tablet computer, or a desktop computer.

[0092] Figure 1 A schematic diagram of the implementation process of a project material text analysis method provided in an embodiment of the present disclosure is shown as follows: Figure 1 As shown, the method includes the following operations:

[0093] Operation S100: segmenting the target project material text to obtain a plurality of segmented texts included in the target project material text.

[0094] The target project material text is the text material uploaded by the project applicant in the application system. The length of the project material text is often large, which puts a lot of pressure on the reviewer. After obtaining the target project material text, in the embodiment of the present disclosure, the target project material text is segmented to obtain multiple segmented texts included in the target project material text. The target project material text can be regarded as a package in multi-instance learning, and each segment contained is an instance. The target project material text is segmented, and the multiple segmented texts obtained are instances. There are many ways to segment the target project material text. As an implementation method, a target project material text contains multiple titles, and the text under a title is a segmented text. The level granularity of the title is not limited, such as a second-level title or a third-level title. Under each title, there is a segmented text that needs to be reviewed and classified. The review classification is, for example, whether the content in the project application material text complies with the relevant regulations of the corresponding project, such as whether the language expression is compliant, whether the project indicators are met, whether the content is complete, etc.

[0095] Alternatively, in another embodiment, the target project material text is segmented based on the following operations, for example, to obtain multiple segmented texts included in the target project material text: determining a ruler text box including a target frame length and a ruler amplitude of the ruler text box corresponding to the target project material text; segmenting the target project material text by moving the ruler text box according to the ruler amplitude to obtain multiple segmented texts included in the target project material text. The target project material text is segmented using the ruler text box. Specifically, the ruler text box including a target frame length corresponding to the target project material text is first determined. The target frame length can be unrestricted, for example, directly set to 500 characters in advance, or dynamically adjusted according to the number of characters in the target project material text, such as setting a basic number of copies and dividing the number of characters in the target project material text by the basic number of copies to obtain the target frame length; determining the ruler amplitude of the ruler text box, that is, the moving stride. The target project material text is segmented by moving the ruler text box on the target project material text to obtain multiple segmented texts included in the target project material text.

[0096] Operation S200: performing text knowledge mining on each segmented text to obtain a segmented text knowledge representation of each segmented text.

[0097] For multiple segmented texts in the target project material text, text knowledge mining is performed on each segmented text to obtain a segmented text knowledge representation for each segmented text. In a feasible design, for example, text knowledge mining is performed on each segmented text based on the following operations to obtain a segmented text knowledge representation for each segmented text: text knowledge mining is performed on each segmented text in a first feature space to obtain a transitional segmented text knowledge representation in the first feature space of each segmented text; the transitional segmented text knowledge representations in the first feature space of each of the multiple segmented texts are spliced ​​to obtain a transitional segmented text knowledge matrix of the target project material text; text knowledge mining is performed on the transitional segmented text knowledge matrix in a second feature space to obtain a segmented text knowledge matrix of the target project material text, the segmented text knowledge matrix includes the segmented text knowledge representation in the second feature space of each segmented text, and the dimension of the second feature space is lower than that of the first feature space.

[0098] In the above content, we can first perform text knowledge mining in the first feature space for each segmented text to obtain the transitional segmented text knowledge representation of the first feature space of each segmented text. The first feature space represents a feature dimension, and text knowledge mining in the first feature space can be achieved through a pre-trained BERT network to help improve the accuracy of text knowledge mining. Text knowledge representation is the feature information obtained after feature mining of the text, which is usually a feature vector, such as a semantic feature vector such as a word vector, context vector, or topic feature vector. The segmented text knowledge representation is the text knowledge representation corresponding to a segmented text, and the transitional segmented text knowledge representation is the segmented text knowledge representation during the processing. Subsequent similar prefix descriptions are understood in the same way.

[0099] Next, the transition segmented text knowledge representations of the first feature space of each of the multiple segmented texts are spliced ​​together to obtain the transition segmented text knowledge matrix of the target project material text. Then, the transition segmented text knowledge matrix is ​​subjected to text knowledge mining in the second feature space to obtain the segmented text knowledge matrix of the target project material text. This segmented text knowledge matrix includes the segmented text knowledge representations of the second feature space of each segmented text, wherein the dimension of the second feature space is lower than that of the first feature space. In specific implementation, text knowledge mining in the second feature space can be achieved through pre-trained basic network operators. The basic network operator is a backbone network operator that constitutes the basic structure of the entire algorithm. Other more complex neural networks can add more layers and modules on this basis to achieve more advanced feature extraction and classification tasks. For the basic network operator, it can be constructed based on a Transformer containing a residual layer to complete knowledge compression and dimensionality reduction. At the same time, the knowledge of each segmented text in the target project material text is integrated to improve the accuracy of text knowledge mining, so that the acquired segmented text knowledge representation can better represent the corresponding segmented text.

[0100] As mentioned above, the extracted transition segmented text knowledge representations are first spliced ​​to obtain the transition segmented text knowledge matrix of the target project material text, and then text knowledge mining is performed on the transition segmented text knowledge matrix. This can complete the knowledge fusion between the segmented texts in the target project material text to improve the accuracy of text knowledge mining, and at the same time complete knowledge compression to improve processing efficiency.

[0101] Operation S300: performing a first item review and estimation on each segmented text based on the segmented text knowledge representation of each segmented text, and obtaining a segmented text review and estimation result for each segmented text.

[0102] After obtaining the segmented text knowledge representation of each segmented text, the first project review estimate is performed on each segmented text respectively through the segmented text knowledge representation of each segmented text to obtain the segmented text review estimate result of each segmented text. The segmented text review estimate result represents the review category corresponding to the corresponding segmented text, and the segmented text review estimate result can also include the segmented estimate support coefficient of the segmented text belonging to each project review category (the support coefficient can be expressed by probability or confidence). In specific implementation, the first project review estimate can be completed based on the segmented dimension classification operator of the text review algorithm. The text review algorithm can be obtained by pre-optimization. The segmented dimension classification operator is a neural network operator that performs classification at the level of text segmentation. The text review algorithm can be based on the BERT architecture.

[0103] Operation S400: for each segmented text, determine the segmented text's eccentricity variable through the segmented text review estimation result of the segmented text, and determine the segmented text eccentricity knowledge representation of the segmented text through the segmented text knowledge representation and the eccentricity variable of the segmented text.

[0104] In operation S400, after obtaining the segmented text review estimation result of each segmented text, the following processing is completed for each segmented text: the eccentricity variable of the segmented text is determined through the segmented text review estimation result of the segmented text, and the segmented text eccentricity knowledge representation of the segmented text is determined through the segmented text knowledge representation of the segmented text and the eccentricity variable. In the embodiment of the present disclosure, the eccentricity variable represents the influence of the segmented text (or the segmented text knowledge representation of the segmented text) on the text review estimation result of the target project material text, and the eccentricity variable can be understood as a weighting coefficient in essence. As described above, the segmented text knowledge representation of the segmented text can be eccentrically adjusted based on the eccentricity variable, that is, weighted calculation is performed to improve the accuracy of the project review estimation of the target project material text through the segmented text knowledge representation of the segmented text.

[0105] In a feasible design, the eccentric variable of the segmented text is determined by the segmented text review estimation result of the segmented text based on the following operations: obtaining the number of project review categories of the first project review estimation, and determining the average distribution variable of the number of project review categories; determining the deviation value between the segmented text review estimation result of the segmented text and the average distribution variable; and taking the deviation value as the eccentric variable of the segmented text.

[0106] In the above, the number of project review categories of the first project review estimation can be obtained first. The project review categories may, for example, be categories of various classification situations from the perspectives of text grammatical errors, index compliance, content completeness, and the like. For example, for text grammatical errors, a binary classification task can be set, such as having grammatical errors or not having grammatical errors, or a ternary classification task can be set, such as not having grammatical errors, having more grammatical errors, and having fewer grammatical errors. For index compliance, a binary classification task can be set, such as index compliance and index non-compliance, or a ternary classification task can be set, such as index non-compliance, index compliance more, and index compliance less. Among them, the project review categories can include one of the above perspectives of text grammatical errors, index compliance, content completeness, and the like, and then multi-classification is performed with respect to this one perspective. At this time, the number of project review categories is the number of classifications corresponding to this perspective. For example, for text grammatical errors, the number of categories is 2 for binary classification. In another feasible embodiment, the project review categories can include multiple of the above perspectives of text grammatical errors, index compliance, content completeness, and the like, and then multi-classification is performed with respect to the multiple perspectives. At this time, the number of project review categories is the combination of the number of classifications corresponding to the multiple perspectives. For example, for text grammatical errors, the number of categories is 4 for binary classification, and for index compliance, the number of categories is 2 for binary classification. The selection of specific review category perspectives and the number of categories corresponding to the perspectives is not limited.

[0107] After determining the project review categories, the average distribution variable corresponding to the number of project review categories is determined. The average distribution variable is the average value of the probability distribution. For ease of understanding, the following is an example. Assuming that the number of project review categories is 2, the average distribution variable is (0.5, 0.5). If the number of project review categories is 3, the average distribution variable is (0.33, 0.33, 0.33). Then, the deviation value between the segmented text review estimation result of the segmented text and the average distribution variable is determined, and the deviation value is taken as the eccentric variable of the segmented text. Alternatively, the K-L distance (K-L Distance) or information entropy between the segmented text review estimation result and the average distribution variable can be determined, and taken as the eccentric variable of the segmented text.

[0108] In a feasible design, for example, the segmented text eccentricity knowledge representation of the segmented text is determined through the segmented text knowledge representation and eccentricity variable of the segmented text based on the following operation: multiply the segmented text knowledge representation of the segmented text by the eccentricity variable of the segmented text to obtain the segmented text eccentricity knowledge representation of the segmented text.

[0109] Operation S500: performing a second project review estimation on the target project material text by using the segmented text eccentricity knowledge representation of the plurality of segmented texts to obtain a text review estimation result of the target project material text.

[0110] In operation S500, after obtaining the segmented text eccentricity knowledge representation of multiple segmented texts in the target project material text, the target project material text is subjected to a second project audit estimation through the segmented text eccentricity knowledge representation of each of the multiple segmented texts in the target project material text, and a text audit estimation result of the target project material text is obtained. The text audit estimation result characterizes the audit category corresponding to the target project material text, and the text audit estimation result may also include the text dimension estimation support coefficient (same as above, which may be expressed as probability or confidence) of the target project material text belonging to each project audit category. In specific implementation, the number of project audit categories estimated in the second project audit may be consistent with the number of project audit categories estimated in the first project audit, for example, both are binary classifications of audit text grammatical errors, and the number of project audit categories is 2. The second project audit estimation can be completed by the text dimension classification operator of the text audit algorithm, and the text audit algorithm can be obtained by pre-optimization. The text dimension classification operator is a neural network operator that performs classification at the level of the entire text (that is, the package in multi-instance classification), and the text audit algorithm can be implemented by BERT. The second project review estimate of the target project material text is performed through the segmented text eccentricity knowledge representation of multiple segmented texts. When the text review estimate result of the target project material text is obtained, the segmented text eccentricity knowledge representation of multiple segmented texts can be specifically subjected to knowledge integration processing. The knowledge integration processing here can be an aggregation processing process at the feature level to obtain an integrated segmented text knowledge representation, and then the second project review estimate of the target project material text is performed through the integrated segmented text knowledge representation to obtain the text review estimate result of the target project material text.

[0111] The following further describes the optimization method of the text audit algorithm. In a feasible design, the optimization method of the text audit algorithm provided by the embodiment of the present disclosure may include the following operations:

[0112] Operation S10 is to obtain a text learning example set for optimizing the text review algorithm. The text learning example set includes a plurality of text learning examples. The text learning examples are training data for optimizing the text review algorithm. The text learning samples are matched with text supervision labels, and the text learning samples include multiple segmented text learning samples, and the segmented text learning samples are matched with segmented text supervision labels; operation S20, through the segmented dimension classification operator, each segmented text learning sample in each text learning sample is respectively subjected to a first project review and estimation, and a segmented review and estimation result of each segmented text learning sample is obtained; operation S30, for each text learning sample, the following processing process is completed: operation S31, based on the segmented text supervision labels of each segmented text learning sample in the text learning sample, a second project review and estimation is performed on the text learning sample through the text dimension classification operator, and a sample review and estimation result of the text learning sample is obtained; operation S32, the algorithm cost value of the text review algorithm is determined through the deviation value of the segmented review and estimation result of each segmented text learning sample and the segmented text supervision label, as well as the deviation value of the sample review and estimation result of the text learning sample and the text supervision label; operation S33, the algorithm parameters of the text review algorithm are corrected through the algorithm cost value to optimize the text review algorithm.

[0113] In the above content, in operation S10, a text learning sample set for optimizing the text audit algorithm is first obtained. The text learning sample set includes multiple text learning samples, and each text learning sample includes multiple segmented text learning samples. The segmented text learning sample is obtained by segmenting the text learning sample. The segmentation method can refer to the aforementioned operation S100 and will not be described here. Among them, each text learning sample matches a text supervision label, and each segmented text learning sample matches a segmented text supervision label. In specific implementation, the segmented text supervision label of the segmented text learning sample is determined as the segmented text supervision label of the segmented text learning sample in the initialization stage. Since there is a possibility that there are segmented text learning samples that are irrelevant to the text supervision label in the text learning sample, the segmented text supervision label of the segmented text learning sample is an uncertain label, that is, a soft label. In operation S20, based on the segmented dimension classification operator of the text audit algorithm, a first project audit estimate is performed on each segmented text learning example in each text learning example, and a segmented audit estimate result of each segmented text learning example is obtained. In a feasible design, for example, the segmented audit estimate result of each segmented text learning example is obtained based on the following operations: the segmented text sample knowledge of each segmented text learning example in each text learning example is obtained; through the segmented dimension classification operator, the segmented text sample knowledge of each segmented text learning example is used to perform a first project audit estimate on each segmented text learning example, and a segmented audit estimate result of each segmented text learning example is obtained.

[0114] In the above content, the following processing process can be completed for each text learning sample: obtain the segmented text sample knowledge of each segmented text learning sample in the text learning sample, such as performing text knowledge mining on each segmented text learning sample to obtain the segmented text sample knowledge of each segmented text learning sample. In the above content, the text knowledge mining can refer to the process of text knowledge mining represented by segmented text knowledge. In this way, through the segmented text sample knowledge of each segmented text learning sample in the text learning sample, the first project review and estimation is performed on each segmented text learning sample in each text learning sample, and the segmented review and estimation result of each segmented text learning sample is obtained. The segmented review and estimation result represents the review category corresponding to the corresponding segmented text learning sample, and the segmented review and estimation result can also include the segmented estimation support coefficient of the corresponding segmented text learning sample corresponding to each project review category.

[0115] In operation S30, for each text learning sample, the following processing is completed: based on the segmented text supervision mark of each segmented text learning sample in the text learning sample, the text learning sample is subjected to a second project review and estimation through a text dimension classification operator to obtain a sample review and estimation result of the text learning sample. In a feasible design, for example, the sample review and estimation result of the text learning sample is obtained based on the following operations: through the segmented text supervision mark of each segmented text learning sample in the text learning sample, the segmented text eccentricity variable of each segmented text learning sample is determined, and through the segmented text sample knowledge and the segmented text eccentricity variable of each segmented text learning sample, the segmented text sample eccentricity knowledge of each segmented text learning sample is determined; through the segmented text sample eccentricity knowledge of multiple segmented text learning samples in the text learning sample, the text learning sample is subjected to a second project review and estimation through a text dimension classification operator to obtain a sample review and estimation result of the text learning sample.

[0116] In the above content, for each segmented text learning sample, the segmented text eccentricity variable of the segmented text learning sample is determined through the segmented text supervision mark and the segmented review estimation result of the segmented text learning sample, and the segmented text eccentricity variable indicates the influence of the segmented text learning sample on the sample review estimation result of the text learning sample. In a feasible design, for example, based on the following operations, the segmented text eccentricity variable of each segmented text learning sample is determined through the segmented text supervision mark of each segmented text learning sample in the text learning sample: for each segmented text learning sample in the text learning sample, the following processing is completed: obtain the number of project review categories estimated by the first project review, and determine the average distribution variable of the number of project review categories; determine the deviation value between the segmented text supervision mark and the average distribution variable of the segmented text learning sample; and use the deviation value as the segmented text eccentricity variable of the segmented text learning sample.

[0117] In the above content, when calculating the segmented text eccentricity variable of the segmented text learning example, the number of project review categories estimated by the first project review can be obtained first, and then the average distribution variable corresponding to the number of project review categories can be determined. Then, the deviation value between the segmented review estimate result of the segmented text learning example and the average distribution variable is obtained, and the deviation value is used as the segmented text eccentricity variable of the segmented text learning example. For details, please refer to the content in operation S400, which will not be repeated here.

[0118] Then, the segmented text sample knowledge of the segmented text learning sample is biased by adjusting the segmented text bias variable. That is, the segmented text bias variable is multiplied by the segmented text sample knowledge to obtain the segmented text sample bias knowledge. Then, based on the segmented text sample bias knowledge of multiple segmented text learning samples in the text learning sample, a second project review estimation is performed on the text learning sample based on the text dimension classification operator to obtain a sample review estimation result of the text learning sample. The sample review estimation result represents the review category corresponding to the text learning sample, and the sample review estimation result may also include the text dimension estimation support coefficient of the text learning sample corresponding to each project review category.

[0119] Operation S32, by means of the deviation value of the segmented audit prediction result and the segmented text supervision mark of each segmented text learning example and the sample audit prediction result of the text learning example and the deviation value of the text supervision mark, determines the algorithm cost value of the text audit algorithm. In a feasible design, operation S32 may include operations 321 to 324: operation 321, for each segmented text learning example, by means of the segmented audit prediction result and the segmented text supervision mark of the segmented text learning example, determines the transition segmented dimension cost of the text audit algorithm; operation 322, accumulates the various transition segmented dimension costs of the text audit algorithm to obtain the segmented dimension cost of the text audit algorithm; operation 323, by means of the deviation value of the sample audit prediction result and the text supervision mark of the text learning example, determines the text dimension cost of the text audit algorithm; operation 324, by means of the first cost eccentricity variable of the segmented dimension cost and the second cost eccentricity variable of the text dimension cost, performs eccentric addition on the segmented dimension cost and the text dimension cost to obtain the algorithm cost value of the text audit algorithm.

[0120] In the above content, the cost function of the transition segment dimension cost is not limited, for example, it is a cross-entropy cost function, and the cost function of the text dimension cost can also be a cross-entropy cost function. The specific values ​​of the first cost eccentricity variable and the second cost eccentricity variable are pre-configured and are not specifically limited.

[0121] Operation S33 corrects the algorithm parameters of the text audit algorithm through the algorithm cost value to optimize the text audit algorithm. When there is a difference between the estimated value of the text audit algorithm and the true value, the cost function returns a non-zero value, indicating the estimated error of the text audit algorithm. In order to minimize the value of the cost function, an optimization algorithm is used to adjust the parameters of the algorithm so that it is as close to the true value as possible. Common optimization algorithms include stochastic gradient descent (SGD), Adagrad, Adadelta, RMSProp, etc. The core idea of ​​these optimization algorithms is to calculate the gradient of the cost function to the parameters of the text audit algorithm and adjust the value of the parameter according to the direction of the gradient to reduce the value of the cost function. Specifically, the optimization algorithm calculates the updated value of each parameter based on the gradient value of the cost function to each parameter, and adds it to the value of the current parameter to achieve parameter update. Each time the parameters are updated, the optimization algorithm controls the amplitude of the parameter update according to the learning rate to avoid overfitting or underfitting. Through the synergy of the cost function and the optimization algorithm, the parameters of the text audit algorithm are continuously adjusted to make them as close to the true value as possible, thereby improving the estimated accuracy of the text audit algorithm. When the text audit algorithm converges, the optimization is stopped and the target text audit algorithm is obtained.

[0122] In the embodiment of the present disclosure, since the segmented text supervision mark of the segmented text learning example is an uncertain mark, the segmented text supervision mark may contain disturbance information, which affects the optimization result of the algorithm. Therefore, in the embodiment of the present disclosure, the segmented text supervision mark of the segmented text learning example can also be updated in real time during optimization to improve the optimization result of the text review algorithm. After operation S20 and before operation S30, the optimization process further includes the following operations:

[0123] Operation A1, determine the core segmented text of each project audit category from the plurality of segmented text learning samples in the text learning sample set based on the segmented audit estimation result of each segmented text learning sample in the text learning sample set. Operation A2, for each project audit category, obtain the basic knowledge representation and the corresponding basic supervision label of the project audit category, and adjust the basic knowledge representation based on the core segmented text of the project audit category to obtain the adjusted basic knowledge representation of the project audit category. Operation A3, for each segmented text learning sample, determine the similarity measurement result between the segmented text knowledge representation of the segmented text learning sample and the adjusted basic knowledge representation of each project audit category, and determine the basic supervision label corresponding to the target adjusted basic knowledge representation with the largest similarity measurement result as the segmented text basic supervision label of the segmented text learning sample. Operation A4, for each segmented text learning sample, adjust the segmented text supervision label of the segmented text learning sample based on the segmented text basic supervision label of the segmented text learning sample to obtain the adjusted segmented text supervision label of the segmented text learning sample. Since each knowledge representation is a vector, the similarity measurement result can be calculated by calculating the feature distance (such as cosine distance) between the segmented text knowledge representation and the adjusted basic knowledge representation of each project audit category. For details, please refer to related technologies.

[0124] After obtaining the segmented audit estimation result of each segmented text learning sample in the text learning sample set in operation S20, in operation A1, the core segmented text of each project audit category is determined from the plurality of segmented text learning samples in the text learning sample set based on the segmented audit estimation result of each segmented text learning sample in the text learning sample set. In the above content, the segmented audit estimation result includes the segmented estimation support coefficient of the segmented text learning sample belonging to each project audit category. For example, based on the following operation, the core segmented text of each project audit category can be determined from the plurality of segmented text learning samples in the text learning sample set: for each project audit category, the following process is completed: determine the target segmented text learning sample belonging to the project audit category from the plurality of segmented text learning samples in the text learning sample set, and the segmented estimation support coefficient is the largest; the preset number of target segmented text learning samples is determined as the preset number of core segmented text of the project audit category. In the above content, the preset number is less than the number of segmented text learning samples.

[0125] In operation A2, for each project review category, the basic knowledge representation of the project review category can be momentum adjusted through the core segmented text of the project review category to obtain the adjusted basic knowledge representation of the project review category. In the above content, for each project review category, the basic knowledge representation and basic supervision mark of the project review category can be initialized in advance. The initialized basic knowledge representation is, for example, an all-zero array. The category of the basic supervision mark can correspond to the project review category estimated by the second project review of the text review algorithm. Momentum adjustment is implemented based on the MomentumUpdate optimization algorithm.

[0126] In a feasible design, when the momentum of the basic knowledge representation is adjusted, in operation A2, the following processing is completed for each project review category: operation A21, through the first core segmented text of the project review category, momentum adjusts the basic knowledge representation to obtain the first transition basic knowledge representation of the project review category; operation A22, through the kth core segmented text of the project review category, momentum adjusts the jth transition basic knowledge representation of the project review category to obtain the kth transition basic knowledge representation of the project review category, where j=k-1, and j is a positive integer less than or equal to S; operation A23, walks through all core segmented texts, that is, traverses all core segmented texts to obtain the Sth transition basic knowledge representation of the project review category, and uses the Sth transition basic knowledge representation of the project review category as the adjusted basic knowledge representation of the project review category.

[0127] Operations A21 to A23 are to adjust the basic knowledge representation once based on each core segmented text respectively to obtain the final adjusted basic knowledge representation. In the embodiment of the present disclosure, the first core segmented text of the project review category is used to momentum adjust the basic knowledge representation to obtain the first transitional basic knowledge representation of the project review category. For example, each momentum adjustment includes: obtaining the core segmented text knowledge representation of the first core segmented text; obtaining the third knowledge representation eccentricity variable of the core segmented text knowledge representation and the fourth knowledge representation eccentricity variable of the basic knowledge representation; through the third knowledge representation eccentricity variable and the fourth knowledge representation eccentricity variable, the core segmented text knowledge representation and the basic knowledge representation are eccentrically added (i.e., eccentric adjustment weighting is performed first, and then the weighted results are added) to obtain the first transitional basic knowledge representation of the project review category. In the above content, the numerical values ​​of the third knowledge representation eccentricity variable and the fourth knowledge representation eccentricity variable are pre-configured and are not specifically limited.

[0128] In another feasible design, when the momentum of the basic knowledge representation is adjusted, in operation A2, the following processing is completed for each project review category: operation A21', obtain the core segmented text knowledge representation of each core segmented text of the project review category, and determine the average segmented text knowledge representation of multiple core segmented text knowledge representations; operation A22', obtain the first knowledge representation eccentricity variable of the average segmented text knowledge representation, and the second knowledge representation eccentricity variable of the basic knowledge representation; operation A23', through the first knowledge representation eccentricity variable and the second knowledge representation eccentricity variable, perform eccentric addition processing on the average segmented text knowledge representation and the basic knowledge representation to obtain the adjusted basic knowledge representation of the project review category. In the above content, the basic knowledge representation is only subjected to one momentum adjustment based on the average segmented text knowledge representation of the core segmented text knowledge representations of multiple core segmented texts to obtain the adjusted basic knowledge representation. The numerical values ​​of the first knowledge representation eccentricity variable and the second knowledge representation eccentricity variable are pre-configured and are not specifically limited.

[0129] In operation A3, the following processing can be completed for each segmented text learning sample: determine the similarity measurement result between the segmented text knowledge representation of the segmented text learning sample and the adjustment basic knowledge representation of each project review category, and determine the basic supervision mark corresponding to the target adjustment basic knowledge representation with the largest similarity measurement result as the segmented text basic supervision mark of the segmented text learning sample.

[0130] In a feasible design, operation A4 may include the following operations: for each segmented text learning example, the following processing is completed: operation A41, obtaining the first label eccentricity variable of the segmented text basic supervision label and the second label eccentricity variable of the segmented text supervision label; operation A42, using the first label eccentricity variable and the second label eccentricity variable, performing eccentric addition processing on the segmented text basic supervision label and the segmented text supervision label to obtain the adjusted segmented text supervision label of the segmented text learning example. In the above content, the segmented text supervision label can be momentum-adjusted by the segmented text basic supervision label. The numerical values ​​of the first label eccentricity variable and the second label eccentricity variable are pre-configured and are not specifically limited.

[0131] In a feasible design, after adjusting the segmented text supervision mark according to operations A1 to A4, the segmented text supervision mark in operation S30 is all passed through the adjusted segmented text supervision mark. In operation S31, based on the adjusted segmented text supervision mark of each segmented text learning sample in the text learning sample, the text learning sample is subjected to a second project review and estimation through the text dimension classification operator to obtain the sample review and estimation result of the text learning sample; in operation S32, the algorithm cost value of the text review algorithm is determined through the segmented review and estimation result of each segmented text learning sample and the deviation value of the adjusted segmented text supervision mark, as well as the deviation value of the sample review and estimation result of the text learning sample and the text supervision mark.

[0132] When the embodiment of the present disclosure performs a project review estimate on the target project material text, the target project material text is first segmented to obtain a plurality of segmented texts included in the target project material text, and then text knowledge mining is performed on each segmented text to obtain a segmented text knowledge representation of each segmented text. Through the segmented text knowledge representation of each segmented text, a first project review estimate is performed on each segmented text to obtain a segmented text review estimate result for each segmented text. Through the segmented text review estimate result of each segmented text, the eccentricity variable of each segmented text is determined, and through the segmented text knowledge representation and eccentricity variable of each segmented text, the eccentricity adjustment is performed on the segmented text knowledge representation to obtain a segmented text eccentricity knowledge representation of each segmented text. Finally, through the segmented text eccentricity knowledge representation of each segmented text, a second project review estimate is performed on the target project material text to obtain a text review estimate result of the target project material text. Since the eccentricity variable used for eccentric adjustment of the knowledge representation of each segmented text indicates the influence of the segmented text on the text review estimation results of the target project material text, the eccentric knowledge representation of each segmented text obtained by eccentric adjustment can better represent the target project material text, making the text review estimation results obtained through the project review estimation of the eccentric knowledge representation of each segmented text more accurate and the evaluation of the target project material text more reliable.

[0133] Based on the foregoing embodiments, the embodiments of the present disclosure provide a project material text analysis device. The various units included in the device, and the various modules included in each unit, can be implemented by a processor in a computer device; of course, they can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0134] Figure 2 A schematic diagram of the structure of a project material text analysis device provided by an embodiment of the present disclosure is shown as follows: Figure 2 As shown, the project material text analysis device 200 includes:

[0135] A text segmentation module 210 is configured to segment the target project material text to obtain a plurality of segmented texts included in the target project material text;

[0136] A knowledge mining module 220 is configured to perform text knowledge mining on each of the segmented texts to obtain a segmented text knowledge representation of each of the segmented texts;

[0137] A first estimation module 230 is configured to perform a first project review estimation on each of the segmented texts using the segmented text knowledge representation of each segmented text, and obtain a segmented text review estimation result for each segmented text;

[0138] an eccentricity calculation module 240 for determining, for each segmented text, an eccentricity variable of the segmented text based on the segmented text review estimation result of the segmented text, and determining a segmented text eccentricity knowledge representation of the segmented text based on the segmented text knowledge representation of the segmented text and the eccentricity variable, wherein the eccentricity variable represents the influence of the segmented text on the text review estimation result of the target project material text;

[0139] The second estimation module 250 is used to perform a second project review estimation on the target project material text by using the segmented text eccentricity knowledge representation of the multiple segmented texts to obtain a text review estimation result of the target project material text.

[0140] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided by the embodiment of the present disclosure can be used to perform the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present disclosure, please refer to the description of the method embodiment of the present disclosure for understanding.

[0141] It should be noted that, in the embodiments of the present disclosure, if the above-mentioned project material text analysis method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present disclosure, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiments of the present disclosure are not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0142] An embodiment of the present disclosure provides a project material text review system, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0143] The present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.

[0144] An embodiment of the present disclosure provides a computer program, including computer-readable codes. When the computer-readable codes are executed in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0145] The present disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0146] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the embodiments, and reference can be made to the similarities or similarities between them. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above-mentioned method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product disclosed herein, please refer to the description of the method embodiments disclosed herein for understanding.

[0147] Figure 3 A hardware entity diagram of a project material text review system provided by an embodiment of the present disclosure, such as Figure 3 As shown, the hardware entity of the project material text review system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.

[0148] The memory 1002 stores computer programs that can be run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001. It can also cache data to be processed or processed by the processor 1001 and various modules in the project material text review system 1000 (for example, image data, audio data, voice communication data and video communication data). It can be implemented through flash memory (FLASH) or random access memory (RAM).

[0149] When the processor 1001 executes the program, the steps of any of the above-mentioned project material text analysis methods are implemented. The processor 1001 generally controls the overall operation of the project material text review system 1000.

[0150] An embodiment of the present disclosure provides a computer storage medium storing one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the project material text analysis method of any of the above embodiments.

[0151] It should be noted here that: the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding. The above processor can be at least one of a target application integrated circuit (Application Specific Integrated Circuit, ASIC), a digital signal processor (Digital Signal Processor, DSP), a digital signal processing device (Digital Signal Processing Device, DSPD), a programmable logic device (Programmable Logic Device, PLD), a field programmable gate array (Field Programmable Gate Array, FPGA), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, and a microprocessor. It is understandable that the electronic device that realizes the above processor function can also be other, and the present disclosure embodiment is not specifically limited.

[0152] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0153] It should be understood that the references to "one embodiment" or "an embodiment" throughout the specification mean that the specific features, structures, or characteristics associated with the embodiment are included in at least one embodiment of the present disclosure. Therefore, the references to "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present disclosure, the order of the serial numbers of the above-mentioned steps / processes does not necessarily indicate the order of execution. The order of execution of the steps / processes should be determined by their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure are for descriptive purposes only and do not represent the advantages or disadvantages of the embodiments. It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further constraints, an element defined by the phrase "comprises a..." does not exclude the existence of other identical elements in the process, method, article or apparatus that includes the element.

[0154] In the several embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0155] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0156] In addition, all functional units in the embodiments of the present disclosure may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0157] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0158] Alternatively, if the above-mentioned integrated unit of the present disclosure is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0159] The above is only an embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the present disclosure, and they should all be covered by the protection scope of the present disclosure.

Claims

1. A method for analyzing project material text, characterized in that: Applied to a project material text review system, the method includes: Segmenting the target project material text to obtain a plurality of segmented texts included in the target project material text; Performing text knowledge mining on each of the segmented texts to obtain a segmented text knowledge representation of each of the segmented texts; specifically comprising: performing text knowledge mining in a first feature space on each of the segmented texts to obtain a transitional segmented text knowledge representation in the first feature space of each of the segmented texts; splicing the transitional segmented text knowledge representations in the first feature space of each of the multiple segmented texts to obtain a transitional segmented text knowledge matrix of the target project material text; performing text knowledge mining in a second feature space on the transitional segmented text knowledge matrix to obtain a segmented text knowledge matrix of the target project material text, the segmented text knowledge matrix including the segmented text knowledge representation in the second feature space of each of the segmented texts, the dimension of the second feature space being lower than the dimension of the first feature space; Performing a first item review and estimation on each of the segmented texts using the segmented text knowledge representation of each segmented text to obtain a segmented text review and estimation result for each segmented text; For each of the segmented texts, determining the eccentricity variable of the segmented text based on the segmented text review estimation result of the segmented text, specifically comprising: obtaining the number of project review categories of the first project review estimation and determining the average distribution variable of the number of project review categories; determining a deviation value between the segmented text review estimation result of the segmented text and the average distribution variable; and using the deviation value as the eccentricity variable of the segmented text; and determining the segmented text eccentric knowledge representation of the segmented text through the segmented text knowledge representation and the eccentric variable of the segmented text, specifically comprising: multiplying the segmented text knowledge representation of the segmented text by the eccentric variable of the segmented text to obtain the segmented text eccentric knowledge representation of the segmented text, wherein the eccentric variable represents the influence of the segmented text on the text review estimation result of the target project material text; The second project review estimation is performed on the target project material text through the segmented text eccentricity knowledge representation of the multiple segmented texts to obtain a text review estimation result of the target project material text.

2. The method according to claim 1, characterized in that The step of performing a second project review and estimation on the target project material text by using the segmented text eccentricity knowledge representation of the plurality of segmented texts to obtain a text review and estimation result of the target project material text includes: Performing knowledge integration processing on the segmented text eccentricity knowledge representations of the plurality of segmented texts to obtain an integrated segmented text knowledge representation; Through the integrated segmented text knowledge representation, a second project review estimate is performed on the target project material text to obtain a text review estimate result of the target project material text.

3. The method according to claim 1, characterized in that The first project review estimate is completed by a segmented dimension classification operator of a text review algorithm, and the second project review estimate is completed by a text dimension classification operator of the text review algorithm; The method further comprises: Obtaining a text learning sample set for optimizing the text audit algorithm, the text learning sample set comprising a plurality of text learning samples, the text learning samples matching text supervision labels, the text learning samples comprising a plurality of segmented text learning samples, the segmented text learning samples matching segmented text supervision labels; By using the segmented dimension classification operator, a first item review and estimation is performed on each segmented text learning example in each of the text learning examples to obtain a segmented review and estimation result of each segmented text learning example; For each of the text learning examples, complete the following processing: Based on the segmented text supervision labels of each segmented text learning example in the text learning example, performing a second project review estimation on the text learning example through the text dimension classification operator to obtain a sample review estimation result of the text learning example; Determine the algorithm cost of the text audit algorithm by using the deviation value between the segmented audit prediction result and the segmented text supervision mark of each segmented text learning example, and the deviation value between the sample audit prediction result and the text supervision mark of the text learning example; The algorithm parameters of the text audit algorithm are modified according to the algorithm cost value to optimize the text audit algorithm.

4. The method according to claim 3, characterized in that Determining the algorithm cost of the text audit algorithm by using the deviation value between the segmented audit prediction result and the segmented text supervision mark of each segmented text learning example, and the deviation value between the sample audit prediction result and the text supervision mark of the text learning example, includes: For each of the segmented text learning examples, determining the transition segmentation dimension cost of the text review algorithm through the segmented review estimation result of the segmented text learning example and the deviation value of the segmented text supervision label; Accumulating the transition segment dimension costs of the text audit algorithm to obtain the segment dimension cost of the text audit algorithm; Determining the text dimension cost of the text audit algorithm by using the sample audit estimation result of the text learning sample and the deviation value of the text supervision mark; By using the first cost eccentricity variable of the segment dimension cost and the second cost eccentricity variable of the text dimension cost, the segment dimension cost and the text dimension cost are subjected to eccentric addition processing to obtain the algorithm cost value of the text audit algorithm; The first item review and estimation is performed on each of the segmented text learning examples in each of the text learning examples by using the segmented dimension classification operator to obtain a segmented review and estimation result of each of the segmented text learning examples, including: Acquiring segmented text example knowledge of each segmented text learning example in each of the text learning examples; By using the segmented dimension classification operator and the segmented text example knowledge of each segmented text learning example, a first project review estimation is performed on each segmented text learning example to obtain a segmented review estimation result of each segmented text learning example; The segmented text supervision labeling of each segmented text learning example in the text learning example, performing a second project review and estimation on the text learning example through the text dimension classification operator, and obtaining a sample review and estimation result of the text learning example, includes: Determining the segmented text eccentricity variable of each segmented text learning example through the segmented text supervision label of each segmented text learning example in the text learning example, and determining the segmented text sample eccentricity knowledge of each segmented text learning example through the segmented text sample knowledge and the segmented text eccentricity variable of each segmented text learning example; By using the segmented text sample eccentricity knowledge of multiple segmented text learning samples in the text learning sample, the second project review estimation is performed on the text learning sample through the text dimension classification operator to obtain the sample review estimation result of the text learning sample.

5. The method according to claim 4, characterized in that The determining of the segmented text eccentricity variable of each segmented text learning example by using the segmented text supervision label of each segmented text learning example in the text learning example comprises: For each of the segmented text learning examples in the text learning examples, the following processing is completed: Obtaining the number of project audit categories estimated in the first project audit, and determining an average distribution variable of the number of project audit categories; Determining a deviation value between the segmented text supervision label of the segmented text learning example and the average distribution variable; The deviation value is used as the segmented text eccentricity variable of the segmented text learning example.

6. The method according to claim 3, characterized in that Before completing the following processing for each of the text learning examples, the method further includes: Determine the core segmented texts of each project review category obtained by the first project review estimate from the multiple segmented text learning examples in the text learning example set based on the segmented review estimate results of each segmented text learning example in the text learning example set; For each of the project review categories, a basic knowledge representation and a corresponding basic supervision mark of the project review category are obtained, and based on the core segmented text of the project review category, the basic knowledge representation is momentum adjusted to obtain an adjusted basic knowledge representation of the project review category; For each of the segmented text learning examples, determining a similarity measurement result between the segmented text knowledge representation of the segmented text learning example and the adjusted basic knowledge representation of each of the project review categories, and determining the basic supervision mark corresponding to the target adjusted basic knowledge representation with the largest similarity measurement result as the segmented text basic supervision mark of the segmented text learning example; For each of the segmented text learning examples, momentum-adjusting the segmented text supervision label of the segmented text learning example using the segmented text basic supervision label of the segmented text learning example to obtain an adjusted segmented text supervision label of the segmented text learning example; The segmented text supervision labeling of each segmented text learning example in the text learning example, performing a second project review and estimation on the text learning example through the text dimension classification operator, and obtaining a sample review and estimation result of the text learning example, includes: Based on the adjusted segmented text supervision label of each segmented text learning example in the text learning example, performing a second project review estimation on the text learning example through the text dimension classification operator to obtain a sample review estimation result of the text learning example; Determining the algorithm cost of the text audit algorithm by using the deviation value between the segmented audit prediction result and the segmented text supervision mark of each segmented text learning example, and the deviation value between the sample audit prediction result and the text supervision mark of the text learning example, includes: The algorithm cost of the text review algorithm is determined by the deviation value of the segmented review prediction result and the adjusted segmented text supervision mark of each segmented text learning sample, as well as the deviation value of the sample review prediction result and the text supervision mark of the text learning sample.

7. The method according to claim 6, characterized in that The segmented review estimation result includes the segmented estimated support coefficient of the segmented text learning sample belonging to each of the project review categories; The segmented review estimation results of each segmented text learning example in the text learning example set are used to determine the core segmented text of each project review category obtained by the first project review estimation from multiple segmented text learning examples in the text learning example set, including: For each of the project review categories described, complete the following process: Determining, from the plurality of segmented text learning examples in the text learning example set, a preset number of target segmented text learning examples that belong to the project review category and have the largest segmented estimated support coefficients; The preset number of target segmented text learning examples are determined as the preset number of core segmented texts of the project review category.

8. The method according to claim 6, characterized in that The number of the core segmented texts is S, where S is greater than 1. The core segmented texts based on the project review category, momentum-adjusting the basic knowledge representation to obtain the adjusted basic knowledge representation of the project review category, include: Momentum-adjust the basic knowledge representation through the first core segmented text of the project review category to obtain a first transition basic knowledge representation of the project review category; Momentum-adjust the jth transition basic knowledge representation of the project review category through the kth core segmented text of the project review category to obtain the kth transition basic knowledge representation of the project review category, where j=k-1, and j is a positive integer less than or equal to S; Wandering through the core segmented text, obtaining the Sth transition basic knowledge representation of the project review category, and using the Sth transition basic knowledge representation of the project review category as the adjusted basic knowledge representation of the project review category; Alternatively, there are multiple core segmented texts, and the core segmented texts based on the project review category are momentum-adjusted to the basic knowledge representation to obtain the adjusted basic knowledge representation of the project review category, including: Obtaining a core segmented text knowledge representation of each core segmented text of the project review category, and determining an average segmented text knowledge representation of a plurality of the core segmented text knowledge representations; Obtaining a first knowledge representation eccentricity variable of the average segmented text knowledge representation and a second knowledge representation eccentricity variable of the basic knowledge representation; Performing eccentric addition processing on the average segmented text knowledge representation and the basic knowledge representation through the first knowledge representation eccentricity variable and the second knowledge representation eccentricity variable to obtain an adjusted basic knowledge representation of the project review category; The step of adjusting the segmented text supervision label of the segmented text learning example by momentum based on the segmented text basic supervision label of the segmented text learning example to obtain the adjusted segmented text supervision label of the segmented text learning example includes: Obtaining a first label eccentricity variable of the segmented text basic supervision label and a second label eccentricity variable of the segmented text supervision label; The segmented text basic supervision label and the segmented text supervision label are subjected to eccentric addition processing by using the first label eccentricity variable and the second label eccentricity variable to obtain the adjusted segmented text supervision label of the segmented text learning example.

9. A project material text analysis device, characterized in that: include: A text segmentation module, configured to segment the target project material text to obtain a plurality of segmented texts included in the target project material text; A knowledge mining module is used to perform text knowledge mining on each of the segmented texts to obtain a segmented text knowledge representation of each of the segmented texts; specifically comprising: performing text knowledge mining in a first feature space on each of the segmented texts to obtain a transitional segmented text knowledge representation in the first feature space of each of the segmented texts; concatenating the transitional segmented text knowledge representations in the first feature space of each of the multiple segmented texts to obtain a transitional segmented text knowledge matrix of the target project material text; performing text knowledge mining in a second feature space on the transitional segmented text knowledge matrix to obtain a segmented text knowledge matrix of the target project material text, the segmented text knowledge matrix including the segmented text knowledge representation in the second feature space of each of the segmented texts, the dimension of the second feature space being lower than the dimension of the first feature space; A first estimation module is configured to perform a first project review estimation on each of the segmented texts based on the segmented text knowledge representation of each segmented text, and obtain a segmented text review estimation result for each segmented text; An eccentricity calculation module is used to determine the eccentricity variable of the segmented text according to the segmented text review estimation result of the segmented text for each segmented text, specifically including: obtaining the number of project review categories of the first project review estimation, and determining the average distribution variable of the number of project review categories; determining the deviation value between the segmented text review estimation result of the segmented text and the average distribution variable; using the deviation value as the eccentricity variable of the segmented text; and determining the segmented text eccentricity knowledge representation of the segmented text according to the segmented text knowledge representation and the eccentricity variable of the segmented text, specifically including: multiplying the segmented text knowledge representation of the segmented text by the eccentricity variable of the segmented text to obtain the segmented text eccentricity knowledge representation of the segmented text, wherein the eccentricity variable represents the influence of the segmented text on the text review estimation result of the target project material text; The second estimation module is used to perform a second project review estimation on the target project material text through the segmented text eccentricity knowledge representation of the multiple segmented texts, and obtain a text review estimation result of the target project material text.

10. A project material text review system, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Detection method and device for AI generated text, medium and equipment

    CN117151074A

  • Multi-business type text auditing method, computer device and computer readable storage medium

    CN117194658A