Text data processing method and system based on mathematical logic and propositional logic
By employing a text data processing method based on mathematical logic and propositional logic, natural language is transformed into computable rule-based entries. Combined with weight calculation and balancing algorithms, this method addresses the insufficient accuracy of existing natural language processing models in high-accuracy domains, achieving efficient and accurate response data generation.
Patent Information
- Application Number
- CN202510276603.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-11-07
AI Technical Summary
Existing natural language processing models are insufficient in accuracy for fields with high accuracy requirements, such as mathematics, medicine, law, and actuarial science, and existing methods are cumbersome and not intuitive enough.
By introducing mathematical and propositional logic, natural language is transformed into computable rule-based entries. Combined with weight calculation and balancing algorithms, the accuracy and rigor of logical operations are ensured, generating highly accurate response data.
It enables efficient natural language processing and generates accurate response data in error-tolerant fields, applicable to fields such as medicine, law, and mathematics, meeting high accuracy requirements.
Smart Images

Figure CN120910233A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of text data processing, and particularly relates to a text data processing method and system based on mathematical logic and propositional logic. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, especially the wide application of natural language processing models such as GPT, artificial intelligence has shown quite comprehensive ability in answering knowledge questions in specified vertical fields. However, in some fields with extremely high accuracy requirements, such as mathematics, medicine, law and actuarial science, the accuracy of the answers of existing models still needs to be further improved. These fields cannot tolerate errors because incorrect answers may have serious consequences. Therefore, it is particularly important to develop a technical method that can ensure accurate answers in these fields.
[0003] At present, natural language processing models such as GPT are mainly based on the principle of "empirical learning", which learns and summarizes rules through a large amount of data to answer user questions. When encountering a type of question that already exists in the database, these models can summarize and operate according to the data and reply. However, when facing a large number of pure mathematical or rule-based operations, the free play characteristics of GPT-type models may lead to answers that do not meet expectations. In order to limit this free play, researchers try to constrain the behavior of the model by providing a large number of artificial instructions, but this method is cumbersome and not intuitive enough. SUMMARY
[0004] To overcome the deficiencies and difficulties of the prior art, the purpose of the present application is to provide a text data processing method and system based on mathematical logic and propositional logic to enable any propositional statement to obtain an accurate answer based on data.
[0005] In a first aspect, the present application provides a text data processing method based on mathematical logic and propositional logic, comprising:
[0006] Obtaining text data corresponding to the data to be processed, assigning weights to the text data, and generating underlying data;
[0007] Based on a pre-constructed logical framework, converting natural language statements in the underlying data into rule-based items that can be operated to generate the base value of each item; wherein the logical framework is constructed based on mathematical logic and propositional logic;
[0008] Introducing the user input logical item set manifests, the positive control item +manifests and the negative control item -manifests into the underlying data for weight balancing processing;
[0009] According to the matching degree of the entry after weight balancing and the basic value thereof, a score of each entry is calculated, and an entry with the highest score is selected as the response data.
[0010] Preferably, the text data corresponding to the to-be-processed data is obtained, and the text data corresponding to the to-be-processed data is obtained.
[0011] The to-be-processed data is obtained, and if the to-be-processed data contains audio or video data, the audio or video data is converted into text data, and the text data is aligned by a large language model.
[0012] Preferably, the text data is assigned a weight, and the bottom layer data is generated, including:
[0013] For each token in the text data, a weight is assigned according to its importance in the global and context, and the text data with the assigned weight is taken as the bottom layer data; wherein the token represents a basic unit of the text data.
[0014] Preferably, the logical framework includes node input definition and operation node definition, wherein:
[0015] Node input definition: each node fixedly inputs all tokens and their corresponding weights; the token represents a basic unit of the text data;
[0016] The following variables are defined:
[0017] var-count: total token quantity;
[0018] pos-count: quantity of weights greater than 1;
[0019] neg-count: quantity of weights less than 1;
[0020] val: token value in the form of weight serialization;
[0021] List S = [var-count, pos-count, neg-count, val];
[0022] S set except var-count is S';
[0023] p': if the input value > 0, return 1, otherwise return 1;
[0024] q': if the input value > 0, return the original value, otherwise return 0;
[0025] e'(a, n): represents q'(n) power of a;
[0026] Operation node definition:
[0027] The operation node is the core of the logical framework, which is used to convert natural language sentences into logical expressions. The following is the formal description of each operation node:
[0028] The formula of the logical negation operation not(S) is: not(S) = (p'(neg-count)) + ((1-sum(val)) x p'(sum(val)) + (0.5 x p'(var-count-pos-count-neg-count-p'(sum(val))));
[0029] The formula of the logical and operation and(var-count, S') is: and(var-count, S') = (e'(0, neg-count)) x (1 / e'(2, var-count)) x (Π n∈val (e'(0, n) + (n x p'(n));
[0030] The formula of the logical or operation or(S) is: or(S) = not(and({not(s') | s' ∈ S}));
[0031] The formula of the logical not operation!(S) is:!(S) = and(var-count, not({not(s') | s' ∈ S'}));
[0032] The formula of the logical implication operation then(var-count, S') is: then(var-count, S') = not(and(var-count, not(S')).
[0033] Preferably, the introduction of the positive control item +manifests and the negative control item -manifests into the underlying data, and the weight balancing processing, comprises:
[0034] After converting the user input data to be processed into the same format as the underlying data, a set of logical entries manifests is obtained from the user input;
[0035] Combine the logical entry set manifests, the positive control item +manifests and the negative control item -manifests with the underlying data;
[0036] Adjust the weight of each entry through the balancing algorithm;
[0037] According to the balancing results of +manifest and -manifest, the corresponding entries in manifest are adjusted again.
[0038] Preferably, the score of each item is calculated according to the matching degree of the item after weight balancing with its base value, the items are sorted according to the scores, the answer data is obtained based on the sorted result list, and the method comprises the following steps:
[0039] The score of each item is calculated according to the matching degree of the item after weight balancing with its base value; the indicators of the score include validity, precision, recall and confidence;
[0040] All the items are sorted according to the scores from high to low to obtain a result list, and the item with the highest score is selected from the result list as the answer data;
[0041] The answer data is polished by combining the user input through a large language model, and the polished answer data is output.
[0042] Preferably, the score of each item is calculated according to the matching degree of the item after weight balancing with its base value, and the method comprises the following steps:
[0043] For each item, the validity, precision, recall and confidence are calculated based on the item result after weight balancing and the base value of the item; wherein:
[0044]
[0045] In the formula, result represents the calculation result of the item, and default represents the base value of the item;
[0046]
[0047] In the formula, |default| represents the var-count corresponding to the base value, and var-count represents the total number of all tokens in the item;
[0048]
[0049] In the formula, intersect represents the intersection number of +manifests and the var of the item, var represents the set of all tokens in the item; -intersect represents the intersection number of -manifests and the negative var of the item; r+intersect represents the intersection number of +manifest and the negative var of the item; r-intersect represents the intersection number of -manifest and the var of the item; and +diff represents the difference set number of +manifest and the var of the item;
[0050]
[0051] In the formula, t-intersect represents the number of intersections of +manifest and taboo-manifest; taboo-manifest represents a user-specified weight node that is not desired to be considered, used to limit the logical operation range; |+manifests| represents the number of +manifests; |-manifests| represents the number of -manifests; taboo-value represents the calculated value of taboo on the logical framework; and taboo represents the user-specified content that is not desired to appear in the output result;
[0052] The calculation formula of the score rating is:
[0053] rating = validity x precision x recall x confidence.
[0054] In a second aspect, the present application provides a text data processing system based on mathematical logic and propositional logic, comprising:
[0055] A text acquisition unit is configured to acquire text data corresponding to the data to be processed, assign weights to the text data, and generate underlying data.
[0056] A logical operation unit is configured to convert natural language sentences in the underlying data into operable rule-making items based on a pre-constructed logical framework, and generate a basic value of each item; wherein the logical framework is constructed based on mathematical logic and propositional logic.
[0057] A weight balancing unit is configured to introduce a user-input logical item set manifests, a positive control item +manifests, and a negative control item -manifests into the underlying data, and perform weight balancing processing.
[0058] A response data output unit is configured to calculate the score of each item according to the matching degree of the balanced item and its basic value, and select the item with the highest score as the response data.
[0059] Through the combination of mathematical logic and propositional logic, the present application realizes efficient processing and accurate answering of natural language, and is particularly suitable for error-intolerant fields. By introducing weight calculation, balancing algorithm, and scoring mechanism, the present application can flexibly respond to user needs, generate high-accuracy response data, and has a wide application prospect and significant technical advantages. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0061] Figure 1 The text data processing method flowchart based on mathematical logic and propositional logic provided by the present application is shown in the figure.
[0062] Figure 2 The text data processing system architecture diagram based on mathematical logic and propositional logic provided by the present application is shown in the figure. DETAILED DESCRIPTION
[0063] In order to better understand the purpose, technical solutions and advantages of the present application, the following will further describe the present application in combination with the drawings and specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the present description.
[0064] In this paper, the phrase "embodiment" means that the specific features, structures or characteristics described in combination with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment that is not mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0065] Although the existing technology (such as GPT and other natural language processing models) has shown quite comprehensive ability in answering knowledge questions in a specified vertical field, the accuracy of the answers is still insufficient in some fields where there is no room for error (such as mathematics, medicine, law, and actuarial science). The main reason is that GPT-type models are based on the principle of "empirical learning", which learns and summarizes rules through a large amount of data. When faced with existing question types in the database, it can summarize and infer according to the data and perform operations and replies. However, when faced with a large number of pure mathematical or rule-based operations, the free play characteristics of GPT-type models may result in answers that do not meet expectations. Although a large number of artificial instructions can be provided to limit the scope of free play, this method is tedious and not intuitive, and it is difficult to meet the high accuracy requirement.
[0066] Therefore, the present application provides a text data processing method and system based on mathematical logic and propositional logic to solve the problem of insufficient accuracy in answering in error-tolerant fields such as mathematics, medicine, law, and actuarial science. By introducing mathematical logic and propositional logic, natural language is converted into computable rule-making items, and combined with weight calculation and balancing algorithm to ensure the accuracy and rigor of logical operation. The following will be expanded and introduced through multiple embodiments.
[0067] Figure 1 The text data processing method based on mathematical logic and propositional logic provided by the present application is shown in the flowchart Figure 1 , which comprises the following steps:
[0068] Step S1, obtaining the text data corresponding to the data to be processed, assigning weights to the text data, and generating bottom layer data.
[0069] Specifically, the data basis unit is text; if the input data contains audio or video, first convert the audio track to text through a general TTS framework (such as Whisper). Use a large language model (LLM) to align the converted text, ensuring the consistency and accuracy of the text.
[0070] Further, according to the importance of each token in the text data in the global and context, a weight is assigned to it. Wherein, token represents the basic unit of text data. The method of assigning weights includes:
[0071] 1) According to the logical framework to calculate in advance:
[0072] Initialize the weight of all tokens to 1.
[0073] This method is simple and direct, and is suitable for the initial construction of the logical framework.
[0074] 2) According to the word frequency:
[0075] According to the frequency of token in the corpus, the weight is assigned.
[0076] High-frequency words may have lower weights, and low-frequency words may have higher weights.
[0077] 3) Use Word2Vec:
[0078] Use word vector models such as Word2Vec to assign weights according to the position of tokens in the vector space.
[0079] This method can capture the semantic information of tokens and is suitable for complex text processing tasks.
[0080] The weight assignment method can be a word-level vector embedding (Word2Vec) or a text-level vector embedding (such as BGE) in addition to the Word2Vec of the present scheme.
[0081] Next, the text data to which the weight is assigned is converted into a vector form to facilitate subsequent logical operations. The vectorized text data serves as precalculation, which is the basis for subsequent logic generation framework and score calculation.
[0082] In step S2, based on the pre-constructed logic framework, the natural language sentence in the underlying data is converted into a rule-making item that can be operated to generate a basic value of each item. The logic framework is constructed based on mathematical logic and propositional logic.
[0083] Specifically, the construction of the logic framework includes node input definition and operation node definition, wherein:
[0084] Node input definition: each node fixedly inputs all tokens and their corresponding weights; token represents the basic unit of text data;
[0085] The following variables are defined:
[0086] var-count: total token quantity;
[0087] pos-count: quantity of weights greater than 1;
[0088] neg-count: quantity of weights less than 1;
[0089] val: token value in the form of weight serialization;
[0090] List S = [var-count, pos-count, neg-count, val];
[0091] S' is the S set excluding var-count;
[0092] p': if the input value > 0, return 1, otherwise return 1;
[0093] q': if the input value > 0, return the original value, otherwise return 0;
[0094] e'(a, n): represents q'(n) power of a;
[0095] Operation node definition:
[0096] The operation node is the core of the logic framework, which is used to convert the natural language sentence into a logical expression. The following is the formal description of each operation node:
[0097] negation not(S), formula: not(S) = (p'(neg-count)) + ((1-sum(val)) * p'(sum(val))) + (0.5 * p'(var-count-pos-count-neg-count-p'(sum(val)));
[0098] logical and and(var-count, S'), formula: and(var-count, S') = (e'(0, neg-count)) * (1 / e'(2, var-count)) * (Pi n∈val (e'(0, n) + (n * p'(n));
[0099] logical or or(S), formula: or(S) = not(and({not(s') | s' element-of S}));
[0100] logical negation!(S), formula:!(S) = and(var-count, not({not(s') | s' element-of S}));
[0101] logical implication then(var-count, S'), formula: then(var-count, S') = not(and(var-count, not(S'))).
[0102] The logical framework of the present application converts natural language sentences into predicate calculus form through the rigor of mathematical logic and the reasoning ability of propositional logic, so that it can perform logical operations.
[0103] It should be noted that the core advantages of the logical framework constructed by the present application include:
[0104] 1. Traditional logical operations cannot directly consider weights. The present application introduces weights into logical operations by defining val (token value in the form of weight sequence) and auxiliary functions (such as p', q', e').
[0105] 2. The logical framework of the present application is designed to be universal and applicable to all natural language sentences, and can handle propositional statements in different fields.
[0106] 3. Through the combination of mathematical logic and propositional logic, natural language is converted into computable rules, ensuring the accuracy and consistency of the operation results.
[0107] In this embodiment, the natural language sentence is converted into an operable rule-making item through the above logical framework. For each item, a default value is generated according to the token, weight and logical operation result contained therein.
[0108] Step S3: Introduce the user input logical item set manifests, positive control item +manifests and negative control item-manifests into the underlying data and perform weight balancing processing.
[0109] In a preferred embodiment of the present application, step S3 specifically includes:
[0110] S31: After converting the user input data to be processed into the same format as the underlying data, the user input logical item set manifests is obtained.
[0111] At the same time, the user can select to input taboo (unwanted output results) and taboo-manifest (weight nodes not to be considered)
[0112] S32: Combine the logical item set manifests, positive control item +manifests and negative control item-manifests with the underlying data.
[0113] Among them, the user or the system can specify the weight of the items in +manifests and-manifests to adjust their importance.
[0114] S33: Adjust the weight of each item through the balancing algorithm.
[0115] Combine manifests, +manifests, -manifests with precalculation (preprocessing data).
[0116] Adjust the weight of each item through the balancing algorithm to meet the user's needs and system logic. The balancing algorithm includes:
[0117] Based on word frequency: adjust the weight according to the frequency of token in the corpus.
[0118] Based on vector space distance: adjust the weight according to the relative distance of token in the vector space.
[0119] Other algorithms: such as TF-IDF, Word2Vec, etc.
[0120] Balancing goals include:
[0121] Positive reinforcement: the weight of the items in +manifests will be enhanced, making them more prominent in the results.
[0122] Negative weakening: the weight of the entry in -manifest is weakened, so that it is suppressed in the result.
[0123] Taboo exclusion: the weight of the entry in taboo-manifest is set to 0, completely excluded from the result.
[0124] S34, according to the balanced results of +manifest and -manifest, the corresponding entries in manifest are adjusted in the second weight.
[0125] After the first balancing, the weights of +manifest and -manifest have been adjusted, but the entry weights in manifest still need further optimization.
[0126] According to the balanced results of +manifest and -manifest, the corresponding entries in manifest are adjusted in the second weight. If an entry has a higher weight in +manifest, it will also be enhanced in manifest. If an entry has a lower weight in -manifest, it will also be weakened in manifest. Through the secondary balancing, the accuracy and user satisfaction of the result are further improved.
[0127] Step S4, according to the matching degree of the entry after weight balancing and its basic value, calculate the score of each entry, select the entry with the highest score as the response data.
[0128] Step S4 can specifically include:
[0129] S41, according to the matching degree of the entry after weight balancing and its basic value, calculate the score of each entry; the indicators of the score include validity, precision, recall and confidence;
[0130] S42, sort all entries according to the score from high to low to get the result list, select the entry with the highest score from the result list as the response data;
[0131] S43, polish the response data by combining the user input through a large language model, and output the polished response data.
[0132] The application realizes efficient processing and accurate answering of natural language by combining mathematical logic and propositional logic, and is particularly suitable for error-intolerant fields (such as medical treatment, law, mathematics and actuarial science). By introducing weight calculation, balanced algorithm and scoring mechanism, the application can flexibly process text data in different scenarios, flexibly respond to user needs, generate high-accuracy response data, and has wide application prospects and significant technical advantages.
[0133] In a preferred embodiment of the application, in S41, the score of each entry is calculated according to the matching degree of the entry after weight balancing and the basic value of the entry, comprising:
[0134] For each entry, the validity, precision, recall and confidence are calculated based on the entry result after weight balancing and the basic value of the entry; wherein:
[0135]
[0136] In the formula, result represents the calculation result of the entry, and default represents the basic value of the entry;
[0137]
[0138] In the formula, |default| represents the var-count corresponding to the basic value, and var-count represents the total number of all tokens in the entry;
[0139]
[0140] In the formula, intersect represents the intersection number of +manifests and the var of the entry, var represents the set of all tokens in the entry; -intersect represents the intersection number of -manifests and the var of the entry; r+intersect represents the intersection number of +manifest and the negative var of the entry; r-intersect represents the intersection number of -manifest and the var of the entry; and +diff represents the difference set number of +manifest and the var of the entry;
[0141]
[0142] In the formula, t-intersect represents the number of intersections of +manifests and taboo-manifests; taboo-manifest represents a user-specified weight node that is not expected to be considered, and is used to limit the logical operation range; |+manifests| represents the number of +manifests; |-manifests| represents the number of -manifests; taboo-value represents the calculated value of taboo on the logical framework; and taboo represents a user-specified content that is not expected to appear in the output result.
[0143] The calculation formula of the score rating is as follows:
[0144] rating=validity*precision*recall*confidence.
[0145] By multiplying the four indexes, the comprehensive score of each item is obtained. The higher the score, the more the item meets the user's demand.
[0146] The text data processing method based on mathematical logic and propositional logic provided by the application ensures that in error-intolerant fields (such as medical treatment and law), the response data completely meets the user's demand and the system logic, and has high accuracy and reliability.
[0147] Figure 2 The text data processing system architecture based on mathematical logic and propositional logic provided by the application is shown in Figure 2 The text data processing system 200 based on mathematical logic and propositional logic comprises:
[0148] A text acquisition unit 201 is configured to acquire text data corresponding to the data to be processed, assign weights to the text data, and generate underlying data.
[0149] A logical operation unit 202 is configured to convert natural language statements in the underlying data into rule-based items that can be operated based on a pre-constructed logical framework, and generate a basic value of each item; wherein the logical framework is constructed based on mathematical logic and propositional logic.
[0150] A weight balancing unit 203 is configured to introduce a user-input logical item set manifests, a positive control item +manifests and a negative control item -manifests into the underlying data, and perform weight balancing processing.
[0151] A response data output unit 204 is configured to calculate the score of each item according to the matching degree of the weight-balanced item and the basic value thereof, and select the item with the highest score as the response data.
[0152] The text data processing system based on mathematical logic and propositional logic provided by the present application executes the text data processing method based on mathematical logic and propositional logic provided by the above-mentioned embodiments through the above-mentioned modules, and the text data processing method based on mathematical logic and propositional logic has been described in detail in the above-mentioned embodiments, which will not be described herein again.
[0153] The above-mentioned embodiments only express several embodiments of the present application, which are described in detail and specifically, but cannot be understood as the limitation of the scope of the present application. It should be pointed out that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
[0154] Finally, it should be pointed out that: the above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that: the technical solutions recorded in the above-mentioned embodiments can still be modified, or some technical features can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for processing text data based on mathematical logic and propositional logic, characterized by, The method comprises the following steps: Obtaining text data corresponding to the to-be-processed data, assigning weights to the text data, and generating underlying data; Based on the pre-constructed logical framework, the natural language statements in the underlying data are converted into rule-based items that can be operated, and the basis value of each item is generated; wherein the logical framework is constructed based on mathematical logic and propositional logic; Introducing the user input logical item set manifests, the positive control item +manifests and the negative control item-manifests into the underlying data, and performing weight balancing processing; According to the matching degree of the balanced items and their basis values, the score of each item is calculated, and the item with the highest score is selected as the response data.
2. The method for processing text data based on mathematical logic and propositional logic according to claim 1, wherein, The method comprises the following steps: Obtaining text data corresponding to the to-be-processed data, assigning weights to the text data, and generating underlying data; 3. The method for processing text data based on mathematical logic and propositional logic according to claim 1, wherein, For each token in the text data, a weight is assigned according to its importance in the global and context, and the text data with the assigned weight is used as the underlying data; wherein token represents the basic unit of text data. The logical framework includes node input definition and operation node definition, wherein:
4. The method for processing text data based on mathematical logic and propositional logic according to claim 1, wherein, Node input definition: each node fixedly inputs all tokens and their corresponding weights; token represents the basic unit of text data; Define the following variables: var-count: total token quantity; pos-count: quantity of weights greater than 1; neg-count: quantity of weights less than 1; val: token value in the form of weight serialization; List S=[var-count, pos-count, neg-count, val]; S' is the S set without var-count; p': if the input value > 0, return 1, otherwise return 1; q': if the input value > 0, return the original value, otherwise return 0; e'(a, n): represents q'(n) power of a; Operation node definition: Operation node is the core of logical framework, which is used to convert natural language statements into logical expressions; the following is the formulaic description of each operation node: Negation operation not(S), formula: not(S)=(p'(neg-count))+((1-sum(val))×p'(sum(val)))+(0.5×p'(var-count-pos-count-neg-count-p'(sum(val)))); Logical or operation or(S), formula: or(S)=not(and({not(s')|s'∈S})); Logical AND operation and (var-count, S'), which is given by the formula: and (var-count, S') = (e'(0, neg-count)) x (1 / e'(2, var-count)) x (Π n∈val (e'(0, n) + (n x p'(n)); Logical non-operation!(S), formula:!(S)=and(var-count,not({not(s')|s'∈S'})); A logic implication operation then(var-count, S') is defined as: then(var-count, S') = not(and(var-count, not(S'))).
5. The method for processing text data based on mathematical logic and propositional logic according to claim 1, wherein, The introducing the positive control item +manifests and the negative control item -manifests into the underlying data and performing weight balancing processing includes: After converting the user inputted to-be-processed data into the same format as the underlying data, a set of logical entries manifests inputted by the user is obtained; Combining the set of logical entries manifests, the positive control item +manifests and the negative control item -manifests with the underlying data; Adjusting the weights of the entries through a balancing algorithm; According to the balancing results of +manifest and -manifest, performing secondary weight adjustment on the corresponding entries in manifests.
6. The text data processing system based on mathematical logic and propositional logic as claimed in claim 5, wherein, According to the matching degree of the entries after weight balancing and their base values, calculating the scores of each entry, sorting the entries according to the scores, and obtaining the response data based on the sorted results, including: According to the matching degree of the entries after weight balancing and their base values, calculating the scores of each entry; the indicators of the scores include validity, precision, recall and confidence; Sorting all the entries according to the scores from high to low to obtain a result list, and selecting the entry with the highest score from the result list as the response data; Through a large language model, the response data is combined with the user input to polish, and the polished response data is output.
7. The text data processing system based on mathematical logic and propositional logic as claimed in claim 6, wherein, According to the matching degree of the entries after weight balancing and their base values, calculating the scores of each entry, including: For each entry, based on the results of the entries after weight balancing and the base values of the entries, the validity, precision, recall and confidence are calculated; wherein: In the formula, result represents the calculation result of the entry, and default represents the base value of the entry. In the formula, |default| represents the var-count corresponding to the base value, and var-count represents the total number of all tokens in the entry. In the formula, intersect represents the intersection number of +manifests and the entry var, var represents the set of all tokens in the entry; -intersect represents the intersection number of -manifests and the negative var of the entry; r+intersect represents the intersection number of +manifest and the negative var of the entry; r-intersect represents the intersection number of -manifest and the var of the entry; +diff represents the difference set number of +manifest and the var of the entry; In the formula, t-intersect represents the number of intersections of +manifests and taboo-manifest; taboo-manifest represents a user-specified weight node that is not desired to be considered, used to limit the logical operation range; |+manifests| represents the number of +manifests; |-manifests| represents the number of -manifests; taboo-value represents the calculated value of taboo on the logical framework; and taboo represents user-specified content that is not desired to appear in the output result. The calculation formula of the score rating is: rating=validity×precision×recall×confidence.
8. A text data processing system applied to the text data processing method based on mathematical logic and propositional logic according to any one of claims 1 to 7, characterized in that, The method comprises the following steps: a text acquisition unit is configured to acquire text data corresponding to the to-be-processed data, assign weights to the text data, and generate underlying data; a logical operation unit is configured to convert natural language statements in the underlying data into operable rule-making items based on a pre-constructed logical framework, and generate a basic value of each item; the logical framework is constructed based on mathematical logic and propositional logic; a weight balancing unit is configured to introduce a user-input logical item set manifests, a positive control item +manifests, and a negative control item -manifests into the underlying data, and perform weight balancing processing; and a response data output unit is configured to calculate the score of each item according to the matching degree of the balanced item and the basic value thereof, and select an item with the highest score as the response data.