Text content rapid classification management method and system based on artificial intelligence
By segmenting language units that are compatible with multiple languages and updating dynamic behavioral profiles, the problem of recognizing mixed languages and emerging words in existing technologies has been solved. This enables accurate classification and rapid adaptation of complex texts, improving the efficiency and accuracy of financial risk control and public opinion monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HUOYANYAN ARTIFICIAL INTELLIGENCE CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies struggle to accurately identify sensitive information when dealing with mixed languages, emerging vocabulary, and deliberately disguised expressions, leading to decreased classification accuracy and timeliness. Furthermore, the high cost of manual intervention makes it impossible to meet the timeliness and accuracy requirements of financial risk control.
A multilingual compatible language unit segmentation method is adopted to identify known and unknown semantic units. The behavioral profile of unknown semantic units is constructed through contextual information, and their meanings and related attributes are dynamically updated and corrected to generate explanatory cues to support classification decisions.
It achieves accurate identification and understanding of complex text information, quickly adapts to the evolution of online language, reduces system performance degradation and manual maintenance costs, and improves the timeliness of sensitive information identification and the accuracy of classification.
Smart Images

Figure CN122019856A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a method and system for rapid classification and management of text content based on artificial intelligence. Background Technology
[0002] In professional fields such as financial risk control and public opinion monitoring, artificial intelligence systems that can quickly and accurately classify and manage massive amounts of text information are core infrastructure supporting efficient business operations. However, with the continuous evolution of internet language and the diversification of information sources, traditional text processing methods often fall short when dealing with mixed language, emerging vocabulary, and deliberately disguised expressions, which directly affects the timeliness of sensitive information identification and the accuracy of classification.
[0003] Specifically, existing systems face challenges when processing text involving code-switching. For example, in social media and online forums, users often mix vocabulary and grammatical structures from two or more languages within the same sentence, making it difficult for the system to determine the dominant language and understand the full meaning behind the mixed structures, thus affecting classification accuracy. Furthermore, in emerging fields such as fintech and cryptocurrency, users create a large number of entirely new jargon or neologisms that blend multiple linguistic elements. These terms have short lifecycles, spread rapidly, and their meanings are highly context-dependent. Due to a lack of training data for these neologisms, existing systems often classify them as meaningless noise, missing potentially significant risk signals.
[0004] To complicate matters further, some users deliberately "transform" text in order to circumvent keyword blocking, using non-standard characters with similar pronunciations or shapes to replace normal words. This practice disrupts the system's preprocessing stage before text classification, resulting in incomplete or distorted text information from the source. This makes it impossible for the system to identify disguised words, leading to misclassification or omission of potential warning information.
[0005] Ultimately, the existing system fell into a vicious cycle, with performance continuously declining. The technical team had to invest a huge amount of manpower to continuously track new internet terms, memes, and circumvention methods, manually label new data, and frequently iterate and train. This passive, "patching" work mode was not only costly but also always lagged behind the speed of evolution of internet language, causing the efficiency advantages brought by the automated system to be completely offset, and failing to meet the timeliness and accuracy required by financial risk control.
[0006] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0007] This application discloses a method and system for rapid classification and management of text content based on artificial intelligence, which aims to solve the problems of declining timeliness and classification accuracy of sensitive information identification when faced with massive amounts of text information, as well as the continuous decline in system performance and the high cost and lag of manual intervention.
[0008] The technical solution of this application is as follows: Firstly, this application discloses a method for rapid classification and management of text content based on artificial intelligence, the method comprising: The text to be classified is obtained, and the text to be classified is segmented into language units that are compatible with multiple languages. The segmented semantic units are matched with a preset set of known semantic units to identify known and unknown semantic units in the text to be classified, while capturing the contextual information of unknown semantic units in the text to be classified. Based on contextual information, co-occurrence association features, syntactic function features, and textual tendency features of unknown semantic units are extracted to construct behavioral profiles of unknown semantic units. When new text containing unknown semantic units is received in the future, the contextual information is updated and the behavioral profiles are adjusted according to the updated contextual information. The adjusted behavioral profile is compared with a pre-set set of known semantic behavioral profiles. Based on the comparison results, the preliminary meaning and related attributes of the unknown semantic units are inferred, and the corresponding confidence scores are generated. When an unknown semantic unit is detected to reappear in newly added text, the initial meaning and related attributes are corrected based on the adjusted behavioral profile to obtain the meaning and related attributes of the unknown semantic unit, and the confidence level is adjusted simultaneously. The system integrates the meanings and related attributes of known semantic units and unknown semantic units in the text to be classified, and determines the degree of participation of the meanings and related attributes of unknown semantic units in the classification decision based on the adjusted confidence level. The system then makes a classification decision for the text to be classified and outputs the classification result. At the same time, it generates explanatory cues, which are used to indicate semantic units whose contribution to the classification result is greater than a preset contribution threshold.
[0009] Secondly, this application also discloses an artificial intelligence-based text content rapid classification management system, the system comprising: The acquisition module is used to acquire the text to be classified, perform multilingual compatible language unit segmentation on the text to be classified, and match the segmented semantic units with a preset set of known semantic units to identify known and unknown semantic units in the text to be classified, while capturing the contextual information of unknown semantic units in the text to be classified. The behavior profile building module is used to extract co-occurrence association features, syntactic function features, and textual tendency features of unknown semantic units based on context information in order to build a behavior profile of the unknown semantic units. When new text containing unknown semantic units is received in the future, the context information is updated and the behavior profile is adjusted according to the updated context information. The meaning inference and confidence generation module is used to compare the adjusted behavioral profile with a preset set of known semantic behavioral profiles, infer the preliminary meaning and related attributes of the unknown semantic units based on the comparison results, and generate the corresponding confidence scores. The meaning correction and confidence adjustment module is used to correct the initial meaning and related attributes based on the adjusted behavioral profile when an unknown semantic unit is detected to reappear in newly added text, so as to obtain the meaning and related attributes of the unknown semantic unit and adjust the confidence level simultaneously. The classification decision and interpretation module integrates the meanings and related attributes of known semantic units and unknown semantic units in the text to be classified. Based on the adjusted confidence level, it determines the degree of participation of the meanings and related attributes of unknown semantic units in the classification decision, makes a classification decision on the text to be classified, and outputs the classification result. At the same time, it generates explanatory clues, which are used to indicate semantic units whose contribution to the classification result is greater than a preset contribution threshold.
[0010] Beneficial Effects: The AI-based rapid text content classification and management method disclosed in this application effectively solves the challenges faced by existing technologies in handling mixed languages, emerging vocabulary, and deliberately disguised expressions. Through multilingual compatible segmentation and dynamic learning and correction of unknown semantic units, this application can accurately identify and understand complex text information that is difficult to handle by traditional methods, overcoming the shortcomings of existing systems in code-switching, new word identification, and handling of disguised vocabulary. Furthermore, by continuously updating behavioral profiles and correcting semantic meanings, this application avoids the performance degradation and high manual maintenance costs caused by data lag in existing systems, achieving rapid adaptation to the evolution of online language. The generated explanatory clues further enhance the credibility and traceability of the classification results, providing a more efficient, accurate, and intelligent text classification and management solution for professional fields such as financial risk control and public opinion monitoring, thereby significantly improving the timeliness of sensitive information identification and the accuracy of classification. Attached Figure Description
[0011] Figure 1 This application provides a flowchart illustrating a method for rapid text content classification and management based on artificial intelligence.
[0012] Figure 2 A flowchart of a text content rapid classification and management system based on artificial intelligence is provided for this application.
[0013] In the diagram: 1. Acquisition module; 2. Behavioral profile construction module; 3. Meaning inference and confidence generation module; 4. Meaning correction and confidence adjustment module; 5. Classification decision and interpretation module. Detailed Implementation
[0014] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0015] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0016] Reference Figure 1 This application proposes a method for rapid classification and management of text content based on artificial intelligence, the method comprising: S1000: Obtain the text to be classified, perform multilingual compatible language unit segmentation on the text to be classified, and match the segmented semantic units with the preset known semantic set to identify known and unknown semantic units in the text to be classified, while capturing the contextual information of unknown semantic units in the text to be classified. S2000: Based on context information, extract the co-occurrence association features, syntactic function features, and textual tendency features of unknown semantic units to construct a behavioral profile of unknown semantic units. When new text containing unknown semantic units is received in the future, update the context information and adjust the behavioral profile according to the updated context information. S3000: Compare the adjusted behavioral profile with the preset set of known semantic behavioral profiles, infer the preliminary meaning and related attributes of the unknown semantic units based on the comparison results, and generate the corresponding confidence scores; S4000: When an unknown semantic unit is detected to reappear in newly added text, the initial meaning and related attributes are corrected based on the adjusted behavioral profile to obtain the meaning and related attributes of the unknown semantic unit, and the confidence level is adjusted simultaneously. S5000: Integrates the meanings and related attributes of known semantic units and unknown semantic units in the text to be classified, and determines the degree of participation of the meanings and related attributes of unknown semantic units in the classification decision based on the adjusted confidence level. It then makes a classification decision for the text to be classified and outputs the classification result. At the same time, it generates explanatory cues, which are used to indicate semantic units whose contribution to the classification result is greater than a preset contribution threshold.
[0017] Here, "semantic unit" refers to the smallest linguistic unit in text with independent or relatively independent meaning, which can be a word, phrase, or even a combination of symbols in a specific context. "Known semantic set" refers to a pre-established database containing identified and labeled vocabulary, phrases, and their related attributes. "Unknown semantic unit" refers to a semantic unit that appears in the text to be classified but is not matched in the known semantic set; these typically represent emerging vocabulary, domain-specific terms, or disguised expressions. "Behavioral profiling" is a comprehensive description of the linguistic behavior characteristics of unknown semantic units in different textual contexts, including their co-occurrence association features, grammatical function features, and textual tendency features. "Confidence" quantifies the reliability of inferring the meaning of unknown semantic units. This method is typically deployed on server clusters or cloud computing platforms with powerful computing capabilities, utilizing deep learning frameworks and natural language processing tools for text analysis and model training.
[0018] The implementation of this application includes the following steps: First, the text to be classified is acquired, and multilingual compatible language unit segmentation is performed on the text to be classified. The segmentation process can adopt rule-based methods, such as segmenting Chinese text based on dictionaries and statistical models, and segmenting English text based on spaces and punctuation marks. Alternatively, machine learning-based methods can be adopted, such as training a multilingual word segmentation model to automatically identify the language in the text and perform corresponding segmentation, for example, using neural network models such as Transformer trained on a large-scale multilingual corpus to process mixed language text. The segmented semantic units are then matched with a preset set of known semantic units to identify known and unknown semantic units in the text to be classified; the matching process can adopt hash lookup or vector similarity-based methods, for example, storing the vector representations of semantic units in the known semantic set, calculating the vector of the segmented semantic unit and comparing it with the known semantic set, and determining that the similarity exceeds a preset threshold as a known semantic unit. At the same time, the context information of the unknown semantic unit in the text to be classified can be captured. The context information can include the N words before and after the unknown semantic unit, the sentence to which it belongs, the paragraph, or even the entire document. For example, all words in the sentence where the unknown semantic unit is located, as well as the content of the two sentences before and after the sentence, can be extracted as its context information.
[0019] Secondly, based on contextual information, co-occurrence association features, grammatical function features, and textual sentiment features of unknown semantic units are extracted to construct behavioral profiles for these units. Co-occurrence association features can be constructed by statistically analyzing the co-occurrence frequency of the unknown semantic unit with surrounding words; grammatical function features can be identified through dependency parsing to determine its grammatical role in the sentence (such as subject, predicate, object, or modifier); and textual sentiment features can be determined as positive, negative, or neutral using a sentiment lexicon or a pre-trained sentiment analysis model. When new text containing unknown semantic units is subsequently received, the contextual information is updated, and the co-occurrence association features, grammatical function features, and textual sentiment features are recalculated or incrementally updated based on the updated contextual information, thereby adjusting the behavioral profiles.
[0020] Next, the adjusted behavioral profile is compared with a pre-defined set of known semantic behavioral profiles. Based on the comparison results, the preliminary meaning and related attributes of the unknown semantic unit are inferred, and a corresponding confidence score is generated. The comparison process can employ metrics such as cosine similarity and Euclidean distance to calculate the similarity between the behavioral profile of the unknown semantic unit and the behavioral profiles of each known semantic unit in the pre-defined set of known semantic behavioral profiles. For example, the N known semantic units with the highest similarity are selected as candidates. Then, the preliminary meaning and related attributes of the unknown semantic unit are inferred based on the meaning and related attributes of the candidate known semantic units. The confidence score can be calculated based on factors such as similarity and candidate consistency. For example, a higher confidence score is assigned when the unknown semantic unit has a significantly higher similarity to a certain candidate than to other candidates.
[0021] Subsequently, when an unknown semantic unit is detected to reappear in newly added text, the initial meaning and related attributes are revised based on the adjusted behavioral profile to obtain the meaning and related attributes of the unknown semantic unit, and the confidence level is adjusted accordingly. For example, if the contextual information in the newly added text better matches the existing inference, the confidence level is increased; if a contradiction occurs, the meaning and related attributes are re-evaluated and revised, and the confidence level is adjusted accordingly.
[0022] Finally, the meanings and related attributes of known semantic units and unknown semantic units in the text to be classified are integrated. Based on the adjusted confidence level, the participation degree of the meanings and related attributes of unknown semantic units in the classification decision is determined. A classification decision is then made for the text to be classified, and the classification result is output. For example, the weight or influence of unknown semantic units in the classification decision can be adjusted with the confidence level: the higher the confidence level, the greater the participation; the lower the confidence level, the smaller the participation. Classification decisions can employ classification models such as Support Vector Machines (SVM), Naive Bayes, and deep neural networks. Simultaneously, explanatory cues are generated. These explanatory cues indicate semantic units whose contribution to the classification result is greater than a preset contribution threshold. For example, identifying several semantic units that have the greatest impact on the final classification result and marking them as explanatory cues helps users understand the classification basis.
[0023] In another embodiment of this application, after extracting the co-occurrence association features, syntactic function features, and textual tendency features of unknown semantic units based on contextual information, the method further includes: S2100: Identify whether rhetorical structures exist in the context information; S2110: When a rhetorical structure is identified, semantic deconstruction of the rhetorical structure is performed to obtain the true textual tendency features of the rhetorical structure. S2120: Based on the real text tendency features, the text tendency features of unknown semantic units are corrected. The corrected text tendency features are used to construct the behavioral profile of unknown semantic units by combining the co-occurrence association features and grammatical function features of unknown semantic units.
[0024] Specifically, identifying the presence of rhetorical structures in contextual information refers to using natural language processing techniques, such as pattern matching, syntactic analysis, semantic role labeling, and deep learning models, to analyze the contextual information of captured unknown semantic units in order to detect whether there are rhetorical devices such as irony, metaphor, hyperbole, and euphemisms. Its purpose is to identify linguistic phenomena where the literal meaning is inconsistent with the actual expressive intent.
[0025] When a rhetorical structure is identified, semantic deconstruction is performed on the rhetorical structure to obtain the true textual tendency characteristics of the rhetorical structure. This refers to conducting deep semantic analysis of the rhetorical structure to reveal the implicit true intention. For example, for ironic statements, the system identifies the negative or critical meaning implied behind the seemingly positive or neutral words. Semantic deconstruction can be achieved by constructing a rhetorical pattern library, using an emotion dictionary, combining common sense knowledge graphs, and training a specialized semantic deconstruction model. Its purpose is to correct the literal tendency bias caused by the rhetorical structure into the true semantic tendency.
[0026] In practical applications, modifying the textual tendency features of unknown semantic units based on real textual tendency features means using the real textual tendency features obtained from semantic deconstruction to adjust or replace the textual tendency features previously extracted from the literal text. For example, if the original extracted textual tendency feature is positive, but semantic deconstruction finds that its real tendency is negative irony, then the textual tendency feature of the unknown semantic unit is modified to negative. The modified textual tendency features are then used to construct a behavioral profile of the unknown semantic unit by combining the co-occurrence association features and grammatical function features of the unknown semantic unit, thereby improving the reliability of the input.
[0027] In another embodiment of this application, S3000 specifically includes: S3100: Compare the adjusted behavior profile with a preset set of known semantic behavior profiles, and identify multiple known semantic units whose similarity to the adjusted behavior profile exceeds a preset similarity threshold; S3200: Check whether there are contradictions in meaning or related attributes among multiple known semantic units; S3300: When a contradiction is identified, analyze the subset of features in the adjusted behavioral profile that makes it similar to different known semantic units; S3400: Generates multidimensional semantic inferences based on feature subsets, which contain multiple preliminary meanings and related attributes of unknown semantic units in different contradictory semantic directions, and assigns initial confidence to the semantic inferences corresponding to each preliminary meaning and related attribute; S3500: When continuously receiving new text containing unknown semantic units, update the adjusted behavioral profile, track changes in feature subsets, and adjust the initial confidence level based on the changes in feature subsets to obtain the confidence level; S3600: When the adjusted behavioral profile continuously presents features with multiple contradictory semantic directions exceeding the preset number of contradictions within a preset time period, maintain multidimensional semantic inference and generate a semantic uncertainty report.
[0028] Specifically, identifying and adjusting multiple known semantic units whose similarity to the behavioral profile exceeds a preset similarity threshold means that the system no longer simply searches for a single best match, but rather identifies all known semantic units that are semantically sufficiently similar to the behavioral profile of the unknown semantic unit. Checking for contradictions in meaning or related attributes among multiple known semantic units means determining a contradiction when these units conflict in meaning or related attributes; for example, if the behavioral profile of an unknown semantic unit is highly similar to both known semantic units representing "positive" emotions and those representing "negative" emotions, then a contradiction exists.
[0029] In practical applications, when a contradiction is identified, the system analyzes the subset of features in the adjusted behavioral profile that leads to similarity with different known semantic units. The aim is to pinpoint the root cause of the contradiction. For example, an unknown semantic unit might be related to "growth" in the financial field due to its co-occurrence association features, but its textual tendency features are related to "risk," thus creating a contradiction. Furthermore, based on the feature subset, a multidimensional semantic inference is generated, containing multiple preliminary meanings and related attributes of the unknown semantic unit in different contradictory semantic directions. An initial confidence level is assigned to each preliminary meaning and related attribute. For example, for the contradiction between "growth" and "risk," the system can generate two preliminary meaning inferences: one leaning towards "positive growth," and the other towards "potential risk," and assign initial confidence levels to each.
[0030] Furthermore, while continuously receiving new text containing unknown semantic units and updating the adjusted behavioral profile, the system tracks changes in feature subsets and adjusts the initial confidence level based on these changes to achieve adaptive correction according to context. For example, if subsequent text further emphasizes the "risk" feature, the confidence level of the "potential risk" inference will be increased. As a preferred implementation, when the adjusted behavioral profile continuously presents features with multiple contradictory semantic directions exceeding a preset number of contradictions within a preset time period, multidimensional semantic inference is maintained, and a semantic uncertainty report is generated to handle unknown semantic units with long-standing semantic ambiguity or polysemy. The system does not force the selection of a single meaning but instead presents its uncertainty through the semantic uncertainty report for human decision-making reference.
[0031] In some preferred embodiments, the following specific example illustrates the situation: In a financial text classification scenario, the system receives a new text containing the unknown semantic unit "gray rhino." First, the behavior profile building module constructs a behavior profile of the "gray rhino" based on contextual information. Then, the meaning inference and confidence generation module compares this behavior profile with a pre-defined set of known semantic behavior profiles.
[0032] During this process, the system identified similarities between the behavioral profile of the "gray rhino" and several known semantic units. For example, its similarity to a known semantic unit representing "significant but predictable risk events" exceeded a preset similarity threshold. Simultaneously, it also showed some similarity to a known semantic unit representing "large mammals in zoos" (possibly due to mentions of zoos or biological backgrounds in some texts). The system then detected a semantic contradiction between these two known semantic units: one is an abstract concept of risk, while the other is a concrete animal entity.
[0033] When a contradiction is identified, the system analyzes the subset of features in the "gray rhino" behavior profile that leads to its similarity to "risk events" (e.g., co-occurrence-related features include "financial crisis," "market volatility," "regulatory gaps," etc.; textual bias features lean towards "negative" and "warning") and the subset of features that leads to its similarity to "animals" (e.g., co-occurrence-related features include "African savanna," "wildlife," "protected area," etc.). Based on these feature subsets, the system generates a multidimensional semantic inference containing two preliminary meanings and related attributes: one inference is "a predictable major risk in the financial field," and the other inference is "a large wild animal." Initial confidence levels are assigned to these two inferences; for example, the initial confidence level for the financial risk inference is 0.7, and the initial confidence level for the animal inference is 0.3.
[0034] Subsequently, the system continues to receive new text containing the term "gray rhino." If the subsequent text is mainly concentrated in financial news and economic analysis reports, and the contextual information of "gray rhino" continuously strengthens its co-occurrence association with words such as "risk" and "crisis," then the system will track the changes in these feature subsets and adjust the confidence level accordingly. This will gradually increase the confidence level of the inference for "predictable major risks in the financial field," for example, to 0.95, while decreasing the confidence level of the inference for "a large wild animal" to 0.05.
[0035] However, if, within a preset timeframe, the system finds that the behavioral profile of the "gray rhino" consistently exhibits contradictory characteristics related to financial risk and animal entities, and the number of contradictions exceeds a preset number (e.g., it frequently appears in certain popular science articles), the system will maintain these two multi-dimensional semantic inferences and generate a semantic uncertainty report. This report will detail how the "gray rhino" in the current corpus simultaneously possesses both "financial risk" and "animal" semantic tendencies, identify the key feature subsets leading to this uncertainty, and the current confidence level for each semantic direction. This provides human decision-makers with comprehensive semantic background information, avoiding misunderstandings arising from a single classification.
[0036] In another embodiment of this application, it is further proposed that after S3600, the following is also included: S3610: Conduct a preliminary assessment of potential risks based on the contradictory semantic direction and the relevant information sources indicated in the semantic uncertainty report; S3611: Select and activate one or more automated risk mitigation processes and / or information collection processes that match potential risks from a pre-defined risk linkage strategy library; S3612: Send an early warning signal. The early warning signal includes the identification information of the unknown semantic unit, the multi-dimensional semantic inference corresponding to the contradictory semantic direction, the semantic uncertainty index, and the suggested follow-up action plan. The semantic uncertainty index is determined by the semantic uncertainty report and is used to characterize the degree of semantic uncertainty. S3613: Initiate the information tracing and verification process based on the key information sources indicated in the semantic uncertainty report.
[0037] Specifically, after identifying semantic uncertainty, the system not only generates a semantic uncertainty report but also assesses the potential risk consequences by combining contradictory semantic directions and relevant information sources. For example, if an unknown semantic unit is simultaneously inferred as both "high-risk investment" and "innovative financial product" in financial text, the system will assess the potential misjudgment risk caused by this contradiction, based on key information sources mentioned in the semantic uncertainty report (such as specific news reports, social media discussions, etc.), such as whether it would lead to incorrect investment advice or regulatory oversight. Potential risks refer to the negative consequences that may be caused by the semantic uncertainty of unknown semantic units, providing a basis for subsequent risk management.
[0038] Furthermore, based on the potential risk type and risk level obtained from the preliminary assessment, the system automatically selects and initiates the corresponding process from the risk linkage strategy library. For example, when the assessment results indicate a high potential risk, the system can activate an automated risk mitigation process, marking text containing the unknown semantic unit as "awaiting manual review," or temporarily restricting its participation in automated decision-making. Simultaneously, it can also activate an information collection process, such as automatically expanding data sources to gather more contextual information about the unknown semantic unit, thereby accelerating the elimination of semantic uncertainty. The risk linkage strategy library contains various preset automated processes to provide a flexible and efficient risk response mechanism.
[0039] In addition, an early warning signal is sent, which includes identification information of the unknown semantic unit, multi-dimensional semantic inference corresponding to the contradictory semantic direction, semantic uncertainty index, and suggested follow-up action plan. The semantic uncertainty index is determined by the semantic uncertainty report and is used to characterize the degree of semantic uncertainty. This means that the system sends key information in the form of an early warning signal to relevant personnel (such as analysts and decision-makers) at the same time as triggering the automated risk mitigation process and / or information collection process. The early warning signal includes at least: identification information of the unknown semantic unit, multi-dimensional semantic inference corresponding to the contradictory semantic direction, semantic uncertainty index determined by the semantic uncertainty report, and suggested follow-up action plan (e.g., "suggesting manual intervention analysis" or "suggesting continuous monitoring"). Among them, the semantic uncertainty index can be used to quantify the degree of semantic uncertainty and provide a risk quantification basis for manual decision-making.
[0040] Meanwhile, the system traces and verifies key sources that cause semantic uncertainty. For example, when a semantic uncertainty report indicates that a specific news source or text within a certain time period is the key source of the contradiction, the system initiates an information tracing and verification process to analyze and verify the reliability, timeliness, and content consistency of the information source, in order to reduce semantic uncertainty from the source and improve the ability to understand unknown semantic units.
[0041] In some preferred embodiments, suppose a new unknown semantic unit, "Green Bond 2.0," appears in a financial news classification system. The system's behavioral profiling module discovers that the behavioral profile of "Green Bond 2.0" is similar to both the known semantic behavioral profiles of "high-yield, high-risk investment" and "sustainable development, low-risk investment" over the past week, and this contradiction persists beyond a preset number of contradictions. At this point, the meaning inference and confidence generation module maintains its multi-dimensional semantic inference of "Green Bond 2.0" and generates a semantic uncertainty report.
[0042] Based on this, the solution in this application will further implement the following steps: First, based on the contradictory semantic directions (high risk and low risk) indicated in the semantic uncertainty report and relevant information sources (e.g., some niche financial forums describe it as high risk, while official media emphasize its sustainability), the system conducts a preliminary assessment of the potential risks of "Green Bonds 2.0". The assessment results show that if it is misclassified, it may lead to investor decision-making errors and even trigger market volatility. Therefore, the potential risk is rated as medium to high risk.
[0043] Secondly, the system selects and activates the automated risk mitigation process and / or information collection process that matches the potential risk from the preset risk linkage strategy library; for example, activating the "information collection process" automatically expands the monitoring scope and frequency of news, reports and social media discussions related to "Green Bonds 2.0"; at the same time, activating the "risk mitigation process" automatically marks all text containing "Green Bonds 2.0" as "awaiting manual review" and suspends its automated classification until the semantics are clear.
[0044] Next, the system sends an early warning signal to the team responsible for financial analysis; the early warning signal includes the identification information of "Green Bond 2.0", its multi-dimensional semantic inference of "high-yield high-risk investment" and "sustainable low-risk investment", semantic uncertainty index (e.g., 0.75), and suggested follow-up action plan (e.g., suggesting that human analysts intervene immediately and refer to additional information collected by the system).
[0045] Finally, based on the key information sources indicated in the semantic uncertainty report (e.g., the report points out that a particular financial blog first proposed "Green Bond 2.0" and assigned it a high-risk attribute), the system initiates an information tracing and verification process to trace the origin, evolution path, and usage of the term in different contexts, in order to understand the reasons for its semantic uncertainty from the source and support the subsequent determination of its meaning and related attributes.
[0046] In another embodiment of this application, after generating explanatory clues, the following steps are further included: S5100: Mark semantic units whose contribution to the classification result is greater than a preset contribution threshold as contributing semantic units, and identify the multidimensional semantic inference and / or semantic uncertainty index corresponding to the contributing semantic units as the first multidimensional semantic inference and / or the first semantic uncertainty index, and when the first semantic uncertainty index exists, obtain the semantic uncertainty report corresponding to the first semantic uncertainty index as the first semantic uncertainty report. S5110: Extract the first semantic inference and its first confidence level that are consistent with the category corresponding to the classification result from the first multidimensional semantic inference; S5120: When the contributing semantic unit is an unknown semantic unit, extract the key feature subset that supports the first semantic inference from the behavioral profile of the contributing semantic unit; S5130: When a first semantic uncertainty indicator is identified, analyze the first semantic uncertainty report to identify the main factors causing the uncertainty and quantify the degree of influence of the main factors on the classification decision; S5140: Based on the role of contributing semantic units in classification decisions, first semantic inference, first confidence level, key feature subsets, main factors and degree of influence, construct hierarchical explanatory cues; S5150: Evaluate the potential impact of hierarchical explanatory cues on human decision-making and obtain the evaluation results; S5160: When the evaluation results are potentially misleading, generate supplementary explanations to guide human decision-making in interpreting the contributing semantic units.
[0047] Specifically, after generating explanatory cues, semantic units whose contribution to the classification result exceeds a preset contribution threshold are first labeled as contributing semantic units. These contributing semantic units are key support for the classification decision. Further, the system identifies the multidimensional semantic inferences and / or semantic uncertainty indicators corresponding to these contributing semantic units, labeling them as first multidimensional semantic inferences and / or first semantic uncertainty indicators, respectively. The first multidimensional semantic inference refers to the multidimensional semantic inference generated by the system based on the behavioral profile of an unknown semantic unit, containing multiple preliminary meanings and related attributes, when there are contradictory meanings. The first semantic uncertainty indicator is determined by a semantic uncertainty report and is used to characterize the degree of semantic uncertainty. When a first semantic uncertainty indicator is identified, the system obtains the corresponding semantic uncertainty report and labels it as a first semantic uncertainty report. This report records in detail the relevant information leading to semantic uncertainty.
[0048] Subsequently, from the first multidimensional semantic inference, the system will extract the first semantic inference and its first confidence level that are consistent with the category corresponding to the classification result. Even if an unknown semantic unit may have multiple potential meanings, the system will focus on the semantic inference that best matches the final classification result and give the corresponding confidence level.
[0049] As a preferred implementation, when a contributing semantic unit is identified as an unknown semantic unit, the system further extracts a subset of key features supporting the first semantic inference from the behavioral profile of the contributing semantic unit. The behavioral profile is constructed based on the co-occurrence association features, grammatical function features, and textual tendency features of the unknown semantic unit. The subset of key features refers to features that play a decisive role in forming a specific semantic inference, such as specific co-occurrence lexical patterns, grammatical structures, or sentiment tendencies.
[0050] Upon identifying the first semantic uncertainty indicator, the system performs an in-depth analysis of the first semantic uncertainty report to identify the main factors leading to the uncertainty. These main factors include insufficient contextual information, high semantic ambiguity, and low frequency of new vocabulary. Simultaneously, the system quantifies the impact of these main factors on the classification decision, for example, by calculating the entropy value of the uncertainty factors on the classification probability distribution.
[0051] Based on this, a hierarchical explanatory cue is constructed according to the role of contributing semantic units in classification decision-making, first semantic inference, first confidence level, key feature subset, main factors, and degree of influence. The hierarchical explanatory cue not only indicates contributing semantic units, but also presents the specific meaning of contributing semantic units, the system's confidence in their meaning, the key behavioral features supporting the meaning, and the source of uncertainty and its specific impact on classification decision-making when uncertainty exists, so as to make the explanation more comprehensive and easier to understand.
[0052] Furthermore, the system evaluates the potential impact of hierarchical explanatory cues on human decision-making to obtain evaluation results; the evaluation is used to determine whether the explanation is clear, accurate and whether there is information that may mislead human decision-makers, and the evaluation can be based on preset explanation quality indicators, including the completeness, consistency and conciseness of the explanation.
[0053] When the evaluation results indicate potential misleading information, the system generates supplementary explanations. These supplementary explanations are used to correct or clarify any misleading information, such as pointing out the limitations of a semantic inference, providing additional background information, or suggesting that human decision-makers focus on specific data points. This guides human decision-makers to interpret contributing semantic units more accurately and avoid making incorrect judgments due to incomplete or ambiguous interpretations.
[0054] The solution proposed in this application, by further integrating multidimensional semantic inference of unknown semantic units, semantic uncertainty indicators and their corresponding semantic uncertainty reports on the basis of generating basic explanatory clues, can provide a more in-depth and comprehensive explanation for classification decisions.
[0055] In some preferred embodiments, it is assumed that a financial institution uses this method to classify news reports for risk. A news report contains a new term, "black swan event," which does not exist in the predefined set of known semantics and is therefore identified as an unknown semantic unit. The system constructs a behavioral profile of the "black swan event" based on contextual information and identifies that it may be related to two contradictory semantic directions: "sudden risk events" and "abnormal market fluctuations," and generates corresponding multi-dimensional semantic inferences and semantic uncertainty reports.
[0056] When a news report is categorized as "high-risk," the system labels "black swan event" as a contributing semantic unit. Subsequently, the system identifies the first multidimensional semantic inference and the first semantic uncertainty index corresponding to "black swan event," and obtains a first semantic uncertainty report. From the first multidimensional semantic inference, the system extracts the first semantic inference of "sudden risk event," which is consistent with the "high-risk" classification result, and assigns its first confidence level, for example, 0.7. Simultaneously, the system extracts a subset of key features supporting the "sudden risk event" inference from the behavioral profile of "black swan events," such as their frequent co-occurrence with words like "crash," "crisis," and "unpredictable."
[0057] Upon identifying the first semantic uncertainty indicator, the system analyzes the first semantic uncertainty report and finds that the main factor causing uncertainty is that "black swan event" is a new term with insufficient historical data, and it is also associated with "abnormal market fluctuations" in some contexts. The system quantifies its impact on classification decisions as moderate.
[0058] Based on the information, the system constructed hierarchical explanatory clues: The first layer: news reports are classified as "high risk", with the main contributing semantic unit being "black swan events".
[0059] The second layer: The initial semantic inference of "black swan event" is "sudden risk event", with a confidence level of 0.7.
[0060] The third layer: Key features supporting this inference include its co-occurrence with words such as "crash" and "crisis".
[0061] The fourth layer contains a moderate degree of semantic uncertainty, mainly due to insufficient historical data and potential semantic ambiguity.
[0062] The system further assesses the potential impact of this hierarchical explanatory cue on human decision-making. If the assessment indicates that, due to the semantic uncertainty of "Black Swan event," human decision-makers may misunderstand its risk level or overlook its potential connection to "abnormal market volatility," the system will generate supplementary explanations. For example, the supplementary explanation might state that although the system infers "Black Swan event" as a sudden risk event, given its novelty and potential connection to "abnormal market volatility," human decision-makers should pay closer attention to real-time market dynamics and make a comprehensive judgment based on other information sources to avoid the risks that may arise from a single interpretation.
[0063] In another embodiment of this application, it is further proposed that, after generating the explanatory clues, the method further includes: S5200: Identify semantic units that are indicated by explanatory cues and whose contribution to the classification results is greater than a preset contribution threshold as contributing semantic units, and determine the semantic evolution cycle of contributing semantic units based on the frequency of occurrence of new texts containing contributing semantic units within a preset historical time period, the rate of change of text tendency features corresponding to contributing semantic units, and / or the difference in behavioral profiles of contributing semantic units at different time points. S5210: During the semantic evolution cycle, continuously receive new text containing contributing semantic units, and update the context information of contributing semantic units based on the new text containing contributing semantic units. S5220: Based on the updated context information, extract the co-occurrence association features, syntactic function features and textual tendency features of contributing semantic units according to the preset time granularity, so as to construct the behavioral profile of contributing semantic units at different time points. S5230: Compare the behavioral profiles of contributing semantic units at different time points to identify the semantic evolution trend and / or magnitude of the contributing semantic units; S5240: When the semantic evolution trend and / or evolution magnitude exceed the preset evolution threshold, adjust the meaning and related attributes of the contributing semantic units, and update the explanatory cues based on the adjusted meaning and related attributes.
[0064] Specifically, contributing semantic units refer to semantic units in text classification decisions whose contribution to the final classification result exceeds a preset contribution threshold. These semantic units are usually the core or key indicators of the text content. The semantic evolution cycle can be determined by comprehensively analyzing the frequency of new texts containing contributing semantic units within a preset historical time period, the rate of change of the text sentiment characteristics corresponding to the contributing semantic units, and the degree of difference between the behavioral profiles constructed by the contributing semantic units at different time points. The role of the semantic evolution cycle is to determine the time window and update rhythm for subsequent actions such as "continuously receiving new texts containing contributing semantic units," "extracting...to construct...behavioral profiles at different time points according to a preset time granularity," and "comparing...behavioral profiles at different time points," thereby matching the monitoring frequency with the semantic change rate of the contributing semantic units. For example, if the frequency of a contributing semantic unit increases significantly in a short period of time, or its text sentiment characteristics (such as sentiment polarity) change drastically, or its behavioral profile differs greatly from the historical profile, it indicates that its semantics are evolving rapidly. In this case, the semantic evolution cycle is set to be shorter for high-frequency monitoring; conversely, if the change is stable, the semantic evolution cycle is set to be longer for routine monitoring.
[0065] During the semantic evolution cycle, the system continuously receives new text containing contributing semantic units and uses the new text to update the context information of the contributing semantic units to ensure that the behavioral profile of the contributing semantic units is constructed based on the latest context.
[0066] In practical applications, based on the updated context information, the system extracts the co-occurrence association features, syntactic function features, and textual tendency features of contributing semantic units according to a preset time granularity (e.g., daily, weekly, or monthly) to construct behavioral profiles of contributing semantic units at different points in time; thus forming a sequence of behavioral profiles to track the dynamic changes of contributing semantic units.
[0067] Furthermore, by comparing the behavioral profiles of contributing semantic units at different points in time, the system identifies their semantic evolution trend (e.g., from positive to negative, or from one domain to another) and / or evolution magnitude (e.g., the degree of semantic drift); the comparison can be achieved by calculating the similarity or distance between behavioral profiles.
[0068] When the identified semantic evolution trend and / or evolution magnitude exceeds the preset evolution threshold, the system determines that the meaning or related attributes of the contributing semantic unit have changed significantly, and adjusts the meaning and related attributes of the contributing semantic unit accordingly to reflect the latest semantic state; at the same time, it updates the explanatory cues based on the adjusted meaning and related attributes to ensure that the explanatory cues can accurately indicate the source of contribution of the current classification result.
[0069] In another embodiment of this application, after determining the semantic evolution period of the contributing semantic unit, the method further includes: S5201: Within a preset sliding time window, when the frequency of occurrence of newly added text contributing semantic units and / or the rate of change of text tendency features corresponding to contributing semantic units exceed the corresponding preset sudden change threshold, the semantic evolution cycle is adjusted to a short-term high-frequency semantic evolution cycle, and new texts containing contributing semantic units are continuously received in a high-frequency manner within the short-term high-frequency semantic evolution cycle. S5202: Within a short-term high-frequency semantic evolution cycle, construct a short-term behavioral profile sequence of contributing semantic units based on newly added text containing contributing semantic units, and calculate the semantic evolution rate and direction of contributing semantic units based on the short-term behavioral profile sequence. S5203: When the semantic evolution rate and direction meet the preset stability conditions, the semantic evolution cycle of the contributing semantic units is adjusted according to the preset stability conditions, and the short-term high-frequency semantic evolution cycle is switched to the adjusted semantic evolution cycle.
[0070] Specifically, the preset sliding time window refers to a dynamically moving time period, such as the most recent 24 hours, 7 days, or 30 days, used to continuously monitor relevant data of contributing semantic units. It does not conflict with the preset historical time period: the preset historical time period is used to comprehensively evaluate and determine the semantic evolution cycle of contributing semantic units in the initialization phase, while the preset sliding time window is used to perform real-time triggered correction of the semantic evolution cycle in the running phase to capture short-term sudden changes.
[0071] When the frequency of occurrence of newly added text contributing to a semantic unit within the sliding time window, such as the amount of newly added text per hour, or the rate of change of its textual tendency features, such as the daily average change of indicators like sentiment polarity and topic preference, exceeds a preset threshold for sudden change, for example, the frequency of occurrence doubles in a short period of time, or the change in textual tendency features exceeds 10%, it indicates that the contributing semantic unit may be undergoing rapid semantic evolution.
[0072] In this context, to capture such changes more promptly, the semantic evolution cycle is adjusted to a short-term, high-frequency semantic evolution cycle. The purpose of the short-term, high-frequency semantic evolution cycle is to shorten the update window of "continuously receiving new text containing contributing semantic units within the semantic evolution cycle" and increase the execution frequency of "extracting...at preset time granularity to construct...behavioral profiles at different time points," thereby obtaining denser behavioral profile sampling and more timely basis for adjusting meaning and related attributes during the rapid semantic evolution phase.
[0073] In this embodiment, the adjusted semantic evolution cycle is the regular cycle. This short-term high-frequency semantic evolution cycle is usually shorter than the regular cycle, for example, shortened from the weekly level to the daily level or even the hourly level. Within this cycle, the system will continuously receive new text containing contributing semantic units in a high-frequency manner, for example, data collection and processing will be performed every few minutes or hours to ensure real-time monitoring of semantic changes.
[0074] Within a short-term, high-frequency semantic evolution cycle, based on the newly received text at high frequency, a short-term behavioral profile sequence of contributing semantic units is constructed; this sequence records the behavioral profiles of contributing semantic units at different points in time within a short period of time, such as hourly or half-day behavioral profile snapshots.
[0075] By analyzing this short-term behavioral profile sequence, the semantic evolution rate and direction of contributing semantic units can be calculated. For example, the speed and direction of semantic vector movement in the semantic space, or the changing trends of its associated vocabulary, grammatical function, and textual tendency features.
[0076] When the semantic evolution rate and direction meet the preset stability conditions, for example, when the movement speed of the semantic vector decreases to below a certain threshold, or when its main features remain relatively stable at multiple consecutive time points, it indicates that the semantic evolution of the contributing semantic unit has become stable. At this time, according to the preset stability conditions, the semantic evolution cycle of the contributing semantic unit will be adjusted again, switching from a short-term high-frequency semantic evolution cycle back to the adjusted regular semantic evolution cycle, so as to restore a more economical and efficient monitoring mode.
[0077] The solution proposed in this application effectively solves the problem of lag in the traditional semantic evolution cycle when faced with rapid and drastic semantic changes by introducing a dynamic response mechanism to sudden semantic changes.
[0078] In another embodiment of this application, after updating the context information of the contribution semantic unit based on the newly added text containing the contribution semantic unit, the following steps are further included: S5211: Based on newly added text containing contributing semantic units received within the semantic evolution cycle, identify multiple influencing factors related to the frequency of occurrence of newly added text containing contributing semantic units and / or the rate of change of text tendency features corresponding to contributing semantic units. The multiple influencing factors include one or more of market sentiment factors, policy guidance factors, industry dynamic factors, and event factors. S5212: For each influencing factor, extract text features corresponding to that influencing factor from the new text containing contributing semantic units received within the semantic evolution cycle. The text features include at least the density of skewed words and / or the frequency of occurrence of preset official terms. S5213: Based on text features, calculate the degree of contribution of each influencing factor to the frequency of occurrence of newly added texts of contributing semantic units and / or the rate of change of text tendency features corresponding to contributing semantic units. S5214: Integrate the contribution levels of various influencing factors to generate a comprehensive evaluation result of the contribution semantic unit. The comprehensive evaluation result is used to characterize the real change trend of the contribution semantic unit under the combined effect of multiple influencing factors and to indicate potential risks.
[0079] Specifically, within the semantic evolution cycle, the system continuously receives new text containing contributing semantic units and updates the context information of these contributing semantic units based on this text. Building upon this, this application further identifies multiple influencing factors related to the frequency of occurrence of new texts contributing semantic units and / or the rate of change in text sentiment characteristics. These multiple influencing factors include one or more of market sentiment factors, policy guidance factors, industry dynamic factors, and event factors. For example, market sentiment factors may refer to the general emotional tendency of the public towards a certain topic; policy guidance factors may refer to policies, regulations, or guidelines issued by the government or regulatory agencies related to the topic; industry dynamic factors may refer to technological developments, changes in the competitive landscape, or business model innovations within a specific industry; and event factors may refer to sudden, significantly influential social, economic, or technological events. The identification of these multiple influencing factors can be performed through a pre-set knowledge base, external data sources (such as news APIs, social media trend analysis), or machine learning-based models.
[0080] For each identified influencing factor, the system extracts text features corresponding to that factor from newly added text containing contributing semantic units received within the semantic evolution cycle. These text features include at least the density of directional words and / or the frequency of occurrence of pre-defined official terms. The density of directional words characterizes the density of words expressing sentiment, attitude, or stance in the text; the frequency of occurrence of pre-defined official terms characterizes the degree of influence of policies or industry norms on semantic units. For example, when a policy-oriented factor is identified, the system extracts the frequency of occurrence and changes of pre-defined official terms related to the policy; when a market sentiment factor is identified, the system extracts the density of words with positive or negative directional tendencies.
[0081] Subsequently, based on the text features, the system calculates the degree of contribution of each influencing factor to the frequency of new text occurrences of the contributing semantic unit and / or the rate of change of the text tendency features corresponding to the contributing semantic unit. The calculation of the degree of contribution can employ correlation analysis, regression models, causal inference models, etc., to quantify the specific impact of each influencing factor on semantic evolution.
[0082] Finally, the system integrates the contribution levels of each influencing factor to generate a comprehensive evaluation result for the contributing semantic unit. This comprehensive evaluation result characterizes the actual change trend of the contributing semantic unit under the combined influence of multiple factors and indicates potential risks. By integrating the contribution levels of multiple factors, the system can reduce the bias caused by a single indicator when the frequency of occurrence and the rate of change of textual tendency features are affected by the interaction or cancellation of multiple factors, and improve the ability to identify actual change trends and potential risks.
[0083] This application's solution overcomes the limitations of relying solely on internal text feature analysis by identifying and quantifying the impact of external influencing factors on the semantic evolution of contributing semantic units. Specifically, by identifying market sentiment factors, policy guidance factors, industry dynamic factors, and event factors, and selectively extracting text features such as the density of trending words and / or the frequency of occurrence of pre-defined official terms, it can capture the driving forces of semantic change from a more macroscopic and multi-dimensional perspective. By calculating the contribution of each influencing factor to the frequency of occurrence of new text and / or the rate of change of text trending features, its specific role in semantic evolution can be quantified, providing a traceable basis for the comprehensive evaluation results. The resulting comprehensive evaluation results not only reflect the "real trend of change" but also indicate potential risks, expanding semantic evolution analysis from "representation of change" to "explanation of driving forces" and "risk indication."
[0084] Through this technical solution, this application provides a deeper and more comprehensive insight into the semantic evolution of contributing semantic units. Based on identifying semantic evolution trends and / or the magnitude of evolution, this solution further reveals the external driving factors that lead to changes in the frequency of new text appearances and / or the rate of change in textual tendency features, thereby improving the accuracy and interpretability of semantic evolution analysis. Especially in risk-sensitive fields such as finance and public opinion monitoring, the comprehensive assessment results can be used to indicate potential risks, supporting more timely risk identification and decision-making, and reducing the possibility of potential losses or missed opportunities.
[0085] In another embodiment of this application, a method is further proposed for integrating the contribution levels of various influencing factors to generate a comprehensive evaluation result of contribution semantic units, which specifically includes: S5214-1: Standardize the contribution of each influencing factor to obtain the standardized contribution. S5214-2: Based on pre-defined empirical rules in the financial field, identify and construct non-linear relationship structures among influencing factors; S5214-3: Transform nonlinear correlation structures into dynamically adjusted weight parameters; S5214-4: Using an iterative optimization method, the comprehensive evaluation result of the contributing semantic unit is calculated based on the occurrence frequency of the new text of the contributing semantic unit and / or the rate of change of the text tendency feature corresponding to the contributing semantic unit, the standardized contribution degree, and the weight parameters. S5214-5: Monitor the deviation between the comprehensive evaluation results and the preset market performance indicators and / or the results of risk events; S5214-6: Adjust the weight parameters based on the deviation to update the overall evaluation results.
[0086] Specifically, standardizing the contribution of each influencing factor refers to unifying the contribution of different influencing factors to the same scale or range, thereby eliminating the impact of differences in dimensions and numerical ranges on subsequent calculations and ensuring the comparability and fairness of each factor in the comprehensive evaluation. For example, methods such as min-max standardization, Z-score standardization, or decimal scaling standardization can be used. The standardized contribution refers to the relative influence of each influencing factor on the change of the semantic unit after unifying its dimensions.
[0087] Furthermore, based on pre-established empirical rules in the financial field, nonlinear correlation structures among influencing factors are identified and constructed. These pre-established empirical rules refer to the knowledge and patterns accumulated in long-term practice within the financial industry regarding the interactions between market behavior, policy effects, industry dynamics, and unforeseen events. For example, the introduction of certain policies may significantly amplify or suppress the impact of market sentiment on specific semantic units. Nonlinear correlation structures refer to relationships between influencing factors that are not simply linear additive but exhibit mutual reinforcement, mutual cancellation, or threshold effects. These nonlinear correlation structures can be identified and modeled using expert knowledge graphs, historical data analysis, or machine learning models (such as neural networks).
[0088] Transforming a nonlinear correlation structure into dynamically adjusted weight parameters means assigning one or a set of weight parameters to each influencing factor based on the nonlinear correlation structure, and enabling these weight parameters to dynamically adjust with changes in the context. For example, when market sentiment factors and policy guidance factors exhibit a strong positive nonlinear correlation, a higher joint weight is assigned to both to characterize their synergistic effect; these weight parameters are used to quantify the importance and influence of each influencing factor in a specific context.
[0089] An iterative optimization approach is employed, calculating the comprehensive evaluation result of contributing semantic units based on the frequency of occurrence of newly added text and / or the rate of change of textual tendency features corresponding to the contributing semantic units, the standardized degree of contribution, and weight parameters. Iterative optimization refers to the process of gradually approaching the optimal solution through repeated calculations and parameter updates, such as using algorithms like gradient descent, genetic algorithms, or particle swarm optimization. In each iteration, the comprehensive evaluation result is calculated based on the latest input and the weight parameters, and the weight parameters are updated by comparing them with actual observation results, thereby improving the evaluation accuracy.
[0090] Monitoring the deviation between the comprehensive assessment results and the preset market performance indicators and / or risk event results refers to comparing the comprehensive assessment results with actual market performance data (e.g., fluctuations in relevant stock prices, changes in industry indices) and / or risk events that have occurred (e.g., credit defaults, market panics) to quantify the accuracy and predictive ability of the comprehensive assessment results; the deviation is the difference between the predicted value and the actual value.
[0091] The system adjusts the weighting parameters based on deviations to update the overall assessment results. This means that when the overall assessment results deviate from actual market performance indicators and / or risk event outcomes, the system automatically adjusts the weighting parameters according to the magnitude and direction of the deviation. For example, if the system underestimates the effect of a certain influencing factor, it increases the weighting parameter of that factor; conversely, it decreases it. This forms a closed-loop feedback mechanism, enabling the overall assessment results to continuously learn and adapt to market changes.
[0092] The solution proposed in this application overcomes the limitations of traditional simple integration methods in dealing with complex and ever-changing influencing factors by introducing standardized processing, identifying nonlinear correlation structures, dynamically adjusting weight parameters, and adopting iterative optimization and feedback correction mechanisms.
[0093] In some preferred embodiments, it is assumed that in the financial field, an emerging technical term, “Green Bond,” is beginning to appear frequently in news reports and social media. The system first identifies key factors influencing its semantic evolution, such as “policy-oriented factors,” “market sentiment factors,” and “industry dynamic factors,” and calculates their initial contribution to the frequency of occurrence and textual shifts in the semantic unit “Green Bond.”
[0094] To more accurately assess the overall trend of "Green Bonds," the system first standardizes these initial contributions to ensure they are comparable on the same scale. Then, based on pre-defined empirical rules in the financial sector (e.g., government support policies for the environmental protection industry typically significantly increase market attention to green financial products), the system identifies a significant non-linear positive correlation between "policy-driven factors" and "market sentiment factors." For example, when policy support reaches a certain threshold, the positive response of market sentiment is amplified.
[0095] The system transforms this non-linear correlation structure into dynamically adjusted weight parameters. For example, during periods when policies explicitly support green finance, the weights associated with "policy guidance factors" and "market sentiment factors" are dynamically increased. Next, the system employs an iterative optimization approach, combining the frequency of new text related to "Green Bonds," the rate of change in textual bias, the standardized contribution level, and these dynamic weight parameters to calculate a comprehensive evaluation result for "Green Bonds." This result indicates the market popularity and risk appetite of "Green Bonds."
[0096] During the calculation process, the system continuously monitors the deviations between the comprehensive assessment results and actual market performance indicators (e.g., green bond issuance volume, related listed company stock price performance) and risk event outcomes (e.g., green bond default events). If a significant deviation is found between the comprehensive assessment results and actual market performance (e.g., the system assesses high interest but actual issuance volume has not increased significantly), the system will automatically adjust the previously set weight parameters based on this deviation, enabling the model to more accurately reflect the actual situation in the next iteration. Through continuous feedback and correction, the system can continuously optimize its assessment of the semantic evolution trend of "green bonds" and more accurately predict its potential market impact and risks.
[0097] Reference Figure 2 In response, this application proposes an artificial intelligence-based rapid text content classification and management system, the system comprising: The acquisition module 1 is used to acquire the text to be classified, perform multilingual compatible language unit segmentation on the text to be classified, and match the segmented semantic units with a preset set of known semantic units to identify known and unknown semantic units in the text to be classified, while capturing the contextual information of unknown semantic units in the text to be classified. Behavioral profiling module 2 is used to extract co-occurrence association features, syntactic function features, and textual tendency features of unknown semantic units based on context information in order to construct behavioral profiling of unknown semantic units. When new text containing unknown semantic units is received in the future, the context information is updated and the behavioral profiling is adjusted according to the updated context information. The meaning inference and confidence generation module 3 is used to compare the adjusted behavioral profile with a preset set of known semantic behavioral profiles, infer the preliminary meaning and related attributes of the unknown semantic units based on the comparison results, and generate the corresponding confidence scores. The meaning correction and confidence adjustment module 4 is used to correct the initial meaning and related attributes based on the adjusted behavioral profile when an unknown semantic unit is detected to reappear in the newly added text, so as to obtain the meaning and related attributes of the unknown semantic unit and adjust the confidence level simultaneously. The classification decision and interpretation module 5 is used to integrate the meaning and related attributes of known semantic units and unknown semantic units in the text to be classified, and to determine the degree of participation of the meaning and related attributes of unknown semantic units in the classification decision based on the adjusted confidence level. It performs classification decision on the text to be classified and outputs the classification result. At the same time, it generates explanatory clues, which are used to indicate semantic units whose contribution to the classification result is greater than a preset contribution threshold.
[0098] The AI-based rapid text content classification management system proposed in this application, through its modular design, achieves dynamic identification of unknown semantic units, construction of behavioral profiles, meaning inference and correction, and adjustment of participation in classification decisions based on confidence levels. This system effectively addresses the challenges posed by the rapid evolution of internet language, emerging vocabulary, and disguised expressions, significantly improving the accuracy and timeliness of text classification and providing interpretable classification results. This reduces the cost of manual intervention and meets the needs of fields such as financial risk control for efficient and accurate information processing.
[0099] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for rapid classification and management of text content based on artificial intelligence, characterized in that, The method includes: The text to be classified is obtained, and the text to be classified is segmented into language units that are compatible with multiple languages. The segmented semantic units are matched with a preset set of known semantic units to identify known and unknown semantic units in the text to be classified, while capturing the contextual information of the unknown semantic units in the text to be classified. Based on the context information, the co-occurrence association features, syntactic function features, and textual tendency features of the unknown semantic unit are extracted to construct a behavioral profile of the unknown semantic unit. When new text containing the unknown semantic unit is subsequently received, the context information is updated, and the behavioral profile is adjusted according to the updated context information. The adjusted behavioral profile is compared with a preset set of known semantic behavioral profiles. Based on the comparison results, the preliminary meaning and related attributes of the unknown semantic unit are inferred, and the corresponding confidence level is generated. When the unknown semantic unit is detected to reappear in the newly added text, the preliminary meaning and related attributes are corrected according to the adjusted behavioral profile to obtain the meaning and related attributes of the unknown semantic unit, and the confidence level is adjusted simultaneously. The text to be classified is integrated with the meanings and related attributes of known semantic units and unknown semantic units. Based on the adjusted confidence level, the degree of participation of the meanings and related attributes of unknown semantic units in the classification decision is determined. The text to be classified is then classified and the classification result is output. At the same time, explanatory clues are generated to indicate semantic units whose contribution to the classification result is greater than a preset contribution threshold.
2. The method for rapid classification and management of text content based on artificial intelligence according to claim 1, characterized in that, Based on the context information, after extracting the co-occurrence association features, syntactic function features, and textual tendency features of the unknown semantic units, the method further includes: Identify whether a rhetorical structure exists in the context information; When the rhetorical structure is identified, the rhetorical structure is semantically deconstructed to obtain the true textual tendency features of the rhetorical structure. Based on the real text tendency features, the text tendency features of the unknown semantic unit are modified. The modified text tendency features are used to construct the behavioral profile of the unknown semantic unit by combining the co-occurrence association features and syntactic function features of the unknown semantic unit.
3. The method for rapid classification and management of text content based on artificial intelligence according to claim 1, characterized in that, The adjusted behavioral profile is compared with a preset set of known semantic behavioral profiles. Based on the comparison results, the preliminary meaning and related attributes of the unknown semantic unit are inferred, and the corresponding confidence level is generated, including: The adjusted behavior profile is compared with a preset set of known semantic behavior profiles to identify multiple known semantic units whose similarity to the adjusted behavior profile exceeds a preset similarity threshold. Check whether there are contradictions in meaning or related attributes among the multiple known semantic units; When the contradiction is identified, the feature subset in the adjusted behavioral profile that causes it to be similar to different known semantic units is analyzed; Based on the feature subset, a multidimensional semantic inference is generated containing multiple preliminary meanings and related attributes of the unknown semantic unit in different contradictory semantic directions, and an initial confidence level is assigned to the semantic inference corresponding to each preliminary meaning and related attribute; While continuously receiving new text containing the unknown semantic units and updating the adjusted behavioral profile, the changes in the feature subset are tracked, and the initial confidence is adjusted according to the changes in the feature subset to obtain the confidence. When the adjusted behavioral profile continuously exhibits features in multiple contradictory semantic directions exceeding a preset number of contradictions within a preset time period, the multidimensional semantic inference is maintained, and a semantic uncertainty report is generated.
4. The method for rapid classification and management of text content based on artificial intelligence according to claim 3, characterized in that, When the adjusted behavioral profile continuously exhibits features in multiple contradictory semantic directions exceeding a preset number of contradictions within a preset time period, after maintaining the multidimensional semantic inference and generating a semantic uncertainty report, the process further includes: Based on the contradictory semantic direction and the relevant information sources indicated in the semantic uncertainty report, a preliminary assessment of the potential risks is conducted; Select and activate one or more automated risk mitigation processes and / or information collection processes that match the potential risks from the preset risk linkage strategy library; Send an early warning signal, which includes the identification information of the unknown semantic unit, the multi-dimensional semantic inference corresponding to the contradictory semantic direction, the semantic uncertainty index, and the suggested follow-up action plan. The semantic uncertainty index is determined by the semantic uncertainty report and is used to characterize the degree of semantic uncertainty. Based on the key information sources indicated in the semantic uncertainty report, the information tracing and verification process is initiated.
5. The method for rapid classification and management of text content based on artificial intelligence according to claim 4, characterized in that, After generating explanatory clues, the following is also included: The semantic units whose contribution to the classification result is greater than a preset contribution threshold are labeled as contributing semantic units, and the multidimensional semantic inference and / or semantic uncertainty index corresponding to the contributing semantic units are identified as the first multidimensional semantic inference and / or the first semantic uncertainty index. When the first semantic uncertainty index exists, the semantic uncertainty report corresponding to the first semantic uncertainty index is obtained as the first semantic uncertainty report. Extract a first semantic inference and its first confidence level that are consistent with the category corresponding to the classification result from the first multidimensional semantic inference; When the contributing semantic unit is an unknown semantic unit, extract the key feature subset that supports the first semantic inference from the behavioral profile of the contributing semantic unit; Upon identifying the first semantic uncertainty index, the first semantic uncertainty report is analyzed to identify the main factors causing the uncertainty, and the degree of influence of the main factors on the classification decision is quantified. Based on the role of the contributing semantic units in the classification decision, the first semantic inference, the first confidence level, the key feature subset, the main factors, and the degree of influence, a hierarchical explanatory clue is constructed. The potential impact of the hierarchical explanatory cues on human decision-making was assessed, and the assessment results were obtained. When the evaluation results are potentially misleading, supplementary explanations are generated to guide human decision-making in interpreting the contribution semantic units.
6. The method for rapid classification and management of text content based on artificial intelligence according to claim 1, characterized in that, After generating the explanatory clues, the method further includes: The semantic units indicated by the explanatory clues that contribute more than a preset contribution threshold to the classification results are identified as contributing semantic units. The semantic evolution cycle of the contributing semantic units is determined based on the frequency of occurrence of new texts containing the contributing semantic units within a preset historical time period, the rate of change of the text tendency features corresponding to the contributing semantic units, and / or the difference in the behavioral profiles of the contributing semantic units at different time points. During the semantic evolution cycle, new text containing the contributing semantic unit is continuously received, and the context information of the contributing semantic unit is updated based on the new text containing the contributing semantic unit. Based on the updated context information, the co-occurrence association features, syntactic function features, and textual tendency features of the contributing semantic units are extracted according to a preset time granularity, so as to construct the behavioral profile of the contributing semantic units at different time points. By comparing the behavioral profiles of the contributing semantic units at different time points, the semantic evolution trend and / or evolution magnitude of the contributing semantic units can be identified. When the semantic evolution trend and / or the magnitude of the evolution exceeds a preset evolution threshold, the meaning and related attributes of the contributing semantic unit are adjusted, and the explanatory clues are updated based on the adjusted meaning and related attributes.
7. The method for rapid classification and management of text content based on artificial intelligence according to claim 6, characterized in that, After determining the semantic evolution period of the contributing semantic unit, the method further includes: Within a preset sliding time window, when the frequency of occurrence of new text in the contributing semantic unit and / or the rate of change of the text tendency feature corresponding to the contributing semantic unit exceed the corresponding preset sudden change threshold, the semantic evolution cycle is adjusted to a short-term high-frequency semantic evolution cycle, and new text containing the contributing semantic unit is continuously received in a high-frequency manner within the short-term high-frequency semantic evolution cycle. Within the short-term high-frequency semantic evolution cycle, a short-term behavioral profile sequence of the contributing semantic unit is constructed based on the newly added text containing the contributing semantic unit, and the semantic evolution rate and direction of the contributing semantic unit are calculated based on the short-term behavioral profile sequence. When the semantic evolution rate and direction meet the preset stability conditions, the semantic evolution cycle of the contributing semantic unit is adjusted according to the preset stability conditions, and the short-term high-frequency semantic evolution cycle is switched to the adjusted semantic evolution cycle.
8. The method for rapid classification and management of text content based on artificial intelligence according to claim 6, characterized in that, After updating the context information of the contribution semantic unit based on the new text containing the contribution semantic unit, the method further includes: Based on the new text containing the contributing semantic unit received within the semantic evolution cycle, identify multiple influencing factors related to the frequency of occurrence of the new text containing the contributing semantic unit and / or the rate of change of the text tendency features corresponding to the contributing semantic unit. The multiple influencing factors include one or more of market sentiment factors, policy guidance factors, industry dynamic factors, and event factors. For each of the aforementioned influencing factors, text features corresponding to the influencing factor are extracted from the newly added text containing the contributing semantic units received within the semantic evolution cycle. The text features include at least the density of skewed words and / or the frequency of occurrence of preset official terms. Based on the text features, calculate the degree of contribution of each of the influencing factors to the frequency of occurrence of new text in the contributing semantic unit and / or the rate of change of the text tendency features corresponding to the contributing semantic unit. By integrating the contribution levels of each of the aforementioned influencing factors, a comprehensive evaluation result of the contribution semantic unit is generated. This comprehensive evaluation result is used to characterize the actual change trend of the contribution semantic unit under the combined effect of multiple influencing factors and to indicate potential risks.
9. The method for rapid classification and management of text content based on artificial intelligence according to claim 8, characterized in that, Integrating the contribution levels of each of the aforementioned influencing factors, a comprehensive evaluation result of the contribution semantic unit is generated, including: The contribution levels of each of the aforementioned influencing factors are standardized to obtain the standardized contribution levels. Based on pre-defined financial field experience rules, the non-linear correlation structure among the influencing factors is identified and constructed; The nonlinear correlation structure is transformed into dynamically adjusted weight parameters; Using an iterative optimization approach, the comprehensive evaluation result of the contributing semantic unit is calculated based on the frequency of occurrence of newly added text in the contributing semantic unit and / or the rate of change of the text tendency features corresponding to the contributing semantic unit, the standardized degree of contribution, and the weight parameters. Monitor the deviation between the comprehensive assessment results and the preset market performance indicators and / or the results of risk events; The weighting parameters are adjusted based on the deviation to update the overall evaluation result.
10. A text content rapid classification and management system based on artificial intelligence, characterized in that, The system includes: The acquisition module is used to acquire the text to be classified, perform multilingual compatible language unit segmentation on the text to be classified, and match the segmented semantic units with a preset set of known semantic units to identify known and unknown semantic units in the text to be classified, while capturing the context information of the unknown semantic units in the text to be classified. The behavior profile building module is used to extract the co-occurrence association features, syntactic function features, and text tendency features of the unknown semantic unit based on the context information, so as to build a behavior profile of the unknown semantic unit. When new text containing the unknown semantic unit is received in the future, the context information is updated and the behavior profile is adjusted according to the updated context information. The meaning inference and confidence generation module is used to compare the adjusted behavior profile with a preset set of known semantic behavior profiles, infer the preliminary meaning and related attributes of the unknown semantic unit based on the comparison result, and generate the corresponding confidence score. The meaning correction and confidence adjustment module is used to correct the preliminary meaning and related attributes based on the adjusted behavioral profile when the unknown semantic unit is detected to reappear in the newly added text, so as to obtain the meaning and related attributes of the unknown semantic unit and adjust the confidence level simultaneously. The classification decision and interpretation module is used to integrate the meanings and related attributes of known semantic units and unknown semantic units in the text to be classified, and to determine the degree of participation of the meanings and related attributes of unknown semantic units in the classification decision based on the adjusted confidence level. The module performs classification decision on the text to be classified and outputs the classification result. At the same time, it generates explanatory clues, which are used to indicate semantic units whose contribution to the classification result is greater than a preset contribution threshold.