Digital feature library construction method for enterprise abnormal behaviors

Through automated analysis and multi-source data processing, a digital feature library of abnormal enterprise behavior is constructed, which solves the problems of manual dependence and insufficient data utilization in existing technologies, realizes the efficiency, intelligence and standardization of the feature library, and improves the accuracy of abnormal behavior identification and regulatory efficiency.

CN120806637APending Publication Date: 2025-10-17BEIJING THUNISOFT INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510946486.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing method of constructing an abnormal behavior feature library for enterprises relies on manual experience, insufficient data utilization, and difficulty in mining abnormal behavior features in massive and complex data. In addition, the feature interpretation is vague and cannot efficiently respond to the update requirements of new regulations or new abnormal behavior patterns.

Method used

By automatically parsing the prescriptive description documents of abnormal enterprise behavior, extracting and structurally defining the characteristic elements of abnormal behavior, combining multi-source data processing and quantitative risk assessment, and building a digital feature library, rapid response and efficient identification of features can be achieved.

Benefits of technology

It has achieved the efficiency, intelligence and standardization of the enterprise abnormal behavior feature library, improved the accuracy of abnormal behavior identification and supervision efficiency, and supported the automated screening of large-scale abnormal behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806637A_ABST
    Figure CN120806637A_ABST
Patent Text Reader

Abstract

The invention discloses a method for constructing a digital feature library of enterprise abnormal behaviors, which comprises the following steps of: analyzing a specified description file of the enterprise abnormal behaviors to determine a digital component set of the digital feature library, the digital constituent element set comprises an abnormal behavior feature element, a digital feature element, a risk level element and a feature interpretation element; extracting enterprise operation mode representation data from the component big data associated with the enterprise, and performing feature engineering processing on the enterprise operation mode representation data to determine abnormal behavior features; extracting enterprise operation condition characterization data from the component big data associated with the enterprise, and performing quantitative conversion processing on the enterprise operation condition characterization data to determine an abnormal behavior digital feature value corresponding to the abnormal behavior feature; and carrying out risk quantile estimation processing on the abnormal behavior digital feature value to determine an abnormal behavior risk level value so as to form a digital feature library of the enterprise abnormal behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, in particular to a method for constructing a digital feature library of enterprise abnormal behavior. BACKGROUND

[0002] In the current rapid development of digital economy, enterprises rely on various components to carry out business activities, generating massive amounts of data. Accurate identification of enterprise abnormal behavior is crucial for maintaining market order and ensuring a fair competitive environment. Building a digital feature library of enterprise abnormal behavior is a key foundation for efficient abnormal behavior identification.

[0003] Currently, existing methods for constructing a feature library of enterprise abnormal behavior mostly use manual rule sorting combined with simple data statistics. This method usually involves industry experts developing judgment rules for abnormal behavior based on experience, extracting a small amount of structured data from enterprise data, calculating feature values through fixed formulas, and then constructing a feature library. This construction process relies on a large amount of manpower, requiring experts to spend a lot of time analyzing legal provisions and industry standards, and manually selecting and processing data. At the same time, in terms of data processing, it is difficult to mine hidden abnormal behavior features in massive complex data due to the use of simple statistical methods. SUMMARY

[0004] To solve the above technical problems, the present application provides a method for constructing a digital feature library of enterprise abnormal behavior to at least alleviate the above technical problems.

[0005] The technical scheme provided by the embodiments of the present application is as follows:

[0006] A method for constructing a digital feature library of enterprise abnormal behavior includes:

[0007] Step 1: Analyze the prescriptive description file of enterprise abnormal behavior to determine the set of digital components of the digital feature library, which includes abnormal behavior feature elements, digital feature elements, risk level elements, and feature interpretation elements.

[0008] Step 2: Extract enterprise operation mode representation data from component big data associated with the enterprise, and perform feature engineering processing on the enterprise operation mode representation data to determine abnormal behavior features.

[0009] Step 3: Extract enterprise operation situation representation data from component big data associated with the enterprise, and perform quantitative conversion processing on the enterprise operation situation representation data to determine abnormal behavior digital feature values corresponding to the abnormal behavior features.

[0010] Step 4: Perform risk quantile estimation processing on the abnormal behavior digital feature values to determine abnormal behavior risk level values.

[0011] Step 5, generating an abnormal behavior feature explanation based on the abnormal behavior feature name, the digital feature value, and the risk level value;

[0012] Step 6, filling the abnormal behavior feature name, the digital feature value, the risk level value, and the abnormal behavior feature explanation into the feature name element, the digital feature element, the risk level element, and the feature explanation element respectively to form a digital feature library of enterprise abnormal behaviors.

[0013] The present application solves the technical bottlenecks of manual dependence, insufficient data utilization, and ambiguous feature explanation in the existing methods through automatic analysis, multi-source data processing, quantitative risk assessment, and structured library construction, and realizes the efficiency, intelligence, and standardization of the construction of the enterprise abnormal behavior feature library, thereby providing technical support for market supervision under the digital economy. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 FIG. 1 is a flowchart of a method for constructing a digital feature library of enterprise abnormal behaviors according to an embodiment of the present application. DETAILED DESCRIPTION

[0015] As shown in FIG. 1, the present application provides a method for constructing a digital feature library of enterprise abnormal behaviors, which includes the following steps: Figure 1

[0016] Step 1, analyzing a prescribed description file of enterprise abnormal behaviors to determine a set of digital constituent elements of the digital feature library, wherein the set of digital constituent elements includes an abnormal behavior feature element, a digital feature element, a risk level element, and a feature explanation element;

[0017] Step 2, extracting enterprise operation mode representation data from component big data associated with the enterprise, and performing feature engineering processing on the enterprise operation mode representation data to determine abnormal behavior features;

[0018] Step 3, extracting enterprise operation situation representation data from component big data associated with the enterprise, and performing quantitative conversion processing on the enterprise operation situation representation data to determine abnormal behavior digital feature values corresponding to the abnormal behavior features;

[0019] Step 4, performing risk quantile estimation processing on the abnormal behavior digital feature values to determine abnormal behavior risk level values;

[0020] Step 5, generating an abnormal behavior feature explanation based on the abnormal behavior feature name, the digital feature value, and the risk level value;

[0021] ​Step 6, fill the abnormal behavior feature name, digitized feature value, risk level value, and abnormal behavior feature explanation into the feature name element, digitized feature element, risk level element, and feature explanation element respectively to form a digitized feature library of enterprise abnormal behavior.

[0022] The technical scheme provided by the present application has the following specific technical benefits:

[0023] The present scheme automatically extracts and structures the abnormal behavior feature element, digitized feature element, and other constituent element sets by analyzing the regulatory description file of enterprise abnormal behavior, reducing the dependence on manual experience. At the same time, through the standardized file analysis process, the updating needs of new regulations or new abnormal behavior patterns can be quickly responded to, ensuring that the feature library is synchronized with the regulatory requirements, and solving the technical problem of lagging rule updates in existing methods.

[0024] The present scheme extracts enterprise operation mode representation data from component big data and performs feature engineering processing, combined with feature engineering techniques such as feature extraction, transformation, selection, etc., to mine abnormal features of enterprise business processes, service calls, and other operation modes from massive unstructured and semi-structured data. Then, the quantitative conversion of operating data is performed to convert text, log, and other multi-source data into quantifiable feature values, solving the problem of insufficient feature mining capability of existing methods and improving the accuracy of abnormal behavior identification.

[0025] The present scheme uses statistical methods to estimate extreme quantile numbers of digitized feature values through risk quantile estimation processing, quantifies the risk level corresponding to different feature values, and avoids the subjectivity of manually setting thresholds. Then, based on the feature name, feature value, and risk level, the feature explanation is generated, automatically associating legal basis, data features, and risk assessment results to form a standardized explanation text, solving the problem of extensive risk assessment and non-standardized explanation in existing methods, and providing a more scientific basis for regulatory decision-making.

[0026] The present scheme fills various features and explanations into standardized elements to form a digitized feature library, realizes the rapid retrieval and calling of features through clear data mapping rules and index creation, and supports large-scale automatic screening of abnormal behaviors through the direct connection of structured feature libraries with AI identification models. The present scheme solves the problem of lack of systematicness of feature libraries in existing methods and difficulty in supporting efficient identification, improving regulatory efficiency and enterprise compliance management capability.

[0027] Optionally, step 1, analyze the regulatory description file of enterprise abnormal behavior to determine the digitized constituent element set of the digitized feature library, including the following steps:

[0028] Step 11, load the enterprise abnormal behavior specification description file through the file reading program, and convert the file content into structured text data;

[0029] Step 12, preprocess the structured text data using a regular expression matching algorithm to generate preprocessed text data;

[0030] Step 13, divide the preprocessed text data into word sequences;

[0031] Step 14, identify the named entities related to enterprise abnormal behavior from the word sequences and construct abnormal behavior feature elements accordingly;

[0032] Step 15, configure digital feature elements, risk level elements, and feature explanation elements corresponding to the abnormal behavior feature elements to form a set of digital constituent elements.

[0033] Preferably, for step 11, an intelligent file reading engine supporting multi-format parsing is constructed, which has built-in parsers for multiple file formats such as PDF, Word, and XML. The engine can automatically identify file types and call corresponding parsing modules. When loading files, stream processing technology is used to read content page by page, and semantic analysis algorithm is used to identify structural elements such as chapter titles, paragraph levels, and list items in the file. Unstructured text is converted into structured text data containing hierarchical relationships and metadata tags. For tables, charts, and other special content in complex documents, OCR technology combined with semantic understanding algorithm is used for recognition, key information is extracted and converted into text form, ensuring that all content is accurately included in the structured data system.

[0034] Preferably, for step 12, an adaptive regular expression engine based on deep learning is built, which automatically learns the feature patterns of enterprise abnormal behavior text through pre-trained models and dynamically generates optimal matching rules. In the preprocessing stage, first use general regular expressions to filter invalid characters (such as HTML tags, special symbols) and redundant information (such as headers and footers, annotations), then configure special matching patterns for legal provisions, professional terms, and other specific content, identify and standardize inconsistent expression methods (such as "suspected of violating the law" and "suspected of violating the rules" are unified as "suspected of violating the law and rules"). At the same time, use context awareness technology to handle fuzzy matching scenarios, improve matching accuracy by analyzing the context relationship of words, and finally generate preprocessed text data with uniform format and clear semantics.

[0035] Preferably, for step 13, a word segmentation component based on the fusion of multimodal information is used, which combines dictionary word segmentation, statistical word segmentation, and deep learning methods to construct a three-level word segmentation architecture. First, preliminary word segmentation is performed based on the dictionary of abnormal corporate behavior to identify common terms and fixed collocations; then, a bidirectional LSTM model is used to analyze the context of the text and disambiguate ambiguous words; finally, a graph neural network is used to capture the semantic associations between words and optimize the word segmentation boundary judgment. For special words such as legal professional terms and industry abbreviations, a knowledge graph is introduced to assist in word segmentation, and entity linking technology is used to ensure accurate segmentation of professional terms to form a word sequence that conforms to business logic.

[0036] Preferably, for step 14, a named entity recognition component based on multimodal knowledge fusion is constructed, which integrates the BERT pre-trained model, the legal knowledge graph, and the industry case library to form a three-dimensional recognition framework. First, the deep semantic features of words are extracted through the BERT model to capture the implicit information in the text; then, the knowledge graph is used for entity linking, mapping the identified words with standard terms in legal provisions and regulatory rules; finally, combined with the annotated data in the historical case library, the recognition model is optimized through transfer learning to improve the recognition ability of new abnormal behaviors. During the recognition process, entity nesting recognition technology is used to process complex expressions. For example, from "financing trade through fictitious trade background", the two entities "fictitious trade background" and "financing trade" are simultaneously identified, and finally a hierarchical and clearly related abnormal behavior feature element system is constructed.

[0037] Preferably, for step 15, a meta-learning-based automatic factor configuration engine is configured. This engine analyzes the semantic features of abnormal behavior characteristic factors and dynamically matches the most appropriate digital features, risk levels, and feature interpretation templates from a predefined feature template library. For digital characteristic factors, feature engineering techniques are used to automatically extract quantifiable indicators (such as transaction frequency and capital flow deviation) and configure data collection interface specifications. Risk level factor configuration utilizes a fuzzy comprehensive evaluation model, combining historical case data and expert experience to automatically set risk thresholds and classification rules. Feature interpretation factors are generated by searching legal text and case databases to generate standardized interpretation texts containing legal basis, behavioral manifestations, and typical cases. Finally, knowledge graph technology is used to semantically associate the four types of factors, forming a set of digital component elements with reasoning capabilities, providing a complete semantic framework for subsequent feature library construction.

[0038] Optionally, the method according to claim 1 is characterized in that step 2, extracting enterprise operation mode characterization data from component big data associated with the enterprise, and performing feature engineering processing on the enterprise operation mode characterization data to determine abnormal behavior characteristics, specifically comprises the following steps:

[0039] Step 21, based on the preset enterprise operation mode data filtering rules, retrieve the data table related to enterprise business process and service call operation mode from the component big data database through database query statement, to extract the original enterprise operation mode characterization data from it;

[0040] Step 22, structure the original enterprise operation mode characterization data to form structured enterprise operation mode characterization data;

[0041] Step 23, principal component analysis of structured enterprise operation mode characterization data, extract principal component features, to determine the abnormal behavior features related to enterprise abnormal behavior.

[0042] Preferably, for step 21, a dynamic adaptive data filtering component is constructed, which uses reinforcement learning algorithm to continuously optimize filtering rules. The preset rules cover the key nodes of enterprise business process (such as order generation, payment settlement, logistics distribution, etc.) and the core indicators of service call operation mode (such as interface call frequency, response time, error rate, etc.). When executing database query, the component uses intelligent index optimization technology to automatically analyze data table structure and query condition, dynamically generates the optimal SQL query statement. For distributed database environment, introduce federal query optimizer, efficiently coordinate the data extraction of multiple data sources, ensure that the data table related to enterprise operation mode can be quickly and accurately retrieved from massive component big data, and the original enterprise operation mode characterization data is extracted, at the same time, the data is preliminarily checked for integrity and de-duplication.

[0043] Preferably, for step 22, an intelligent heterogeneous data structuring engine is constructed, which integrates natural language processing and knowledge graph technology. For text type original data, such as business process description document, use named entity recognition and relationship extraction algorithm to convert free text into structured data with semantic relationship; for unstructured log data, such as service call log, through pattern matching and clustering analysis, identify fixed data format template, automatically fill in to form structured record. For numerical and time series data, use data type inference algorithm to automatically identify data attributes and add corresponding metadata tags. At the same time, use the enterprise business domain knowledge in knowledge graph for semantic verification of the structuring process, to ensure that the data structure conforms to the enterprise business logic, finally form standardized, unified and semantically clear structured enterprise operation mode characterization data.

[0044] Preferably, for step 23, a deep learning-based principal component analysis enhancement framework is constructed, which is based on traditional principal component analysis (PCA) and combines autoencoder and generative adversarial network (GAN) techniques. First, the autoencoder is used to reduce the dimension of the structured data, and by learning the compressed representation of the data, the most representative feature dimension is extracted. Then the discriminator of GAN is introduced to evaluate the principal component features and judge their ability to distinguish normal and abnormal behaviors. Through adversarial training, the principal component features generated by the autoencoder are continuously optimized so that they can better capture abnormal behavior signals in the enterprise operation mode. At the same time, an attention mechanism is used to assign weights to different principal component features according to their relevance to historical abnormal behavior cases, highlighting the key features that contribute to abnormal behavior identification, and finally accurately determining the abnormal behavior features closely related to the enterprise abnormal behavior to provide core data support for subsequent abnormal behavior analysis.

[0045] Optionally, step 3, extracting enterprise operation situation representation data from component big data associated with the enterprise, and quantitatively converting the enterprise operation situation representation data to determine the abnormal behavior digital feature value corresponding to the abnormal behavior feature, specifically including the following steps:

[0046] Step 31, according to the pre-set enterprise operation data screening condition, retrieving the operation data table related to enterprise revenue, cost, profit, market share, and transaction size from the database of component big data through database operation language to extract the original enterprise operation situation representation data therefrom;

[0047] Step 32, preprocessing the original enterprise operation situation representation data to obtain preprocessed enterprise operation situation representation data;

[0048] Step 33, based on the constructed quantitative conversion library, quantitatively converting the preprocessed enterprise operation situation representation data to generate abnormal behavior digital feature values.

[0049] Preferably, for step 31, an intelligent operation data retrieval component is constructed, which can automatically convert the screening conditions described in natural language by business personnel into efficient database query statements based on natural language processing technology. The component has an enterprise operation data semantic model built-in, which can understand the business meaning and associated relationship of operation indicators such as revenue, cost, and profit, and automatically associate relevant data tables during retrieval. For distributed and heterogeneous component big data environments, the component uses a federated query optimization strategy to intelligently schedule query requests to the optimal data source node and improves query efficiency through a data index preheating mechanism. At the same time, the component has real-time data sensing capability, which can dynamically identify newly added operation data tables and update the retrieval range to ensure that the original representation data related to the enterprise operation situation can be extracted comprehensively and accurately.

[0050] Preferably, for step 32, an adaptive business data preprocessing engine is configured, which integrates multidimensional data cleaning and standardization techniques. First, through an anomaly detection algorithm, outliers, missing values and incorrect data in the original data are identified based on statistical methods and machine learning models, and intelligent repair strategies are adopted for different types of data problems (such as Kalman filter interpolation for time series data, and business rule-based filling for classification data). Then, data standardization processing is performed, and dynamic feature scaling technology is used to automatically select appropriate standardization methods (such as Z-score standardization, Min-Max standardization) according to the business characteristics of different business indicators. At the same time, the engine has a built-in business data quality evaluation system, which sets up multiple evaluation indicators such as data integrity, accuracy and consistency to score the quality of the preprocessed data, ensuring that the data quality meets the quantitative conversion requirements.

[0051] Preferably, for step 33, a knowledge-driven intelligent quantitative conversion component is created, which semantically associates enterprise abnormal behavior characteristics with business data indicators. The quantitative conversion library adopts a hierarchical structure configuration, including a basic conversion rule layer, an industry characteristic adaptation layer and an abnormal behavior mapping layer. The basic conversion rule layer provides general business data quantitative methods (such as converting the market share described in text into a numerical percentage); the industry characteristic adaptation layer customizes conversion strategies for different industry characteristics (such as the cost structure of manufacturing industry and the user growth pattern of internet companies); the abnormal behavior mapping layer establishes a mapping relationship between business indicators and abnormal behavior characteristics based on historical cases and expert knowledge (such as the correlation rules between revenue fluctuation amplitude and abnormal transactions). The component uses a deep learning model to dynamically optimize the conversion rules, continuously learns new abnormal behavior patterns, and continuously improves the accuracy and relevance of quantitative conversion, ultimately generating digital feature values that accurately reflect enterprise abnormal behavior characteristics.

[0052] Optionally, step 4, the abnormal behavior digital feature value is subjected to risk quantile estimation processing to determine the abnormal behavior risk level value, which specifically includes the following steps:

[0053] Step 41, the abnormal behavior digital feature value is subjected to multiple sampling simulations, the quantile of the sample is calculated after each sampling, a quantile distribution data set is constructed, and the distribution of the quantile is estimated;

[0054] Step 42, the extreme quantile in the abnormal behavior digital feature value is estimated by a generalized extreme value distribution model to optimize the distribution of the estimated quantile;

[0055] Step 43, the optimized distribution of the estimated quantile is compared with the pre-set safe, low, medium and high four-grade risk level standards, and quantile regression is performed on the comparison result to determine the abnormal behavior risk level value.

[0056] Preferably, for step 41, an adaptive Monte Carlo sampling component is constructed, which employs a sequential Monte Carlo algorithm (SMC) combined with importance sampling techniques to dynamically sample the abnormal behavior digital feature values. The component pre-analyzes the probability distribution of the feature values through kernel density estimation (KDE), automatically identifies the thick-tailed and high-variability regions in the data, and increases the sampling density in these regions. After each sampling, the quantile of the sample is efficiently calculated using a fast quantile calculation algorithm (such as the T-Digest algorithm), which significantly reduces the computational complexity while maintaining accuracy through data compression techniques. As the number of samplings increases, the component continuously updates the distribution estimate of the quantile using an online expectation-maximization (EM) algorithm, constructing a dynamically expanding quantile distribution dataset. Through the Bayesian inference framework, the historical sampling results are integrated to achieve progressive learning and optimization of the quantile distribution, providing a reliable foundation for subsequent extreme quantile estimation.

[0057] Preferably, for step 42, a deep learning-based generalized extreme value distribution (GEV) fusion model is configured, which combines traditional statistical theory with neural network technology. First, the moment estimation method is used to preliminarily fit the parameters of the abnormal behavior digital feature values, determining the location parameter, scale parameter, and shape parameter of the GEV model. Then, a deep residual network (ResNet) is introduced to modify the GEV model, with the input being the quantile distribution dataset generated in step 41 and the output being the optimized distribution parameters. During model training, an adversarial training mechanism is adopted, with the discriminator of the generative adversarial network (GAN) evaluating the fitting effect of the GEV model on the extreme quantile, guiding the generator to adjust the parameters to improve the ability to capture tail data. At the same time, the variational inference technique is used to quantify the parameter uncertainty, constructing a robust extreme quantile estimate. This method breaks through the dependence of traditional GEV models on the independent and identically distributed assumption, and can more accurately depict the extreme risk in the abnormal behavior digital feature values.

[0058] Preferably, for step 43, a multi-scale quantile regression decision component is constructed. This component employs a two-layer regression architecture to achieve accurate risk classification. The first layer is a preliminary classifier based on a quantile regression forest (QRF). This uses the optimized quantile distribution as input features and constructs multiple regression trees using the random forest algorithm to non-parametrically estimate the boundaries of each risk level. The second layer is a deep quantile regression network (DQRN). This network employs an attention mechanism to weightedly fuse the QRF output, focusing on the impact of the tail region of the distribution on risk classification. The component incorporates an adaptive threshold adjustment mechanism that dynamically calibrates the quantile thresholds for the four risk levels using a historical case library and expert knowledge. During the comparison process, the Wasserstein distance between the quantile distribution and each risk level standard is calculated using optimal transport theory. The resulting distance metric serves as the input to the regression model, ultimately outputting a probabilistic risk level determination. By incorporating the principle of structural risk minimization, this component effectively balances model complexity and generalization, ensuring the accuracy and stability of risk classification.

[0059] Optionally, step 5, generating an abnormal behavior feature explanation based on the abnormal behavior feature name, the digital feature value, and the risk level value, specifically includes the following steps:

[0060] Step 51: Based on the constructed pattern sequence generator, the abnormal behavior feature name, digital feature value, and risk level value are encoded to generate a behavior pattern sequence;

[0061] Step 52: Based on the constructed behavior feature interpreter, the behavior pattern sequence is mapped to the constructed interpretation template to generate an abnormal behavior feature interpretation.

[0062] The method according to claim 6 is characterized in that step 51, based on the constructed pattern sequence generator, encoding the abnormal behavior feature name, digital feature value, and risk level value to generate a behavior pattern sequence, specifically comprises the following steps:

[0063] Step 511: vectorize the abnormal behavior feature name, digital feature value, and risk level value to generate a multimodal semantic enhancement vector;

[0064] Step 512: Perform temporal modeling and sequence generation processing on the multimodal semantic enhancement vector to generate a behavior pattern sequence.

[0065] Preferably, for step 511, a cross-modal knowledge fusion vectorization engine is constructed, which realizes multi-modal semantic enhancement using a three-layer encoding architecture. The first layer uses a BERT model pre-trained based on a legal corpus for text encoding of abnormal behavior feature names, extracts key legal terms and business concepts in the feature names through an attention mechanism, and generates high-dimensional semantic vectors. The second layer uses a variational autoencoder (VAE) to model the probability distribution of digital feature values, maps the numerical values to the latent space while preserving their probability distribution characteristics, and enhances the representation of abnormal values. The third layer configures a symbol encoder based on a knowledge graph to map the four risk levels to a knowledge network composed of legal provisions, regulatory requirements and industry standards, generating structured semantic vectors. Finally, a tensor fusion network is used to nonlinearly fuse the encoding results of the three layers, dynamically adjust the weights of each modality using a gating mechanism, and form a multi-modal semantic enhancement vector containing text semantics, numerical distribution and domain knowledge.

[0066] Preferably, for step 512, a spatio-temporal attention sequence generation network is configured, which combines the advantages of the Transformer architecture and the long short-term memory network (LSTM). First, a spatio-temporal attention mechanism is used to capture the temporal dependence between multi-modal vectors. In the time dimension, the evolution law of different features in the business process is learned through self-attention mechanism, and in the spatial dimension, the semantic association between features is mined through cross-modal attention mechanism. Then, a bidirectional LSTM is used to model the time sequence of the attention-weighted vector sequence, and extract the potential behavior patterns hidden in the data. To ensure the logic and coherence of the generated sequence, a sequence generation strategy based on Monte Carlo tree search is introduced. At each step of generation, multiple possible sequence extension paths are evaluated by simulation and deduction, and the optimal path is selected to expand the sequence. At the same time, through an adversarial training mechanism, the discriminator is used to evaluate the similarity between the generated sequence and the real behavior pattern, and the parameters of the generator are constantly optimized, finally generating a behavior pattern sequence that conforms to the business logic.

[0067] Optionally, step 52, based on the constructed behavior feature interpreter, maps the behavior pattern sequence to the constructed explanation template to generate an abnormal behavior feature explanation, including the following steps:

[0068] Step 521, convert the behavior pattern sequence to graph structure data based on the GNN semantic parsing module;

[0069] Step 522, perform legal and industry knowledge enhancement processing on the graph structure data to generate semantic enhanced graph structure data;

[0070] Step 523, map the semantic enhanced graph structure data to the constructed explanation template to generate an abnormal behavior feature explanation.

[0071] Preferably, for step 521, a spatio-temporal evolution graph neural network (STEGNN) is constructed, which converts the behavior pattern sequence into a semantic correlation graph using a dynamic graph construction strategy. First, the self-attention mechanism is used to capture the dependency between different time step features in the sequence, and a semantic similarity matrix between the features is calculated. Then, a dynamic adjacency graph is constructed based on the similarity matrix, with nodes representing feature vectors in the sequence and edge weights representing the correlation strength between features. To preserve the temporal information of the sequence, a time embedding mechanism is introduced to add a timestamp encoding to each node. In the process of graph convolution, a spatio-temporal message passing algorithm is used to simultaneously aggregate information in the spatial dimension (feature correlation) and the temporal dimension (evolution relationship). Through the multi-head attention mechanism, the network can focus on different semantic substructures in the sequence to generate multi-scale graph representations. Finally, the behavior pattern sequence is converted into structured graph data containing node features, edge weights, and global graph features, providing a good representation basis for subsequent knowledge enhancement.

[0072] Preferably, for step 52, a knowledge graph guided graph enhancement network (KGGN) is configured, which deeply integrates pre-constructed legal knowledge graphs and industry knowledge graphs with graph structure data. First, an entity alignment module is configured, and a knowledge embedding algorithm based on TransE is used to map the nodes in the graph structure data to the entities in the knowledge graph. Then, through the relationship reasoning module, the triple information (entity-relation-entity) in the knowledge graph is used to add new edge connections and attributes to the graph structure data. For example, if the knowledge graph contains the relationship "fake invoice - legal consequences - administrative punishment", and the graph structure data contains the "abnormal invoice feature" node, the knowledge transfer is used to add legal consequence-related semantic information to this node. To avoid knowledge noise, a confidence propagation mechanism is introduced to assign weights to different knowledge based on the reliability of the knowledge source. At the same time, a graph attention mechanism is used to enable the network to selectively focus on the most relevant knowledge to the current abnormal behavior. The final semantic enhanced graph structure data not only retains the feature information of the original sequence, but also incorporates prior knowledge in the legal and industry domains.

[0073] Preferably, for step 523, a multi-modal graph-to-text generation component (MG2T) is constructed, which converts graph structure data into natural language explanations using a hierarchical mapping strategy. First, a graph parser is configured to extract key subgraphs from the graph structure data through a graph attention mechanism, identifying the core features of abnormal behavior and relevant legal basis. Then a template matching engine is constructed, which uses an optimal matching algorithm based on dynamic programming to match the extracted key subgraphs with a pre-defined explanation template library. Each template in the template library corresponds to a pattern of abnormal behavior, containing a structured explanation framework and variable parameter slots. In the matching process, semantic similarity calculations in the knowledge graph are used to ensure that the most suitable explanation template for the current abnormal behavior is found. Finally, a parameter filling module is used to map the node features and edge relationships in the graph structure data into the parameter slots of the template, generating natural language explanations. To improve the readability and professionalism of the explanations, a legal text generation model is introduced to polish the initially generated explanations, ensuring accurate terminology and clear logic. Throughout the process, reinforcement learning is used to optimize the mapping strategy, with a reward function guiding the component to generate more accurate and comprehensive explanations of abnormal behavior characteristics.

[0074] Optionally, step 6, the abnormal behavior characteristic name, the digital characteristic value, the risk level value, and the abnormal behavior characteristic explanation are filled into the characteristic name element, the digital characteristic element, the risk level element, and the characteristic explanation element, respectively, to form a digital characteristic library of enterprise abnormal behavior, specifically including the following steps:

[0075] Step 61, based on the constructed data mapping rules, the corresponding relationships between the abnormal behavior characteristic name and the characteristic name element, the digital characteristic value and the digital characteristic element, the risk level value and the risk level element, and the abnormal behavior characteristic explanation and the characteristic explanation element are determined.

[0076] Step 62, based on the corresponding relationship, the abnormal behavior characteristic name, the digital characteristic value, the risk level value, and the abnormal behavior characteristic explanation are filled into the characteristic name element, the digital characteristic element, the risk level element, and the characteristic explanation element, respectively, through a database operation statement, and a retrieval index is created to form a digital characteristic library of enterprise abnormal behavior.

[0077] Preferably, for step 61, a dynamic semantic alignment engine is constructed, which implements the automatic generation and optimization of data mapping rules using a knowledge-guided multi-modal matching algorithm. First, a pre-trained legal language model (such as LegalBERT) is used to perform semantic analysis on the abnormal behavior feature names, extracting legal entities, action verbs, and modification relationships. At the same time, the feature name elements are modeled using an ontology, constructing a knowledge graph containing business terms, legal concepts, and data structures. Through a graph embedding-based entity alignment algorithm (such as TransH), the semantic similarity between feature names and elements is calculated, establishing a preliminary mapping relationship. For digitized feature values and digitized feature elements, a type-aware mapping network is developed, combining statistical feature analysis and data distribution fitting techniques to automatically identify the corresponding relationships between value types, value ranges, and business meanings. For risk level values and risk level elements, a decision tree-based mapping rule generator is configured, generating multi-condition matching rules based on the definition thresholds and business logic of risk levels. For abnormal behavior feature explanations and feature explanation elements, a text structure analysis algorithm is used to identify key arguments, evidence chains, and conclusions in the explanations, and perform semantic alignment with the preset fields of the elements. Throughout the process, mapping weights are dynamically adjusted through reinforcement learning, and mapping rules are continuously optimized based on actual application feedback.

[0078] Preferably, for step 62, an adaptive data injection and index optimization component is configured, which implements efficient data filling and index construction using a batch parallel processing architecture. First, a transaction consistency guarantee layer is configured to ensure atomicity and integrity of data writing based on the two-phase commit protocol, preventing partial data loss or inconsistency. For structured digitized feature values and risk level values, precompiled SQL insert statements are used for batch writing, and the database's vectorized execution engine is used to improve writing performance. For unstructured abnormal behavior feature explanations, a text storage module based on word segmentation indexing is constructed, and an inverted index technology is used to achieve fast text retrieval. During data filling, data quality is monitored in real time, and mapping results are checked for legality through pre-set verification rules, and data that does not meet the requirements is automatically repaired or marked. In the index creation stage, a multi-dimensional index optimizer is developed to automatically select the optimal index type (such as B-tree index, hash index, spatial index) based on query pattern prediction analysis. For feature name elements and feature explanation elements, a semantic index based on word vectors is constructed, and an approximate nearest neighbor search (ANN) algorithm is used to speed up semantic queries. To improve the spatial efficiency and query performance of the index, an incremental index merging strategy is introduced to periodically reorganize and optimize the index. The entire component uses an adaptive load balancing mechanism to dynamically adjust resource allocation for data processing and index construction, ensuring efficient and stable operation in large-scale data scenarios.

[0079] The above descriptions are only the preferred embodiments of the present application, and are not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for constructing a digital feature library of abnormal enterprise behavior, characterized in that: include: Step 1: Parse the prescriptive description file of the enterprise's abnormal behavior to determine a digital component set of a digital feature library, wherein the digital component set includes abnormal behavior feature elements, digital feature elements, risk level elements, and feature explanation elements; Step 2: extract enterprise operation mode representation data from the component big data associated with the enterprise, and perform feature engineering on the enterprise operation mode representation data to determine abnormal behavior characteristics; Step 3: extracting enterprise operating status representation data from the component big data associated with the enterprise, and performing quantitative conversion processing on the enterprise operating status representation data to determine the abnormal behavior digital feature value corresponding to the abnormal behavior feature; Step 4: Perform risk quantile estimation on the digital characteristic value of abnormal behavior to determine the abnormal behavior risk level value; Step 5: Generate an abnormal behavior feature explanation based on the abnormal behavior feature name, digital feature value, and risk level value; Step 6: Fill the abnormal behavior feature name, digital feature value, risk level value, and abnormal behavior feature explanation into the feature name element, digital feature element, risk level element, and feature explanation element respectively to form a digital feature library of the enterprise's abnormal behavior.

2. The method according to claim 1, characterized in that Step 1: Parse the prescriptive description file of the enterprise's abnormal behavior to determine the digital component set of the digital feature library, which specifically includes the following steps: Step 11: Load the prescriptive description file of the enterprise's abnormal behavior through a file reading program and convert the file content into structured text data; Step 12: Preprocess the structured text data using a regular expression matching algorithm to generate preprocessed text data; Step 13: Segment the preprocessed text data into word sequences; Step 14: Identify named entities related to the enterprise's abnormal behavior from the word sequence and construct abnormal behavior feature elements based on them; Step 15: Configure digital feature elements, risk level elements, and feature explanation elements corresponding to the abnormal behavior feature elements to form a digital component element set.

3. The method according to claim 1, characterized in that The method according to claim 1 is characterized in that step 2, extracting enterprise operation mode characterization data from component big data associated with the enterprise, and performing feature engineering processing on the enterprise operation mode characterization data to determine abnormal behavior characteristics, specifically comprises the following steps: Step 21: Based on the preset enterprise operation mode data screening rules, a database query statement is used to retrieve data tables related to the enterprise business process and service call operation mode from the component big data database to extract the original enterprise operation mode representation data; Step 22: Structuring the original enterprise operation mode representation data to form structured enterprise operation mode representation data; Step 23: Perform principal component analysis on the structured enterprise operation mode representation data to extract principal component features to determine abnormal behavior features related to abnormal enterprise behavior.

4. The method according to claim 1, wherein Step 3: Extract enterprise operating status representation data from the component big data associated with the enterprise, and perform quantitative conversion processing on the enterprise operating status representation data to determine the abnormal behavior digital feature value corresponding to the abnormal behavior feature, specifically including the following steps: Step 31: Based on pre-set enterprise business data screening conditions, using the database operation language, retrieve business data tables related to enterprise revenue, cost, profit, market share, and transaction scale from the component big data database to extract original enterprise business performance representative data; Step 32: pre-processing the original enterprise operating condition representation data to obtain pre-processed enterprise operating condition representation data; Step 33: Based on the constructed quantitative conversion library, perform quantitative conversion on the pre-processed enterprise operating condition representation data to generate digital characteristic values ​​of abnormal behavior.

5. The method according to claim 1, wherein Step 4: Perform risk quantile estimation on the digital feature value of abnormal behavior to determine the abnormal behavior risk level value, which specifically includes the following steps: Step 41: Perform multiple sampling simulations on the digital characteristic values ​​of abnormal behavior, calculate the quantile of the sample after each sampling, construct a quantile distribution data set, and then estimate the distribution of the quantiles; Step 42: Estimate the extreme quantiles in the digital characteristic values ​​of the abnormal behavior using a generalized extreme value distribution model to optimize the distribution of the estimated quantiles; Step 43: Compare the distribution of the optimized estimated quantiles with the pre-set four risk level standards of safe, low, medium, and high, and perform quantile regression on the comparison results to determine the abnormal behavior risk level value.

6. The method according to claim 1, characterized in that Step 5: Generate an abnormal behavior feature explanation based on the abnormal behavior feature name, digital feature value, and risk level value, specifically including the following steps: Step 51: Based on the constructed pattern sequence generator, the abnormal behavior feature name, digital feature value, and risk level value are encoded to generate a behavior pattern sequence; Step 52: Based on the constructed behavior feature interpreter, the behavior pattern sequence is mapped to the constructed interpretation template to generate an abnormal behavior feature interpretation.

7. The method according to claim 6, characterized in that Step 51: Based on the constructed pattern sequence generator, the abnormal behavior feature name, digital feature value, and risk level value are encoded to generate a behavior pattern sequence, which specifically includes the following steps: Step 511: vectorize the abnormal behavior feature name, digital feature value, and risk level value to generate a multimodal semantic enhancement vector; Step 512: Perform temporal modeling and sequence generation processing on the multimodal semantic enhancement vector to generate a behavior pattern sequence.

8. The method according to claim 6, characterized in that Step 52: Based on the constructed behavior feature interpreter, the behavior pattern sequence is mapped to the constructed interpretation template to generate an abnormal behavior feature interpretation, which specifically includes the following steps: Step 521: Based on the GNN semantic parsing module, convert the behavior pattern sequence into graph structure data; Step 522: Enhance the graph structure data with legal and industry knowledge to generate semantically enhanced graph structure data. Step 523: Map the semantically enhanced graph structure data to the constructed explanation template to generate an abnormal behavior feature explanation.

9. The method according to claim 1, characterized in that Step 6: Fill the abnormal behavior feature name, digital feature value, risk level value, and abnormal behavior feature explanation into the feature name element, digital feature element, risk level element, and feature explanation element respectively to form a digital feature library of abnormal behavior of the enterprise, which specifically includes the following steps: Step 61: Based on the constructed data mapping rules, clarify the correspondence between the abnormal behavior feature name and the feature name element, the digital feature value and the digital feature element, the risk level value and the risk level element, and the abnormal behavior feature explanation and the feature explanation element; Step 62: Through database operation statements, based on the corresponding relationship, the abnormal behavior feature name, digital feature value, risk level value, and abnormal behavior feature explanation are respectively filled into the feature name element, digital feature element, risk level element, and feature explanation element, and a retrieval index is created to form a digital feature library of the enterprise's abnormal behavior.