An AI-driven big data intelligent parsing and analysis method and system

Through the AI-driven big data intelligent analysis and analysis method, the problem that traditional methods are difficult to adapt to the diversity and complexity of big data is solved, efficient and accurate user behavior decision-making support and strategy formulation are achieved, and information communication efficiency and decision-making quality are improved.

CN119808794BActive Publication Date: 2025-07-01SHENYANG QINGAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510290084.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-01
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional big data analysis processing methods are difficult to adapt to the diversity and complexity of big data, resulting in inefficient intelligent resolution and ineffective support for user behavior activity decisions.

Method used

Using AI-driven big data intelligent analysis and analysis method, by obtaining the large data set of user behavior activity description text, behavior text semantic feature analysis, embedded and filtered feature selection, behavior activity feature classification and local interpretability AI intelligent analysis analysis is carried out, and the results of user behavior decision interpretation factor analysis are generated, and logical reasons are visualized.

Benefits of technology

It improves the efficiency and accuracy of intelligent analysis and analysis of big data, can quickly respond to market changes, formulate targeted and timely strategies, reduce the probability of decision-making errors, promote effective communication among decision-makers, and improve the efficiency of information dissemination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808794B_ABST
    Figure CN119808794B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of big data processing, and particularly relates to a big data intelligent parsing and analysis method and system driven by AI. The method includes the following steps: obtaining a big data set of user behavior activity description texts and performing behavioral text semantic feature analysis, and at the same time performing embedded and filtering feature selection processing to obtain a user behavior semantic embedded feature subset and a user behavior semantic filtering feature subset; performing repeated representative feature screening processing and behavioral activity feature classification according to the user behavior semantic embedded feature subset and the user behavior semantic filtering feature subset to obtain a behavioral semantic representative feature subset corresponding to each user behavior activity, and performing local interpretable AI intelligent parsing and analysis and logical reason visualization display processing on it to obtain the interpretable logical reasons behind the decisions of each user behavior activity. The present invention can improve the efficiency and accuracy of big data parsing and analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and particularly to a big data intelligent parsing and analysis method and system driven by AI. Background Art

[0002] In the digital age, the explosive growth of big data has posed unprecedented challenges and opportunities to all industries. The generation and storage of massive user behavior activity data provide rich basic data guarantees for the optimization of big data services. Through automated data integration technology, data from different sources (such as databases, sensors, text files, social media, etc.) are effectively aggregated and cleaned to ensure data quality and consistency. This process includes steps such as removing duplicate data, filling in missing values, and formatting, laying a solid foundation for subsequent analysis. At the same time, by applying advanced machine learning and deep learning algorithms, large-scale user behavior activity data is efficiently modeled and analyzed. These algorithms can automatically identify complex patterns in big data and capture non-linear relationships, thereby improving the accuracy of the big data intelligent parsing and analysis process. However, traditional big data parsing and processing methods often rely on rules and statistical models and are difficult to adapt to the diversity and complexity of big data, resulting in low efficiency of big data intelligent parsing and thus being unable to effectively support the activity decision-making of user behavior. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a big data intelligent parsing and analysis method and system driven by AI to solve at least one of the above technical problems.

[0004] To achieve the above object, a big data intelligent parsing and analysis method driven by AI includes the following steps:

[0005] Step S1: Obtain a big data set of user behavior activity description texts, and perform behavioral text semantic feature analysis on the big data set of user behavior activity description texts to obtain a user behavior activity description text semantic feature set;

[0006] Step S2: Perform embedded and filter-based feature selection processing on the user behavior activity description text semantic feature set to obtain a user behavior semantic embedded feature subset and a user behavior semantic filter-based feature subset; perform repeated representative feature screening processing based on the user behavior semantic embedded feature subset and the user behavior semantic filter-based feature subset to obtain a user behavior semantic comprehensive representative feature set;

[0007] Step S3: Classify the behavioral activity features of the comprehensive representative feature set of user behavior semantics to obtain the subset of representative features of behavior semantics corresponding to each user behavior activity; perform local interpretable AI intelligent analysis on the subset of representative features of behavior semantics corresponding to each user behavior activity to obtain the interpretable factor analysis results of user behavior decisions corresponding to each user behavior activity;

[0008] Step S4: Perform logical reason visualization display processing on the interpretable factor analysis results of user behavior decisions corresponding to each user behavior activity to obtain the interpretable logical reasons behind the decisions of each user behavior activity.

[0009] Further, Step S1 includes the following steps:

[0010] Step S11: Obtain a large dataset of user behavior activity description texts, where the large dataset of user behavior activity description texts includes user behavior activity text record data from user behavior data sources related to social media, mobile applications, online platforms, and e-commerce websites;

[0011] Step S12: Perform sensitive elimination processing on the large dataset of user behavior activity description texts to obtain a large dataset of user behavior activity texts with sensitive words eliminated; perform stop word removal and text normalization processing on the large dataset of user behavior activity texts with sensitive words eliminated to obtain a standardized large dataset of user behavior activity texts;

[0012] Step S13: Perform semantic atomic unit extraction processing on the standardized large dataset of user behavior activity texts to obtain a set of atomic units of semantic components of user behavior activity texts;

[0013] Step S14: Perform text semantic relationship mining analysis on each basic text semantic component atomic unit in the set of atomic units of semantic components of user behavior activity texts to obtain the user behavior activity text semantic relationships between each basic text semantic component atomic unit; construct a semantic relationship map for each corresponding basic text semantic component atomic unit based on the user behavior activity text semantic relationships between each basic text semantic component atomic unit to generate a user behavior activity text semantic relationship map;

[0014] Step S15: Perform semantic feature recognition analysis on the corresponding user behavior activity description texts in the standardized large dataset of user behavior activity texts based on the user behavior activity text semantic relationship map to obtain the semantic feature set of user behavior activity description texts.

[0015] Further, the construction of the semantic relationship map for each corresponding basic text semantic component atomic unit based on the user behavior activity text semantic relationships between each basic text semantic component atomic unit in Step S14 includes the following steps:

[0016] For each basic semantic component atom unit in the set of basic semantic component atom units of the user behavior activity text semantics, perform text frequency and inverse document frequency statistical calculations to obtain the text frequency and inverse document frequency corresponding to each basic semantic component atom unit of the text;

[0017] Based on the text frequency and inverse document frequency corresponding to each basic semantic component atom unit of the text, perform semantic relationship weight distribution calculations on the semantic relationships of the user behavior activity text between each basic semantic component atom unit to obtain the text semantic relationship distribution weights between each basic semantic component atom unit of the text;

[0018] According to the text semantic relationship distribution weights between each basic semantic component atom unit of the text, perform semantic relationship weight mapping determination on the semantic relationships of the user behavior activity text between each basic semantic component atom unit to obtain the user activity text semantic relationship weight mapping rules between each basic semantic component atom unit of the text;

[0019] Based on the user activity text semantic relationship weight mapping rules between each basic semantic component atom unit of the text, construct a semantic relationship map for the semantic relationships of the user behavior activity text between each basic semantic component atom unit and each corresponding basic semantic component atom unit of the text to generate a user behavior activity text semantic relationship map.

[0020] Further, step S2 includes the following steps:

[0021] Step S21: Perform embedded feature selection processing on the user behavior activity description text semantic feature set to obtain a user behavior semantic embedded feature subset;

[0022] Step S22: Perform filter-based feature selection processing on the user behavior activity description text semantic feature set to obtain a user behavior semantic filter-based feature subset;

[0023] Step S23: Perform duplicate feature screening processing according to the user behavior semantic embedded feature subset and the user behavior semantic filter-based feature subset to obtain a user behavior semantic duplicate feature set;

[0024] Step S24: Use the behavior semantic representativeness calculation formula to perform representativeness scoring calculations on each user behavior semantic sub-feature in the user behavior semantic duplicate feature set to obtain the feature representativeness scoring values corresponding to each user behavior semantic sub-feature;

[0025] Among them, the behavior semantic representativeness calculation formula is specifically:

[0026] ;

[0027] In the formula, is the th user behavior semantic sub - feature corresponding feature representativeness score value, is the th user behavior semantic sub - feature, is the th user behavior semantic sub - feature, and are both item - index parameters of user behavior semantic sub - features, is the total number of user behavior semantic sub - features, is the normalization constant, is the lower limit of the integration time range, is the upper limit of the integration time range, is the time - variable parameter, is the weight parameter corresponding to the th user behavior semantic sub - feature, is the exponential function, is the th feature center position corresponding to the user behavior semantic sub - feature in the user behavior semantic repeated feature set, is the th feature dispersion degree corresponding to the user behavior semantic sub - feature in the user behavior semantic repeated feature set, is the th semantic feature similarity measurement value between the th user behavior semantic sub - feature and the th user behavior semantic sub - feature at the time point is the correction coefficient of the feature representativeness score value;

[0028] Step S25: Compare and judge the feature representativeness score value corresponding to each user behavior semantic sub - feature according to a preset representativeness score threshold. When the feature representativeness score value is greater than or equal to the preset representativeness score threshold, mark the corresponding user behavior semantic sub - feature as a representative feature; when the feature representativeness score value is less than the preset representativeness score threshold, mark the corresponding user behavior semantic sub - feature as a non - representative feature; screen out and aggregate the user behavior semantic sub - features corresponding to the representative features marked in the user behavior semantic repeated feature set to obtain the user behavior semantic comprehensive representative feature set.

[0029] Furthermore, step S21 includes the following steps:

[0030] Perform feature importance evaluation and analysis on each user behavior semantic sub - feature in the user behavior activity description text semantic feature set to obtain the feature importance degree corresponding to each user behavior semantic sub - feature;

[0031] Perform feature sparse quantization calculation on each user behavior semantic sub - feature in the user behavior activity description text semantic feature set to obtain the feature sparse degree corresponding to each user behavior semantic sub - feature;

[0032] Perform L1 - regularization statistical analysis on each corresponding user behavior semantic sub - feature in the user behavior activity description text semantic feature set according to the feature importance degree and the feature sparse degree corresponding to each user behavior semantic sub - feature to obtain the L1 - regularization strength parameter of the user behavior semantic feature;

[0033] Perform L1 - regularization embedded feature selection on each corresponding user behavior semantic sub - feature in the user behavior activity description text semantic feature set based on the L1 - regularization strength parameter of the user behavior semantic feature to obtain the user behavior semantic embedded feature subset.

[0034] Further, step S22 includes the following steps:

[0035] Perform statistical calculation of the number of occurrences of activity descriptions on each user behavior semantic sub - feature in the user behavior activity description text semantic feature set to obtain the number of occurrences of user behavior activity descriptions corresponding to each user behavior semantic sub - feature;

[0036] Perform feature frequency statistical analysis on each corresponding user behavior semantic sub - feature according to the number of occurrences of user behavior activity descriptions corresponding to each user behavior semantic sub - feature to obtain the user behavior semantic feature frequency corresponding to each user behavior semantic sub - feature;

[0037] Perform feature correlation evaluation and analysis on each user behavior semantic sub - feature in the user behavior activity description text semantic feature set to obtain the correlation coefficient between each user behavior semantic sub - feature and the user behavior activity description text; set the significance level of the user behavior activity according to the correlation coefficient between each user behavior semantic sub - feature and the user behavior activity description text to obtain the significance level of the user behavior activity semantic feature;

[0038] Perform sub - feature chi - square test calculation on each corresponding user behavior semantic sub - feature based on the significance level of the user behavior activity semantic feature and the user behavior semantic feature frequency corresponding to each user behavior semantic sub - feature to obtain the semantic feature chi - square value corresponding to each user behavior semantic sub - feature;

[0039] Perform a filtering feature selection process on each user behavior semantic sub - feature corresponding to the semantic feature chi - square value in the user behavior activity description text semantic feature set to obtain a user behavior semantic filtering feature subset.

[0040] Further, step S3 includes the following steps:

[0041] Step S31: Extract context information for each user behavior semantic representative sub - feature in the user behavior semantic comprehensive representative feature set to obtain the user behavior semantic context information corresponding to each user behavior semantic representative sub - feature;

[0042] Step S32: Perform sub - feature activity attribute extraction processing on each user behavior semantic representative sub - feature based on the user behavior semantic context information corresponding to it to obtain the user behavior activity attribute corresponding to each user behavior semantic representative sub - feature;

[0043] Step S33: Classify the behavior activity features of each user behavior semantic representative sub - feature in the user behavior semantic comprehensive representative feature set based on the user behavior activity attribute corresponding to it to obtain the behavior semantic representative feature subsets corresponding to each user behavior activity;

[0044] Step S34: Perform a local interpretability evaluation and analysis on the behavior semantic representative feature subsets corresponding to each user behavior activity to obtain the user behavior decision local interpretability factors corresponding to each user behavior activity;

[0045] Step S35: Perform an AI - driven intelligent parsing and analysis on the user behavior decision local interpretability factors corresponding to each user behavior activity to obtain the user behavior decision interpretability factor analysis results corresponding to each user behavior activity.

[0046] Further, step S34 includes the following steps:

[0047] Step S341: Perform a user behavior decision semantic association analysis on each behavior semantic representative sub - feature in the behavior semantic representative feature subsets corresponding to each user behavior activity to obtain the semantic association relationship between the behavior semantic representative sub - features corresponding to each user behavior activity and the user behavior decision;

[0048] Step S342: Perform a behavior decision local influencing factor analysis on the behavior semantic representative feature subsets corresponding to each user behavior activity based on the semantic association relationship between the behavior semantic representative sub - features corresponding to each user behavior activity and the user behavior decision to obtain the candidate set of user behavior decision local influencing factors corresponding to each user behavior activity;

[0049] Step S343: Perform user behavior performance evaluation and analysis on each behavioral semantic representative sub - feature within the behavioral semantic representative feature subset corresponding to each user behavior activity, to obtain the user behavior activity performance scores for each behavioral semantic representative sub - feature corresponding to each user behavior activity;

[0050] Step S344: Based on the user behavior activity performance scores for each behavioral semantic representative sub - feature corresponding to each user behavior activity, design a local interpretability framework for the behavioral semantic representative feature subset corresponding to each user behavior activity, to generate a local interpretability evaluation and analysis model for each user behavior activity;

[0051] Step S345: According to the local interpretability evaluation and analysis model corresponding to each user behavior activity, perform local interpretability evaluation and analysis on the candidate set of local influencing factors of user behavior decision corresponding to each user behavior activity, to obtain the local interpretability factors of user behavior decision corresponding to each user behavior activity.

[0052] Further, step S4 includes the following steps:

[0053] Step S41: Perform factor correlation degree statistical analysis on the corresponding user behavior decision interpretability factors within the user behavior decision interpretability factor analysis results corresponding to each user behavior activity, to obtain the correlation degrees between the decision interpretability factors corresponding to each user behavior activity;

[0054] Step S42: Based on the correlation degrees between the decision interpretability factors corresponding to each user behavior activity, construct a decision logic relationship network for the corresponding user behavior decision interpretability factors, to generate the user behavior decision logic relationship network structure corresponding to each user behavior activity;

[0055] Step S43: Based on the user behavior decision logic relationship network structure corresponding to each user behavior activity, perform logical reason visualization display processing on the user behavior decision interpretability factor analysis results corresponding to each user behavior activity, to obtain the interpretable logical reasons behind each user behavior activity decision.

[0056] Further, the present invention also provides a big - data intelligent analysis system driven by AI, used to execute the big - data intelligent analysis method driven by AI as described above. The big - data intelligent analysis system driven by AI includes:

[0057] A user behavior semantic feature analysis module, used to obtain a big - data set of user behavior activity description texts, and perform behavioral text semantic feature analysis on the big - data set of user behavior activity description texts, so as to obtain a user behavior activity description text semantic feature set;

[0058] The user behavior semantic repetition representative screening module is used to perform embedded and filtering feature selection processing on the semantic feature set of the user behavior activity description text, so as to obtain the user behavior semantic embedded feature subset and the user behavior semantic filtering feature subset; perform repeated representative feature screening processing according to the user behavior semantic embedded feature subset and the user behavior semantic filtering feature subset, so as to obtain the user behavior semantic comprehensive representative feature set;

[0059] The behavior activity AI-driven intelligent analysis module is used to classify the behavior activity features of the user behavior semantic comprehensive representative feature set to obtain the behavior semantic representative feature subset corresponding to each user behavior activity; perform local interpretable AI intelligent analysis on the behavior semantic representative feature subset corresponding to each user behavior activity to obtain the user behavior decision interpretable factor analysis result corresponding to each user behavior activity;

[0060] The behavior decision analysis reason visualization module is used to perform logical reason visualization display processing on the user behavior decision interpretable factor analysis result corresponding to each user behavior activity, so as to obtain the interpretable logical reason behind the decision of each user behavior activity.

[0061] The beneficial effects of the present invention:

[0062] 1. The big data intelligent parsing and analysis method driven by AI proposed by the present invention, compared with the prior art, the beneficial effects of the present application are as follows: By collecting large datasets of text records related to user behavior activities from multiple channels such as social media, mobile applications, online platforms, and e-commerce websites, it provides rich materials for subsequent analysis. User interaction records on social media, user operation logs within mobile applications, comments and feedback on online platforms, and purchase and browsing records on e-commerce websites are all true reflections of user behavior activities. By aggregating these data, the user's needs, preferences, and behavior patterns can be comprehensively understood, and thus support can be provided for the subsequent decision-making analysis process of user behavior activities. The huge amount of data also lays a foundation for the application of algorithms such as data mining, machine learning, and deep learning, making the analysis results more accurate and reliable. At the same time, by performing behavioral text semantic feature analysis on the large dataset of user behavior activity description texts to analyze and identify the semantic features in the user behavior activity description texts, a key feature set reflecting the user's needs and preferences can be effectively extracted. These features can not only help better understand the actual needs of users but also provide a scientific basis for the subsequent processing process. Secondly, by performing embedded feature selection processing on the user behavior activity description text semantic feature set, aiming to automatically identify those features that have the greatest impact on subsequent analysis through L1 regularization technology, embedded feature selection not only depends on the independence or importance score of features but also combines the model training process, enabling feature selection and model construction to be carried out simultaneously, thereby improving the efficiency and effectiveness of feature selection, and can also significantly reduce redundant and irrelevant features, reducing the complexity of text semantic features, thus providing a reliable basis for subsequent analysis and decision-making. It also processes the semantic feature set of user behavior activity description texts by adopting a filter-based feature selection method. Different from embedded feature selection, filter-based feature selection is usually carried out before model training. By performing independent statistical analysis on each feature in the feature set and using statistics (such as chi-square test, information gain, or mutual information, etc.) to evaluate the relationship between the feature and the target variable, the features most relevant to user behavior are screened out. The advantage of this process is its high computational efficiency and independence from the training of any machine learning model, enabling it to be quickly executed even on large-scale datasets, which can effectively remove noise features, thus providing a clearer and more concise feature set for the subsequent processing process.By performing repeated feature screening based on the user behavior semantic embedded feature subset and the user behavior semantic filtering feature subset, the aim is to further screen out the repetitive features in the feature set. By cross-comparing these two feature subsets, it is possible to identify which features appear in different selection methods and screen them out for later use. By setting a reasonable threshold, it is possible to effectively screen out the key features that can accurately reflect user behavior, while excluding features that contribute little to the model. Such a selection mechanism can ensure the simplicity and effectiveness of the final feature set, while reducing the complexity of model training. This not only improves the accuracy of subsequent data analysis, but also provides a more scientific and reliable basis for the subsequent big data intelligent analysis and processing process. Then, by classifying the behavior activity features of the user behavior semantic comprehensive representative feature set, the value of this step lies in that it can structure the massive user behavior information, enabling different types of behaviors to be effectively classified. The classification process not only helps to identify the commonalities and individualities of user behavior, but also reveals the potential relationships between behaviors. This analysis can provide a basis for user stratification, thus providing basic data guarantee for the subsequent analysis and processing process. By also performing local interpretability evaluation and analysis on the behavior semantic representative feature subsets corresponding to each user behavior activity, and conducting AI-driven intelligent analysis, it is possible to further deepen the understanding of user behavior. This process converts the interpretability factors of user behavior decisions into actionable analysis results through machine learning and data analysis techniques. The key to this analysis is that it can provide clear insights into user behavior for the subsequent processing process. This AI-driven analysis process can quickly respond to market changes, timely adjust strategies to meet the changing needs of users, and formulate more targeted and timely strategies to ensure a competitive advantage at each stage of user decision-making, thereby improving the efficiency of the big data intelligent analysis process. Finally, by performing logical reason visualization display processing on the analysis results of the user behavior decision interpretability factors corresponding to each user behavior activity, the key to this step is that it not only provides an intuitive understanding of user behavior decisions, but also promotes effective communication with decision-makers. The visualization display of logical reasons converts complex data and relationships into easy-to-understand graphics, greatly reducing the difficulty of information transmission. Through the graphical method, decision-makers can quickly grasp the key factors and their relationships, facilitating decision-making and action. Compared with traditional text reports, this visualization method is more able to attract the attention of the audience and improve the dissemination efficiency of information. It can also help decision-makers consider more dimensions of factors when formulating strategies. For example, in industries such as e-commerce and finance, decision-making often involves trade-offs in multiple aspects. Being able to clearly see the relationships between various factors helps to comprehensively evaluate the corresponding impact results. This systematic way of thinking promotes a more rational decision-making process, reduces the probability of decision-making errors, and thus can effectively support the activity decisions of user behavior.

[0063] 2. The big data intelligent parsing and analysis system driven by AI proposed by the present invention is generally composed of a user behavior semantic feature analysis module, a user behavior semantic repetition representative screening module, a behavior activity AI-driven intelligent parsing module, and a behavior decision parsing reason visualization module. It can implement any of the big data intelligent parsing and analysis methods driven by AI described in the present invention. For the operation between computer programs running on each module to implement the big data intelligent parsing and analysis method driven by AI, the internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower input, and can quickly and effectively provide a more accurate and efficient big data intelligent parsing and analysis process driven by AI, thus simplifying the operation process of the big data intelligent parsing and analysis system driven by AI. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0065] Figure 1 It is a schematic flow chart of the steps of the big data intelligent parsing and analysis method driven by AI of the present invention;

[0066] Figure 2 is Figure 1 a detailed schematic flow chart of step S1 in

[0067] Figure 3 is Figure 1 a detailed schematic flow chart of step S2 in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.

[0069] To achieve the above object, please refer to Figures 1 to 3 The present invention provides a big data intelligent parsing and analysis method driven by AI. In the embodiments of the present invention, please refer to Figure 1 shown in

[0070] Step S1: Obtain a big data set of user behavior activity description texts, and perform behavior text semantic feature analysis on the big data set of user behavior activity description texts to obtain a user behavior activity description text semantic feature set;

[0071] In an embodiment of the present invention, descriptive text information data related to user behavior activities is collected in real time from social media platforms (such as Twitter, Facebook, Weibo), mobile applications (such as in-app logs), online platforms (such as forums, blogs), and e-commerce websites (such as user reviews and feedback on e-commerce platforms). The user behavior activity text data on various platforms is batch-crawled through API interfaces or data crawler technologies. These data include the text descriptions of the user's posting records, comment content, purchase behavior, browsing records, likes, shares, and other behaviors on the platform, thereby obtaining a large dataset of user behavior activity description texts. By performing processing to eliminate sensitive words, removing stop words, and normalizing the text on the large dataset of user behavior activity description texts collected previously, using a unified word segmentation tool, such as the Jieba tokenizer or SpaCy, to perform word segmentation, formatting special characters and punctuation marks in the text uniformly, ensuring the consistency and standardization of the text in subsequent processing, and using the word segmentation tool in natural language processing technology to split the text into lexical units, performing lexical-level segmentation using Jieba or SpaCy to ensure that the granularity of semantic units is fine enough, combining dependency parsing technology, extracting semantic components such as nouns and verbs in the text, constructing a set of semantic atomic units. At the same time, by mining and analyzing the semantic relationships between each basic semantic component atomic unit in the set of semantic component atomic units extracted previously, aiming to discover potential semantic associations between the basic semantic component atomic units of the text, and regarding each basic semantic component atomic unit as a node in the graph, while the semantic relationship is used as the edge between the nodes. Additionally, when constructing the graph, the weight of the edge is obtained from the similarity of the semantic relationship mining results, that is, the edges between pairs of semantic components with a higher relationship strength have a greater edge weight, thereby transforming the text data and semantic relationships into a graphical structure, ensuring that the semantic relationship structure of user behavior activities is clear and operable. Then, by combining the previously constructed semantic relationship graph, the semantic features of the corresponding user behavior activity description text are identified and analyzed. Through the relationship graph stored in the graph database, combined with community detection algorithms in graph theory (such as the Louvain algorithm or spectral clustering), the text is grouped to identify user groups with similar behaviors in a specific semantic domain. During the feature extraction process, a machine learning model (such as support vector machine SVM or random forest) is used to classify the nodes in the graph, extracting the main semantic features related to user behavior activities. Through feature engineering techniques, the feature set is further screened and optimized, thereby obtaining a set of semantic feature sets that are helpful for further user behavior analysis and prediction, and finally obtaining the semantic feature set of user behavior activity description texts.

[0072] Step S2: Perform embedded and filtered feature selection processing on the semantic feature set of the user behavior activity description text to obtain a user behavior semantic embedded feature subset and a user behavior semantic filtered feature subset; perform repeated representative feature screening processing based on the user behavior semantic embedded feature subset and the user behavior semantic filtered feature subset to obtain a user behavior semantic comprehensive representative feature set;

[0073] In the embodiment of the present invention, by using a pre-trained deep learning model, such as BERT or Word2Vec, the corresponding user behavior semantic sub-features in the semantic feature set of the user behavior activity description text are vectorized. The user behavior semantic sub-features are converted into dense vector representations in a high-dimensional feature space. By analyzing these vectors, the relationship strength between each feature and the target variable can be calculated, and the L1 regularization method (Lasso regression) is used to perform feature selection, retaining the features with higher correlation with the user behavior performance, and removing redundant and irrelevant features, thereby obtaining a user behavior semantic embedded feature subset. Also, a filtered feature selection is performed on the corresponding user behavior semantic sub-features in the semantic feature set of the user behavior activity description text. The chi-square test statistical method is used to evaluate the independence between each user behavior semantic sub-feature and the user behavior result, and by calculating the chi-square value of each user behavior semantic sub-feature, the features with chi-square values higher than the preset threshold are screened out, thereby obtaining a user behavior semantic filtered feature subset. At the same time, by screening the repeated features of the previously discriminated user behavior semantic embedded feature subset and the filtered feature subset, at this time, the intersection operation is performed on the two feature subsets, and the features with obvious overlap in user behavior semantics are further screened out and de-duplicated. Also, a representative score calculation is performed on each user behavior semantic sub-feature in the previously screened semantic repeated feature set to ensure that the representative score of each sub-feature is accurately calculated. Then, by comparing and judging the feature representative score values calculated previously using a preset representative score threshold, a reasonable score threshold is set. This value is based on the analysis results of the feature performance in the previous steps to ensure that it can effectively distinguish representative features from non-representative features. Subsequently, the feature representative score values corresponding to each user behavior semantic sub-feature are traversed, and through conditional judgment statements, representative features (if the feature representative score value is greater than or equal to the preset representative score threshold) and non-representative features (if the feature representative score value is less than the preset representative score threshold) are marked, and by screening out the user behavior semantic sub-features marked as representative features and integrating them in a dataset, a user behavior semantic comprehensive representative feature set is finally obtained.

[0074] Step S3: Classify the behavioral activity features of the user behavior semantic comprehensive representative feature set to obtain the behavioral semantic representative feature subsets corresponding to each user behavioral activity; perform local interpretable AI intelligent analysis on the behavioral semantic representative feature subsets corresponding to each user behavioral activity to obtain the interpretable factor analysis results of the user behavior decision corresponding to each user behavioral activity;

[0075] In the embodiments of the present invention, for each representative sub-feature of user behavior semantics in the comprehensively representative feature set of user behavior semantics obtained from previous analysis, context information is extracted. These sub-features include user clicks, search terms, purchase records, etc. Using natural language processing (NLP) techniques, methods such as word segmentation, part-of-speech tagging, and named entity recognition are employed to extract context information related to these sub-features. For example, for the statement "the user clicked on the link of product A", keywords such as "user", "clicked", "product A", "link" can be extracted, and then context information about this behavior is generated, including information such as time, location, device type, etc. By combining the previously extracted context information of user behavior semantics, the extraction of sub-feature activity attributes is performed for each corresponding representative sub-feature of user behavior semantics. Taking the behavior of "the user clicks on the link of product A" as an example, the activity attributes can include "product type", "click frequency", "user group", etc., thus summarizing the activity attributes of each sub-feature. At the same time, by combining the user behavior activity attributes corresponding to each representative sub-feature obtained from previous mining and analysis, the classification processing of behavior activity features is performed for each corresponding representative sub-feature of user behavior semantics in the comprehensively representative feature set of user behavior semantics. In this step, a machine learning model (such as support vector machine SVM or random forest) can be used to classify user behavior activities. The activity attributes of each sub-feature are used as input features. After the model is trained, it can identify different categories of user behavior activities, such as "browsing", "purchasing", "commenting", etc., thereby classifying to obtain the behavior semantics representative feature subsets corresponding to each user behavior activity.Then, through the evaluation and analysis of local interpretability for each user behavior activity and its corresponding subset of representative behavioral semantic features, in this process, local interpretable model-agnostic interpretation methods (such as LIME or SHAP) are adopted to identify the key factors affecting user decisions. By generating local models and analyzing the relationship between user behavior activity features and decision results, interpretable factors for each decision can be provided. For example, if a user chooses to purchase a product, LIME can help identify the roles played by features such as "discount" and "product evaluation" in the decision. This process needs to be combined with data visualization tools (such as Matplotlib or Tableau) to display the results of interpretability analysis, making the local interpretable factors of the decision clear at a glance, that is, the degree of influence of the corresponding local influencing factors in the user decision-making process. And through AI-driven intelligent parsing and analysis of the local interpretable factors of user behavior decisions obtained after previous evaluation and analysis, this step will be combined with deep learning techniques (such as deep neural networks) for factor parsing. By using the neural network model to conduct more in-depth feature learning on the local interpretable factors, through the model's non-linear mapping of the relationships between factors, more complex decision-making logics are extracted. After training the model, corresponding decision interpretable factor parsing results can be generated. For example, the model will point out that a certain user is more influenced by "social media influence" and "similar product evaluation" when making a purchase. This process relies on deep learning frameworks (such as TensorFlow or PyTorch) and parsing tools to ensure that the final parsing results are highly credible and practical, and finally obtain the parsing results of user behavior decision interpretable factors corresponding to each user behavior activity.

[0076] Step S4: Perform logical reason visualization display processing on the parsing results of user behavior decision interpretable factors corresponding to each user behavior activity to obtain the interpretable logical reasons behind each user behavior activity decision.

[0077] In the embodiments of the present invention, through the use of the Pearson correlation coefficient or the Spearman rank correlation coefficient, a statistical analysis of the correlation degree of the user behavior decision interpretability factors corresponding to each user behavior activity is carried out in the analysis results of the user behavior decision interpretability factors, so as to input the collected user behavior decision interpretability factors into statistical software or data analysis tools, such as the pandas library in Python or the psych package in R language, and use the correlation calculation function to calculate to obtain the correlation degree matrix between the factors. This matrix clearly shows the correlation degree between the decision interpretability factors corresponding to each user behavior activity, and through combining the correlation degree between the decision interpretability factors corresponding to each user behavior activity obtained by the previous quantitative calculation, a connection construction of the logical relationship network of the corresponding user behavior decision interpretability factors is carried out. By using the concepts of nodes and edges in graph theory, the user behavior decision interpretability factors are regarded as nodes, and the correlation degree between the factors is used as the weight of the edge. Then, with the help of network analysis tools, such as Gephi or NetworkX, a directed graph containing all relevant factors is established, and the weight of the edge is set to reflect the intensity of its correlation degree. Next, by applying network analysis algorithms (such as the Dijkstra algorithm or the PageRank algorithm), important factors and their mutual relationships are identified, and the topological structure of the decision logical relationship network is generated. Then, through combining the decision logical relationship network structure obtained by the previous connection construction, a visual display of the logical reasons for the corresponding user behavior decision interpretability factor analysis results is carried out, so as to more intuitively understand the decision logical reasons behind the user behavior activities. By selecting key factors based on the constructed decision logical relationship network and presenting their data in a graphical way, an interactive visualization graph can be created using chart tools (such as D3.js or Plotly) to display each factor and the relationships and influences between them. By setting the interaction function, the user can click on a certain factor to view its specific impact on the decision and the logical path, and then analyze the importance of this factor in the decision-making process and its logical explanation reasons, thereby generating a corresponding decision logical display report. This report will include a detailed analysis of the relationships between the factors, an influence assessment, and a comprehensive explanation of their impact on the user behavior decision, and finally obtain the interpretable logical reasons behind each user behavior activity decision.

[0078] Further, as an embodiment of the present invention, referring to Figure 2 as shown, it is Figure 1 a detailed step flow schematic diagram of step S1 in

[0079] Step S11: Obtain a large dataset of user behavior activity description texts, where the large dataset of user behavior activity description texts includes user behavior activity text record data from user behavior data sources related to social media, mobile applications, online platforms, and e-commerce websites;

[0080] In an embodiment of the present invention, descriptive text information data related to user behavior activities is obtained by real-time collection from social media platforms (such as Twitter, Facebook, Weibo), mobile applications (such as in-app logs), online platforms (such as forums, blogs), and e-commerce websites (such as user comments and feedback on e-commerce platforms). The user behavior activity text data on various platforms is batch-crawled through API interfaces or data scraping technologies to ensure coverage of a wide range of user activity scenarios. These data include text descriptions of user posting records, comment content, purchase behaviors, browsing records, likes, shares, etc. on the platforms. To ensure the extensiveness and comprehensiveness of the data, the Requests library in the Python programming language is used for API calls, or the Scrapy framework is used for web data scraping, and the data is organized into a large dataset, and finally a large dataset of user behavior activity description texts is obtained.

[0081] Step S12: Perform sensitive information elimination processing on the large dataset of user behavior activity description texts to obtain a large dataset of user behavior activity text with sensitive words eliminated; perform stop word removal and text normalization processing on the large dataset of user behavior activity text with sensitive words eliminated to obtain a standardized large dataset of user behavior activity texts;

[0082] In an embodiment of the present invention, by performing elimination processing of sensitive words on the previously collected large dataset of user behavior activity description texts, the sensitive word list is updated regularly through a standard thesaurus or set manually. Regular expressions are used to match sensitive words and replace them with harmless characters. This process is implemented using the re module in Python to ensure that sensitive information in the text data is removed, thereby obtaining a large dataset of user behavior activity text with sensitive words eliminated. Then, stop word removal is performed on the large dataset of user behavior activity text with sensitive words eliminated. The stop word list includes common words without substantial meaning (such as "de", "shi", "zai", etc.). These words are not helpful for semantic analysis. Therefore, the stop word list in the NLTK library is used for removal, and text normalization processing is performed on the large dataset of text after stop word removal. By using a unified word segmentation tool, such as the Jieba word segmenter or SpaCy, word segmentation is performed, and special characters and punctuation marks in the text are uniformly formatted to ensure the consistency and standardization of the text in subsequent processing, and finally a standardized large dataset of user behavior activity texts is obtained.

[0083] Step S13: extracting semantic atomic units from the standardized user behavior activity text big data set to obtain a set of semantic component atomic units of the user behavior activity text;

[0084] In an embodiment of the present invention, semantic atomic units are extracted from a previously standardized large data set of standardized user behavior activity text, so as to split the text into vocabulary units by using a word segmentation tool in natural language processing technology, and use Jieba word segmentation or SpaCy to perform vocabulary-level segmentation to ensure that the granularity of the semantic units is sufficiently refined. Next, in combination with dependency syntactic analysis technology, semantic components such as nouns and verbs in the text are extracted to construct a set of semantic atomic units. This process is achieved by using the dependency parsing function in StanfordNLP or Spacy, ensuring that each semantic component unit is accurately extracted from the text, forming a set of atomic units with vocabulary as the core, and finally obtaining a set of semantic component atomic units of the user behavior activity text.

[0085] Step S14: performing text semantic relationship mining and analysis on each text basic semantic component atomic unit in the user behavior activity text semantic component atomic unit set to obtain the user behavior activity text semantic relationship between each text basic semantic component atomic unit; constructing a semantic relationship graph for each corresponding text basic semantic component atomic unit based on the user behavior activity text semantic relationship between each text basic semantic component atomic unit to generate a user behavior activity text semantic relationship graph;

[0086] In the embodiments of the present invention, by mining and analyzing the semantic relationships between each text basic semantic component atomic unit within the set of text semantic component atomic units of the user behavior activity text extracted previously, the aim is to discover the potential semantic associations between the text basic semantic component atomic units. The analysis method adopted is relationship mining based on a semantic model, such as word embedding models like Word2Vec or GloVe. By calculating the vector space similarity between different text basic semantic component atomic units to extract semantic relationships. In specific operations, first, the semantic component atomic units are mapped into the vector space through a pre-trained word embedding model (such as Word2Vec or FastText). Then, the cosine similarity between each text basic semantic component atomic unit is calculated to determine the strength of their semantic relationships. The implicit relationships between different semantic units in the text can be represented numerically, thereby obtaining the user behavior activity text semantic relationships between each text basic semantic component atomic unit. At the same time, by combining the user behavior activity text semantic relationships between each text basic semantic component atomic unit obtained from the previous analysis, a semantic relationship graph is constructed for the corresponding text basic semantic component atomic units. Each basic semantic component atomic unit is regarded as a node in the graph, and the semantic relationship is used as the edge between the nodes. When constructing the graph, the weight of the edge is obtained from the similarity of the semantic relationship mining results, that is, the edges between semantic component pairs with higher relationship strength have greater weights. This process is stored and managed using graph database technology, using graph database systems such as Neo4j, converting the text data and semantic relationships into a graphical structure for modeling to ensure that the semantic relationship structure of the user behavior activity is clear and operable, and finally generating a user behavior activity text semantic relationship graph.

[0087] Step S15: Based on the user behavior activity text semantic relationship graph, perform semantic feature recognition and analysis on the corresponding user behavior activity description text in the standardized user behavior activity text big data set to obtain the user behavior activity description text semantic feature set.

[0088] In an embodiment of the present invention, by combining the user behavior activity text semantic relationship graph constructed previously, semantic feature recognition and analysis are performed on the corresponding user behavior activity description text in the standardized user behavior activity text big data set after standardization. Through the relationship graph stored in the graph database and combining community detection algorithms in graph theory (such as the Louvain algorithm or spectral clustering), the text is clustered to identify user groups with similar behaviors in a specific semantic domain. During the feature extraction process, a machine learning model (such as support vector machine SVM or random forest) is used to classify the nodes in the graph, and the main semantic features related to user behavior activities are extracted. Through feature engineering techniques, the feature set is further screened and optimized, so as to obtain a set of semantic feature sets that are helpful for further user behavior analysis and prediction, and finally a user behavior activity description text semantic feature set is obtained.

[0089] Further, the construction of the semantic relationship graph for each text basic semantic component atomic unit corresponding to the user behavior activity text semantic relationship between each text basic semantic component atomic unit in step S14 includes the following steps:

[0090] Perform text frequency and inverse document frequency statistical calculations on each text basic semantic component atomic unit in the set of user behavior activity text semantic component atomic units to obtain the text frequency and inverse document frequency corresponding to each text basic semantic component atomic unit;

[0091] In an embodiment of the present invention, by performing text frequency and inverse document frequency statistical calculations on each text basic semantic component atomic unit in the set of user behavior activity text semantic component atomic units extracted previously, a word segmentation tool in natural language processing technology (such as NLTK, SpaCy or Jieba) is used to segment the text, and each basic semantic component atomic unit is extracted. Then, by constructing a document-term matrix (DTM), the term frequency TF(t, d) (Term Frequency, TF) of each atomic unit in the document is calculated, where t represents the occurrence time of the atomic unit in the document and d is the number of occurrences of the atomic unit in the document. The inverse document frequency (Inverse Document Frequency, IDF) is calculated by calculating the number of documents containing a specific atomic unit, and then applying the formula , where is the total number of documents, is the number of documents containing this atomic unit, and finally the text frequency and inverse document frequency corresponding to each text basic semantic component atomic unit are obtained.

[0092] Preferably, based on the text frequency and inverse document frequency corresponding to each atomic unit of the basic semantic components of the text, a semantic relationship weight distribution calculation is performed on the text semantic relationships of the user behavior activity between each atomic unit of the basic semantic components of the text, so as to obtain the text semantic relationship distribution weights between each atomic unit of the basic semantic components of the text;

[0093] In the embodiment of the present invention, by combining the text frequency and inverse document frequency corresponding to each atomic unit of the basic semantic components of the text obtained previously, the text semantic relationships of the user behavior activity between each atomic unit of the basic semantic components of the text are further analyzed, so as to adopt the weighted cosine similarity or correlation coefficient method to quantify the similarity between atomic units, thereby distributing semantic relationship weights. The specific operations include constructing a similarity matrix, where the rows and columns of the matrix respectively represent different atomic units of the basic semantic components. For each pair of atomic units, the cosine similarity is calculated through their TF-IDF values to obtain a similarity value between 0 and 1 as the semantic relationship weight between this pair of atomic units. Matrix calculations can be performed using the linear algebra module in the NumPy library, and these weights are integrated into a weight distribution list to form the semantic relationship mapping between each atomic unit, and finally the text semantic relationship distribution weights between each atomic unit of the basic semantic components of the text are obtained.

[0094] Preferably, based on the text semantic relationship distribution weights between each atomic unit of the basic semantic components of the text, a semantic relationship weight mapping determination is performed on the text semantic relationships of the user behavior activity between each atomic unit of the basic semantic components of the text, so as to obtain the user activity text semantic relationship weight mapping rules between each atomic unit of the basic semantic components of the text;

[0095] In the embodiment of the present invention, by combining the text semantic relationship distribution weights between each atomic unit of the basic semantic components of the text obtained previously, a weight mapping determination is performed on the text semantic relationships of the user behavior activity between each atomic unit of the basic semantic components of the text obtained previously, so as to determine the relationship mapping rules between each atomic unit of the basic semantic components in the user behavior activity text. By constructing a weighted directed graph, where the nodes represent the atomic units of the basic semantic components and the weights of the edges represent the intensity of the semantic relationships between them. In the specific implementation, the Dijkstra algorithm in graph theory is used to calculate the shortest path, thereby revealing the core connections between each atomic unit in the user behavior activity. At the same time, a graph visualization tool (such as NetworkX and Matplotlib) is used to visualize the relationship graph, making the relationships between atomic units clear at a glance, and finally the user activity text semantic relationship weight mapping rules between each atomic unit of the basic semantic components of the text are obtained through mapping analysis.

[0096] Preferably, based on the user activity text semantic relationship weight mapping rules between each atomic unit of the basic text semantic components, a semantic relationship map is constructed for the user behavior activity text semantic relationship between each atomic unit of the basic text semantic components and each corresponding atomic unit of the basic text semantic components, so as to generate a user behavior activity text semantic relationship map.

[0097] In the embodiment of the present invention, by combining the user activity text semantic relationship weight mapping rules determined previously, a connection construction of the semantic relationship map is carried out for the user behavior activity text semantic relationship between the atomic units of the basic text semantic components and each corresponding atomic unit of the basic text semantic components, so as to realize in-depth analysis of each atomic unit of the basic semantic component and its semantic relationship. Taking the atomic unit of the basic text semantic component as a node and the user behavior activity text semantic relationship as an edge and assigning the corresponding user activity text semantic relationship weight to the edge, a multi-level semantic relationship map is constructed. This map not only shows the direct connections between the atomic units of the basic semantic components, but also includes indirect connections and hierarchical relationships. The graph database technology (such as Neo4j) can be used to store and manage these relationships, and the query language of the graph database (such as Cypher) can be used for efficient retrieval and analysis. In addition, by combining machine learning algorithms (such as clustering analysis or classification algorithms), the user behavior patterns are identified and classified, and then the structure and content of the semantic relationship map are further optimized, and finally a user behavior activity text semantic relationship map is constructed and generated.

[0098] Further, as an embodiment of the present invention, referring to Figure 3 shown, for Figure 1 the detailed step flow diagram of step S2 in

[0099] Step S21: Perform embedded feature selection processing on the user behavior activity description text semantic feature set to obtain a user behavior semantic embedded feature subset;

[0100] In the embodiment of the present invention, by using a pre-trained deep learning model, such as BERT or Word2Vec, the corresponding user behavior semantic sub-features in the user behavior activity description text semantic feature set are vectorized. The user behavior semantic sub-features are converted into dense vector representations in a high-dimensional feature space. By analyzing these vectors, the relationship strength between each feature and the target variable can be calculated, and the L1 regularization method (Lasso regression) is used to perform feature selection, retaining the features with higher correlation with the user behavior performance, removing redundant and irrelevant features. The generated user behavior semantic embedded feature subset contains sub-features that have a significant impact on the user behavior activity, and finally a user behavior semantic embedded feature subset is obtained.

[0101] Step S22: performing filtering feature selection processing on the semantic feature set of the user behavior activity description text to obtain a user behavior semantic filtering feature subset;

[0102] In an embodiment of the present invention, a filtering feature selection is performed on the corresponding user behavior semantic sub-features in the semantic feature set of the user behavior activity description text, so as to adopt statistical methods such as information gain, chi-square test or mutual information to evaluate the independence between each user behavior semantic sub-feature and the user behavior result, and by calculating the chi-square value of each user behavior semantic sub-feature, the features with chi-square values ​​higher than a preset threshold are screened out to form a user behavior semantic filtering feature subset. During the filtering process, the tools used, such as the feature selection module of Scikit-learn, can effectively help to quickly calculate the feature score and filter. The goal of this step is to obtain features with strong discrimination ability in different user behavior scenarios, and finally obtain the user behavior semantic filtering feature subset.

[0103] Step S23: performing repeated feature screening processing according to the user behavior semantic embedded feature subset and the user behavior semantic filtered feature subset to obtain a user behavior semantic repeated feature set;

[0104] In an embodiment of the present invention, duplicate features are screened for the previously identified and screened user behavior semantic embedded feature subset and the filtered feature subset. At this time, the two feature subsets are subjected to an intersection operation to identify duplicate features. On this basis, those features that have obvious overlaps in user behavior semantics are further screened out. The Pandas library in Python can be used to perform data frame operations to efficiently process the intersection of feature sets and deduplicate them, and finally obtain a duplicate feature set of user behavior semantics.

[0105] Step S24: using the behavior semantic representativeness calculation formula to perform representativeness score calculation on each user behavior semantic sub-feature in the user behavior semantic repeated feature set, so as to obtain a feature representativeness score value corresponding to each user behavior semantic sub-feature;

[0106] In an embodiment of the present invention, a suitable behavior semantic representativeness calculation formula is formed by combining the corresponding user behavior semantic sub-features in the user behavior semantic repetitive feature set, normalization constants, lower limit of integration time range, upper limit of integration time range, time variable parameters, weight parameters, exponential functions, feature center positions, feature dispersion, semantic feature similarity measurement values ​​and related parameters to perform representative score calculation on each user behavior semantic sub-feature in the user behavior semantic repetitive feature set, so as to ensure that the representative score of each sub-feature is accurately calculated, and finally obtain the feature representativeness score value corresponding to each user behavior semantic sub-feature.

[0107] Among them, the specific calculation formula for the representativeness of behavioral semantics is as follows:

[0108] ;

[0109] In the formula, is the characteristic representativeness score value corresponding to the th sub - feature of user behavioral semantics, is the th sub - feature of user behavioral semantics, is the th sub - feature of user behavioral semantics, is the and are both item - index parameters of user behavioral semantics sub - features, is the total number of user behavioral semantics sub - features, is the normalization constant, is the lower limit of the integration time range, is the upper limit of the integration time range, is the time - variable parameter, is the weight parameter corresponding to the th sub - feature of user behavioral semantics, is the exponential function, is the characteristic center position corresponding to the th sub - feature of user behavioral semantics in the user behavioral semantics repeated feature set, is the characteristic dispersion degree corresponding to the th sub - feature of user behavioral semantics in the user behavioral semantics repeated feature set, is the semantic feature similarity measurement value between the th sub - feature of user behavioral semantics and the th sub - feature of user behavioral semantics at the time point , is the correction coefficient of the characteristic representativeness score value;

[0110] The present invention has obtained a representative calculation formula for behavioral semantics through the use of a specific mathematical model and verification, which is used to calculate the representative score for each user behavioral semantics sub-feature within the user behavioral semantics repetition feature set. By using a normalization constant, this representative calculation formula for behavioral semantics can ensure that the scores of all features are within the same range, thereby avoiding the influence of the order of magnitude difference between features on the final result. This process helps to improve the comparability of scores, making the feature selection process more stable and accurate. By using the weight parameter of the user behavioral semantics sub-feature, it is allowed to weight the score according to the importance of the feature in different contexts. By setting different weights, the degree of emphasis on the feature can be flexibly adjusted, so as to better adapt to the diversity and complexity of user behavior. At the same time, by considering the central position and dispersion degree of the feature, the formula can reflect the clustering situation of the feature and its position in the overall feature space, which enables a more accurate evaluation of the similarity between features, thereby avoiding the selection of redundant features and improving the compactness and effectiveness of the feature set. Considering the similarity between features through the semantic feature similarity metric value can help identify similar features, thereby reducing redundant features. This mechanism helps to improve the efficiency of feature selection, ensuring that only features with unique information are retained, and further enhancing the expressiveness of the model. By introducing a time variable, the dynamic changes of user behavior can be analyzed, which is crucial for capturing the temporal features of user behavior, enabling the model to better adapt to behavior changes in different time periods, and thus improving the accuracy of prediction. In addition, the introduction of a correction coefficient can fine-tune the feature score. Considering the uncertainty or other potential factors in the calculation process, this adjustment can improve the reliability of the score, making the finally obtained comprehensive representative feature set more in line with the actual situation. By comprehensively considering multiple factors such as the weight, central position, dispersion degree, similarity, and time change of the feature, this formula can effectively identify and evaluate the representativeness of user behavioral semantics features. This mechanism not only improves the accuracy and effectiveness of feature selection, but also enhances the prediction ability and reliability of the user behavior analysis model, thereby achieving a deeper understanding of user behavior and providing strong support for personalized services and optimized decision-making. To sum up, this formula fully considers the th user behavioral semantics sub-feature corresponding feature representative score value ,the th user behavioral semantics sub-feature ,the th user behavioral semantics sub-feature ,the item index parameter of the user behavioral semantics sub-feature and ,the total number of user behavioral semantics sub-features ,the normalization constant ,the lower limit of the integration time range , the upper limit of the integration time range , the time variable parameter , the weight parameter corresponding to the th user behavior semantic sub - feature , the exponential function , the feature center position corresponding to the th user behavior semantic sub - feature in the user behavior semantic repetition feature set , the feature dispersion degree corresponding to the th user behavior semantic sub - feature in the user behavior semantic repetition feature set , the semantic feature similarity measurement value between the th user behavior semantic sub - feature and the th user behavior semantic sub - feature at the time point , the correction coefficient of the feature representativeness score value , according to the th user behavior semantic sub - feature corresponding feature representativeness score value and the mutual correlation relationships among the above - mentioned parameters constitute a functional relationship , this formula can implement the calculation process of the representativeness score for each user behavior semantic sub - feature in the user behavior semantic repetition feature set. At the same time, through the introduction of the correction coefficient of the feature representativeness score value, it can be adjusted according to the error situation in the calculation process, so as to improve the accuracy and applicability of the behavior semantic representativeness calculation formula.

[0111] Step S25: Compare and judge the feature representativeness score value corresponding to each user behavior semantic sub - feature according to the preset representativeness score threshold. When the feature representativeness score value is greater than or equal to the preset representativeness score threshold, mark the corresponding user behavior semantic sub - feature as a representative feature; when the feature representativeness score value is less than the preset representativeness score threshold, mark the corresponding user behavior semantic sub - feature as a non - representative feature; screen out and aggregate the user behavior semantic sub - features corresponding to the representative features marked in the user behavior semantic repetition feature set to obtain the user behavior semantic comprehensive representative feature set.

[0112] In an embodiment of the present invention, by using a preset representative scoring threshold to compare and judge the feature representative scoring values corresponding to the user behavior semantic sub-features obtained by previous quantization calculations, a reasonable scoring threshold is set. This value is based on the analysis results of the feature performance in the previous steps to ensure that it can effectively distinguish representative features from non-representative features. Subsequently, the feature representative scoring values corresponding to each user behavior semantic sub-feature are traversed, and through conditional judgment statements, representative features (if the feature representative scoring value is greater than or equal to the preset representative scoring threshold) and non-representative features (if the feature representative scoring value is less than the preset representative scoring threshold) are marked. At the same time, by screening out the user behavior semantic sub-features marked as representative features and integrating them into a dataset, a user behavior semantic comprehensive representative feature set is finally obtained.

[0113] Further, step S21 includes the following steps:

[0114] Perform feature importance evaluation and analysis on each user behavior semantic sub-feature in the user behavior activity description text semantic feature set to obtain the feature importance degree corresponding to each user behavior semantic sub-feature;

[0115] In an embodiment of the present invention, by performing importance evaluation and analysis on each user behavior semantic sub-feature in the user behavior activity description text semantic feature set obtained by previous analysis, by using ensemble learning algorithms such as random forest and gradient boosting tree, by constructing multiple decision trees, calculate the information gain or the reduction of Gini impurity of each feature in the tree model. The feature importance of each user behavior semantic sub-feature can be determined by calculating the frequency of feature selection and the contribution to the model prediction result, so as to quantitatively calculate the feature importance degree of each sub-feature, and finally obtain the feature importance degree corresponding to each user behavior semantic sub-feature.

[0116] Preferably, perform feature sparse quantization calculation on each user behavior semantic sub-feature in the user behavior activity description text semantic feature set to obtain the feature sparsity degree corresponding to each user behavior semantic sub-feature;

[0117] In an embodiment of the present invention, by performing quantization calculation of the feature sparsity degree on each user behavior semantic sub-feature in the user behavior activity description text semantic feature set obtained by previous analysis, by using the Lasso regression method to analyze the sparsity of each feature in the user behavior activity description text. In specific implementation, select an appropriate regularization parameter to limit the feature coefficients through the minimization objective function of the model, so as to quantitatively evaluate the feature sparsity degree. The sparsity degree can be realized by calculating the ratio of the number of non-zero feature coefficients to the total number of features, and finally obtain the feature sparsity degree corresponding to each user behavior semantic sub-feature.

[0118] Preferably, perform L1 regularization statistical analysis on each user behavior semantic sub - feature in the user behavior activity description text semantic feature set according to the feature importance degree and feature sparsity degree corresponding to each user behavior semantic sub - feature, so as to obtain the L1 regularization strength parameter of the user behavior semantic feature;

[0119] In the embodiment of the present invention, perform statistical analysis of the L1 regularization strength on the corresponding user behavior semantic sub - features in the user behavior activity description text semantic feature set by combining the feature importance degree and feature sparsity degree corresponding to the user behavior semantic sub - features obtained by previous quantization calculations. This step constructs a comprehensive model including feature importance and sparsity degree, and performs modeling in the form of L1 regularization, aiming to improve the effect of feature selection. In implementation, first determine the weights of feature importance and sparsity degree, select significant user behavior semantic sub - features by setting appropriate thresholds, then adjust the contribution of features using regularization parameters, and perform linear regression analysis, so as to statistically calculate the corresponding L1 regularization strength parameter, and finally obtain the L1 regularization strength parameter of the user behavior semantic feature.

[0120] Preferably, perform L1 regularization embedded feature selection on each user behavior semantic sub - feature in the user behavior activity description text semantic feature set based on the L1 regularization strength parameter of the user behavior semantic feature, so as to obtain the user behavior semantic embedded feature subset.

[0121] In the embodiment of the present invention, perform L1 regularization embedded feature selection on each user behavior semantic sub - feature in the user behavior activity description text semantic feature set by combining the L1 regularization strength parameter of the user behavior semantic feature obtained by previous quantization calculations. In specific implementation, adopt the Lasso regression model, and process the semantic feature set of the user behavior activity description text through the corresponding L1 regularization strength parameter of the user behavior semantic feature. During the model training process, the coefficients of the features are gradually shrunk to zero, so as to realize the selection and elimination of features. By setting thresholds, features strongly related to user behavior can be effectively screened out, and a subset containing important user behavior semantic features is formed, and finally the user behavior semantic embedded feature subset is selected.

[0122] Further, step S22 includes the following steps:

[0123] Perform statistical calculation on the occurrence times of activity descriptions for each user behavior semantic sub - feature in the user behavior activity description text semantic feature set, so as to obtain the occurrence times of user behavior activity descriptions corresponding to each user behavior semantic sub - feature;

[0124] In the embodiments of the present invention, by extracting each corresponding user behavior semantic sub-feature from the user behavior activity description text semantic feature set, and by using natural language processing (NLP) tools such as NLTK or SpaCy to preprocess the user behavior semantic sub-feature in the corresponding user behavior activity description text, including word segmentation, stop word removal, and lemmatization. Then, using a statistical analysis library (such as Pandas) to count the frequency of each user behavior semantic sub-feature appearing in the text, and recording the occurrence times of each sub-feature through a dictionary or set structure to ensure that all data is accurately recorded and organized, and finally obtaining the occurrence times of the user behavior activity description corresponding to each user behavior semantic sub-feature.

[0125] Preferably, perform feature frequency statistical analysis on each user behavior semantic sub-feature according to the occurrence times of the user behavior activity description corresponding to each user behavior semantic sub-feature, to obtain the user behavior semantic feature frequency corresponding to each user behavior semantic sub-feature;

[0126] In the embodiments of the present invention, perform statistical calculation of the feature frequency on each user behavior semantic sub-feature by combining the occurrence times of the user behavior activity description corresponding to each user behavior semantic sub-feature obtained from the previous statistical analysis, which is achieved by dividing the occurrence times of each sub-feature by the total text length (or total occurrence times), and using numerical calculation libraries such as NumPy or Pandas for mathematical calculations, and finally obtaining the user behavior semantic feature frequency corresponding to each user behavior semantic sub-feature.

[0127] Preferably, perform feature correlation evaluation analysis on each user behavior semantic sub-feature in the user behavior activity description text semantic feature set to obtain the correlation coefficient between each user behavior semantic sub-feature and the user behavior activity description text; set the significance level of the user behavior activity according to the correlation coefficient between each user behavior semantic sub-feature and the user behavior activity description text, so as to obtain the significance level of the user behavior activity semantic feature;

[0128] In an embodiment of the present invention, by using the Pearson correlation coefficient or the Spearman rank correlation coefficient, the statistical calculation of the feature correlation coefficient is performed between each user behavior semantic sub - feature in the user behavior activity description text semantic feature set obtained from the previous analysis and the user behavior activity description text, so as to calculate the correlation coefficients between all user behavior semantic sub - features and the text content through the SciPy library, and display the correlation through a visualization tool (such as Matplotlib or Seaborn), thereby obtaining the correlation coefficient between each user behavior semantic sub - feature and the user behavior activity description text. At the same time, by performing a significance test on the correlation coefficient between the user behavior semantic sub - feature and the user behavior activity description text obtained from the previous quantitative calculation, the feature significance is evaluated by using a t - test or ANOVA analysis, so as to determine the corresponding significance level (for example, 0.05), and finally obtain the user behavior activity semantic feature significance level.

[0129] Preferably, based on the user behavior activity semantic feature significance level and the user behavior semantic feature frequency corresponding to each user behavior semantic sub - feature, a chi - square test calculation is performed on each corresponding user behavior semantic sub - feature to obtain the semantic feature chi - square value corresponding to each user behavior semantic sub - feature;

[0130] In an embodiment of the present invention, by combining the user behavior activity semantic feature significance level obtained from the previous analysis and the user behavior semantic feature frequency corresponding to the user behavior semantic sub - feature, a chi - square test of the sub - feature is performed on the corresponding user behavior semantic sub - feature. By constructing a contingency table for the chi - square test, recording the observed frequency and expected frequency of each sub - feature, and using the chi2_contingency function in SciPy for the chi - square test, the corresponding chi - square value is calculated. This process ensures that each user behavior semantic sub - feature undergoes a strict statistical test, and the obtained chi - square value will be recorded, and finally the semantic feature chi - square value corresponding to each user behavior semantic sub - feature is obtained.

[0131] Preferably, based on the semantic feature chi - square value corresponding to each user behavior semantic sub - feature, a filter - based feature selection process is performed on each corresponding user behavior semantic sub - feature in the user behavior activity description text semantic feature set to obtain a user behavior semantic filter - based feature subset.

[0132] In an embodiment of the present invention, a filtered feature selection is performed on each corresponding user behavior semantic sub-feature in the semantic feature set of the user behavior activity description text by combining the semantic feature chi-square value corresponding to each user behavior semantic sub-feature obtained by previous quantitative calculation, so as to set a chi-square threshold in advance. If the chi-square value is greater than or equal to the set chi-square threshold, the corresponding user behavior semantic sub-feature is filtered out to obtain sub-features with higher significance. This process uses Pandas to perform filtering operations on the data frame, and finally obtains the user behavior semantic filtered feature subset.

[0133] Further, step S3 includes the following steps:

[0134] Step S31: extracting context information from each user behavior semantics representative sub-feature in the user behavior semantics comprehensive representative feature set to obtain user behavior semantics context information corresponding to each user behavior semantics representative sub-feature;

[0135] In an embodiment of the present invention, context information is extracted from each user behavior semantic representative sub-feature in the user behavior semantic comprehensive representative feature set obtained by previous analysis. These sub-features include user clicks, search terms, purchase records, etc., and natural language processing (NLP) technology is used to extract context information related to these sub-features using methods such as word segmentation, part-of-speech tagging and named entity recognition. For example, for "the user clicked on the link of product A", keywords such as "user", "click", "product A" and "link" can be extracted, and then context information about the behavior is generated, including time, place, device type and other information. This process uses tools such as spaCy or NLTK, which can efficiently process text data and extract meaningful information, and finally extract the user behavior semantic context information corresponding to each user behavior semantic representative sub-feature.

[0136] Step S32: extracting sub-feature activity attributes from each corresponding user behavior semantic representative sub-feature based on the user behavior semantic context information corresponding to each user behavior semantic representative sub-feature, and obtaining the user behavior activity attributes corresponding to each user behavior semantic representative sub-feature;

[0137] In an embodiment of the present invention, by combining the user behavior semantic context information corresponding to each previously extracted user behavior semantic representative sub - feature, the extraction of the sub - feature activity attributes is performed on each corresponding user behavior semantic representative sub - feature. Taking the behavior of "a user clicks on the link of product A" as an example, the activity attributes may include "product type", "click frequency", "user group", etc. Then, clustering analysis and classification algorithms (such as K - means or decision tree) are used to analyze the extracted context information to summarize the activity attributes of each sub - feature. By comparing historical data and combining user portraits, it is ensured that the extracted activity attributes are representative and reliable. This process can also use tools such as Scikit - learn or TensorFlow to implement data mining and feature engineering, and finally obtain the user behavior activity attributes corresponding to each user behavior semantic representative sub - feature.

[0138] Step S33: Based on the user behavior activity attributes corresponding to each user behavior semantic representative sub - feature, perform behavior activity feature classification on each corresponding user behavior semantic representative sub - feature in the user behavior semantic comprehensive representative feature set, so as to obtain the behavior semantic representative feature subsets corresponding to each user behavior activity.

[0139] In an embodiment of the present invention, through combining the user behavior activity attributes corresponding to each previously mined and analyzed user behavior semantic representative sub - feature, the classification process of behavior activity features is performed on each corresponding user behavior semantic representative sub - feature in the user behavior semantic comprehensive representative feature set. In this step, a machine learning model (such as support vector machine SVM or random forest) can be used to classify user behavior activities. The activity attributes of each sub - feature are used as input features. After the model is trained, it can identify the categories of different user behavior activities, such as "browsing", "purchasing", "commenting", etc. To ensure classification accuracy, a cross - validation method is used to evaluate the performance of the model, and indicators such as accuracy, recall rate, and F1 - score are used to optimize the model. Finally, the behavior semantic representative feature subsets corresponding to each user behavior activity are obtained through classification.

[0140] Step S34: Perform a local interpretability evaluation and analysis on the behavior semantic representative feature subsets corresponding to each user behavior activity to obtain the user behavior decision local interpretability factors corresponding to each user behavior activity.

[0141] In the embodiments of the present invention, by performing a local interpretability evaluation and analysis on each user behavior activity and its corresponding subset of representative behavioral semantic features, in this process, a local interpretable model-agnostic interpretation method (such as LIME or SHAP) is used to identify the key factors affecting user decisions. By generating a local model to analyze the relationship between user behavior activity features and decision results, interpretable factors for each decision can be provided. For example, if a user chooses to purchase a product, LIME can help identify the roles played by features such as "discount" and "product evaluation" in the decision-making process. This process needs to be combined with data visualization tools (such as Matplotlib or Tableau) to display the results of the interpretability analysis, making the local interpretable factors of the decision clear at a glance, that is, the degree of influence of the corresponding local influencing factors in the user decision-making process, and finally obtaining the local interpretable factors of user behavior decisions corresponding to each user behavior activity.

[0142] Step S35: Perform AI-driven intelligent parsing and analysis on the local interpretable factors of user behavior decisions corresponding to each user behavior activity to obtain the parsing results of the interpretable factors of user behavior decisions corresponding to each user behavior activity.

[0143] In the embodiments of the present invention, through AI-driven intelligent parsing and analysis of the local interpretable factors of user behavior decisions corresponding to each user behavior activity obtained after previous evaluation and analysis, this step will combine deep learning techniques (such as deep neural networks) for factor parsing. By using the neural network model to perform more in-depth feature learning on the local interpretable factors, through the model's non-linear mapping of the relationships between factors, more complex decision-making logics are extracted. After training the model, corresponding parsing results of interpretable factors for decisions can be generated. For example, the model will indicate that a certain user is more influenced by "social media influence" and "evaluation of similar products" when making a purchase. This process relies on deep learning frameworks (such as TensorFlow or PyTorch) and parsing tools to ensure that the final parsing results are highly reliable and practical, and finally obtaining the parsing results of the interpretable factors of user behavior decisions corresponding to each user behavior activity.

[0144] Furthermore, step S34 includes the following steps:

[0145] Step S341: Perform user behavior decision semantic association analysis on each behavioral semantic representative sub-feature within the subset of representative behavioral semantic features corresponding to each user behavior activity to obtain the semantic association relationship between the representative sub-features of behavioral semantics corresponding to each user behavior activity and user behavior decisions;

[0146] In an embodiment of the present invention, semantic association analysis of behavioral decisions is performed on each behavioral semantic representative sub - feature within the subset of behavioral semantic representative features corresponding to each previously screened user behavioral activity to extract user behavioral data, including browsing history, purchase records, search queries, etc. And by using natural language processing (NLP) techniques, these user behavioral data are converted into behavioral semantic features. For example, through keyword extraction and sentiment analysis, the potential intentions behind user behaviors are identified. At the same time, statistical methods such as correlation analysis or regression analysis are applied to explore the association relationship between behavioral semantic features and user decisions. This process can be implemented using the Python programming language and its related data analysis libraries (such as Pandas and Scikit - learn), and finally, the semantic association relationship between the behavioral semantic representative sub - features corresponding to each user behavioral activity and user behavioral decisions is obtained.

[0147] Step S342: Perform an analysis of local influencing factors of behavioral decisions on the subset of behavioral semantic representative features corresponding to each user behavioral activity based on the semantic association relationship between the behavioral semantic representative sub - features corresponding to each user behavioral activity and user behavioral decisions, to obtain a candidate set of local influencing factors of user behavioral decisions corresponding to each user behavioral activity;

[0148] In an embodiment of the present invention, by combining the semantic association relationship between the behavioral semantic representative sub - features corresponding to each user behavioral activity and user behavioral decisions obtained from the previous analysis, an identification analysis of local influencing factors of behavioral decisions is performed on the corresponding subset of behavioral semantic representative features, so as to analyze the influence degree of each behavioral feature on the decision result by constructing a decision tree model or using techniques such as SHAP (SHapley Additive exPlanations) values. At this time, the tool selection can include the R language or the Shap library in Python to provide interpretable results. In specific operations, first, the user behavioral activity data set is divided into a training set and a test set to train the model and evaluate the local influencing factors of each feature, thereby forming a corresponding candidate set of local influencing factors, and finally, a candidate set of local influencing factors of user behavioral decisions corresponding to each user behavioral activity is obtained.

[0149] Step S343: Perform an evaluation analysis of user behavioral performance on each behavioral semantic representative sub - feature within the subset of behavioral semantic representative features corresponding to each user behavioral activity, to obtain a user behavioral activity performance score for each behavioral semantic representative sub - feature corresponding to each user behavioral activity;

[0150] In an embodiment of the present invention, by evaluating and analyzing the user behavior performance for each behavioral semantic representative sub - feature within the subset of behavioral semantic representative features corresponding to each previously screened user behavior activity, and by adopting multi - dimensional evaluation metrics such as conversion rate, user retention rate, and user satisfaction, etc., and by setting a series of criteria, applying these metrics to different user behavior activities. For example, using A / B testing to evaluate the impact of different strategies on user behavior performance, and at the same time applying statistical methods such as analysis of variance (ANOVA) or t - test, which can be used to compare the behavioral performance scores of different user groups to reflect the effectiveness and influence of each behavioral semantic representative sub - feature, and finally obtaining the user behavior activity performance scores for each behavioral semantic representative sub - feature corresponding to each user behavior activity.

[0151] Step S344: Design a local interpretability framework for the subset of behavioral semantic representative features corresponding to each user behavior activity based on the user behavior activity performance scores for each behavioral semantic representative sub - feature corresponding to each user behavior activity, so as to generate a local interpretability evaluation and analysis model for each user behavior activity;

[0152] In an embodiment of the present invention, by combining the user behavior activity performance scores for each behavioral semantic representative sub - feature corresponding to each user behavior activity obtained from the previous evaluation and calculation, design a local interpretability model framework for the corresponding subset of behavioral semantic representative features. By integrating interpretable tools such as LIME (Local Interpretable Model - agnostic Explanations) or SHAP, and combining the corresponding user behavior activity performance scores, conduct an interpretability evaluation for the subset of behavioral semantic representative features corresponding to each user behavior activity, so as to construct a local model that fits around each user behavior activity to better understand the impact of features on decisions, and by using graphical tools (such as Matplotlib or Seaborn) to display the contributions of each feature, that is, the user behavior activity performance scores, for easy analysis and interpretation. This framework will generate a corresponding local interpretability evaluation and analysis model, aiming to make the decision - making process transparent and easy to understand, and finally design and generate a local interpretability evaluation and analysis model for each user behavior activity.

[0153] Step S345: Conduct a local interpretability evaluation and analysis on the candidate set of local influencing factors of user behavior decisions corresponding to each user behavior activity according to the local interpretability evaluation and analysis model corresponding to each user behavior activity, so as to obtain the local interpretability factors of user behavior decisions corresponding to each user behavior activity.

[0154] In the embodiments of the present invention, through the local interpretability evaluation and analysis models corresponding to each user behavior activity generated by combining previous designs, a local interpretability evaluation and analysis is carried out on the candidate set of local influencing factors corresponding to the user behavior decision, so as to quantitatively analyze each local influencing factor in the candidate set by using the above LIME or SHAP method within the local interpretability framework to determine its role and importance in different user behavior activities. Specifically, regression analysis is used to calibrate the prediction results of the local model, and further analyze the performance of each local influencing factor in the process of influencing user decisions. In terms of operation, relevant data visualization tools in Python are used to generate a factor importance ranking diagram to intuitively display the contribution degree of each local influencing factor to the user behavior decision, and finally, the local interpretability factors corresponding to each user behavior activity are obtained.

[0155] Further, step S4 includes the following steps:

[0156] Step S41: Conduct a statistical analysis of the factor correlation degree on the user behavior decision interpretability factors corresponding to each user behavior activity in the user behavior decision interpretability factor analysis results, so as to obtain the correlation degree between the decision interpretability factors corresponding to each user behavior activity;

[0157] In the embodiments of the present invention, a statistical analysis of the correlation degree is carried out on the user behavior decision interpretability factors corresponding to each user behavior activity by using the Pearson correlation coefficient or the Spearman rank correlation coefficient, so as to input the collected user behavior decision interpretability factors into statistical software or data analysis tools, such as the pandas library in Python or the psych package in the R language, and use the correlation calculation function for calculation to obtain the correlation matrix of factors. This matrix clearly shows the strength and direction between the decision interpretability factors corresponding to each user behavior activity. Further, a visualization tool (such as Matplotlib or Tableau) is used to display the correlation matrix in the form of a heat map to facilitate observing the potential relationships and mutual influences between factors, and finally, the correlation degree between the decision interpretability factors corresponding to each user behavior activity is obtained.

[0158] Step S42: Based on the correlation degree between the decision interpretability factors corresponding to each user behavior activity, construct a decision logic relationship network for the corresponding user behavior decision interpretability factors, so as to generate the user behavior decision logic relationship network structure corresponding to each user behavior activity;

[0159] In the embodiments of the present invention, by combining the correlation degrees between the decision interpretability factors corresponding to each user behavior activity obtained from the previous quantitative calculation, a logical relationship network is constructed for connecting the corresponding user behavior decision interpretability factors. By using the concepts of nodes and edges in graph theory, the user behavior decision interpretability factors are regarded as nodes, and the correlation degrees between the factors are used as the weights of the edges. With the help of network analysis tools such as Gephi or NetworkX, a directed graph containing all relevant factors is established, and the weights of the edges are set to reflect the intensity of their correlation degrees. Then, by applying network analysis algorithms (such as the Dijkstra algorithm or the PageRank algorithm), important factors and their mutual relationships are identified, and the topological structure of the decision logical relationship network is generated. This network structure not only reflects the direct relationships between the factors but also reveals the potential influence paths of the factors in the decision-making process. Finally, the user behavior decision logical relationship network structures corresponding to each user behavior activity are constructed and generated.

[0160] Step S43: Based on the user behavior decision logical relationship network structures corresponding to each user behavior activity, a logical reason visualization display process is performed on the analysis results of the user behavior decision interpretability factors corresponding to each user behavior activity, so as to obtain the interpretable logical reasons behind each user behavior activity decision.

[0161] In the embodiments of the present invention, by combining the user behavior decision logical relationship network structures corresponding to each user behavior activity obtained from the previous connection construction, a visual display of the logical reasons is performed on the analysis results of the corresponding user behavior decision interpretability factors, so as to more intuitively understand the decision logical reasons behind the user behavior activities. By selecting key factors based on the constructed decision logical relationship network and presenting their data in a graphical manner, interactive visualization graphs can be created using chart tools (such as D3.js or Plotly) to display each factor and the relationships and influences between them. By setting interactive functions, users can click on a certain factor to view its specific influence on the decision and the logical path, and then analyze the importance of the factor in the decision-making process and its logical interpretation reasons, thereby generating a corresponding decision logical display report. The report will include a detailed analysis of the relationships between the factors, an influence assessment, and a comprehensive interpretation of their impact on the user behavior decision. Finally, the interpretable logical reasons behind each user behavior activity decision are obtained.

[0162] Furthermore, the present invention also provides a big data intelligent analysis system driven by AI for executing the big data intelligent analysis method driven by AI as described above. The big data intelligent analysis system driven by AI includes:

[0163] The user behavior semantic feature analysis module is used to obtain a large dataset of user behavior activity description texts and perform behavioral text semantic feature analysis on the large dataset of user behavior activity description texts, so as to obtain a set of user behavior activity description text semantic features;

[0164] The user behavior semantic repetition representative screening module is used to perform embedded and filtering feature selection processing on the set of user behavior activity description text semantic features to obtain a user behavior semantic embedded feature subset and a user behavior semantic filtering feature subset; perform repeated representative feature screening processing based on the user behavior semantic embedded feature subset and the user behavior semantic filtering feature subset, so as to obtain a set of user behavior semantic comprehensive representative features;

[0165] The behavior activity AI-driven intelligent analysis module is used to classify the behavior activity features of the set of user behavior semantic comprehensive representative features to obtain a subset of behavior semantic representative features corresponding to each user behavior activity; perform local interpretable AI intelligent analysis on the subset of behavior semantic representative features corresponding to each user behavior activity to obtain the analysis results of user behavior decision interpretable factors corresponding to each user behavior activity;

[0166] The behavior decision analysis reason visualization module is used to perform logical reason visualization display processing on the analysis results of user behavior decision interpretable factors corresponding to each user behavior activity to obtain the interpretable logical reasons behind the decisions of each user behavior activity.

[0167] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A big data intelligent parsing and analysis method based on AI, characterized in that: The following steps are involved: Step S1: obtaining a large dataset of user behavior activity description texts, and performing behavior text semantic feature analysis on the large dataset of user behavior activity description texts to obtain a user behavior activity description text semantic feature set; Step S2: performing embedded and filtered feature selection processing on the semantic feature set of the user behavior activity description text to obtain a user behavior semantic embedded feature subset and a user behavior semantic filtered feature subset; performing repeated representative feature screening processing on the user behavior semantic embedded feature subset and the user behavior semantic filtered feature subset to obtain a user behavior semantic comprehensive representative feature set; Step S3: classifying the behavior activity features of the user behavior semantics comprehensive representative feature set to obtain a behavior semantics representative feature subset corresponding to each user behavior activity; Perform local interpretable AI intelligent analysis on the behavioral semantic representative feature subsets corresponding to each user behavior activity to obtain the interpretable factor analysis results of user behavior decisions corresponding to each user behavior activity; Step S4: Perform a logical reason visualization processing on the user behavior decision explainability factor analysis results corresponding to each user behavior activity to obtain the explainable logical reasons behind each user behavior activity decision.

2. The AI-driven big data intelligent parsing and analysis method according to claim 1 is characterized in that: Step S1 includes the following steps: Step S11: obtaining a large data set of user behavior activity description texts, wherein the large data set of user behavior activity description texts includes user behavior activity text record data from user behavior data sources related to social media, mobile applications, online platforms, and e-commerce websites; Step S12: performing sensitivity elimination processing on the user behavior activity description text big data set to obtain the user behavior activity text sensitive word elimination big data set; performing stop word removal and text normalization processing on the user behavior activity text sensitive word elimination big data set to obtain the standardized user behavior activity text big data set; Step S13: extracting semantic atomic units from the standardized user behavior activity text big data set to obtain a set of semantic component atomic units of the user behavior activity text; Step S14: performing text semantic relationship mining and analysis on each text basic semantic component atomic unit in the user behavior activity text semantic component atomic unit set to obtain the user behavior activity text semantic relationship between each text basic semantic component atomic unit; constructing a semantic relationship graph for each corresponding text basic semantic component atomic unit based on the user behavior activity text semantic relationship between each text basic semantic component atomic unit to generate a user behavior activity text semantic relationship graph; Step S15: Based on the user behavior activity text semantic relationship graph, semantic feature recognition analysis is performed on the corresponding user behavior activity description text in the standardized user behavior activity text big data set to obtain a semantic feature set of the user behavior activity description text.

3. The AI-driven big data intelligent parsing and analysis method according to claim 2 is characterized in that: The step S14 of constructing a semantic relationship graph for each corresponding text basic semantic component atomic unit based on the user behavior activity text semantic relationship between each text basic semantic component atomic unit includes the following steps: Perform text frequency and reverse document frequency statistics calculation on each text basic semantic component atomic unit in the user behavior activity text semantic component atomic unit set, so as to obtain the text frequency and reverse document frequency corresponding to each text basic semantic component atomic unit; Based on the text frequency and reverse document frequency corresponding to each text basic semantic component atomic unit, the semantic relationship weight distribution calculation of the user behavior activity text semantic relationship between each text basic semantic component atomic unit is performed to obtain the text semantic relationship distribution weight between each text basic semantic component atomic unit; According to the text semantic relationship allocation weights between the basic semantic component atomic units of each text, the user behavior activity text semantic relationship between the basic semantic component atomic units of each text is determined by semantic relationship weight mapping, so as to obtain the user activity text semantic relationship weight mapping rules between the basic semantic component atomic units of each text; Based on the user activity text semantic relationship weight mapping rules between each text basic semantic component atomic unit, the user behavior activity text semantic relationship between each text basic semantic component atomic unit and each corresponding text basic semantic component atomic unit are constructed to generate a user behavior activity text semantic relationship map.

4. The AI-driven big data intelligent parsing and analysis method according to claim 1 is characterized in that: Step S2 includes the following steps: Step S21: performing embedded feature selection processing on the semantic feature set of the user behavior activity description text to obtain a user behavior semantic embedded feature subset; Step S22: performing filtering feature selection processing on the semantic feature set of the user behavior activity description text to obtain a user behavior semantic filtering feature subset; Step S23: performing repeated feature screening processing according to the user behavior semantic embedded feature subset and the user behavior semantic filtered feature subset to obtain a user behavior semantic repeated feature set; Step S24: using the behavior semantic representativeness calculation formula to perform representativeness score calculation on each user behavior semantic sub-feature in the user behavior semantic repeated feature set, so as to obtain a feature representativeness score value corresponding to each user behavior semantic sub-feature; The specific calculation formula for the representativeness of behavioral semantics is: ; In the formula, For the User behavior semantic sub-features The corresponding feature representative score value, For the User behavior semantic sub-features, For the User behavior semantic sub-features, and are all item index parameters of user behavior semantic sub-features, is the total number of semantic sub-features of user behavior, is the normalization constant, is the lower limit of the integration time range, is the upper limit of the integration time range, is the time variable parameter, For the The weight parameter corresponding to the semantic sub-feature of user behavior, is an exponential function, For the The feature center position corresponding to each user behavior semantic sub-feature in the user behavior semantic repeated feature set, For the The degree of feature dispersion corresponding to each user behavior semantic sub-feature in the user behavior semantic repeated feature set, For the The semantic sub-features of user behavior and User behavior semantic sub-features at time point The semantic feature similarity measure at is the correction factor of the characteristic representativeness score; Step S25: Compare and judge the feature representativeness score value corresponding to each user behavior semantic sub-feature according to the preset representativeness score threshold; when the feature representativeness score value is greater than or equal to the preset representativeness score threshold, the corresponding user behavior semantic sub-feature is marked as a representative feature; when the feature representativeness score value is less than the preset representativeness score threshold, the corresponding user behavior semantic sub-feature is marked as a non-representative feature; the user behavior semantic sub-features marked as representative features in the user behavior semantic repeated feature set are screened out and combined to obtain a comprehensive representative feature set of user behavior semantics.

5. The AI-driven big data intelligent parsing and analysis method according to claim 4 is characterized in that: Step S21 includes the following steps: Perform feature importance evaluation and analysis on each user behavior semantic sub-feature in the semantic feature set of the user behavior activity description text to obtain the feature importance degree corresponding to each user behavior semantic sub-feature; Perform feature sparsity quantification calculation on each user behavior semantic sub-feature in the semantic feature set of the user behavior activity description text to obtain the feature sparsity degree corresponding to each user behavior semantic sub-feature; According to the feature importance and feature sparsity of each user behavior semantic sub-feature, L1 regularization statistical analysis is performed on each user behavior semantic sub-feature corresponding to the user behavior activity description text semantic feature set to obtain the user behavior semantic feature L1 regularization strength parameter; Based on the L1 regularization strength parameter of the user behavior semantic feature, L1 regularized embedded feature selection is performed on each corresponding user behavior semantic sub-feature in the semantic feature set of the user behavior activity description text to obtain the user behavior semantic embedded feature subset.

6. The AI-driven big data intelligent parsing and analysis method according to claim 4 is characterized in that: Step S22 includes the following steps: Performing statistical calculations on the number of occurrences of activity descriptions for each user behavior semantic sub-feature in the user behavior activity description text semantic feature set, and obtaining the number of occurrences of user behavior activity descriptions corresponding to each user behavior semantic sub-feature; Perform feature frequency statistical analysis on each corresponding user behavior semantic sub-feature according to the number of occurrences of the user behavior activity description corresponding to each user behavior semantic sub-feature, and obtain the user behavior semantic feature frequency corresponding to each user behavior semantic sub-feature; Performing feature correlation evaluation and analysis on each user behavior semantic sub-feature in the user behavior activity description text semantic feature set to obtain the correlation coefficient between each user behavior semantic sub-feature and the user behavior activity description text; setting the user behavior activity significance level according to the correlation coefficient between each user behavior semantic sub-feature and the user behavior activity description text to obtain the user behavior activity semantic feature significance level; Based on the significance level of the semantic feature of the user behavior activity and the user behavior semantic feature frequency corresponding to each user behavior semantic sub-feature, a sub-feature chi-square test is performed on each corresponding user behavior semantic sub-feature to obtain the semantic feature chi-square value corresponding to each user behavior semantic sub-feature; Based on the semantic feature chi-square value corresponding to each user behavior semantic sub-feature, a filtering feature selection process is performed on each corresponding user behavior semantic sub-feature in the semantic feature set of the user behavior activity description text to obtain a user behavior semantic filtering feature subset.

7. The AI-driven big data intelligent parsing and analysis method according to claim 1 is characterized in that: Step S3 includes the following steps: Step S31: extracting context information from each user behavior semantics representative sub-feature in the user behavior semantics comprehensive representative feature set to obtain user behavior semantics context information corresponding to each user behavior semantics representative sub-feature; Step S32: extracting sub-feature activity attributes from each corresponding user behavior semantic representative sub-feature based on the user behavior semantic context information corresponding to each user behavior semantic representative sub-feature, and obtaining the user behavior activity attributes corresponding to each user behavior semantic representative sub-feature; Step S33: classifying each corresponding user behavior semantics representative sub-feature in the user behavior semantics comprehensive representative feature set based on the user behavior activity attribute corresponding to each user behavior semantics representative sub-feature, so as to obtain a behavior semantics representative feature subset corresponding to each user behavior activity; Step S34: performing local interpretability evaluation and analysis on the behavioral semantic representative feature subset corresponding to each user behavior activity to obtain a local interpretability factor of the user behavior decision corresponding to each user behavior activity; Step S35: Perform AI-driven intelligent parsing analysis on the local explainability factors of user behavior decisions corresponding to each user behavior activity to obtain the parsing results of the explainability factors of user behavior decisions corresponding to each user behavior activity.

8. The AI-driven big data intelligent parsing and analysis method according to claim 7 is characterized in that: Step S34 includes the following steps: Step S341: Performing user behavior decision semantic association analysis on each behavior semantic representative sub-feature in the behavior semantic representative feature subset corresponding to each user behavior activity, to obtain the semantic association relationship between the behavior semantic representative sub-feature corresponding to each user behavior activity and the user behavior decision; Step S342: Based on the semantic association relationship between the behavioral semantic representative sub-features corresponding to each user behavior activity and the user behavior decision, a behavior decision local influencing factor analysis is performed on the behavior semantic representative feature subset corresponding to each user behavior activity to obtain a user behavior decision local influencing factor candidate set corresponding to each user behavior activity; Step S343: performing user behavior performance evaluation and analysis on each behavior semantics representative sub-feature in the behavior semantics representative feature subset corresponding to each user behavior activity, and obtaining a user behavior activity performance score corresponding to each behavior semantics representative sub-feature of each user behavior activity; Step S344: Based on the user behavior activity performance score corresponding to each behavior semantic representative sub-feature of each user behavior activity, a local interpretability framework design is performed on the behavior semantic representative feature subset corresponding to each user behavior activity to generate a local interpretability evaluation and analysis model corresponding to each user behavior activity; Step S345: Perform local interpretability evaluation and analysis on the candidate set of local influencing factors of user behavior decisions corresponding to each user behavior activity according to the local interpretability evaluation and analysis model corresponding to each user behavior activity, and obtain the local interpretability factor of user behavior decision corresponding to each user behavior activity.

9. The AI-driven big data intelligent parsing and analysis method according to claim 1 is characterized in that: Step S4 includes the following steps: Step S41: performing factor correlation statistical analysis on the corresponding user behavior decision interpretability factors in the user behavior decision interpretability factor analysis results corresponding to each user behavior activity, and obtaining the correlation between the decision interpretability factors corresponding to each user behavior activity; Step S42: constructing a decision logic relationship network for the corresponding user behavior decision interpretability factors based on the correlation between the decision interpretability factors corresponding to each user behavior activity, so as to generate a user behavior decision logic relationship network structure corresponding to each user behavior activity; Step S43: Based on the user behavior decision logical relationship network structure corresponding to each user behavior activity, the user behavior decision interpretability factor analysis results corresponding to each user behavior activity are processed for logical reason visualization to obtain the interpretable logical reasons behind each user behavior activity decision.

10. An AI-driven big data intelligent parsing and analysis system, characterized in that: Used to execute the AI-driven big data intelligent parsing and analysis method as claimed in claim 1, the AI-driven big data intelligent parsing and analysis system comprises: A user behavior semantic feature analysis module is used to obtain a large data set of user behavior activity description texts, and perform behavior text semantic feature analysis on the large data set of user behavior activity description texts, thereby obtaining a user behavior activity description text semantic feature set; The user behavior semantic repeated representative screening module is used to perform embedded and filtered feature selection processing on the semantic feature set of the user behavior activity description text to obtain the user behavior semantic embedded feature subset and the user behavior semantic filtered feature subset; perform repeated representative feature screening processing based on the user behavior semantic embedded feature subset and the user behavior semantic filtered feature subset, so as to obtain the user behavior semantic comprehensive representative feature set; The behavior activity AI-driven intelligent analysis module is used to classify the behavior activity features of the user behavior semantics comprehensive representative feature set to obtain the behavior semantics representative feature subset corresponding to each user behavior activity; perform local interpretable AI intelligent analysis on the behavior semantics representative feature subset corresponding to each user behavior activity to obtain the user behavior decision interpretability factor analysis result corresponding to each user behavior activity; The behavior decision analysis reason visualization module is used to visualize the logical reasons of the user behavior decision explainability factor analysis results corresponding to each user behavior activity, so as to obtain the explainable logical reasons behind each user behavior activity decision.

Citation Information

Patent Citations

  • Semiconductor data analysis and visualization method and device based on large model

    CN119149708A

  • Method for AI model maternal and infant assistant from question answering to accurate recommendation

    CN119539921A