Audit management method based on multi-source data fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明所解决的技术问题在于提供一种基于多源数据融合的审计管理方法,能够通过采集审计用户的操作行为日志构建动态用户画像,并基于相似用户群体的协同过滤预测生成个性化的驾驶舱布局,解决现有技术无法适应不同岗位个性化需求以及无法预测用户潜在关注点的问题
[0031] Based on the normalized attention probability list, the cockpit area is divided into three display areas: core, secondary, and auxiliary, with a probability threshold range configured for each area. Then, display areas are allocated according to the probability value range of each dimension. Within the same area, the layout order and area proportion of charts are determined by probability, with higher probability charts occupying larger areas and appearing at the front. Finally, corresponding components are called from the chart component library, dynamically rendered using a grid layout or flexible layout algorithm, and pushed to the front end for display. The advantage of this step is that it achieves a complete mapping from predicted probability to physical layout. Existing cockpits often use a fixed grid arrangement, which cannot reflect the differences in the importance of data. This solution, through area division and area allocation, makes the most important data visually prominent, conforming to visual perception principles. For example, the distribution chart of the projects with the highest attention probability is placed in the core area and occupies a larger area, making it easily visible to users, while auxiliary data with lower probabilities are placed in the corner areas. This dynamic layout ensures the prominent display of key information, makes full use of screen space, improves user experience and data acquisition efficiency, and the entire layout process is fully automated, requiring no manual intervention.
Smart Images

Figure CN122550124A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audit data display and management technology, specifically to an audit management system and method based on multi-source data fusion. Background Technology
[0002] Current audit dashboard technologies primarily focus on basic data visualization and simple rule-based alerts. According to industry surveys, over 70% of large enterprises have invested heavily in digital transformation over the past two years, but less than 30% have truly achieved intelligent business operations and efficient management loops. Many managers are still struggling with issues such as outdated reports, data silos, and insufficient analytical capabilities. The core function of existing audit dashboards typically involves displaying key indicators such as the number of audit projects, issue distribution, and rectification effectiveness using fixed chart templates, resulting in a standardized interface that is indistinguishable from the user's. While this design initially met the basic needs of audit data aggregation, its limitations have become increasingly apparent as audit operations become more complex and management levels become more differentiated.
[0003] The main shortcomings of existing technologies are as follows. First, the data collection dimensions are too broad. Traditional audit dashboards typically only record page views or simple clicks, failing to distinguish between quick browsing and in-depth analysis, let alone capture the subtle differences in user attention implied by fine-grained operations such as dragging and adjusting charts, hovering to view details, or drilling down to explore details. This leaves the system's understanding of the user's true intent at a superficial level, making it difficult to build accurate user behavior profiles. Second, personalized recommendations are lacking. Existing dashboard layouts are mostly statically configured or rely on manual drag-and-drop adjustments by users. The system lacks proactive learning and predictive capabilities, and cannot dynamically adjust the displayed content based on user job characteristics and behavioral habits. When audit leaders need to focus on new areas or cross-departmental data, they can only manually search or rely on technical personnel to modify configurations, resulting in low efficiency and an inability to respond promptly to management needs.
[0004] A deeper technological bias exists within the traditional approach to user profiling. Current mainstream technologies pursue precise characterization of individual users' historical behavior, believing that the closer the profile is to the user's individual characteristics, the better, and all recommendations should be based on the user's own historical data. This technological approach implicitly assumes that a user's future needs can be found in their past behavior. However, in audit management practice, this assumption has significant flaws. Audit leaders' focus is dynamically shifting. With adjustments in national policies, changes in organizational business, or job rotations, leaders may need to focus on entirely new data areas, and these new needs are not recorded in their personal historical behavior logs. If the concept of precise profiling is adhered to, and predictions are made solely based on the user's own data, the system will be unable to cope with this shift in needs, leading to the failure of personalized recommendations. Furthermore, existing dashboards generally lack mechanisms for transferring collective wisdom. Each user's learning experience is confined to the individual level, unable to achieve knowledge sharing and experience transfer through the system. When a leader discovers the important value of a certain data dimension, other leaders in similar positions cannot benefit from it, creating information silos within the organization. Summary of the Invention
[0005] The technical problem solved by this invention is to provide an audit management method based on multi-source data fusion, which can construct dynamic user profiles by collecting the operation behavior logs of audit users, and generate personalized cockpit layouts based on collaborative filtering prediction of similar user groups, thus solving the problems that existing technologies cannot adapt to the personalized needs of different positions and cannot predict users' potential concerns.
[0006] The basic solution provided by this invention is an audit management method based on multi-source data fusion, comprising the following steps: S1. Obtain the historical operation behavior logs of each audit user on the front-end dashboard interface. The historical operation behavior logs include the user ID, operation timestamp, operation type and operation object of each audit user. The operation type includes one or more of clicking, dragging, pausing and drilling down. The operation object includes the visual chart components in the dashboard and the corresponding data dimensions. S2. Parse the historical operation behavior logs and construct a dynamic user profile for each audit user based on the parsing results. The dynamic user profile includes static attribute tags and dynamic behavior tags for the audit user. The static attribute tags are generated based on the basic information associated with the user identifier. The dynamic behavior tags are generated based on frequency statistics and sequence analysis of operation type and operation object, and are used to represent the user's attention weight to different data dimensions. The dynamic user profile is associated with the user identifier. S3. Respond to the login request of the target audit user, obtain the user identifier of the currently logged-in target audit user and retrieve its corresponding dynamic user profile. Calculate the similarity between its dynamic user profile and the dynamic user profiles of other audit users, and filter out several other audit users whose similarity to the target audit user exceeds a preset threshold to form a similar user group. S4. Obtain the group behavior preference data of the similar user group within a preset historical time period, and predict the probability of the target audit user's attention to different data dimensions in the current login session using a collaborative filtering prediction model based on user similarity. S5. Based on the attention probability, call the corresponding visualization chart component from the preset chart component library, and dynamically generate the personalized cockpit layout for the target audit user according to the probability value; wherein, the data dimension with the highest attention probability and its corresponding chart component are placed in the core display area of the cockpit.
[0007] The principle and advantages of this invention are as follows: By collecting historical logs of user operations on the cockpit interface, including user identifier, operation timestamp, operation type, and operation object, a dynamic profile of each user is constructed. This profile integrates the user's static attributes, such as department and position, with dynamic behavioral tags derived from operation frequency and sequence analysis, enabling the quantification of the user's attention to different data dimensions. When a user logs into the system, the method calculates the similarity between their profile and other users, filters out similar user groups, and uses a collaborative filtering model to predict the data dimensions and probabilities that the user may pay attention to in the current session. Finally, based on the probability, a personalized cockpit layout is dynamically generated, placing the most likely charts to be paid attention to in the core area.
[0008] This method represents a significant technological breakthrough by incorporating collaborative information from similar user groups to aid in predicting the target user's interests. This seems to deviate from the traditional pursuit of highly accurate user profiling. In traditional user profiling, engineers typically strive to make the profile as close as possible to the individual user's true characteristics, believing that the more accurate the profile, the better, and all recommendations and predictions should be strictly based on the user's historical behavioral data. This approach implicitly assumes that a user's future needs can be found in their past behavior. However, this assumption has limitations in audit management scenarios. Audit leaders' concerns are not static. With adjustments in national policies, changes in organizational business, or job rotations, leaders may need to focus on entirely new data areas, and these new needs may be completely unrecorded in their personal historical behavior logs. If the concept of precise profiling is rigidly adhered to, relying solely on the user's own data, the system will be unable to cope with this shift in needs, leading to predictive failures.
[0009] This method cleverly overcomes the aforementioned technical biases by introducing collaborative analysis of similar user groups. It doesn't abandon the pursuit of individual accuracy, but rather recognizes that when individual data is sparse or needs shift, drawing on the experience of similar users can fill information gaps. This design brings unexpected technical effects. First, it solves the cold start problem. For newly appointed audit leaders, the system lacks their personal behavioral data, making it impossible to generate a personalized layout using traditional methods. However, this method can find similar senior leaders based on their static attributes, such as their area of responsibility, and use their behavioral preferences to generate an initial layout for the new leader, providing a dashboard interface tailored to their job requirements from their first login. Second, it enables the discovery of implicit needs across domains. A leader who has long been responsible for financial audits may never have recorded any interest in engineering audits in their personal behavior logs. However, if similar leaders at the same level have recently frequently viewed engineering audit data, the system can predict that this leader may also need to focus on this area, thus proactively pushing relevant charts to their dashboard. This forward-looking prediction goes beyond the scope of individual historical data and is an effect that cannot be achieved by simply relying on precise profiling.
[0010] From a technical perspective, this method endows the audit dashboard with the ability to transfer collective intelligence. It doesn't simply statistically replay individual behaviors, but rather transfers the experiential knowledge of a group to another through similarity calculation and collaborative filtering. In large enterprises and institutions, audit leaders often share similar job responsibilities and management pain points, and their information needs have inherent commonalities. This method leverages this commonality, enabling the system to learn by analogy. When a leader discovers value in a certain data dimension, the system can automatically disseminate that insight to other similar leaders. This mechanism breaks down information silos and promotes the flow of tacit knowledge within the organization. Furthermore, this method enhances the system's adaptability. Over time, changes in the behavior of similar user groups continuously influence the prediction results for target users, allowing the dashboard layout to dynamically track the shift in the organization's overall focus, always maintaining a high degree of alignment with current audit priorities. This two-way reinforcement mechanism—where the group drives the individual, and the individual feeds back into the group—is a technical advantage that traditional fixed profiling or single-user profiling methods lack.
[0011] Furthermore, S1 includes the following steps: S11. Embed a behavior collection script in the front-end cockpit interface. The behavior collection script listens for interaction events between the user and the visual chart component. The interaction events include mouse click events, mouse drag events, mouse hover events, and chart drill-down events. S12. When an interaction event is detected, extract the event type of the interaction event and identify the operation type according to the event type: if a mouse click event is detected, the operation type is identified as click; if a mouse drag event is detected and the drag distance exceeds a preset threshold, the operation type is identified as drag; if a mouse hover event is detected and the hover duration exceeds a preset duration threshold, the operation type is identified as stay; if a chart drill-down event is detected, the operation type is identified as drill-down. S13. Obtain the component identifier of the target chart component affected by the current interactive event, and query the preset chart-dimension mapping table according to the component identifier to determine the data dimension corresponding to the target chart component. S14. Obtain the user identifier of the current audit user, assemble the user identifier, operation timestamp, identified operation type and determined data dimension into a structured log record, and transmit it to the backend server for storage.
[0012] A behavior tracking script is embedded in the front-end dashboard interface to monitor user interaction events in real time. The script can distinguish between four different operation types: click, drag, pause, and drill-down. Dragging requires a distance exceeding a threshold to be considered valid, and pause requires a duration exceeding a threshold to be considered engaged. This granular design accurately captures conscious user behavior rather than accidental actions. Simultaneously, by querying a pre-defined chart dimension mapping table using chart component identifiers, the system can precisely determine the data dimension corresponding to the user's operation. For example, if a user clicks on an item type distribution chart, the system knows they are engaged with the category classification under the item dimension. Finally, user identifiers, time, operation type, and dimension are assembled into structured logs for storage. The advantage of this step is that it solves the problem of coarse logging in traditional technologies. In existing technologies, server logs typically only record the URL of the page accessed, unable to know what specific actions the user performed on the page, let alone distinguish between quick browsing and in-depth analysis. This method, through front-end tracking and refined identification rules, can collect high-quality, fine-grained behavioral data, laying a solid foundation for building accurate user profiles and ensuring the authenticity and effectiveness of the profiles.
[0013] Furthermore, S2 includes the following steps: S21. Obtain historical operation behavior logs, group them according to user identifiers, and construct an initial profile object for each user identifier; S22. Obtain the basic information associated with each user identifier, and extract one or more of the user's department, position, rank, and area of responsibility from the basic information as static attribute tags for the user and store them in the initial profile object. S23. Iterate through all historical operation behavior logs under the same user ID, count the frequency of operation for each data dimension, and assign different weight coefficients according to the operation type, and calculate the initial attention weight for each data dimension; among them, click operation is assigned the first weight coefficient, drag operation is assigned the second weight coefficient, stay operation is assigned the third weight coefficient, drill-down operation is assigned the fourth weight coefficient, and the fourth weight coefficient > the third weight coefficient > the second weight coefficient > the first weight coefficient. S24. Sort all historical operation behavior logs under the same user ID by timestamp to generate the user's operation sequence; perform sequence pattern mining on the operation sequence to identify frequently occurring adjacent operation pairs and continuous operation paths as sequence features. S25. The initial attention weight is fused with the sequence features to generate the final attention weight for each data dimension. The data dimension and its corresponding final attention weight are stored as dynamic behavior labels in the initial profile object to form a complete dynamic user profile.
[0014] First, the collected logs are grouped according to user identifiers to create an initial profile for each user. Then, static tags such as department and job title are extracted from the user's basic information, reflecting the user's fixed identity characteristics. More crucially, dynamic behavior tags are constructed by statistically analyzing the frequency of operations on each data dimension and assigning different weights based on operation type; for example, drill-down has a higher weight than click, as drill-down represents a stronger exploration intent. Furthermore, this method incorporates sequence analysis, sorting user operation timestamps to uncover frequently occurring pairs of adjacent operations. For instance, users often view project distributions before question lists, revealing business analysis logic. Finally, the initial weights from frequency statistics are merged with sequence features to generate the final attention weight for each dimension. Compared to the simple method of only counting clicks, this solution not only knows what users viewed but also how they viewed it and the relationships between dimensions, making the profile more comprehensive and accurate. For example, frequency alone might suggest a user is most interested in project distributions, but sequence analysis might reveal that they view project distributions only to subsequently view specific questions, thus underestimating the true needs of the question dimension. The merged calculation corrects this bias.
[0015] Furthermore, in S23, the first weighting coefficient is 1, the second weighting coefficient is 1.5, the third weighting coefficient is 2, and the fourth weighting coefficient is 3; In step S24, a prefix projection frequent pattern mining algorithm or a sliding window-based co-occurrence frequency statistics method is used to identify two data dimensions that appear consecutively within a time window as adjacent operation pairs, and the confidence level of each adjacent operation pair is calculated. In S25, the fusion calculation is performed using the following formula: Final attention weight = initial attention weight × (1 + α × confidence of adjacent operation pairs); Wherein, α is a preset fusion coefficient, with a value range of 0.3 to 0.7; the confidence of adjacent operation pairs refers to the average confidence of all adjacent operation pairs that appear as the latter item in the user's operation sequence.
[0016] The document clarifies the weight coefficients for four operations: click, drag, pause, and drill-down, with coefficients of 1, 1.5, 2, and 3 respectively. This weight allocation reflects the differences in attention represented by different operations; pause receives more attention than click, and drill-down receives more attention than pause, making the weight settings interpretable and engineering-feasible. Regarding sequence analysis, it clarifies that a prefix projection frequent pattern mining algorithm or a sliding window co-occurrence frequency statistical method can be used to identify adjacent operation pairs and calculate the confidence score for each operation pair. A specific formula for fusion calculation is also provided: the final attention weight equals the initial attention weight multiplied by one, plus alpha multiplied by the confidence score of the adjacent operation pair, where alpha ranges from 0.3 to 0.7. The advantage of this step is that it transforms abstract concepts into executable algorithms and parameters. Existing technologies often remain at the theoretical level, lacking concrete implementation methods. By providing specific weight values, specific mining algorithms, and specific fusion formulas, those skilled in the art can reproduce the complete technical solution based on the specification, satisfying the patent law's requirement for full disclosure and providing a clear comparative basis for subsequent infringement determination.
[0017] Furthermore, the confidence level of adjacent operation pairs in S25 is calculated as follows: For any two data dimensions A and B, the confidence level of a vector operation on A→B is defined as follows: :
[0018] in This indicates the number of times dimension B appears after dimension A in the user's action sequence. This represents the total number of times dimension A appears; The confidence level of adjacent operation pairs refers to the average confidence level of all adjacent operation pairs appearing as subsequent items in the user's operation sequence. For target data dimension i, its average confidence level AvgConf(i) is calculated in the following way:
[0019] Where X traverses all data dimensions that form adjacent operation pairs X→i with i, and M is the total number of X.
[0020] For any two data dimensions A and B, the confidence level equals the number of times B occurs immediately after A, divided by the total number of times A occurs. This is essentially an estimate of conditional probability, reflecting the strength of the operation transition from A to B. For the target dimension i, its average confidence level equals the sum of the confidence levels of all adjacent operation pairs pointing to i, divided by the number of such operation pairs. The advantage of this calculation method is its clear logic and ease of programming implementation. For example, in an auditing system, if dimension A is the distribution of project types and dimension B is the list of issues, then a high confidence level means that a user is highly likely to continue viewing the issue list after viewing the project types, revealing the inherent relationship between the two business dimensions. By calculating the average confidence level, the prevalence of each dimension in the user's operation path can be quantified. Compared to complex methods such as graph neural networks, the statistical method used in this solution is simple to calculate, has strong interpretability, is easier to deploy and maintain in engineering practice, and can also improve the accuracy of user profiles.
[0021] Furthermore, S3 includes the following steps: S31. Read the dynamic user profile of the target audit user and the dynamic user profiles of all other audit users except the target audit user from the profile database. Each dynamic user profile includes a set of static attribute tags and a set of dynamic behavior tags. The set of dynamic behavior tags consists of several data dimensions and their corresponding final attention weights. S32. Convert the dynamic user profile of each audited user into a feature vector. The feature vector is composed of static feature components and dynamic feature components. The static feature components are obtained by one-hot encoding or embedding vectorization of static attribute labels. The dynamic feature components are constructed with all data dimensions as the dimension space and the final attention weight of each data dimension as the component value. For data dimensions that the user has not operated, the component value is set to 0. S33. Use the Pearson correlation coefficient to calculate the similarity between the feature vector of the target audit user and the feature vector of each other audit user; S34. Sort all other audit users in descending order of similarity to the target audit user, and select users whose similarity exceeds the preset similarity threshold, or select the top N users in terms of similarity, as the similar user group. S35. When the number of users whose similarity exceeds a preset similarity threshold is less than the preset minimum group size, the similarity threshold is dynamically reduced until the number of filtered users reaches the preset minimum group size, and the filtered results with the adjusted threshold are used as the similar user group.
[0022] After converting user profiles into feature vectors, similarity is calculated. Specifically, user profiles are read from the profile database, static attribute labels are converted into vectors using one-hot encoding, and dynamic behavior labels are weighted vectors across all data dimensions. These two vectors are then concatenated to form a complete user feature vector. The Pearson correlation coefficient is used to calculate the similarity between the target user and other users, and similar user groups are selected based on similarity ranking. When high-quality similar users are insufficient, a dynamic threshold adjustment mechanism is designed to automatically lower the threshold until a minimum group size is reached. This step solves the cold start and sparsity problems. In existing technologies, if the target user is unique, a fixed threshold may not be sufficient to select enough similar users, making subsequent predictions impossible. The dynamic threshold adjustment in this scheme ensures that the system can output a usable similar user group under any circumstances, enhancing the system's robustness. Furthermore, by integrating static attributes and dynamic behaviors into the same feature vector for similarity calculation, both user identity characteristics and actual behavioral habits are considered, making the identification of similar users more comprehensive and accurate.
[0023] Furthermore, S32 includes the following steps: S321. Suppose that there are K possible values for a static attribute label, and the static feature vector of the u-th audit user is... ; S322. Let the set constructed from all data dimensions be... M represents the total number of dimensions, and the u-th user's view on the j-th dimension. The final attention weight is Then the dynamic feature vector is ;in This indicates that the u-th audit user has access to the first data dimension. The final attention weight, This indicates that the u-th audit user has access to the second data dimension. The final attention weight, This represents the u-th audit user's view on the m-th data dimension. The final attention weight; S323. Perform concatenation of static and dynamic feature components:
[0024] S33 uses the Pearson correlation coefficient formula to calculate the similarity between the target audit user and other audit users:
[0025] Where n is the dimension of the feature vector, i.e., K+M. Feature vector of the target audit user The i-th component, Feature vectors for other audit users The i-th component, Feature vector representing the target audit user The average of all components, Feature vectors representing other audit users The average of all components.
[0026] A user's static feature vector is a K-dimensional zero-one vector. The dynamic feature vector uses all M data dimensions as its space, with each component representing the user's final attention weight for that dimension; unprocessed dimension components are zero. The two are then concatenated to obtain a K+M dimensional composite feature vector. Similarity is calculated using the Pearson correlation coefficient formula, measuring correlation by calculating the deviation of each dimension of the two vectors from their mean. The advantage of this step is that it provides a complete and executable mathematical solution. The Pearson correlation coefficient can eliminate the influence of different user rating scales; for example, some users may have high weights across all dimensions, while others may have low weights. The Pearson correlation coefficient, through centering, more accurately reflects the consistency of two users' preference patterns. Unifying the vectorization of static and dynamic features allows all users to be compared within the same mathematical framework, laying a solid mathematical foundation for subsequent collaborative filtering predictions.
[0027] Furthermore, S4 includes the following steps: S41. Determine the set of data dimensions for which the probability of attention needs to be predicted. ; S42. Read the similar user groups from the user profile database. Each similar user Dynamic user profiles, extracting each similar user For the data dimension set The final attention weight for each data dimension i in the data. The final attention weight is generated and stored by S25; S43, for the data dimension set For each data dimension i, calculate the predicted probability of target user u's attention to that data dimension. :
[0028] S44. Calculated predicted probability of attention Normalization is performed to obtain the normalized probability of attention. : S45. Normalize the attention probability The final attention probability of the target audit user for different data dimensions in the current login session is used as the basis for sorting the data according to the probability values from high to low, and a attention probability list is generated.
[0029] First, determine the set of data dimensions to be predicted; this can be all dimensions or dimensions that the target user has not yet interacted with. Then, extract the final attention weights for these dimensions from the profiles of similar user groups. Next, calculate a weighted average using the collaborative filtering formula, where the weights represent the similarity between users. After obtaining the raw predicted probabilities, normalization is performed to ensure the sum of the probabilities of all dimensions is one, ultimately generating a ranked list of attention probabilities. The advantage of this step is that it successfully applies the classic collaborative filtering algorithm to the personalized scenario of the audit dashboard. Compared to methods that directly recommend popular dimensions, this solution fully considers the preferences of similar users and can discover data dimensions that the target user may be potentially interested in but have not yet interacted with. For example, a leader in charge of engineering audits may never have paid attention to procurement audit-related dimensions, but if many people in their similar user group are interested, the system can predict that they may also need this data, thus achieving cross-domain intelligent recommendation. Normalization makes the probability values comparable, providing standardized input for subsequent layout decisions.
[0030] Furthermore, S5 includes the following steps: S51. Obtain the normalized attention probability list of the target audit user for different data dimensions. , where d represents the data dimension and p represents the corresponding normalized attention probability; S52. According to the preset cockpit layout template, the cockpit display area is divided into multiple display areas, including a core display area, a secondary display area and an auxiliary display area, and a region weight threshold range is configured for each display area; wherein, the core display area corresponds to the highest probability threshold range, the secondary display area corresponds to the second highest probability threshold range, and the auxiliary display area corresponds to the lowest probability threshold range. S53. Traverse the list of attention probabilities L. For each data dimension d and its probability p, determine the display area type to be assigned to the data dimension based on the probability threshold range to which p belongs. S54. For each display area, all data dimensions assigned to that area are sorted from high to low according to their probability values, and the layout order and area ratio of the visualization chart components corresponding to each data dimension in the area are determined according to the sorting results; among them, the data dimension with the higher probability value has a larger area ratio occupied by its chart components in the area, and the higher the arrangement order. S55. Based on the layout order and area ratio determined in step S54, call the visualization chart component corresponding to each data dimension from the preset chart component library, and dynamically render and generate chart component instances in the corresponding display area according to the preset grid layout algorithm or flexible layout algorithm. S56. Push the generated complete cockpit layout to the front-end browser for rendering and display, and monitor the user's operation behavior in real time during the user's interaction with the cockpit for incremental updates of dynamic user profiles.
[0031] Based on the normalized attention probability list, the cockpit area is divided into three display areas: core, secondary, and auxiliary, with a probability threshold range configured for each area. Then, display areas are allocated according to the probability value range of each dimension. Within the same area, the layout order and area proportion of charts are determined by probability, with higher probability charts occupying larger areas and appearing at the front. Finally, corresponding components are called from the chart component library, dynamically rendered using a grid layout or flexible layout algorithm, and pushed to the front end for display. The advantage of this step is that it achieves a complete mapping from predicted probability to physical layout. Existing cockpits often use a fixed grid arrangement, which cannot reflect the differences in the importance of data. This solution, through area division and area allocation, makes the most important data visually prominent, conforming to visual perception principles. For example, the distribution chart of the projects with the highest attention probability is placed in the core area and occupies a larger area, making it easily visible to users, while auxiliary data with lower probabilities are placed in the corner areas. This dynamic layout ensures the prominent display of key information, makes full use of screen space, improves user experience and data acquisition efficiency, and the entire layout process is fully automated, requiring no manual intervention. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of an embodiment of the audit management method based on multi-source data fusion of the present invention. Detailed Implementation
[0033] The following detailed description illustrates the specific implementation method: The basic implementation examples are as follows: Figure 1 As shown: The audit management method based on multi-source data fusion includes the following steps: S1. Obtain the historical operation behavior logs of each audit user on the front-end dashboard interface. The historical operation behavior logs include the user ID, operation timestamp, operation type, and operation object of each audit user. The operation type includes one or more of clicking, dragging, pausing, and drilling down. The operation object includes the visualization chart components in the dashboard and the corresponding data dimensions.
[0034] S1 includes the following steps: S11. Embed a behavior collection script in the front-end cockpit interface. The behavior collection script listens for interaction events between the user and the visual chart component. The interaction events include mouse click events, mouse drag events, mouse hover events, and chart drill-down events. S12. When an interaction event is detected, extract the event type of the interaction event and identify the operation type according to the event type: if a mouse click event is detected, the operation type is identified as click; if a mouse drag event is detected and the drag distance exceeds a preset threshold, the operation type is identified as drag; if a mouse hover event is detected and the hover duration exceeds a preset duration threshold, the operation type is identified as stay; if a chart drill-down event is detected, the operation type is identified as drill-down. S13. Obtain the component identifier of the target chart component affected by the current interactive event, and query the preset chart-dimension mapping table according to the component identifier to determine the data dimension corresponding to the target chart component. S14. Obtain the user identifier of the current audit user, assemble the user identifier, operation timestamp, identified operation type and determined data dimension into a structured log record, and transmit it to the backend server for storage.
[0035] Specifically, in this embodiment, the front-end dashboard interface is developed using the Vue framework, and the behavior collection script is implemented by listening to DOM events. Assume the audit department has three leaders: Mr. Li, head of the Engineering Audit Department; Mr. Wang, head of the Financial Audit Department; and Mr. Zhang, head of the Comprehensive Audit Department. After the script is embedded in the front end, when Mr. Li logs into the dashboard and performs operations, the script captures the following interaction events: Mr. Li first clicks the "Project Type Distribution" chart, then drags and adjusts the size of the "Problem Trend Chart," then hovers the mouse over the "Rectification Completion Rate" chart for 3 seconds, and finally double-clicks the "Detailed Problems of a Certain Unit" chart to perform a drill-down operation. The script identifies the operation type based on the event type: a click event is identified as a click, a drag distance exceeding 5 pixels is identified as a drag, a hover duration exceeding 2 seconds is identified as a pause, and a double-click drill-down event is identified as a drill-down. Simultaneously, the script queries the preset chart-dimension mapping table based on the component identifier of each chart to determine that the data dimension corresponding to "Project Type Distribution" is the project type under the project dimension, the data dimension corresponding to "Problem Trend Chart" is the problem quantity trend under the problem dimension, the data dimension corresponding to "Rectification Completion Rate" is the rectification status under the rectification dimension, and the data dimension corresponding to "Problem Details of a Certain Unit" is the unit details under the problem dimension. The script assembles Mr. Li's user identifier L001, operation timestamp, identified operation type, and determined data dimensions into structured log records, and transmits them to the backend server for storage via HTTP requests. The table below shows some of the log records generated by Mr. Li in one session:
[0036] S2. Analyze the historical operation behavior logs and construct a dynamic user profile for each audit user based on the analysis results. The dynamic user profile includes static attribute tags and dynamic behavior tags for the audit user. The static attribute tags are generated based on the basic information associated with the user identifier. The dynamic behavior tags are generated based on frequency statistics and sequence analysis of operation type and operation object, and are used to represent the user's attention weight to different data dimensions. The dynamic user profile is then associated with the user identifier.
[0037] S2 includes the following steps: S21. Obtain historical operation behavior logs, group them according to user identifiers, and construct an initial profile object for each user identifier; S22. Obtain the basic information associated with each user identifier, and extract one or more of the user's department, position, rank, and area of responsibility from the basic information as static attribute tags for the user and store them in the initial profile object. S23. Iterate through all historical operation behavior logs under the same user ID, count the frequency of operation for each data dimension, and assign different weight coefficients according to the operation type, and calculate the initial attention weight for each data dimension; among them, click operation is assigned the first weight coefficient, drag operation is assigned the second weight coefficient, stay operation is assigned the third weight coefficient, drill-down operation is assigned the fourth weight coefficient, and the fourth weight coefficient > the third weight coefficient > the second weight coefficient > the first weight coefficient. S24. Sort all historical operation behavior logs under the same user ID by timestamp to generate the user's operation sequence; perform sequence pattern mining on the operation sequence to identify frequently occurring adjacent operation pairs and continuous operation paths as sequence features. S25. The initial attention weight is fused with the sequence features to generate the final attention weight for each data dimension. The data dimension and its corresponding final attention weight are stored as dynamic behavior labels in the initial profile object to form a complete dynamic user profile.
[0038] In S23, the first weighting coefficient is 1, the second weighting coefficient is 1.5, the third weighting coefficient is 2, and the fourth weighting coefficient is 3. In step S24, a prefix projection frequent pattern mining algorithm or a sliding window-based co-occurrence frequency statistics method is used to identify two data dimensions that appear consecutively within a time window as adjacent operation pairs, and the confidence level of each adjacent operation pair is calculated. In S25, the fusion calculation is performed using the following formula: Final attention weight = initial attention weight × (1 + α × confidence of adjacent operation pairs); Wherein, α is a preset fusion coefficient, with a value range of 0.3 to 0.7; the confidence of adjacent operation pairs refers to the average confidence of all adjacent operation pairs that appear as the latter item in the user's operation sequence.
[0039] The confidence level of adjacent operation pairs in S25 is calculated as follows: For any two data dimensions A and B, the confidence level of a vector operation on A→B is defined as follows: :
[0040] in This indicates the number of times dimension B appears after dimension A in the user's action sequence. This represents the total number of times dimension A appears; The confidence level of adjacent operation pairs refers to the average confidence level of all adjacent operation pairs appearing as subsequent items in the user's operation sequence. For target data dimension i, its average confidence level AvgConf(i) is calculated in the following way:
[0041] Where X traverses all data dimensions that form adjacent operation pairs X→i with i, and M is the total number of X.
[0042] Specifically, the backend server processes the logs collected the previous day in batches every morning at midnight. Taking Mr. Li as an example, firstly, all of Mr. Li's log records are read from the database, grouped by user identifier L001, and an initial profile object is created for Mr. Li. Then, Mr. Li's basic information is obtained from the organizational structure system, including his department as Engineering Audit Department, his position as Director, his rank as Senior, and his area of responsibility as Engineering Audit. This information is stored as static attribute tags in the profile object. Next, all of Mr. Li's log records are traversed, the frequency of operation on each data dimension is counted, and weights are assigned according to the operation type: click weight 1, drag weight 1.5, dwell weight 2, drill-down weight 3. The initial attention weight of each dimension is calculated. For example, if the project type dimension is clicked 2 times and dragged 1 time, the initial attention weight is 2×1+1×1.5=3.5; if the issue quantity trend dimension is clicked 1 time, dragged 2 times, and drill-down 1 time, the initial attention weight is 1×1 +2×1.5 +1×3 =7. The calculation results of all dimensions are summarized in the following table:
[0043] Next, Mr. Li's log records are sorted by timestamp to generate an operation sequence: [Project Type, Problem Quantity Trend, Rectification Status, Unit Details]. A sliding window method (window size 2) is used to identify adjacent operation pairs, resulting in three adjacent operation pairs: Project Type → Problem Quantity Trend, Problem Quantity Trend → Rectification Status, and Rectification Status → Unit Details. Calculate the confidence level for each operation pair: Total number of occurrences of project type N(project type) = 3, number of occurrences of issue quantity trend immediately following project type N(project type → issue quantity trend) = 1, so Conf(project type → issue quantity trend) = 1 / 3 ≈ 0.333; Total number of occurrences of issue quantity trend N(issue quantity trend) = 4, number of occurrences of rectification status immediately following issue quantity trend N(issue quantity trend → rectification status) = 1, so Conf(issue quantity trend → rectification status) = 1 / 4 = 0.25; Total number of occurrences of rectification status N(rectification status) = 3, number of occurrences of unit details immediately following rectification status N(rectification status → unit details) = 1, so Conf(rectification status → unit details) = 1 / 3 ≈ 0.333. For each dimension appearing as a consequent, calculate its average confidence level: the average confidence level of the issue quantity trend as a consequent is Conf(Project Type → Issue Quantity Trend) = 0.333; the average confidence level of the rectification status as a consequent is Conf(Issue Quantity Trend → Rectification Status) = 0.25; the average confidence level of the unit details as a consequent is Conf(Rectification Status → Unit Details) = 0.333. Taking the fusion coefficient α = 0.5, calculate the final attention weight for each dimension according to the formula: Final Attention Weight = Initial Attention Weight × (1 + α × Average Confidence Level): Final weight of issue quantity trend = 7 × (1 + 0.5 × 0.333) = 7 × 1.1665 ≈ 8.17; final weight of rectification status = 6 × (1 + 0.5 × 0.25) = 6 × 1.125 = 6.75; final weight of unit details = 6 × (1 + 0.5 × 0.333) = 6 × 1.1665 ≈ 7. Since the project type appears as the first item and not as a subsequent item, its average confidence level is considered to be 0, and its final weight remains unchanged at 3.5. The data dimensions and their final attention weights are stored as dynamic behavioral tags in Mr. Li's user profile object, forming a complete dynamic user profile.
[0044] S3. Respond to the login request of the target audit user, obtain the user identifier of the currently logged-in target audit user and retrieve its corresponding dynamic user profile. Calculate the similarity between its dynamic user profile and the dynamic user profiles of other audit users, and filter out several other audit users whose similarity to the target audit user exceeds a preset threshold to form a similar user group.
[0045] S3 includes the following steps: S31. Read the dynamic user profile of the target audit user and the dynamic user profiles of all other audit users except the target audit user from the profile database. Each dynamic user profile includes a set of static attribute tags and a set of dynamic behavior tags. The set of dynamic behavior tags consists of several data dimensions and their corresponding final attention weights. S32. Convert the dynamic user profile of each audited user into a feature vector. The feature vector is composed of static feature components and dynamic feature components. The static feature components are obtained by one-hot encoding or embedding vectorization of static attribute labels. The dynamic feature components are constructed with all data dimensions as the dimension space and the final attention weight of each data dimension as the component value. For data dimensions that the user has not operated, the component value is set to 0. S33. Use the Pearson correlation coefficient to calculate the similarity between the feature vector of the target audit user and the feature vector of each other audit user; S34. Sort all other audit users in descending order of similarity to the target audit user, and select users whose similarity exceeds the preset similarity threshold, or select the top N users in terms of similarity, as the similar user group. S35. When the number of users whose similarity exceeds a preset similarity threshold is less than the preset minimum group size, the similarity threshold is dynamically reduced until the number of filtered users reaches the preset minimum group size, and the filtered results with the adjusted threshold are used as the similar user group.
[0046] S32 includes the following steps: S321. Suppose that there are K possible values for a static attribute label, and the static feature vector of the u-th audit user is... ; S322. Let the set constructed from all data dimensions be... M represents the total number of dimensions, and the u-th user's view on the j-th dimension. The final attention weight is Then the dynamic feature vector is ;in This indicates that the u-th audit user has access to the first data dimension. The final attention weight, This indicates that the u-th audit user has access to the second data dimension. The final attention weight, This represents the u-th audit user's view on the m-th data dimension. The final attention weight; S323. Perform concatenation of static and dynamic feature components:
[0047] S33 uses the Pearson correlation coefficient formula to calculate the similarity between the target audit user and other audit users:
[0048] Where n is the dimension of the feature vector, i.e., K+M. Feature vector of the target audit user The i-th component, Feature vectors for other audit users The i-th component, Feature vector representing the target audit user The average of all components, Feature vectors representing other audit users The average of all components.
[0049] Specifically, when Mr. Li logs into the system again, the front end sends a login request, and the back end retrieves his dynamic user profile from the profile database based on the user identifier L001. Simultaneously, it reads the dynamic user profiles of all other audit users besides Mr. Li from the database, assuming there are 10 users in the system. Taking Mr. Wang as an example, his static attribute tags include department: Finance and Audit Department; position: Director; rank: Senior; and area of responsibility: Finance and Audit. His dynamic behavior tags include the final attention weights for each data dimension. Each user's static attribute tags are one-hot encoded. Assuming there are 20 possible values for all static attributes, the static feature vector is a 20-dimensional 0 / 1 vector. There are 15 data dimensions in total, and the dynamic feature vector is a 15-dimensional real-number vector, with each component corresponding to the final attention weight for that dimension. Dimensions not manipulated are 0. The static and dynamic feature vectors are concatenated to obtain a 35-dimensional comprehensive feature vector. The Pearson correlation coefficient is used to calculate the similarity of Mr. Li's feature vector with each other. For example, Mr. Li's feature vector is... Mr. Wang's feature vector is The similarity sim(L,W) was calculated to be 0.82. The similarity to the other 8 users was calculated sequentially, with results of 0.75, 0.68, 0.91, 0.43, 0.57, 0.62, 0.38, and 0.71 respectively. Setting a similarity threshold of 0.6, users with a similarity exceeding 0.6 were selected, excluding Mr. Li himself, including Mr. Wang (0.82), another Mr. Zhang (0.91), Mr. Liu (0.75), and Mr. Zhao (0.71), totaling 4 users, forming a similar user group. If the number of selected users is less than the preset minimum group size of 3, this embodiment meets the requirement; if not, the system will automatically lower the threshold by 0.05 and re-filter until the minimum group size is reached.
[0050] S4. Obtain the group behavior preference data of the similar user group within a preset historical time period, and predict the probability of the target audit user's attention to different data dimensions in the current login session using a collaborative filtering prediction model based on user similarity.
[0051] S4 includes the following steps: S41. Determine the set of data dimensions for which the probability of attention needs to be predicted. ; S42. Read the similar user groups from the user profile database. Each similar user Dynamic user profiles, extracting each similar user For the data dimension set The final attention weight for each data dimension i in the data. The final attention weight is generated and stored by S25; S43, for the data dimension set For each data dimension i, calculate the predicted probability of target user u's attention to that data dimension. :
[0052] S44. Calculated predicted probability of attention Normalization is performed to obtain the normalized probability of attention. : S45. Normalize the attention probability The final attention probability of the target audit user for different data dimensions in the current login session is used as the basis for sorting the data according to the probability values from high to low, and a attention probability list is generated.
[0053] Specifically, the set of data dimensions for which the probability of attention needs to be predicted is determined to be all 15 data dimensions. Dynamic user profiles of four users from similar user groups are read from the user profile database, and their final attention weights for each data dimension are extracted. For example, for the "procurement audit" dimension, Mr. Wang's final attention weight is 0.8, Mr. Zhang's is 0.6, Mr. Liu's is 0.4, and Mr. Zhao's is 0.7. Mr. Li's similarity to these four users is 0.82, 0.91, 0.75, and 0.71, respectively. The predicted probability of Mr. Li's attention to the "Procurement Audit" dimension is calculated using the collaborative filtering formula: Numerator = 0.82×0.8 + 0.91×0.6 + 0.75×0.4 + 0.71×0.7 = 0.656 + 0.546 + 0.3 + 0.497 = 1.999, Denominator = 0.82+0.91+0.75+0.71=3.19, P=1.999 / 3.19≈0.627. Similarly, the predicted probability of attention for the other 14 dimensions is calculated to obtain the original probability vector. Then, the original probabilities of all dimensions are normalized, i.e., the original probability of each dimension is divided by the sum of the original probabilities of all dimensions to obtain the normalized probability of attention. For example, if the sum of the original probabilities of all dimensions is 5.2, then the normalized probability of "Procurement Audit" = 0.627 / 5.2≈0.121. Sort by normalized probability from high to low to generate a list of probabilities of interest, as shown in the table below:
[0054] S5. Based on the attention probability, call the corresponding visualization chart component from the preset chart component library, and dynamically generate the personalized cockpit layout for the target audit user according to the probability value; wherein, the data dimension with the highest attention probability and its corresponding chart component are placed in the core display area of the cockpit.
[0055] S5 includes the following steps: S51. Obtain the normalized attention probability list of the target audit user for different data dimensions. , where d represents the data dimension and p represents the corresponding normalized attention probability; S52. According to the preset cockpit layout template, the cockpit display area is divided into multiple display areas, including a core display area, a secondary display area and an auxiliary display area, and a region weight threshold range is configured for each display area; wherein, the core display area corresponds to the highest probability threshold range, the secondary display area corresponds to the second highest probability threshold range, and the auxiliary display area corresponds to the lowest probability threshold range. S53. Traverse the list of attention probabilities L. For each data dimension d and its probability p, determine the display area type to be assigned to the data dimension based on the probability threshold range to which p belongs. S54. For each display area, all data dimensions assigned to that area are sorted from high to low according to their probability values, and the layout order and area ratio of the visualization chart components corresponding to each data dimension in the area are determined according to the sorting results; among them, the data dimension with the higher probability value has a larger area ratio occupied by its chart components in the area, and the higher the arrangement order. S55. Based on the layout order and area ratio determined in step S54, call the visualization chart component corresponding to each data dimension from the preset chart component library, and dynamically render and generate chart component instances in the corresponding display area according to the preset grid layout algorithm or flexible layout algorithm. S56. Push the generated complete cockpit layout to the front-end browser for rendering and display, and monitor the user's operation behavior in real time during the user's interaction with the cockpit for incremental updates of dynamic user profiles.
[0056] Specifically, obtain the normalized attention probability list mentioned above. The cockpit layout template pre-divides the display area into three parts: the core display area is located at the top center of the screen, occupying approximately 60% of the area; the secondary display area is located below the core area, occupying approximately 30% of the area; and the auxiliary display area is located in the right sidebar, occupying approximately 10% of the area. Set probability threshold ranges: the core area corresponds to a probability ≥ 0.12, the secondary area corresponds to a probability of 0.05 ≤ probability < 0.12, and the auxiliary area corresponds to a probability < 0.05. Traverse the probability list and allocate the three dimensions of project type, issue quantity trend, and procurement audit (probabilities of 0.185, 0.152, and 0.121 respectively) to the core display area; allocate dimensions with probabilities between 0.05 and 0.12, such as rectification status and unit details, to the secondary display area; and allocate the remaining dimensions with probabilities below 0.05 to the auxiliary display area. Within each area, sort the allocated dimensions from highest to lowest probability and determine the area ratio occupied by the charts. For example, if there are three charts within the core area, with a total area of S, then the area of the project type chart is approximately 0.185 / (0.185+0.152+0.121)×S ≈ 0.185 / 0.458×S ≈ 0.404S, the area of the issue quantity trend chart is approximately 0.152 / 0.458×S ≈ 0.332S, and the area of the procurement audit chart is approximately 0.121 / 0.458×S ≈ 0.264S. These charts are arranged from left to right within the core area according to this area ratio. The system calls the corresponding chart components from the preset chart component library, such as a pie chart for project type, a line chart for issue quantity trend, and a bar chart for procurement audit. A grid layout algorithm is used to dynamically render and generate chart instances within the corresponding display areas. Finally, the complete cockpit layout is pushed to the front-end browser in JSON format, and the front-end parses and renders it for Mr. Li. Simultaneously, the system continues to monitor Mr. Li's actions in the new session in real time and uses the newly added behavior logs for subsequent incremental updates to his dynamic user profile.
[0057] The above are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. An audit management method based on multi-source data fusion, characterized in that: Includes the following steps: S1. Obtain the historical operation behavior logs of each audit user on the front-end dashboard interface. The historical operation behavior logs include the user ID, operation timestamp, operation type and operation object of each audit user. The operation type includes one or more of clicking, dragging, pausing and drilling down. The operation object includes the visual chart components in the dashboard and the corresponding data dimensions. S2. Parse the historical operation behavior logs and construct a dynamic user profile for each audit user based on the parsing results. The dynamic user profile includes static attribute tags and dynamic behavior tags for the audit user. The static attribute tags are generated based on the basic information associated with the user identifier. The dynamic behavior tags are generated based on frequency statistics and sequence analysis of operation type and operation object, and are used to represent the user's attention weight to different data dimensions. The dynamic user profile is associated with the user identifier. S3. Respond to the login request of the target audit user, obtain the user identifier of the currently logged-in target audit user and retrieve its corresponding dynamic user profile. Calculate the similarity between its dynamic user profile and the dynamic user profiles of other audit users, and filter out several other audit users whose similarity to the target audit user exceeds a preset threshold to form a similar user group. S4. Obtain the group behavior preference data of the similar user group within a preset historical time period, and predict the probability of the target audit user's attention to different data dimensions in the current login session using a collaborative filtering prediction model based on user similarity. S5. Based on the attention probability, call the corresponding visualization chart component from the preset chart component library, and dynamically generate the personalized cockpit layout for the target audit user according to the probability value; wherein, the data dimension with the highest attention probability and its corresponding chart component are placed in the core display area of the cockpit.
2. The audit management method based on multi-source data fusion according to claim 1, characterized in that: S1 includes the following steps: S11. Embed a behavior collection script in the front-end cockpit interface. The behavior collection script listens for interaction events between the user and the visual chart component. The interaction events include mouse click events, mouse drag events, mouse hover events, and chart drill-down events. S12. When an interaction event is detected, extract the event type of the interaction event and identify the operation type according to the event type: if a mouse click event is detected, the operation type is identified as click; if a mouse drag event is detected and the drag distance exceeds a preset threshold, the operation type is identified as drag; if a mouse hover event is detected and the hover duration exceeds a preset duration threshold, the operation type is identified as stay; if a chart drill-down event is detected, the operation type is identified as drill-down. S13. Obtain the component identifier of the target chart component affected by the current interactive event, and query the preset chart-dimension mapping table according to the component identifier to determine the data dimension corresponding to the target chart component. S14. Obtain the user identifier of the current audit user, assemble the user identifier, operation timestamp, identified operation type and determined data dimension into a structured log record, and transmit it to the backend server for storage.
3. The audit management method based on multi-source data fusion according to claim 2, characterized in that: S2 includes the following steps: S21. Obtain historical operation behavior logs, group them according to user identifiers, and construct an initial profile object for each user identifier; S22. Obtain the basic information associated with each user identifier, and extract one or more of the user's department, position, rank, and area of responsibility from the basic information, and store them as the user's static attribute tags in the initial profile object. S23. Iterate through all historical operation behavior logs under the same user ID, count the frequency of operation for each data dimension, and assign different weight coefficients according to the operation type, and calculate the initial attention weight for each data dimension; among them, click operation is assigned the first weight coefficient, drag operation is assigned the second weight coefficient, stay operation is assigned the third weight coefficient, drill-down operation is assigned the fourth weight coefficient, and the fourth weight coefficient > the third weight coefficient > the second weight coefficient > the first weight coefficient. S24. Sort all historical operation behavior logs under the same user ID by timestamp to generate the user's operation sequence; perform sequence pattern mining on the operation sequence to identify frequently occurring adjacent operation pairs and continuous operation paths as sequence features. S25. The initial attention weight is fused with the sequence features to generate the final attention weight for each data dimension. The data dimension and its corresponding final attention weight are stored as dynamic behavior labels in the initial profile object to form a complete dynamic user profile.
4. The audit management method based on multi-source data fusion according to claim 3, characterized in that: In S23, the first weighting coefficient is 1, the second weighting coefficient is 1.5, the third weighting coefficient is 2, and the fourth weighting coefficient is 3. In step S24, a prefix projection frequent pattern mining algorithm or a sliding window-based co-occurrence frequency statistics method is used to identify two data dimensions that appear consecutively within a time window as adjacent operation pairs, and the confidence level of each adjacent operation pair is calculated. In S25, the fusion calculation is performed using the following formula: Final attention weight = initial attention weight × (1 + α × confidence of adjacent operation pairs); Wherein, α is a preset fusion coefficient, with a value range of 0.3 to 0.7; the confidence of adjacent operation pairs refers to the average confidence of all adjacent operation pairs that appear as the latter item in the user's operation sequence.
5. The audit management method based on multi-source data fusion according to claim 4, characterized in that: The confidence level of adjacent operation pairs in S25 is calculated as follows: For any two data dimensions A and B, the confidence level of a vector operation on A→B is defined as follows: : in This indicates the number of times dimension B appears after dimension A in the user's action sequence. This represents the total number of times dimension A appears; The confidence level of adjacent operation pairs refers to the average confidence level of all adjacent operation pairs appearing as subsequent items in the user's operation sequence. For target data dimension i, its average confidence level AvgConf(i) is calculated in the following way: Where X traverses all data dimensions that form adjacent operation pairs X→i with i, and M is the total number of X.
6. The audit management method based on multi-source data fusion according to claim 5, characterized in that: S3 includes the following steps: S31. Read the dynamic user profile of the target audit user and the dynamic user profiles of all other audit users except the target audit user from the profile database. Each dynamic user profile includes a set of static attribute tags and a set of dynamic behavior tags. The set of dynamic behavior tags consists of several data dimensions and their corresponding final attention weights. S32. Convert the dynamic user profile of each audited user into a feature vector. The feature vector is composed of static feature components and dynamic feature components. The static feature components are obtained by one-hot encoding or embedding vectorization of static attribute labels. The dynamic feature components are constructed with all data dimensions as the dimension space and the final attention weight of each data dimension as the component value. For data dimensions that the user has not operated, the component value is set to 0. S33. Use the Pearson correlation coefficient to calculate the similarity between the feature vector of the target audit user and the feature vector of each other audit user; S34. Sort all other audit users in descending order of similarity to the target audit user, and select users whose similarity exceeds the preset similarity threshold, or select the top N users in terms of similarity, as the similar user group. S35. When the number of users whose similarity exceeds a preset similarity threshold is less than the preset minimum group size, the similarity threshold is dynamically reduced until the number of filtered users reaches the preset minimum group size, and the filtered results with the adjusted threshold are used as the similar user group.
7. The audit management method based on multi-source data fusion according to claim 6, characterized in that: S32 includes the following steps: S321. Suppose that there are K possible values for a static attribute label, and the static feature vector of the u-th audit user is... ; S322. Let the set constructed from all data dimensions be... M represents the total number of dimensions, and the u-th user's view on the j-th dimension. The final attention weight is Then the dynamic feature vector is ,in This indicates that the u-th audit user has access to the first data dimension. The final attention weight, This indicates that the u-th audit user has access to the second data dimension. The final attention weight, This represents the u-th audit user's view on the m-th data dimension. The final attention weight; S323. Perform concatenation of static and dynamic feature components: S33 uses the Pearson correlation coefficient formula to calculate the similarity between the target audit user and other audit users: Where n is the dimension of the feature vector, i.e., K+M. Feature vector of the target audit user The i-th component, Feature vectors for other audit users The i-th component, Feature vector representing the target audit user The average of all components, Feature vectors representing other audit users The average of all components.
8. The audit management method based on multi-source data fusion according to claim 7, characterized in that: S4 includes the following steps: S41. Determine the set of data dimensions for which the probability of attention needs to be predicted. ; S42. Read the similar user groups from the user profile database. Each similar user Dynamic user profiles, extracting each similar user For the data dimension set The final attention weight for each data dimension i in the data. The final attention weight is generated and stored by S25; S43, for the data dimension set For each data dimension i, calculate the predicted probability of target user u's attention to that data dimension. : S44. Calculated predicted probability of attention Normalization is performed to obtain the normalized probability of attention. : S45. Normalize the attention probability The final attention probability of the target audit user for different data dimensions in the current login session is used as the basis for sorting the data according to the probability values from high to low, and a attention probability list is generated.
9. The audit management method based on multi-source data fusion according to claim 8, characterized in that: S5 includes the following steps: S51. Obtain the normalized attention probability list of the target audit user for different data dimensions. , where d represents the data dimension and p represents the corresponding normalized attention probability; S52. According to the preset cockpit layout template, the cockpit display area is divided into multiple display areas, including a core display area, a secondary display area and an auxiliary display area, and a region weight threshold range is configured for each display area; wherein, the core display area corresponds to the highest probability threshold range, the secondary display area corresponds to the second highest probability threshold range, and the auxiliary display area corresponds to the lowest probability threshold range. S53. Traverse the list of attention probabilities L. For each data dimension d and its probability p, determine the display area type to be assigned to the data dimension based on the probability threshold range to which p belongs. S54. For each display area, all data dimensions assigned to that area are sorted from high to low according to their probability values, and the layout order and area ratio of the visualization chart components corresponding to each data dimension in the area are determined according to the sorting results; among them, the data dimension with the higher probability value has a larger area ratio occupied by its chart components in the area, and the higher the arrangement order. S55. Based on the layout order and area ratio determined in step S54, call the visualization chart component corresponding to each data dimension from the preset chart component library, and dynamically render and generate chart component instances in the corresponding display area according to the preset grid layout algorithm or flexible layout algorithm. S56. Push the generated complete cockpit layout to the front-end browser for rendering and display, and monitor the user's operation behavior in real time during the user's interaction with the cockpit for incremental updates of dynamic user profiles.