Enterprise Digital Internal Control Management System Based on Big Data Analytics

The enterprise digital internal control management system, which utilizes big data analytics, has solved the problem of discrepancies between employee communication behavior and actual output, enabling more accurate efficiency assessment and evaluation, and improving the comparability and fairness of assessments across cycles and departments.

CN120806755BActive Publication Date: 2025-11-14HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511309298.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-11-14
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

In existing technologies, employees' communication behavior patterns fluctuate significantly before and after key milestones, leading to discrepancies between proxy indicators and actual outputs. This results in a systematic bias in efficiency assessment, affecting the accuracy of performance evaluation and the credibility of data-driven management.

Method used

By constructing an enterprise digital intelligent internal control management system based on big data analysis, the system uses a data acquisition module to extract event streams of communication behaviors and actual outputs from the enterprise's business database, applies a symmetric kernel smoothing function for data processing, calculates the decoupling rate sequence and smoothing scale parameters, and combines bias strength and phase stability indicators to generate a report dependence index. This is then used for debiasing processing to generate debiasing assessment results.

Benefits of technology

It significantly improves the comparability and fairness of efficiency assessments across cycles, positions, and departments, reduces misjudgments caused by communication volume substituting for efficiency, supports personalized incentives and person-job matching, and restores the temporal consistency between proxy indicators and actual output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806755B_ABST
    Figure CN120806755B_ABST
Patent Text Reader

Abstract

This invention relates to the field of enterprise digitalization technology and discloses an enterprise digitalization internal control management system based on big data analysis, comprising: extracting proxy event streams and actual event streams from a business database; constructing a time window set, including the window center time and the window half-width; applying symmetric kernel smoothing to the two types of event streams to obtain proxy event intensity sequences and actual event intensity sequences, and calculating the decoupling rate sequence; determining smoothing scale parameters based on the decoupling rate sequence and the time window set; dividing the time window into a pre-window region, a post-window region, and a symmetric baseline region, and calculating the pre-window mean, post-window mean, and baseline mean respectively; calculating a bias intensity index based on the three means; calculating a phase stability index based on the peak value of the decoupling rate sequence and the window center time; generating a report dependency index by combining the bias intensity index and the phase stability index; and performing de-biasing processing on the proxy event streams and actual event streams accordingly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise digitalization technology, and more specifically, to an enterprise digitalization internal control management system based on big data analysis. Background Technology

[0002] As enterprise management becomes increasingly digitalized, employees' daily behaviors are widely recorded in various systems, such as instant messaging tools, email systems, work order systems, and attendance and schedule management platforms. To simplify measurement, companies often directly use easily obtainable metrics such as the number of communications and message response speed as approximations of efficiency. However, when these approximations are used directly as performance evaluation criteria, employees tend to concentrate their communication activities at key moments, thus creating a false impression of high efficiency on the data level.

[0003] As fixed milestones such as periodic performance reviews, shift handovers, or routine reports approach, employees' communication behavior patterns exhibit significant shifts: before these milestones, communication is artificially accumulated and released in a concentrated manner; after these milestones, communication activity decreases, but actual work output increases. This fluctuation in the time series causes proxy indicators, represented by communication volume, to become out of sync with actual output, resulting in a systemic deviation between the two. This leads to the following problems:

[0004] 1. Employees' true efficiency is concealed or distorted;

[0005] 2. Efficiency results between different periods cannot be directly compared;

[0006] 3. It is difficult to establish a fair comparison of efficiency levels between departments and positions.

[0007] The aforementioned systemic biases not only lead managers to draw incorrect conclusions in performance appraisals but also affect the rationality of personnel-job matching, training investment, and incentive allocation. More importantly, if such biases persist, they will weaken the credibility of data-driven management and reduce the organization's reliance on and willingness to use efficiency assessment tools. Therefore, accurately characterizing and correcting the systemic biases between proxy indicators and actual outputs has become a pressing technical problem that needs to be addressed in enterprise efficiency assessment methods. Summary of the Invention

[0008] This invention provides an enterprise digital intelligent internal control management system based on big data analysis, which solves the technical problems mentioned in the background.

[0009] This invention provides an enterprise digital intelligent internal control management system based on big data analysis, comprising:

[0010] The data acquisition module extracts proxy event streams reflecting communication behaviors and implementation event streams reflecting actual outputs from the enterprise business database, taking employees as the unit; it constructs a set of time windows, each containing the center time of the window;

[0011] The data processing module applies a symmetric kernel smoothing function to the proxy event stream and the actual event stream respectively to obtain the proxy event intensity sequence and the actual event intensity sequence. It determines the decoupling rate sequence based on the ratio of the proxy event intensity sequence and the actual event intensity sequence, and determines the smoothing scale parameter of the symmetric kernel smoothing function based on the decoupling rate sequence and the time window set.

[0012] The index calculation module divides each time window into a pre-window region, a post-window region, and a symmetrical baseline region. It calculates the pre-window mean of the decoupling rate sequence in the pre-window region, the post-window mean in the post-window region, and the baseline mean in the symmetrical baseline region. Based on the pre-window mean, post-window mean, and baseline mean, it calculates the bias strength index, which reflects the degree of deviation of the decoupling rate sequence from the baseline region in the pre-window and post-window regions. Based on the peak value of the decoupling rate sequence and the corresponding center time of the time window, it calculates the phase stability index, which reflects the stability of the position of the peak value of the decoupling rate sequence within the time window.

[0013] The bias removal module, combined with the bias strength index and the phase stability index, generates a report dependency index that characterizes the pattern of employees adjusting proxy events around the time window. Based on the report dependency index, the employees' proxy event flow and actual event flow are debiased, and the debiased proxy event flow and actual event flow are input into the preset assessment model to obtain the debiased assessment results.

[0014] Furthermore, a symmetric kernel smoothing function is applied to the proxy event stream and the actual event stream respectively to obtain the proxy event intensity sequence and the actual event intensity sequence. The decoupling rate sequence is determined based on the ratio of the proxy event intensity sequence to the actual event intensity sequence, including:

[0015] The decoupling rate sequence is obtained by performing a ratio operation on the proxy event intensity sequence and the actual event intensity sequence at each time point on the time axis.

[0016] Furthermore, the smoothing scale parameters of the symmetric kernel smoothing function are determined based on the decoupling rate sequence and the time window set, including:

[0017] The average half-width of the time window set is used as the alignment reference; candidate intervals for smoothing scales are determined within multiples of the alignment reference; within the candidate intervals, a one-dimensional search is performed with the square of the difference between the first zero-crossing position of the first autocorrelation function of the decoupling rate sequence and the alignment reference as the objective function to obtain the smoothing scale parameter that minimizes the objective function, and this smoothing scale parameter is used for the symmetric kernel smoothing function.

[0018] Furthermore, for each time window, the system is divided into a pre-window region, a post-window region, and a symmetrical baseline region. The mean of the decoupling rate sequence in the pre-window region, the post-window region, and the baseline region are calculated, including:

[0019] Use half of the time window as the window's half width;

[0020] The front area of ​​the window is obtained by extending the center of the window forward by half the width of the window.

[0021] The back area of ​​the window is obtained by extending the center of the window backward by half the width of the window.

[0022] A symmetrical baseline area is formed by extending one half-width of the window at equal intervals on both sides of the time window;

[0023] On the time axis, the mean before the window, the mean after the window, and the baseline mean are obtained by averaging the decoupling rate sequence over time in the pre-window, post-window, and symmetrical baseline regions, respectively.

[0024] Furthermore, the bias intensity index is calculated based on the pre-window mean, post-window mean, and baseline mean, including:

[0025] The ratio of the mean before the window to the baseline mean is used as the ratio before the window.

[0026] The ratio of the mean after the window to the baseline mean is called the post-window ratio.

[0027] The single-window offset is obtained by taking the logarithm of the product of the ratio before and after the window.

[0028] The average of all single-window offsets corresponding to the time window set is calculated to obtain the offset strength index.

[0029] Furthermore, the phase stability index is calculated based on the peak value of the decoupling rate sequence and the corresponding center moment of the time window, including:

[0030] The peak value of the decoupling rate sequence is defined as the moment corresponding to the peak value within the time window.

[0031] The relative position angle is obtained by normalizing the time difference between the peak time and the corresponding window center time.

[0032] The phase stability index is obtained by aggregating all relative position angles corresponding to the time window set using the circular statistical method.

[0033] Furthermore, by combining the bias strength index and the phase stability index, a reporting dependence index is generated that reflects the degree to which employee behavior depends on the time window, including:

[0034] The larger value between the bias strength index and 0 is taken as the non-negative bias.

[0035] The difference between 1 and the phase stability index is taken as the phase concentration.

[0036] The product of the non-negative bias and the phase concentration is used as the reporting dependency index.

[0037] Furthermore, based on the report dependency index, bias-free processing is performed on the employee's proxy event flow and implementation event flow, including:

[0038] If the report relies on an index exceeding the limit, a debiasing process will be performed, including:

[0039] A time weighting function is constructed for the time window; the weight strength parameter of the time weighting function is specifically calculated as follows: the square of the difference between the mean of the decoupling rate sequence before the window and the mean of the decoupling rate sequence after the window is calculated, and the square of the difference between all the means corresponding to the time window set is minimized to obtain the weight strength parameter applied to the time weighting function.

[0040] The proxy event flow and the actual event flow are weighted first and second by a time weighting function and a symmetric kernel smoothing function, respectively, to obtain a weighted proxy event intensity sequence and a weighted actual event intensity sequence. Within a preset assessment period, the weighted proxy event intensity sequence and the weighted actual event intensity sequence are used as inputs to a preset assessment model to obtain the debiased assessment results.

[0041] The beneficial effects of this invention are as follows: By implementing symmetric kernel smoothing on the proxy event flow and the actual event flow within a time window and constructing a decoupling rate, combined with three-part statistics of the pre-window area, post-window area and symmetric baseline area, and circular statistical phase aggregation based on peak position, the bias intensity and phase stability are quantitatively obtained and a report dependence index is formed; when the index exceeds the limit, the bias is removed and weighted by a time weight function and further smoothing, and the corrected assessment results are output, thereby identifying and eliminating the systematic bias caused by communication backlog and output lag around reporting nodes, restoring the temporal consistency between proxy indicators and actual output, significantly improving the comparability and fairness across cycles, positions and departments, reducing misjudgments caused by substituting efficiency for communication volume, and supporting robust assessment and early warning at the employee level. Attached Figure Description

[0042] Figure 1 This is a module diagram of the enterprise digital intelligent internal control management system based on big data analysis according to the present invention. Detailed Implementation

[0043] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.

[0044] like Figure 1 As shown, the enterprise digital intelligent internal control management system based on big data analytics includes:

[0045] The data acquisition module extracts proxy event streams reflecting communication behaviors and implementation event streams reflecting actual outputs from the enterprise business database, taking employees as the unit; it constructs a set of time windows, each containing the center time of the window;

[0046] The data processing module applies a symmetric kernel smoothing function to the proxy event stream and the actual event stream respectively to obtain the proxy event intensity sequence and the actual event intensity sequence. It determines the decoupling rate sequence based on the ratio of the proxy event intensity sequence and the actual event intensity sequence, and determines the smoothing scale parameter of the symmetric kernel smoothing function based on the decoupling rate sequence and the time window set.

[0047] The index calculation module divides each time window into a pre-window region, a post-window region, and a symmetrical baseline region. It calculates the pre-window mean of the decoupling rate sequence in the pre-window region, the post-window mean in the post-window region, and the baseline mean in the symmetrical baseline region. Based on the pre-window mean, post-window mean, and baseline mean, it calculates the bias strength index, which reflects the degree of deviation of the decoupling rate sequence from the baseline region in the pre-window and post-window regions. Based on the peak value of the decoupling rate sequence and the corresponding center time of the time window, it calculates the phase stability index, which reflects the stability of the position of the peak value of the decoupling rate sequence within the time window.

[0048] The bias removal module, combined with the bias strength index and the phase stability index, generates a report dependency index that characterizes the pattern of employees adjusting proxy events around the time window. Based on the report dependency index, the employees' proxy event flow and actual event flow are debiased, and the debiased proxy event flow and actual event flow are input into the preset assessment model to obtain the debiased assessment results.

[0049] It should be noted that "employee-based" means that all data extraction, organization, and analysis are carried out around a single employee, avoiding the confusion of event data from different employees, ensuring that subsequent efficiency evaluations can accurately pinpoint individuals, and meeting the management needs of matching people to positions and providing personalized incentives.

[0050] It's important to note that a proxy event flow refers to a set of discrete events that directly reflect employee communication and interaction behaviors. It's called a proxy because companies often use the frequency of communication as a substitute indicator for efficiency (i.e., a proxy indicator). However, such indicators are easily manipulated by employees, for example, by concentrating communication at a window. Therefore, it's necessary to capture the temporal characteristics of communication behaviors through this event flow.

[0051] It should be noted that the proxy event stream is taken from the enterprise's existing business database, requiring no additional data collection equipment. Specific sources include:

[0052] Instant Messaging (IM) System Database: Stores records of instant messages sent / received by employees, including message sending time, receiving time, message recipients, etc.

[0053] Email system database: Stores records of emails sent / received by employees, including email sending time, receiving time, email subject, recipient / sender, etc.

[0054] Meeting Management System Database: Stores records of employee participation in meetings, including meeting start time, end time, and participation status (whether they attended, whether they were late or left early), etc.

[0055] Collaborative office system database: Stores communication records of employees initiating / responding to collaborative tasks, such as the time when task comments were sent and the time when others were mentioned.

[0056] The proxy event stream is a set of discrete-time events, where each communication action corresponds to a unique timestamp. For example, employee A sends an IM message at 9:05 on May 20, 2024; employee A sends an email at 10:10 on May 20, 2024; all such events are arranged in chronological order to form a discrete event sequence based on employees.

[0057] It's important to note that a performance event flow refers to a set of discrete events that directly reflect an employee's actual work output. It's called "performance" because these events directly correspond to value-creating behaviors recognized by the company. Examples include completing work orders and delivering results. The performance event flow serves as a benchmark for judging whether communication behaviors match actual efficiency. If there are many communication behaviors but few performance events, it indicates communication manipulation; if the communication behaviors and performance events are sequential, it means the communication is not divorced from actual output.

[0058] The actual event flow is taken from the company's existing business database. The specific source varies slightly depending on the job type. Common sources include:

[0059] Work order management system database: Stores records of employees processing work orders, with a focus on extracting work order completion time (work order completion time), work order type, and work order completion quality (e.g., whether it passed acceptance on the first try).

[0060] Code Management System Database (Technical Positions): Stores records of code submissions by employees, including submission time, amount of code submitted, and code approval time.

[0061] Deliverables System Database (Design and R&D positions): Stores records of deliverables uploaded by employees, including upload time, deliverable type (e.g., design drawings, test reports), and deliverable acceptance time.

[0062] Project management system database: Stores records of employees completing project tasks, including task start time, task completion time, task completion progress (whether it is 100% completed), etc.

[0063] Customer service system database (customer service position): Stores records of employees handling customer inquiries / complaints, including service start time, service end time, customer satisfaction rating, etc.

[0064] The implementation event flow is a set of discrete-time events, with each actual output behavior corresponding to a unique timestamp. For example, employee A completes a work order closure at 11:30 on May 20, 2024; employee A submits code once at 15:20 on May 20, 2024; and the events are arranged in chronological order as a discrete event sequence based on the employee.

[0065] It's important to note that a time window set refers to a collection of multiple key time intervals. Each key time interval corresponds to a fixed point in enterprise management that may trigger manipulation of employee behavior (such as reporting meetings or performance review deadlines). Constructing a time window set allows for the analysis of changes in employee communication and output behavior before and after key points. By comparing behavioral differences before, after, and outside the window, manipulation patterns of accumulated communication before the window and concentrated output after the window can be accurately identified.

[0066] It should be noted that the window center time refers to the core reference time of each key time interval, that is, the starting time of the key node that triggers employee behavior manipulation.

[0067] The window center time is determined entirely based on the company's existing management rules and historical records, without any subjective setting. Specific sources include:

[0068] Reporting points: such as the departmental meeting at 9:00 am every Monday and the monthly performance review report at 4:00 pm on the last working day of each month, the start time of the meeting / report is the window center time;

[0069] Shift handover points: For example, the handover between the morning shift (8:00-16:00) and the afternoon shift (16:00-24:00) of the customer service position, the handover start time (16:00) is the window center time;

[0070] Task deadlines: such as the deadlines for project phase tasks (e.g., submitting design drafts before 24:00 on May 20th), the deadline is the time shown in the center of the window;

[0071] Cross-departmental collaboration nodes: such as the cross-departmental communication meeting held every Wednesday at 2:00 pm, the meeting start time is the window center time.

[0072] For example, if employee A is required to attend a departmental report meeting every Monday at 9:00 AM, and this report is held on a fixed schedule (once a week), then there are 4 windows in the time window set (assuming the evaluation period is 4 weeks). The center time of each window is 9:00 AM on Monday of the first week, 9:00 AM on Monday of the second week, 9:00 AM on Monday of the third week, and 9:00 AM on Monday of the fourth week.

[0073] It should be noted that the window half-width refers to the length of time extending forward and backward from the center time of the window, that is, half the length of each key time interval. Together with the center time of the window, it forms a complete time window interval (window interval = center time - half-width to center time + half-width). The window half-width is used to define the time range within which employees may manipulate behavior. That is, the size of the window half-width must match the actual duration of the sustained impact of the key nodes to avoid analytical bias caused by the range being too large or too small.

[0074] The window half-width is determined by data-driven calculations, based entirely on actual records of similar historical nodes for each employee, as detailed below:

[0075] Step 1: Extract the historical continuous records of similar key nodes for employees over the past 6 months (or longer). For example, for the weekly Monday morning 9:00 reporting meeting, extract the actual duration of the past 24 reporting meetings (e.g., one reporting meeting started at 9:00 and ended at 9:40, lasting 40 minutes; another started at 9:00 and ended at 10:00, lasting 60 minutes).

[0076] Step 2: Sort the extracted historical durations and take the median as the half-width of the window for this type. Choosing the median instead of the mean avoids interference from extreme values ​​(such as a report lasting 2 hours due to special circumstances) and ensures that the half-width reflects the duration of the critical impact in most cases.

[0077] For example, the duration of Employee A's past 24 Monday reporting meetings were 40 minutes, 45 minutes, 50 minutes, and 60 minutes respectively (24 data points in total). After sorting, the 12th and 13th data points are 45 minutes and 50 minutes respectively. Therefore, the median is (45+50) / 2 = 47.5 minutes, meaning the half-width of each reporting window is 47.5 minutes. Combining this with the center time of the window (e.g., 9:00 AM on Monday of the first week), the complete window interval is from 9:00 - 47.5 minutes = 8:12:30 to 9:00 + 47.5 minutes = 9:47:30. This interval is the core range for analyzing Employee A's behavior before and after this reporting session.

[0078] In one embodiment of the present invention, a symmetric kernel smoothing function is applied to the proxy event stream and the actual event stream respectively to obtain a proxy event intensity sequence and an actual event intensity sequence. A decoupling rate sequence is determined based on the ratio of the proxy event intensity sequence to the actual event intensity sequence, including:

[0079] The decoupling rate sequence is obtained by performing a ratio operation on the proxy event intensity sequence and the actual event intensity sequence at each time point on the time axis.

[0080] It should be noted that symmetric kernel smoothing functions are mathematical tools used to transform discrete data into continuous data, and their core characteristic is time symmetry. That is, for a given target time, historical data within the same time interval before and after it are assigned equal weights, avoiding analytical bias caused by weighting favoring one side (such as focusing only on past data). In this invention, the symmetric kernel function typically used is the Gaussian kernel function (commonly used in the industry for smoothing event streams, combining smoothing effect with computational convenience), to eliminate random fluctuations in discrete event streams and extract trend features.

[0081] It's important to note that both the original proxy event flow and the actual event flow are discrete sets of time points. For example, an employee might send one message at 9:05, a second message at 9:20, and complete a work order at 10:10. This type of discrete data exhibits strong randomness and large fluctuations, making it impossible to directly observe the overall activity trend of communication or output within a specific time period. For instance, it's not intuitive to determine whether communication is more concentrated between 9:00-10:00 and 10:00-11:00. Symmetric kernel smoothing transforms discrete time points into a continuous intensity sequence, thereby highlighting the temporal trend characteristics.

[0082] Proxy event stream: A set of timestamps reflecting communication behaviors, mathematically represented as:

[0083]

[0084] Where δ(·) represents the Dirac function, t i (c)Let s represent the time of the i-th communication event, where s is a time variable;

[0085] Implementation event flow: A set of timestamps reflecting actual output, mathematically represented as:

[0086]

[0087] Among them, t j (r) Indicates the time of the j-th actual event;

[0088] Taking a proxy event stream as an example, for any target time t, its proxy event intensity C σ (t) is obtained by weighting the contributions of all communication events before and after this moment using a symmetric kernel, and the formula is:

[0089]

[0090] Among them, K σ (·) denotes a symmetric kernel function, which is a Gaussian kernel; where the standard deviation of the Gaussian kernel is the smoothing scale parameter, and the mean of the Gaussian kernel is 0 by default.

[0091] Integration is used to transform the time distribution of discrete events into continuous intensity values ​​using a symmetric kernel function;

[0092] Similarly, the intensity sequence r of the actual events is calculated. σ (t);

[0093] The proxy event intensity sequence is a continuous numerical sequence. The value at each time t represents the average occurrence density of communication events per unit time at that time. The higher the value, the more active the communication is near that time.

[0094] The intensity sequence of actual events is a continuous numerical sequence. The value of each time t represents the average occurrence density of actual events per unit time at that time. The higher the value, the more active the actual output is near that time.

[0095] It should be noted that the degree of decoupling at each moment is quantified by the relative relationship between communication intensity and implementation intensity. That is, if the ratio of the intensity of proxy events to the intensity of implementation events is too high, it indicates that the communication input far exceeds the output; if the ratio is balanced, it indicates that communication and output are matched.

[0096] Iterate through all moments on the continuous timeline. For each moment t, calculate the ratio of the surrogate event strength to the actual event strength to obtain the decoupling rate at that moment. To avoid calculation errors when the actual event strength is 0, the document requires adding a machine-precision constant ε (e.g., ε=10) to the denominator. -12 (This is only for numerical stability and does not affect the overall trend), the formula is: .

[0097] In detail, the decoupling rate sequence is a continuous numerical sequence covering the entire analysis period, with the value at each moment corresponding to the degree of decoupling between communication and output at that moment;

[0098] When H(t) > 1, the communication intensity is higher than the implementation intensity, and there is a decoupling between communication and output. The larger the value, the more severe the decoupling.

[0099] When H(t) = 1, the communication intensity matches the implementation intensity, with no obvious decoupling;

[0100] When H(t) < 1, the intensity of practice is higher than the intensity of communication, and the communication input is insufficient.

[0101] In one embodiment of the present invention, determining the smoothing scale parameter of the symmetric kernel smoothing function based on the decoupling rate sequence and the time window set includes:

[0102] The average half-width of the time window set is used as the alignment reference; candidate intervals for smoothing scales are determined within multiples of the alignment reference; within the candidate intervals, a one-dimensional search is performed with the square of the difference between the first zero-crossing position of the first autocorrelation function of the decoupling rate sequence and the alignment reference as the objective function to obtain the smoothing scale parameter that minimizes the objective function, and this smoothing scale parameter is used for the symmetric kernel smoothing function.

[0103] It should be noted that symmetric kernel smoothing is used to transform discrete event streams into continuous intensity sequences, and the smoothing scale parameter directly determines the smoothing effect. If the parameter is too small, the intensity sequence will retain too much random fluctuation and will not reflect the true trend; if the parameter is too large, the intensity sequence will be overly blurred and lose the decoupling features before and after the window.

[0104] It should be noted that for all windows in the time window set, the half-width of each window is extracted. This is the median duration of the sustained impact of each key node, such as a half-width of 47.5 minutes for a briefing and 30 minutes for a handover meeting. The arithmetic mean of all window half-widths is calculated, and this mean is defined as the alignment benchmark.

[0105] It's important to note that the time window set corresponds to key nodes in employee behavior manipulation (such as reporting and handover), and the window half-width reflects the duration of the key node's sustained impact on employee behavior. The smoothing scale parameter must be aligned with this duration. If the smoothing scale is much smaller than the window half-width, it will cause frequent fluctuations in the intensity sequence, making it impossible to capture the overall decoupling trend before and after the window; if the smoothing scale is much larger than the window half-width, it will cause over-smoothing of the intensity sequence, masking the decoupling differences before and after the window. Therefore, using the average window half-width as the alignment benchmark ensures that the smoothing scale matches the impact duration of the key node.

[0106] For example, if the time window set contains 3 windows with half-widths of 47.5 minutes, 30 minutes, and 42.5 minutes respectively, then the average half-width of the windows is (47.5+30+42.5) / 3=40 minutes, and this 40 minutes is the alignment reference.

[0107] In detail, the candidate interval for smoothing is defined within a multiple of the alignment benchmark. Typically, a range of 0.5 to 2 times the alignment benchmark is used (a reasonable range commonly used in the industry, balancing search efficiency and parameter coverage). That is, the lower limit of the candidate interval is 0.5 × the alignment benchmark, and the upper limit is 2 × the alignment benchmark.

[0108] Searching directly for the smoothing scale parameter across all positive numbers would result in extremely low search efficiency and could lead to anomalous parameters completely unrelated to the window features (e.g., a smoothing scale of 10 hours, much larger than the window half-width of 40 minutes). By limiting the candidate interval to multiples of the benchmark, we can ensure that all parameters within the interval are related to the window's influence duration. This narrows the search range, improves computational efficiency, avoids interference from irrelevant parameters, and guarantees the validity of subsequent search results.

[0109] For example, if the alignment baseline is 40 minutes, then the lower limit of the candidate interval is 0.5 × 40 = 20 minutes, and the upper limit is 2 × 40 = 80 minutes, that is, the candidate interval for smoothing scale is 20 minutes to 80 minutes.

[0110] It should be noted that the first-order autocorrelation function of the decoupling rate sequence reflects the similarity of the decoupling rate sequence at different time intervals. If the autocorrelation value of the sequence is positive at a certain time interval, it means that the decoupling rate trends before and after that interval are similar; if the autocorrelation value is negative, it means that the trends are opposite; if the autocorrelation value is 0, it means that the trends are unrelated.

[0111] The first zero-crossing point of the first-order autocorrelation function: This refers to the time interval during which the autocorrelation function first equals 0 as it transitions from positive to negative. This point reflects the effective trend period of the decoupling rate sequence. The closer the zero-crossing point is to the alignment reference, the better it matches the trend period of the sequence with the duration of the window's influence, accurately capturing the decoupling changes before and after the window.

[0112] For each candidate value of the smoothing scale within the candidate interval, the symmetric kernel smoothing function is substituted to smooth the proxy event flow and the actual event flow, and the corresponding decoupling rate sequence is calculated.

[0113] Calculate the objective function value for each candidate value: For the decoupling rate sequence corresponding to each candidate value, calculate the first zero-crossing position of its first-order autocorrelation function, then calculate the difference between this zero-crossing position and the alignment benchmark, and define the square of the difference as the objective function value. That is, the smaller the objective function value, the better the trend period of the decoupling rate sequence corresponding to the candidate value matches the duration of the window influence.

[0114] Traverse all candidate values ​​within the candidate interval, find the candidate value that minimizes the objective function value, and determine the candidate value as the optimal smoothing scale parameter.

[0115] The decoupling rate sequence identifies the communication and output decoupling before and after the window, while the optimal smoothing scale parameter must both eliminate random fluctuations and preserve the decoupling characteristics before and after the window. A one-dimensional search by minimizing the objective function ensures that the trend period of the final decoupling rate sequence accurately matches the duration of the window's influence, avoiding misjudgments of decoupling characteristics due to poor smoothing. For example, if the difference between the zero-crossing position corresponding to a candidate value and the alignment benchmark is small, it indicates that the candidate value clearly shows the trend of an increasing decoupling rate before the window and a decreasing decoupling rate after the window.

[0116] For example, if the candidate interval is 20 minutes to 80 minutes, then for a candidate value of 35 minutes:

[0117] After substituting into the smoothing function, the first zero-crossing position of the first-order autocorrelation of the decoupling rate sequence is calculated to be 38 minutes.

[0118] The alignment baseline is 40 minutes, the difference is 38-40=-2 minutes, and the objective function value is (-2). 2 =4;

[0119] For candidate values, 40 minutes:

[0120] The corresponding zero-crossing point is 39 minutes, the difference is -1 minute, and the objective function value is 1 (smaller);

[0121] Ultimately, 40 minutes was determined to be the optimal smoothing scale parameter because it minimizes the objective function value and best matches the sequence trend period with the window influence duration.

[0122] In one embodiment of the present invention, for each time window, a pre-window region, a post-window region, and a symmetrical baseline region are respectively divided. The mean of the decoupling rate sequence in the pre-window region, the mean in the post-window region, and the baseline mean in the symmetrical baseline region are calculated, including:

[0123] Use half of the time window as the window's half width;

[0124] The front area of ​​the window is obtained by extending the center of the window forward by half the width of the window.

[0125] The back area of ​​the window is obtained by extending the center of the window backward by half the width of the window.

[0126] A symmetrical baseline area is formed by extending one half-width of the window at equal intervals on both sides of the time window;

[0127] On the time axis, the mean before the window, the mean after the window, and the baseline mean are obtained by averaging the decoupling rate sequence over time in the pre-window, post-window, and symmetrical baseline regions, respectively.

[0128] It should be noted that each time window includes the window center time and the window half-width. The division of the three intervals is based on the window center time and the window half-width to ensure that the interval range matches the range of influence of the time window.

[0129] In detail, taking the center time of the time window as the reference, and extending it in the negative direction of the time axis (i.e., earlier times) by half a window width, the resulting time interval is the front area of ​​the window.

[0130] The purpose of employee manipulation is to concentrate communication activities before critical moments, and the half-width of the window reflects the duration of the sustained impact of critical moments on employee behavior. Extending forward by half a width precisely covers the time period when employees begin preparing and accumulating communication activities. For example, if a meeting starts at 9:00 AM, with a half-width of 47.5 minutes, the window's front area is 8:12:30 AM to 9:00 AM, covering the 47.5 minutes of communication accumulation before the meeting. This interval can be used to identify high-incidence areas of communication manipulation. For example, if the center time of a certain window is 9:00 AM, and the half-width is 47.5 minutes, then the window's front area is 9:00 AM minus 47.5 minutes to 9:00 AM, i.e., 8:12:30 AM to 9:00 AM.

[0131] In detail, taking the center time of the time window as the reference, extending half a window width in the positive direction of the time axis (i.e., a later time) forms the time interval that is called the back area of ​​the window.

[0132] After employees accumulate communication at the window, they concentrate on processing actual output behind the window. For example, focusing on completing work orders after a briefing leads to reduced communication and increased output behind the window. Extending the window by one and a half widths covers the time period when employees shift from communication manipulation to actual output. For example, if a briefing starts at 9:00 AM, with a half width of 47.5 minutes, the output period behind the window is 9:00 AM to 9:47:30 AM, covering the 47.5-minute peak output period after the briefing. This interval is used to define the areas where communication declines and output increases. For example, if the center time of a window is 9:00 AM and the half width is 47.5 minutes, then the output period behind the window is 9:00 AM to 9:00 AM plus 47.5 minutes, i.e., 9:00 AM to 9:47:30 AM.

[0133] In detail, using the start of the pre-window zone and the end of the post-window zone as benchmarks, a window half-width is extended in the negative and positive directions of the time axis, respectively. The two extended time intervals together constitute a symmetrical baseline zone. Both the pre-window and post-window zones are affected by key nodes (manipulation exists), while the symmetrical baseline zone is a normal period far from key nodes and free from manipulation. The baseline zone is divided into a pre-window baseline segment (extending forward by one half-width from the start of the pre-window zone) and a post-window baseline segment (extending backward by one half-width from the end of the post-window zone), symmetrically distributed on the time axis with the pre-window and post-window zones to avoid benchmark deviations in the decoupling rate due to differences in time period location; the duration of each baseline segment is one half-width, exactly the same as the pre-window and post-window zones (both half-width), ensuring that the subsequently calculated baseline mean is statistically comparable to the pre-window and post-window mean. This interval is used to establish a normal decoupling level benchmark. By comparing the difference between the pre-window and post-window mean and the baseline mean, the decoupling deviation caused by key nodes can be quantified. For example, if the pre-window mean is much higher than the baseline mean, it indicates that decoupling before the window is exacerbated by manipulation. For example, the center time of a certain window is 9:00, the half-width of the window is 47.5 minutes, the front area of ​​the window is 8:12:30-9:00, and the back area of ​​the window is 9:00-9:47:30.

[0134] In detail, on the time axis, for each pre-window region, post-window region, and symmetrical baseline region, the time average of the decoupling rate sequence within that interval is calculated, resulting in the pre-window mean, post-window mean, and baseline mean.

[0135] The decoupling rate sequence is a continuous numerical sequence (reflecting the degree of decoupling at each moment), while time averaging is the core method to convert the decoupling rate at all moments within an interval into a single quantified value. The calculation logic is as follows: the integral value of the decoupling rate sequence within the interval is divided by the interval duration, thus reflecting the average decoupling level within the entire interval.

[0136] The mean before the window reflects the average degree of decoupling in the area before the window (a high-incidence area of ​​communication manipulation). If the mean before the window is high, it means that the intensity of communication before the window is much higher than the intensity of actual action, and the decoupling is aggravated by manipulation.

[0137] The mean after the window reflects the average degree of decoupling in the area after the window (the output concentration area). If the mean after the window is low, it indicates that the communication intensity after the window has decreased and the implementation intensity has increased, and the decoupling has been alleviated due to the increase in output.

[0138] The baseline mean reflects the normal average degree of decoupling in the window-free interference area and serves as a benchmark value for comparing the deviation between the mean before and after the window: if the mean before the window is much higher than the baseline mean, it indicates that there is additional deviation caused by manipulation in the decoupling before the window; if the mean after the window is much lower than the baseline mean, it indicates that the decoupling after the window is due to additional mitigation caused by concentrated output.

[0139] In one embodiment of the present invention, the bias intensity index is calculated based on the pre-window mean, the post-window mean, and the baseline mean, including:

[0140] The ratio of the mean before the window to the baseline mean is used as the ratio before the window.

[0141] The ratio of the mean after the window to the baseline mean is called the post-window ratio.

[0142] The single-window offset is obtained by taking the logarithm of the product of the ratio before and after the window.

[0143] The average of all single-window offsets corresponding to the time window set is calculated to obtain the offset strength index.

[0144] It should be noted that the baseline mean is the average decoupling rate of the symmetrical baseline area, representing the normal decoupling level of employees when there is no window interference; the window mean is the average decoupling rate of the area in front of the window, representing the decoupling level of employees in front of the window affected by manipulative behavior. By dividing the window mean by the baseline mean, the decoupling level in front of the window can be converted into a multiple relationship relative to the normal level, intuitively reflecting the degree of deviation of the decoupling in front of the window.

[0145] If the ratio before the window is greater than 1, it means that the decoupling rate before the window is higher than the normal level. The larger the ratio, the more serious the decoupling caused by the accumulation of communication before the window (e.g., if the ratio is 3, it means that the decoupling before the window is 3 times the normal level).

[0146] If the ratio in front of the window is equal to 1, it means that the unhooking in front of the window is consistent with the normal level and there is no additional bias.

[0147] If the ratio before the window is less than 1, it means that the unhooking before the window is below the normal level.

[0148] It should be noted that the post-window mean is the average decoupling rate in the post-window region, representing the level of decoupling after employees shift to actual output. The deviation of post-window decoupling from the normal level can be quantified by the ratio of the post-window mean to the baseline mean.

[0149] If the ratio after the window is less than 1, it means that the decoupling rate after the window is lower than the normal level. The smaller the ratio, the more obvious the decoupling relief caused by the concentration of output after the window (e.g., if the ratio is 0.25, it means that the decoupling after the window is only 25% of the normal level).

[0150] If the ratio after the window is equal to 1, it means that the unhooking after the window is consistent with the normal level and there is no additional bias.

[0151] If the ratio after the window is greater than 1, it indicates that the unhooking after the window is higher than the normal level.

[0152] It should be noted that employees' manipulative behaviors around the window exhibit synergy. The accumulation of communication before the window (previous ratio > 1) and the concentration of output after the window (later ratio < 1) are two stages of the same manipulative behavior, and their synergistic effect needs to be synthesized through multiplication.

[0153] If typical manipulation behavior exists, the ratio before the window is >1 and the ratio after the window is <1, and their product is usually close to 1 but slightly deviates from it;

[0154] The deviation of the product is converted into an offset that can be directly judged as positive or negative. If the product > 1, the logarithmic result is positive, indicating that there is a positive bias in the single window (higher at the front and lower at the back) (the manipulation behavior causes asymmetry in the unhooking); if the product = 1, the logarithmic result is 0, indicating that there is no bias; if the product < 1, the logarithmic result is negative, indicating that there is a negative bias.

[0155] It should be noted that the time window set usually includes multiple windows (e.g., one report per week, four windows if the evaluation period is four weeks). The offset of a single window may be affected by accidental factors (e.g., an unexpected communication need before a certain week's window). Accidental biases need to be eliminated by averaging to reflect the overall degree of employee bias around all windows throughout the entire evaluation period.

[0156] If the bias intensity index is >0, it means that employees have a decoupling bias with the front higher than the back in most windows. The larger the value, the more serious the overall bias (e.g., if the bias intensity index = 0.1, it means that there is an average of 10% positive bias in all windows).

[0157] If the bias strength index = 0, it means that the biases of all windows cancel each other out or there is no bias, and the employee behavior is less affected by the windows.

[0158] If the bias strength index is less than 0, it indicates that there is a reverse bias.

[0159] The bias strength index quantifies the overall degree to which employee behavior is manipulated by the window.

[0160] In one embodiment of the present invention, the phase stability index is calculated based on the peak value of the decoupling rate sequence and the center time of the corresponding time window, including:

[0161] The peak value of the decoupling rate sequence is defined as the moment corresponding to the peak value within the time window.

[0162] The relative position angle is obtained by normalizing the time difference between the peak time and the corresponding window center time.

[0163] The phase stability index is obtained by aggregating all relative position angles corresponding to the time window set using the circular statistical method.

[0164] It's important to note that for each time window, the maximum value of the decoupling rate sequence within the window's coverage area is extracted, and the moment corresponding to this maximum value is defined as the peak moment. The peak of the decoupling rate sequence represents the moment when communication and output are most severely decoupled within that window; that is, for employees engaging in manipulative behavior, the peak moment typically corresponds to the time point before the window when communication is most concentrated. For example, multiple communication messages might be sent in concentrated bursts 30 minutes before a reporting meeting. Locating the peak moment helps identify the core time points of employee manipulative behavior around the window.

[0165] It should be noted that while the time difference itself is linear data (e.g., -30 minutes, -25 minutes), the time difference between different windows cannot be directly used for consistency analysis. Because the window coverage interval is periodically repetitive (e.g., the weekly reporting window), the linear time difference needs to be converted into periodic angular data to be suitable for circular statistical methods.

[0166] Using the total duration of the window coverage interval (twice the window half-width, i.e., the duration from center time - half-width to center time + half-width) as the denominator, normalize the time difference to the range [-1, 1], then multiply by π to convert it to an angle [-π, π], and finally adjust it to the range [0, 2π) by adding π to the angle. The formula can be simplified to:

[0167] Relative position angle = 2π × [(peak time - window center time) ÷ (2 × window half width)] + π;

[0168] This ensures that the closer the peak times are to the same relative position (e.g., both are 30 minutes before the window), the closer the corresponding relative position angles are.

[0169] For example, the window center time = 9:00, the peak time = 8:40, the time difference = 8:40 - 9:00 = -20 minutes; the window half-width = 47.5 minutes, 2 times the half-width = 95 minutes; then the relative position angle = 2π × [(-20) ÷ 95] + π ≈ 2π × (-0.2105) + π ≈ -1.322 + 3.1416 ≈ 1.8196 radians (approximately 104.2 degrees), that is, the relative position angle of the window ≈ 1.8196 radians.

[0170] It should be noted that the circular statistical method is used to aggregate the relative position angles of all windows in the time window set: first, each angle is converted into a complex number (the modulus of the complex number is 1, and the argument is the relative position angle); then, the arithmetic mean of all complex numbers is calculated (the sum of complex numbers divided by the total number of windows); finally, the modulus of the average is subtracted from 1, and the result is defined as the phase stability index.

[0171] The relative position angle is a periodic data in the range [0, 2π). Conventional linear statistics (such as arithmetic mean) will cause deviation (e.g., the conventional mean of 350 degrees and 10 degrees is 180 degrees, but it should actually be closer to 0 degrees). Circular statistics must be used.

[0172] Angles are transformed into points on the unit circle of the complex plane, and the positions of these points can intuitively reflect the distribution of angles; for example, angle θ corresponds to the complex number cosθ + i×sinθ.

[0173] The modulus of the complex mean represents the degree of concentration of all points on the unit circle. The closer the modulus is to 1, the more concentrated the points are (the more consistent the angles); the closer the modulus is to 0, the more dispersed the points are (the more chaotic the angles).

[0174] Subtract the modulus of the complex mean from 1 to make the index range [0,1]. The closer the phase stability index is to 0, the more concentrated the relative position angles of different windows are (the more consistent the peak operation time of employees in all windows, such as always being 30 minutes before the report); the closer the phase stability index is to 1, the more dispersed the angles are (the peak operation time is irregular and has a large influence from accidental factors).

[0175] Phase stability index quantifies the temporal consistency of employee manipulation behavior. Combined with bias strength index, it can fully determine the dependence of employee behavior on the window (e.g., high bias strength and low phase stability indicate that the employee has regular and serious manipulation behavior).

[0176] For example, the time window set contains four windows with relative position angles of 1.8196, 1.798, 1.832, and 1.805 radians, respectively, and their corresponding complex numbers are:

[0177] Angle 1.8196: cos1.8196 + isin1.8196 ≈ -0.245 + 0.9694i;

[0178] Angle 1.798: cos1.798 + isin1.798 ≈ -0.218 + 0.9760i;

[0179] Angle 1.832: cos1.832 + isin1.832 ≈ -0.261 + 0.9653i;

[0180] Angle 1.805: cos1.805 + isin1.805 ≈ -0.226 + 0.9741i;

[0181] The sum of complex numbers ≈ (-0.245-0.218-0.261-0.226) + (0.9694+0.9760+0.9653+0.9741)i ≈ -0.95+3.8848i;

[0182] The complex mean is approximately (-0.95 + 3.8848i) ÷ 4, which is approximately -0.2375 + 0.9712i.

[0183] The modulus of the mean is approximately 0.9998;

[0184] Where i represents the imaginary unit, i 2 Equals -1;

[0185] The phase stability index = 1 - 0.9998 = 0.0002, indicating that the relative position angles of the four windows are extremely concentrated, and the peak time of employee operation is highly consistent.

[0186] In one embodiment of the present invention, a reporting dependence index reflecting the degree to which employee behavior depends on a time window is generated by combining a bias strength index and a phase stability index, including:

[0187] The larger value between the bias strength index and 0 is taken as the non-negative bias.

[0188] The difference between 1 and the phase stability index is taken as the phase concentration.

[0189] The product of the non-negative bias and the phase concentration is used as the reporting dependency index.

[0190] It should be noted that the bias strength index represents the overall deviation of employees from decoupling around the time window, and its value may be positive, negative, or zero.

[0191] When the bias strength index is >0, it indicates that employees have a positive bias in most windows, with high decoupling in front of the window and low decoupling behind the window (i.e., typical communication manipulation behavior). This value directly reflects the effective magnitude of the bias and should be retained in its entirety.

[0192] When the bias strength index is ≤0, it indicates that employees either have no obvious bias or exhibit a reverse bias (index <0) with low decoupling before the window and high decoupling after the window. Since reverse bias has no practical management significance—that is, employees will not reduce communication before the window and increase communication after the window to cater to the index, and such situations are extremely rare—the bias strength index ≤0 needs to be corrected to 0 to ensure that the non-negative bias only reflects the actual positive communication manipulation amplitude and avoids interference from invalid or reverse data.

[0193] It should be noted that the phase stability index has a value range of [0,1], and is the degree of dispersion of the decoupling peak time calculated by the circular statistical method.

[0194] The closer the phase stability index is to 0, the more concentrated the decoupling peak times are in different windows (such as communication backlog peaks appearing 30 minutes before each reporting meeting), and the more regular the employees' time response to the window.

[0195] The closer the phase stability index is to 1, the more dispersed the decoupling peak times are in different windows (such as peaks occurring 20 minutes before a certain report and 40 minutes before another report), and the more chaotic the employees' time response to the window is.

[0196] The dispersion is transformed into the concentration by using the 1-phase stability index:

[0197] The phase concentration range is also [0,1], and it is inversely related to the phase stability index.

[0198] The closer the phase concentration is to 1, the closer the corresponding phase stability index is to 0, indicating that the decoupling peak time is more concentrated and the temporal regularity of employee manipulation behavior is stronger.

[0199] The closer the phase concentration is to 0, the closer the corresponding phase stability index is to 1, indicating that the decoupling peak times are more dispersed and the temporal regularity of employee manipulation behavior is weaker.

[0200] It should be noted that report dependency indicates that employees regularly adjust their behavior around a time window, leading to a decoupling between communication and output. Two conditions must be met simultaneously:

[0201] 1. The decoupling amplitude is large enough (high non-negative bias).

[0202] 2. The decoupling time is sufficiently regular (high phase concentration).

[0203] Multiplication allows for the coordinated assessment of two types of conditions, ensuring that the exponent only increases significantly when both are high, as detailed below:

[0204] Scenario 1: High non-negative bias (e.g., 0.1108) and high phase concentration (e.g., 0.9998), the product result is approximately 0.1108 × 0.9998 ≈ 0.1107 (high exponent). This indicates that employees exhibit window dependency with "large amplitude and strong regularity," requiring priority for bias removal.

[0205] Scenario 2: High non-negative bias (e.g., 0.1108) but low phase concentration (e.g., 0.2), the product result is approximately 0.1108 × 0.2 ≈ 0.0222 (low exponent). This indicates a large decoupling amplitude but irregular timing, which may be due to accidental factors (e.g., a sudden communication need in a certain week). It does not belong to purposeful window dependency and does not require bias removal.

[0206] Scenario 3: Low non-negative bias (e.g., 0) and high phase concentration (e.g., 0.9998), product result = 0 × 0.9998 = 0 (low exponent). This indicates strong time regularity but no decoupling amplitude; employees have a fixed behavioral rhythm but are not manipulated, so no bias removal is needed.

[0207] Scenario 4: Low non-negative bias (e.g., 0) and low phase concentration (e.g., 0.2), the product result = 0 × 0.2 = 0 (low exponent). This indicates no decoupling amplitude and no regularity, with no window dependency whatsoever.

[0208] In one embodiment of the present invention, bias removal processing is performed on the employee's proxy event flow and implementation event flow based on the report dependency index, including:

[0209] If the report relies on an index exceeding the limit, a debiasing process will be performed, including:

[0210] A time weighting function is constructed for the time window; the weight strength parameter of the time weighting function is specifically calculated as follows: the square of the difference between the mean of the decoupling rate sequence before the window and the mean of the decoupling rate sequence after the window is calculated, and the square of the difference between all the means corresponding to the time window set is minimized to obtain the weight strength parameter applied to the time weighting function.

[0211] The proxy event flow and the actual event flow are weighted first and second by a time weighting function and a symmetric kernel smoothing function, respectively, to obtain a weighted proxy event intensity sequence and a weighted actual event intensity sequence. Within a preset assessment period, the weighted proxy event intensity sequence and the weighted actual event intensity sequence are used as inputs to a preset assessment model to obtain the debiased assessment results.

[0212] It should be noted that the report dependency index is compared with the enterprise's preset debiasing trigger threshold: if the report dependency index is greater than the debiasing trigger threshold, it is determined that the employee's behavior is dependent on the time window beyond the acceptable range, and subsequent debiasing processing is performed; if the index is less than or equal to the debiasing trigger threshold, it is determined that there is no significant dependency, and the original proxy / implementation event intensity sequence is used directly for assessment without debiasing.

[0213] The report dependency index ranges from [0,1], reflecting the severity of window dependency. Companies need to set differentiated thresholds based on the job characteristics of different positions (e.g., functional positions have high communication frequency, technical positions have a high proportion of hands-on work), rather than using a uniform standard.

[0214] For example, the threshold for functional positions can be set to 0.3 (a slight dependence is allowed for communication-related jobs), and the threshold for technical positions can be set to 0.2 (a strict de-biasing is required for practical jobs).

[0215] Setting a threshold for bias correction is crucial to distinguish between normal behavioral fluctuations and abnormal manipulation. Exceeding the threshold indicates that the decoupling of employee communication and output is due to regular human manipulation, rather than random fluctuations. In such cases, bias correction is necessary to avoid wasting resources.

[0216] It should be noted that employees' manipulative behaviors are concentrated before and after the time window (communication accumulates before the window, and output is concentrated after the window). The core function of the time weighting function is to assign lower weights to events in the high-incidence area of ​​manipulation near the window (to weaken the impact of manipulative behavior on the assessment) and assign higher weights to events in the normal behavior area outside the window (to retain the true behavioral characteristics), thereby achieving neutralization.

[0217] It should be noted that a window-sensitive exponential weighting is used, constructed based on the center time and half-width of each time window. The formula for the time weighting function is as follows:

[0218]

[0219] Where K is the total number of windows in the time window set; ψ k (t) is the weight kernel (Gaussian kernel) for the k-th time window. , τ k Indicates the time at the center of the window;

[0220] Δ k λ is the half-width of the window; λ is the weight strength parameter, that is, the larger λ is, the more the weights near the window decrease, and the stronger the debiasing is; the smaller λ is, the weaker the debiasing is. Its value needs to be determined by data-driven methods.

[0221] λ needs to be determined by minimizing the difference between the mean decoupling rates before and after the window, specifically:

[0222] For the current candidate value of λ, substituting it into the time weighting function yields the weighted decoupling rate sequence:

[0223]

[0224] Then calculate the pre-window mean and post-window mean of the weighted decoupling rate sequence for each window, and calculate the square of the difference between the pre-window mean and the post-window mean;

[0225] The total target value is obtained by summing the squares of the differences between all windows in the time window set.

[0226] Within the range of λ≥0 (λ is non-negative to ensure that the weight does not increase), traverse the candidate values ​​and calculate the total objective value, find the λ that minimizes the total objective value, which is the final weight strength parameter;

[0227] Specifically, balance the decoupling rates before and after the weighted window:

[0228] If employees engage in manipulative behavior, the original decoupling rate sequence satisfies the condition that the mean before the window is greater than the mean after the window, and the square of the mean difference is larger.

[0229] By adjusting λ, the weights near the window are reduced, the influence of manipulation is weakened, and the squared mean difference gradually decreases.

[0230] When the total target value is minimized, the difference in decoupling rate before and after the window is minimized, and the debiasing effect is optimal.

[0231] It should be noted that the first weighting is used to correct bias. By reducing the weight of events in high-incidence areas of manipulation near the window, the bias caused by the backlog of communication before the window and the concentration of output after the window is directly neutralized.

[0232] The second weighting is used to optimize data quality. Even after time-weighting, the event stream may still have random fluctuations (such as a sudden communication message at a non-window time). The fluctuations are eliminated by smoothing with a symmetric kernel to ensure that the intensity sequence reflects the true communication and output trends.

[0233] In detail, the preset assessment cycle is the efficiency evaluation cycle set by the enterprise according to management needs. It is usually weekly, monthly or quarterly. It needs to match the cycle of the time window set (e.g., if the window is a weekly report meeting, the assessment cycle is set to weekly) to ensure that the bias removal process covers the complete window-dependent behavior cycle and avoids the bias caused by the cycle being too short and not being completely eliminated.

[0234] In detail, the pre-set assessment model is the core model used by enterprises to calculate employee efficiency. Its inputs are a weighted sequence of proxy event intensity and a weighted sequence of actual event intensity, and the output is a debiased efficiency score. The model format is consistent with the enterprise's existing assessment logic, only replacing the original intensity sequence with a weighted intensity sequence.

[0235] For example, the original efficiency model for an enterprise is: Efficiency = Total number of actual events in the period / Total number of proxy events in the period. Then, the debiased model is: Debiased efficiency = Weighted total number of actual events in the period / Weighted total number of proxy events in the period.

[0236] Among them, the total weighted implementation events are the integral value of the weighted implementation event intensity sequence within the assessment period, and the total weighted proxy events are the integral value of the weighted proxy event intensity sequence within the assessment period.

[0237] In detail, the efficiency score calculated by substituting the weighted proxy event intensity sequence and the weighted implementation event intensity sequence into the preset assessment model within the preset assessment period is the debiased assessment result. This debiased assessment result represents the true efficiency after eliminating window dependency bias.

[0238] If the original performance evaluation results of employees are underestimated due to window dependence (e.g., there is more output after the window but the original statistics are not corrected), the results after debiasing will increase significantly.

[0239] If the original results are overestimated due to accumulated communication (e.g., a lot of communication at the window but little output), the results after debiasing will return to the true level.

[0240] Ultimately, this will enable comparable efficiency across cycles and job roles, resolving the technical issue of masking true efficiency.

[0241] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.

Claims

1. An enterprise digital intelligent internal control management system based on big data analysis, characterized in that: include: The data acquisition module extracts proxy event streams reflecting communication behaviors and implementation event streams reflecting actual outputs from the enterprise business database, taking employees as the unit; it constructs a set of time windows, each containing the center time of the window; The data processing module applies a symmetric kernel smoothing function to the proxy event stream and the actual event stream respectively to obtain the proxy event intensity sequence and the actual event intensity sequence. It determines the decoupling rate sequence based on the ratio of the proxy event intensity sequence and the actual event intensity sequence, and determines the smoothing scale parameter of the symmetric kernel smoothing function based on the decoupling rate sequence and the time window set. The indicator calculation module divides each time window into a pre-window zone, a post-window zone, and a symmetrical baseline zone, and calculates the pre-window mean of the decoupling rate sequence in the pre-window zone, the post-window mean in the post-window zone, and the baseline mean in the symmetrical baseline zone. The bias intensity index, which reflects the degree of deviation of the decoupling rate sequence from the baseline region in the pre-window and post-window regions, is calculated based on the mean before the window, the mean after the window, and the baseline mean. The phase stability index, which reflects the stability of the position of the decoupling rate sequence peak within the time window, is calculated based on the peak value of the decoupling rate sequence and the corresponding center time of the time window. The bias removal module, combined with the bias strength index and the phase stability index, generates a report dependency index that characterizes the pattern of employees adjusting proxy events around the time window. Based on the report dependency index, the employees' proxy event flow and actual event flow are debiased, and the debiased proxy event flow and actual event flow are input into the preset assessment model to obtain the debiased assessment results.

2. The enterprise digital intelligent internal control management system based on big data analysis according to claim 1, characterized in that, A symmetric kernel smoothing function is applied to the surrogate event stream and the actual event stream respectively to obtain the surrogate event intensity sequence and the actual event intensity sequence. The decoupling rate sequence is determined based on the ratio of the surrogate event intensity sequence to the actual event intensity sequence, including: The decoupling rate sequence is obtained by performing a ratio operation on the proxy event intensity sequence and the actual event intensity sequence at each time point on the time axis.

3. The enterprise digital intelligent internal control management system based on big data analysis according to claim 2, characterized in that, The smoothing scale parameters of the symmetric kernel smoothing function are determined based on the decoupling rate sequence and the time window set, including: The average half-width of the time window set is used as the alignment reference; candidate intervals for smoothing scales are determined within multiples of the alignment reference; within the candidate intervals, a one-dimensional search is performed with the square of the difference between the first zero-crossing position of the first autocorrelation function of the decoupling rate sequence and the alignment reference as the objective function to obtain the smoothing scale parameter that minimizes the objective function, and this smoothing scale parameter is used for the symmetric kernel smoothing function.

4. The enterprise digital intelligent internal control management system based on big data analysis according to claim 3, characterized in that, For each time window, the system is divided into a pre-window region, a post-window region, and a symmetrical baseline region. The mean of the decoupling rate sequence in the pre-window region, the post-window region, and the baseline region are calculated, including: Use half of the time window as the window's half width; The front area of ​​the window is obtained by extending the center of the window forward by half the width of the window. The back area of ​​the window is obtained by extending the center of the window backward by half the width of the window. A symmetrical baseline area is formed by extending one half-width of the window at equal intervals on both sides of the time window; On the time axis, the mean before the window, the mean after the window, and the baseline mean are obtained by averaging the decoupling rate sequence over time in the pre-window, post-window, and symmetrical baseline regions, respectively.

5. The enterprise digital intelligent internal control management system based on big data analysis according to claim 4, characterized in that, The bias intensity index is calculated based on the pre-window mean, post-window mean, and baseline mean, including: The ratio of the mean before the window to the baseline mean is used as the ratio before the window. The ratio of the mean after the window to the baseline mean is called the post-window ratio. The single-window offset is obtained by taking the logarithm of the product of the ratio before and after the window. The average of all single-window offsets corresponding to the time window set is calculated to obtain the offset strength index.

6. The enterprise digital intelligent internal control management system based on big data analysis according to claim 5, characterized in that, Phase stability indices are calculated based on the peak value of the decoupling rate sequence and the center moment of the corresponding time window, including: The peak value of the decoupling rate sequence is defined as the moment corresponding to the peak value within the time window. The relative position angle is obtained by normalizing the time difference between the peak time and the corresponding window center time. The phase stability index is obtained by aggregating all relative position angles corresponding to the time window set using the circular statistical method.

7. The enterprise digital intelligent internal control management system based on big data analysis according to claim 6, characterized in that, By combining the bias strength index and the phase stability index, a reporting dependence index is generated that reflects the degree to which employee behavior depends on time windows, including: The larger value between the bias strength index and 0 is taken as the non-negative bias. The difference between 1 and the phase stability index is taken as the phase concentration. The product of the non-negative bias and the phase concentration is used as the reporting dependency index.

8. The enterprise digital intelligent internal control management system based on big data analysis according to claim 7, characterized in that, The employee's proxy event flow and implementation event flow are debiased based on the report dependency index, including: If the report relies on an index exceeding the limit, a debiasing process will be performed, including: A time weighting function is constructed for the time window; the weight strength parameter of the time weighting function is specifically calculated as follows: the square of the difference between the mean of the decoupling rate sequence before the window and the mean of the decoupling rate sequence after the window is calculated, and the square of the difference between all the means corresponding to the time window set is minimized to obtain the weight strength parameter applied to the time weighting function. The proxy event flow and the actual event flow are weighted first and second by a time weighting function and a symmetric kernel smoothing function, respectively, to obtain a weighted proxy event intensity sequence and a weighted actual event intensity sequence. Within a preset assessment period, the weighted proxy event intensity sequence and the weighted actual event intensity sequence are used as inputs to a preset assessment model to obtain the debiased assessment results.

Citation Information

Patent Citations

  • Systems and / or methods for event stream deviation detection

    CN102902699A

  • Project intelligent management method and system based on enterprise digital information

    CN120198073A