Data security monitoring and analysis method for bidding informatization platform
By using end-to-end data collection and scenario segmentation, abnormal behaviors of bidding platforms are identified, an abnormal behavior relationship graph is constructed, and path sorting is dynamically adjusted. This solves the problem of difficulty in identifying abnormal behaviors in existing technologies and achieves efficient data security monitoring and early warning.
Patent Information
- Application Number
- CN202511374534.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing technologies ignore abnormal behavior in user uploads, making it difficult to identify anomalies in bidding data and respond quickly to complex abnormal behaviors.
By collecting data from the bidding platform across the entire process, user behavior is divided into multiple monitoring points based on the operation object, operation quantity, and operation type. Scene segmentation and abnormal behavior identification are performed, an abnormal behavior relationship graph is constructed, and the path sorting priority is dynamically adjusted to achieve efficient early warning of abnormal behavior.
This improved the accuracy of identifying abnormal behavior and the timeliness of early warnings, ensuring the precision and efficiency of data security monitoring and analysis on the bidding platform.
Smart Images

Figure CN120880786A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information monitoring technology, specifically a data security monitoring and analysis method for a bidding and tendering information platform. Background Technology
[0002] Tendering and bidding, or the process by which a buyer publishes its procurement needs to the public, includes credit sales by manufacturing companies to corporate clients, i.e., product credit sales. In the process of product credit sales, the credit provider is typically a material supplier, product manufacturer, or wholesaler, while the buyer is the beneficiary, and these are various corporate clients or agents.
[0003] For example, Chinese Patent Publication No. CN114971796A discloses a bidding system based on a cloud service platform. The system comprises: a transaction platform (for receiving bidding applications from mobile users and verifying their identities); a public service platform (for verifying user identity information, generating bidding transaction data, and managing bidding); and an administrative supervision platform (for receiving bidding data from mobile users, synchronously signing it in the cloud, and generating bidding data).
[0004] For example, Chinese Patent Publication No. CN119027223A discloses a bidding information management method and system based on an electronic bidding transaction platform, including the following steps: S01, importing readable data from multiple data sources into a database; S02, managing the relevant metadata of the corresponding project components in the imported database data based on the project components stored in the database, to form a metadata repository matching the data volume of the database; S03, matching the metadata repository with the data import module, and generating difference information between mismatched information and multiple identifiable database entities located under the metadata repository. This invention simplifies the process of filling out bidding templates and intelligently processes bidding templates through the bidding document.
[0005] In existing technologies, the transaction entities in the bidding and tendering scenario are described, and the bidding and tendering data is processed by using encrypted authorization authentication. The information under the bidding and tendering text is checked to complete the overlay processing of the bidding and tendering data. However, existing technologies ignore abnormal behavior of user upload behavior, making it difficult to identify anomalies in the received data and unable to respond quickly to complex abnormal behaviors. Summary of the Invention
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a data security monitoring and analysis method for a bidding and tendering information platform, comprising: S1, collecting bidding and tendering platform data across the entire chain, dividing user behavior into multiple monitoring points based on the operation objects, operation quantity and operation type of the bidding and tendering platform data in a continuous time period, and obtaining a user behavior sequence composed of each monitoring point.
[0007] S2, segment the scenario based on the user behavior implementation scenario, and use the time distribution of user behavior processing as the main body to determine the abnormal behavior relative to the historical behavior baseline under each scenario segmentation dimension.
[0008] S3, obtain the correlation values between each abnormal behavior, and then merge the abnormal behavior into multiple correlation sets according to the correlation values, and determine the behavior correlation result between each correlation set and the user operation object.
[0009] S4 uses the behavioral association results to trace the source of data association, and connects the behavioral association results in the form of the distribution of monitoring points to determine the association path under different scene segmentation dimensions.
[0010] S5 performs iterative verification of the associated paths under all scene segmentation dimensions, updates the path list of each associated path combination based on the intersection of the iterative verification, and uses the updated path list as the output warning target.
[0011] The beneficial effects of this invention are as follows: First, this invention monitors bidding platform data according to the operation object, operation quantity, and operation type, and uses user behavior sequences deployed under multiple monitoring points to avoid the problem of incorrect user behavior identification under a single dimension; then, it performs scene segmentation on user behavior, identifies abnormal behaviors in the current user behavior sequence relative to the historical behavior baseline using multiple scene segmentation dimensions, improves the adaptability of the historical behavior baseline through state transition path modeling and cross-validation, and achieves more accurate initial screening of abnormal behaviors by combining comprehensive risk scores.
[0012] Second, this invention abstracts abnormal behavior into graph nodes, calculates correlation values through text similarity and behavior similarity, constructs an abnormal behavior relationship graph, and extracts a feature pool of operation objects to realize the correlation mining of abnormal behavior. The feature pool sorting mechanism further improves the focus of the correlation set and enhances the correlation and accuracy of isolated analysis of abnormal behavior.
[0013] Third, this invention generates associated paths by sorting monitoring points according to global timestamps, and dynamically adjusts the path sorting priority by verifying the positive correlation between the number of monitoring points and the value range. The positive correlation between the value range and the number of monitoring points enables dynamic adjustment of path priority, ensuring that highly concentrated abnormal paths are responded to first. Finally, the intersection of paths is calculated through subset cross-validation, and high-frequency intersection paths are merged by combining frequency thresholds and their priority is moved up. This completes the merging analysis under each scenario segmentation dimension, realizes dynamic updating of the path list, and improves the timeliness and accuracy of early warning. Attached Figure Description
[0014] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0015] Figure 1 This is a flowchart illustrating a data security monitoring and analysis method for a bidding and tendering information platform.
[0016] Figure 2 This is a flowchart illustrating step S2 of a data security monitoring and analysis method for a bidding and tendering information platform.
[0017] Figure 3 This is a flowchart illustrating step S3 of a data security monitoring and analysis method for a bidding and tendering information platform.
[0018] Figure 4 This is a flowchart illustrating step S4 of a data security monitoring and analysis method for a bidding and tendering information platform.
[0019] Figure 5 This is a flowchart illustrating step S5 of a data security monitoring and analysis method for a bidding and tendering information platform. Detailed Implementation
[0020] The embodiments of the present invention are described in detail below. The embodiments described below are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. Where specific techniques or conditions are not specified in the embodiments, they shall be performed in accordance with the techniques or conditions described in the literature in the art or in accordance with the product manual.
[0021] See Figure 1 A data security monitoring and analysis method for a bidding and tendering information platform includes: S1, collecting bidding and tendering platform data across the entire chain, dividing user behavior into multiple monitoring points based on the operation objects, operation quantity, and operation type of the bidding and tendering platform data in a continuous time period, and obtaining a user behavior sequence composed of each monitoring point.
[0022] S2, segment the scenario based on the user behavior implementation scenario, and use the time distribution of user behavior processing as the main body to determine the abnormal behavior relative to the historical behavior baseline under each scenario segmentation dimension.
[0023] S3, obtain the correlation values between each abnormal behavior, and then merge the abnormal behavior into multiple correlation sets according to the correlation values, and determine the behavior correlation result between each correlation set and the user operation object.
[0024] S4 uses the behavioral association results to trace the source of data association, and connects the behavioral association results in the form of the distribution of monitoring points to determine the association path under different scene segmentation dimensions.
[0025] S5 performs iterative verification of the associated paths under all scene segmentation dimensions, updates the path list of each associated path combination based on the intersection of the iterative verification, and uses the updated path list as the output warning target.
[0026] When collecting data from the bidding platform, the data chain process, upload time, and relevant data of the bid text and file attributes are used to identify user behavior when uploading on the bidding platform. By analyzing the user's operation type, operation object, operation time, operation quantity, and the text data involved in the operation, the rationality of the user's uploaded bidding documents is comprehensively analyzed. The system monitors whether there are specific cases of bid rigging or collusion in the currently received bidding documents, so as to determine whether each user has abnormal behavior when uploading relevant bid documents.
[0027] The data collected at this time includes, but is not limited to, business operation logs, file metadata, and network traffic metadata; the business operation logs describe the username, timestamp, operation type (upload, modify, delete, etc.), operation object (such as file ID, file name, size, etc.), upload IP, encryption certificate ID used, and upload result.
[0028] File metadata represents the specific text of the uploaded tender document; batch processing of the text contained in the file, as well as file attributes such as creator, last modifier, creation time, and last modification time, is used to determine whether there are any abnormal associations between tender documents from multiple different users under the same tender document.
[0029] Network traffic data includes source IP, destination IP, port, protocol, data volume transmitted, and timestamp. This information represents the upload status under business operations and is used to identify abnormal uploads or user behavior involving large-volume abnormal access. This data is combined to form a complete user behavior report, explaining the specific actions the user took during the upload process.
[0030] The above operation time represents the timestamp when the user action was performed, and the operation quantity represents the number of files involved in each operation. Furthermore, at this time, based on the total number of files received by the system in a certain time period, it will analyze whether some of the tender documents have abnormal behavior or whether there is excessive text similarity. These user actions will be classified into nodes to determine the changes in user behavior under the time series of multiple time period combinations.
[0031] The implementation of step S1 also includes: according to the time window corresponding to each monitoring point, traversing all user behaviors, determining the context information of user behaviors within each time window, the context information represents the process of successful login → uploading files → downloading files → logging out, and subsequent re-authentication → modifying files → logging out, etc. The context information is used to mark the overall process of the user uploading the tender documents, and to mark the corresponding user, time and IP to assist in the subsequent anomaly identification of each user behavior data.
[0032] Based on the contextual information of user behavior, discrete user behaviors are aggregated in chronological order to form a continuous user behavior sequence; it is then determined whether the timestamp of each user behavior in the user behavior sequence corresponds to the timestamp in the tender document, and the result is marked in the user behavior sequence.
[0033] Once acquired, this user behavior data will be converted into a unified, standardized format to facilitate the subsequent review of inconsistent or abnormal behavioral information regarding user-uploaded and processed events.
[0034] The timestamps are used to determine whether the timestamps of the last modification and storage of the bid documents correspond to the timestamps specified in the bidding documents, i.e., whether the time requirements are met or whether there are any timeouts in the submission or modification information. This part of the processing is equivalent to data preprocessing, which verifies the relevant operations performed by users when uploading bid documents in the data obtained from the bidding platform, and checks their timestamps to determine whether there are any inconsistencies in timestamp markings.
[0035] Preferably, the aforementioned monitoring points represent data indexes corresponding to each user behavior, and the number of monitoring points is equal to the number of tender documents currently being processed. That is, the monitoring points will represent the number of operations performed by the user in a continuous time period, and monitor the uploading, modification, and deletion of each tender document. The monitoring points are placed at the operation object, operation quantity, and operation type in the continuous time period to illustrate the corresponding tender documents in the operation object, operation quantity, and operation type, as well as the data formats that need to be paid attention to in each tender document.
[0036] In one embodiment of the present invention, when performing scene segmentation, a comprehensive verification is performed based on the data chain flow and upload time of the user-uploaded bid documents. The data chain flow verification verifies the data of the bid documents under the creation time, modification records, version iterations, etc., to illustrate the process of the current file upload, modification, and deletion corresponding to the historical user behavior. The upload time emphasizes the time anomalies of the bid document submission, and the upload time of the bid documents in the platform is centrally analyzed. By comparing the time distribution of the bid documents uploaded in the past, abnormal concentrated submissions are identified, and it is checked whether the file upload distribution conforms to the normal operating rhythm. The user's upload behavior pattern is identified according to the form of its upload time distribution to identify abnormal behaviors relative to the historical behavior baseline, such as abnormal behaviors of bid rigging and collusion.
[0037] like Figure 2 As shown, the implementation of step S2 includes: S21, treating each scene segmentation dimension as a processing node, verifying the user behavior under each processing node relative to the historical behavior baseline of historical data, which is set according to the time distribution of historical uploaded tender documents.
[0038] For example, analyzing the time distribution of historical uploaded bid documents helps determine the usual upload time periods by analyzing the time distribution of bid documents uploaded by a particular user or all users of the same company. Then, by merging the behavior of multiple users, it can be determined whether the bid documents are similar in time distribution, thus enabling anomaly detection in the time distribution dimension.
[0039] Preferably, the scenario segmentation dimension represents the segmentation based on the time period in which the user's behavior occurs, the user's geographical location, and the bidding type corresponding to the user's behavior; the time period in which the user's behavior occurs is segmented by quarter to illustrate the time distribution of bidding relative to each quarter; the geographical location represents the relative situation of users bidding in different regions and provinces; and the bidding type represents the classification of bidding documents into different types such as entrusted bidding documents, construction project bidding documents, procurement bidding documents, and information system bidding documents, in order to conduct security monitoring of the bidding platform.
[0040] In step S21, when verifying the historical behavior baseline of each user behavior under the processing node relative to the historical data, the implementation method includes: S211, dividing the time interval according to the upload time of the user behavior, taking each time interval as a state interval, determining the operation type of the user behavior in each state interval, dividing the time interval into equal sizes, taking one hour as a state interval, and counting the operation types of all users in the corresponding time period.
[0041] As for the statistical operation types, they record the time periods during which all users tend to upload bid documents, as well as the time periods corresponding to other operations such as modifying and deleting bid documents. This facilitates the subsequent segmentation of the relationship between bid documents from multiple users under the same tender document, and the analysis of the relevant content of a single user and multiple tender documents in the bid documents.
[0042] S212, in response to the operation type of user behavior, record the time point of user behavior change, and according to the operation object and operation quantity at the time of change, perform state identification on user behavior in continuous state interval, and determine the state transition path of user behavior under the corresponding operation type.
[0043] S213, using the optimal state interval corresponding to each operation type in the state transition path, the frequency of executing that operation type in the optimal state interval is used as the historical behavior baseline for output.
[0044] At this point, the user behavior sequence aggregated from multiple users is divided into multiple state intervals, with each state interval corresponding to multiple states, such as A1, A2, A3, A4, etc. These multiple states represent the number of files uploaded and processed by users within this time period. These numbers are converted into the total frequency of file processing by the corresponding user within a time period. The value of each frequency corresponds to a value range in states A1, A2, A3, A4, etc. At this point, the highest frequency of submissions in a single time period in historical data is used as the upper limit, and the average of the lowest frequency of submissions in historical data is used as the lower limit. The interval between the upper and lower limits is divided into four equal parts, forming the value ranges of four states, such as A1, A2, A3, A4. The states are then set according to the maximum and minimum values in each hourly segment throughout the day. Then, the probability value of the transition between adjacent state intervals in the A1→A2 state is calculated. The probability value is expressed as the frequency of the transition between adjacent state intervals in the A1→A2 state divided by the frequency of all state transitions. Finally, the states marked in all state intervals are connected to obtain the state transition path under the corresponding operation type. If a portion of a time interval exceeds the upper and lower limits, it is labeled as a state described in the form of A5, A6, etc., to complete the labeling of relative states within all state intervals.
[0045] The probability values of adjacent state intervals are iterated continuously until the probability values represented by each pair of state intervals stabilize after statistical analysis of data over multiple consecutive days. This means the average probability over multiple days within the same time interval is stable. The corresponding average probability is then taken as the optimal state interval for that state interval, and the frequency of the optimal state interval is used as the historical behavior baseline for output. The output optimal state interval now represents the optimal transition between two adjacent state intervals, such as A1→A2. The frequency of states like A1→A2 in historical data is then statistically analyzed, and the weighted average of these frequencies is used as the historical behavior baseline. The weights used here are based on the ratio of the frequency of each state interval within the optimal state interval to the sum of the frequencies of all optimal state intervals. If the probability values of state intervals still fluctuate after incorporating data from multiple days, the judgment criterion is adjusted to: the average probability fluctuation of the same state transition path within 7 consecutive days is ≤3%. The corresponding state interval is then considered the optimal state interval, thus obtaining the optimal state interval relative to the hourly division of state intervals.
[0046] Finally, each state interval is compared to obtain the historical behavior baseline of the uploaded bid documents within a 24-hour period. This refers to the processing of bid documents uploaded to the bidding platform. The deletion and modification operations are processed in the same way as the upload. It should be noted that the upload refers to the initial submission of the bid document, while the modification refers to subsequent modifications to the bid document, in order to identify the operation status of the bid document in the data chain process.
[0047] The process of finding the optimal state interval can be illustrated by an example. The optimal state interval from 10:00 to 11:00 can be identified by the number of bid submissions within 10:00 and the number of bid submissions within 11:00. This will give us a relatively stable historical behavior baseline that originally belonged to 10:00. Then, we can determine the optimal state interval from 11:00 to 12:00 and complete the processing under each of the divided state intervals.
[0048] S22, determine the degree of deviation between user behavior and historical behavior baseline under each processing node, and set initial anomaly scores for behavior data sets corresponding to multiple degrees of deviation.
[0049] At this point, the deviation of the current user behavior from the historical behavior baseline is calculated by the operations of multiple users within a time interval. This illustrates the degree of deviation of the user behavior in the current scene segmentation dimension from the historical behavior baseline. Based on the deviation value, multiple behavior data sets are combined. Each behavior data set is assigned an initial anomaly score based on the deviation value. This score is set based on the deviation value. For example, a Z-Score is used to set a standard score, which is obtained by subtracting the frequency of the current user behavior in a single time interval from the historical behavior baseline and dividing by the standard deviation of the frequency of the state interval corresponding to the historical behavior baseline in the historical data.
[0050] S23, perform cross-validation on the behavioral data set corresponding to each degree of deviation, and if the validation results are inconsistent, trigger a consistency anomaly flag.
[0051] Step S23, when performing cross-validation, includes: identity and operation consistency verification, file attribute association verification, and traffic and operation matching verification. Identity and operation consistency verification mainly verifies whether the operation IP is consistent with the authentication IP to exclude account theft. File attribute association verification is used to compare the creator, modification time, and MAC address of files uploaded by different users to identify shared devices or collaborative editing behavior and determine whether they belong to the same person for editing and processing. Traffic and operation matching verification is used to check whether the upload traffic is consistent with the file size to prevent fragmented transmission or hidden data.
[0052] If there are inconsistencies in the current bid submission process under these three verification methods, a consistency anomaly flag will be triggered to indicate that there is a problem with the uploaded files.
[0053] For example, identity and operation consistency verification mainly compares the consistency of the geographical location and network attributes of the operation IP and the authentication IP. It sets a risk score based on the impossibility of IP address changes, with higher risk for greater geographical distance and shorter time difference.
[0054] Within the same city, different IP segments may be using dynamic IPs or different network exits, such as switching between WiFi / 4G, which is common under normal circumstances. The current situation is considered low risk and is assigned a score of 10.
[0055] Within the same province but different cities, there may be account sharing or operations performed while traveling. In such cases, a medium risk level will be set, and a score of 30 will be used to explain the situation. If the time interval between operations is less than 1 hour, a score of 50 will be set to explain the relative unreasonableness of the situation.
[0056] If the situation involves different provinces or countries, it is highly likely that the account has been stolen or maliciously shared. In this case, it will be set to high risk and represented by a score of 70. At the same time, if the operation time interval is very short, such as within 10 minutes, it will be set to 100 points to indicate that there is a serious inconsistency in the cross-validation.
[0057] File attribute association verification compares metadata from different tender documents, such as creator, last saver, last printr, company information, MAC address, and editing duration. Risk scores are assigned based on the accuracy and uniqueness of the metadata match. The more unique and accurate the match, the higher the risk. File attribute association verification primarily targets different users; for portions submitted by the same user, the current portion is split into multiple datasets and processed separately to verify the association between text attributes.
[0058] After verifying the usernames of the creator and the last saver, it indicates that the file was created by the same person. If multiple files match this, they are directly judged as the highest risk and set with a score of 100.
[0059] When verifying computer name and MAC address, if the file originates from the same physical device and the current user's file matches other user files, it is considered high risk and is described as 90 points.
[0060] When verifying company information, if the company names embedded in the file metadata are the same, but the submitter's account belongs to a different company, it is considered high risk and will be described as 80 points.
[0061] When verifying editing time, if the total editing time of multiple files is exactly the same or very close, it may be due to batch generation by a script. These data are considered medium risk, and the risk score is set to 60 points (range 0-100).
[0062] Traffic and operation matching verification primarily compares the actual amount of data transmitted as recorded at the network layer with the file size declared at the application layer. Risk scores are assigned based on the significance and direction of the traffic discrepancy; upload traffic significantly exceeding the file size is an extremely high-risk signal.
[0063] If the actual upload traffic is significantly greater than the file size, it is highly likely that other data was included while uploading the file, resulting in the leakage of source data or communication. The current situation is considered high-risk, and is given a score of 90 for explanation.
[0064] If the actual upload traffic is less than the file size, the upload may fail, be uploaded in chunks, or there may be abnormal traffic monitoring. This is a technical issue, and the current situation is considered low-risk. We will use 30 points to explain this.
[0065] If the same file is uploaded in small batches multiple times, it may be attempting data fragmentation to avoid triggering an alarm due to a large volume of data at once. The current situation is a relatively abnormal pattern and is considered medium risk, which is explained with a score of 70.
[0066] It should be noted that the above content is only used to describe what content is used to describe the abnormal situation when the consistency anomaly flag appears. These scores will identify the label when each user behavior occurs according to the logical rules set in the database in advance, and set the corresponding scores and labels according to the anomalies corresponding to these labels. Then, the scores of the consistency anomaly flag are used to complete the setting of the comprehensive risk score, and the abnormal behaviors in the current user behavior are output.
[0067] S24: Using the initial anomaly score and consistency anomaly flag in the behavior data set, set a comprehensive risk score for user behavior, and output the abnormal behavior according to the value of the comprehensive risk score.
[0068] At this point, the consistency anomaly indicates that a score will be set according to the inconsistent parts of the cross-validation. If there are no inconsistent parts, the score will be set to 0; if there are, the set score will be weighted and summed with the initial anomaly score. The set weights will be set according to four weights: identity and operation consistency verification, file attribute association verification, traffic and operation matching verification, and deviation degree. For example, weights of 0.3, 0.2, 0.2, and 0.3 can be used. Alternatively, a machine learning aggregation method can be used to cluster based on the consistency anomaly indicator and the deviation degree, and output the probability of taking the corresponding value. The probability is used as the score, and the sum of the scores is regarded as the comprehensive risk score. This machine learning aggregation method clusters the parts with obvious abnormal behavior in the current data and historical data and identifies the corresponding probability value. The probability value can be based on the ratio of frequency to describe the comprehensive risk score relative to the abnormal behavior.
[0069] When outputting abnormal behavior based on the comprehensive risk score, the comprehensive risk score is divided into low risk (0-30 points), which only needs to be marked with logs and will not be directly analyzed and processed in the subsequent process; medium risk (31-70 points), which generates an alarm and pushes it to the subsequent processing; and high risk (71-100 points), which requires automatic forced intervention on the account and analysis of whether there is batch generation of relevant data text.
[0070] It should be noted that the comprehensive risk score represents a set of labels representing anomalous behavior. That is, during cross-validation and deviation calculation, each anomalous behavior is labeled according to the calculated deviation value and the cross-validation part, indicating the type of the current anomalous behavior. When outputting each anomalous behavior, it includes not only the comprehensive risk score but also the multiple labels used to calculate the comprehensive risk score, to illustrate the overall situation of the anomalous behavior.
[0071] In one embodiment of the present invention, such as Figure 3 As shown, the implementation of step S3 includes: S31, receiving the input abnormal behavior, abstracting the abnormal behavior into graph nodes, and the attributes of each graph node representing the data information containing specific user behavior, such as file ID, user ID, comprehensive risk score, and bidding text corresponding to the abnormal behavior.
[0072] S32, calculate the correlation value between any two graph nodes, and construct the abnormal behavior relationship graph using the correlation value between graph nodes as the edge;
[0073] In the current step, the correlation value between any two graph nodes is determined by locating the similarity of the bidding text; that is, the implementation of step S32 includes: S321, for any two graph nodes, the text similarity and behavior similarity between the graph nodes are obtained respectively.
[0074] S322: When the current graph nodes are not from the same user, the similarity of the bidding text is used to represent the text similarity between the graph nodes; when the graph nodes are from the same user, the historical bidding text of the same user is extracted, and the similarity between the historical bidding text and the current bidding text is used to represent the text similarity between the graph nodes.
[0075] When characterizing the associated values of graph nodes, it is also necessary to obtain the comprehensive risk score of the graph nodes. The abnormal behavior represented by the comprehensive risk score is combined with the comprehensive risk score in the cross-validation part to form a behavior vector. At this time, the data related to the abnormal behavior is standardized to calculate the similarity between multiple graph nodes.
[0076] S323: Based on the abnormal behavior corresponding to each graph node, the abnormal behavior is transformed into a behavior vector according to the label corresponding to the abnormal behavior. The cosine similarity between any two graph nodes with respect to the behavior vector is used as the behavior similarity between the graph nodes.
[0077] S324, Set association values based on text similarity and behavior similarity between graph nodes.
[0078] It should be noted that text similarity is calculated using NLP methods to assess the similarity of content related to project implementation plans, implementation schedules, and safeguard measures within the bidding text. This similarity can be based on converting the text into word vectors and using the cosine similarity calculated from the word vectors as the text similarity. Behavioral similarity is calculated in the same way, continuously selecting anomalous behaviors in the current scenario until the correlation values between all anomalous behaviors have been calculated.
[0079] Preferably, when setting the association value based on text similarity and behavior similarity, it can be set based on the weighted sum of text similarity and behavior similarity. In this case, the weight can be set in the form of 0.5, or the weight can be set as the ratio of text similarity and behavior similarity to the average value of text similarity and behavior similarity in historical data. If the weight is set in the form of a ratio, the association value needs to be normalized after the association value calculation is completed to prevent it from being too large and causing the value range of subsequent processes to exceed the limit.
[0080] After calculating the associated values, the graph nodes where the abnormal behavior is located and the edges corresponding to the associated values can be connected one by one to eventually form an abnormal behavior relationship graph.
[0081] S33. Based on the abnormal behavior relationship graph, perform association mining on abnormal behaviors to obtain the association set corresponding to the abnormal behaviors.
[0082] When selecting the association set and mining the associations of anomalous behaviors, the Louvain algorithm can be used. The goal of the Louvain algorithm is to maximize the modularity of the entire graph, which is an indicator of the quality of community partitioning. By partitioning into communities, the partitioned communities are used as the association set for the output of anomalous behaviors.
[0083] The following steps will then be performed: A1, assign each graph node to an independent community, with the number of communities equal to the number of graph nodes initially.
[0084] A2, in turn, attempts to move each graph node from its original community to the community of one of its neighboring nodes.
[0085] A3. After each move, calculate the change in modularity Q of the entire graph, ΔQ. Modularity measures how tightly connected within a community is compared to random connections. The higher the Q value, the better the community division.
[0086] A4. Move the graph node to the community that will increase ΔQ the most. If no move produces a positive ΔQ, the graph node remains in its original community.
[0087] A5, repeat A2-A4 until the movement of any graph node can no longer increase the modularity Q, at which point a local optimum is obtained.
[0088] A6 aggregates the communities of local optima into a new super node, resulting in a new graph. The weight of the edge between the super nodes in the new graph is the sum of the weights of all edges between the corresponding two communities in the original graph; the edge weights within a community are converted into the self-loop weights of the super node.
[0089] A7. Repeat the content described in A2-A6 on this new graph until the modularity can no longer be increased. At this point, stop the calculation and finally aggregate a super node as an output association set to obtain multiple association sets related to abnormal behaviors.
[0090] The modularity can be calculated as follows: ;in, This represents the weight between graph node i and graph node j, where i and j represent the graph node indices within the aggregated community at this point. and Let i and j represent the degrees of graph node i and graph node j, respectively. Let represent the sum of the weights of all edges connected to graph node i. This represents the sum of the weights of all edges connected to graph node j, indicating the connection strength of the graph node. The higher the degree of a graph node, the more important it is. This represents the sum of the weights of all edges, which is the sum of the weights of the edges between all nodes in the entire abnormal behavior graph. This represents the Kronecker delta function, which is 1 if graph node i and graph node j belong to the same community, and 0 otherwise. and Let represent the community indices of graph node i and graph node j, respectively, to illustrate their relative positions under the clustering partition. Here, the modularity continuously searches for graph nodes within the community. After traversing each pair of nodes, the sum of these sums is used as the modularity to indicate the clustering pattern of the current anomalous behavior.
[0091] S34, extract the operation object for each abnormal behavior in the association set, and generate behavior association results about multiple operation objects.
[0092] When extracting the operation object in step S34, the tender documents under multiple abnormal behaviors are described according to the file ID, file name, size and other contents corresponding to the operation object to obtain the association results of different file descriptions.
[0093] The implementation process of step S34 is further represented as follows: S341, in response to the operation objects in the association set, a feature pool is set for each operation object according to each association set. The operation objects include, but are not limited to, file identifiers, metadata, and source information. The file identifier mainly reflects the file ID and file name; the metadata mainly reflects the file size, creation time, and modification time; the source information mainly reflects the upload IP address, device fingerprint, and user identity. The data described here is used to point to the specific object of the abnormal behavior when abnormal behavior occurs, and these objects will reflect the corresponding problem set.
[0094] S342, Collect the frequency of occurrence of each element in the feature pool, sort the elements in the feature pool according to the frequency of occurrence, and form a behavior association list; use the behavior association list as the output behavior association result.
[0095] The behavioral association results will be represented as a set of data results in the form of association set ID + similarity of bid document content + frequency of occurrence of corresponding bid document ID + IP, etc., indicating the existence of data source and the specific manifestation of the current abnormal behavior.
[0096] For example, if a certain IP address appears most frequently in the abnormal behaviors aggregated in the current association set, this IP address will be used as the output behavior association result, and the data related to this IP will be recorded in the current association set to facilitate subsequent analysis of the temporal distribution of this data.
[0097] In one embodiment of the present invention, when performing data association, the behavioral data set contained in the result of the association analysis of abnormal behavior is mapped to the initially extracted monitoring points, and these monitoring points are connected to illustrate the similar parts of the abnormal behavior in the entire processing flow. These similar parts are then combined according to different scene segmentation dimensions to determine the association path of the state change of a table detection point.
[0098] like Figure 4 As shown, the implementation of step S4 includes: S41, traversing each operation object in the behavior association results, extracting the monitoring point corresponding to each operation object, and sorting all monitoring points in ascending order according to the global timestamp; at this time, by tracing the operation object, the monitoring point directly corresponding to each operation object in the behavior association results is found. These monitoring points will represent the user behavior identified in the initial situation, to illustrate the specific distribution of user behavior under abnormal behavior analysis.
[0099] S42, connect all monitoring points in the order in which they were identified to determine the associated paths belonging to the same operation object.
[0100] When connecting all monitoring points, the connection will be made according to the results analyzed in step S3. The abnormal behaviors represented by multiple monitoring points will be connected one by one according to the content aggregated by the association set. Multiple bid documents with high correlation in the association set will be connected to illustrate the association path through the combination of similar bid documents under each scenario segmentation dimension.
[0101] S43. Based on the weight of the monitoring points on the associated path, determine the value range of each associated path, and sort the associated paths from largest to smallest according to the value range to output the associated paths under the corresponding scene segmentation dimensions.
[0102] The monitoring point weights described here are determined based on the correlation values calculated between various abnormal behaviors during the association set processing. As for the value range of each association path, it represents the value range of all monitoring point correlation values on this association path. The smaller this value range, the more concentrated the correlation values are on this path, indicating that the problematic user behavior shows a strong consistency, and the abnormal behavior will show a high correlation with the identified operation object. When this value range is larger, it indicates that the abnormal behavior on the aggregated association path is more dispersed, indicating that the abnormal behavior is related to multiple data contained in the current association path, and is not an abnormal behavior targeting a single operation object.
[0103] It should be noted that the value range of the above-mentioned associated path represents the difference between the upper and lower limits of the associated values on the current associated path.
[0104] The implementation method for the value range of the associated path in step S43 also includes: when the value range of the associated path is positively correlated with the number of monitoring points on the associated path, the sorting position of the current associated path is raised by one position; otherwise, the sorting position of the current associated path is maintained.
[0105] At this point, the correlation between the range of values of the associated path and the number of monitoring points is analyzed. The significance test is used, that is, after the data of the associated path is normalized, the Pearson correlation coefficient is calculated with the historical data, the t-statistic is used to observe the change of the t value, and the p value obtained by the two-tailed test is used to indicate whether the current range of values of the associated path and the number of monitoring points are positively correlated.
[0106] The process involves comparing the data from the current correlation path with other correlation paths or historical data with the same number of monitoring points, and then calculating the Pearson correlation coefficient for each correlation path. This is done by using the correlation values of each monitoring point as the basis for the calculation.
[0107] The t-statistic is then expressed as: ;in, This represents the Pearson correlation coefficient, indicating the Pearson correlation coefficient of the current association path; This represents the sample size, indicating the number of monitoring points on the current association path. The p-value is calculated by statistically analyzing the distribution of t-values under different sample sizes. When the p-value is less than 0.05, it indicates a significant relationship between the calculated association value and the sample size corresponding to the monitoring point. This suggests a positive correlation between the value range of the corresponding association path and the number of monitoring points on the association path, requiring an increased monitoring priority for the corresponding data on the current association path, thus raising its ranking and increasing attention to related issues.
[0108] The p-value for significance testing is then expressed as: ;in, This represents the degrees of freedom, which take the value n-2; This represents the cumulative distribution function of the t-distribution; this function is usually obtained directly through numerical integration or statistical software to illustrate the values of the t-distribution under different conditions.
[0109] The associated paths under different scene segmentation dimensions described in step S4 will form a path list according to the value range of the associated paths. Each path list represents data under a scene segmentation dimension. The order of these data in the path list represents the priority of the associated paths during monitoring. This content will serve as the main target for subsequent early warnings, making it easier for back-end staff to verify the upload status of bidding documents and promptly process any data with related anomalies.
[0110] In one embodiment of the present invention, step S5 mainly involves using the associated path in a cyclical analysis scenario to verify the currently acquired bidding platform data in a cyclical manner using multiple subsets, finding the intersection of the cyclical processing and whether the intersection is abnormal behavior, in order to find abnormal behavior in multiple dimensions of analysis, and regard this abnormal behavior as the target data under the current security monitoring.
[0111] like Figure 5 As shown, step S5 is implemented as follows: S51, a scene segmentation dimension is considered as a subset, and the associated paths under all subsets are combined into a path list; the subsets formed at this time are related to the specific content of the scene segmentation dimension, including time period subsets, geographical location subsets, and bidding type subsets. At this time, the associated paths under each subset are initially combined into a whole path list. This list is equivalent to a list that has not undergone intersection loop verification. After completing the intersection verification, the priority of the data with intersection is adjusted, and the sorting of this part of the data is adjusted upwards.
[0112] S52, perform pairwise cross-validation on the association path combinations of all subsets, and calculate the intersection of association paths for each pair of subsets.
[0113] S53 verifies the frequency of intersections of associated paths. When the frequency of intersections exceeds a preset frequency threshold, the corresponding intersections are merged and added to the path list as new paths. If the merged path already exists in the path list, its sorting position is moved up. The priority of the intersection paths is adjusted by moving them up, or an initial priority is configured for each sorting position. The priority of the data corresponding to each associated path in the path list is updated according to the number of moves up. The current path list is updated sequentially, and the portion that has completed sorting adjustments and data merging is output as a warning target.
[0114] The following examples illustrate this: the path list for the time period subset Q1: [Path A (Entrusted Bidding), Path B (Construction Project Bidding)]; the path list for the geographical location subset Beijing: [Path A (Entrusted Bidding), Path C (Procurement Bidding)]; the initial path list after merging: [Path A, Path B, Path C] (not deduplicated, only preliminary combination).
[0115] During frequency verification, the number of times each intersection path appears in all subset combinations is counted. This number represents the frequency of the corresponding intersection path. If the frequency exceeds a preset threshold, it is marked as an abnormal path and merged into the path list.
[0116] If the intersection path does not exist in the main list, it will be added and its initial priority will be set, such as the lowest priority. If the intersection path already exists, its priority will be moved up one position each time based on the frequency. When outputting warning targets, the path with higher priority will be output first to describe the main problem part in the current scenario.
[0117] As for the aforementioned preset frequency threshold, the frequency distribution of the intersection of the corresponding related paths in historical data is used as the quartile as the preset frequency threshold at this time. That is, after sorting the frequency of the intersection of related paths from smallest to largest, the upper limit of the frequency in the top 75% of the data is used as the preset frequency threshold at this time. This is to indicate the obviously abnormal parts compared to the normal situation, so as to facilitate the subsequent verification and processing by the staff.
[0118] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention, which are still covered within the protection scope of the present invention.
Claims
1. A data security monitoring and analysis method for a bidding and tendering information platform, characterized in that, include: S1 collects data from the bidding platform across the entire chain. Based on the operation objects, operation quantity, and operation type of the bidding platform data in a continuous time period, user behavior is divided into multiple monitoring points, and the user behavior sequence composed of each monitoring point is obtained. S2, segment the scenario based on the user behavior implementation scenario, and use the time distribution of user behavior processing as the main body to determine the abnormal behavior relative to the historical behavior baseline under each scenario segmentation dimension; S3, obtain the correlation values between each abnormal behavior, and then merge the abnormal behavior into multiple correlation sets according to the correlation values, and determine the behavior correlation result between each correlation set and the user operation object; S4. Use the behavior association results to trace the source of data association, connect the behavior association results in the form of the distribution of monitoring points, and determine the association path under different scene segmentation dimensions. S5 performs iterative verification of the associated paths under all scene segmentation dimensions, updates the path list of each associated path combination based on the intersection of the iterative verification, and uses the updated path list as the output warning target.
2. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 1, characterized in that, The implementation of step S1 also includes: According to the time window corresponding to each monitoring point, all user behaviors are traversed to determine the context information of user behaviors within each time window; Based on the contextual information of user behavior, discrete user behaviors are aggregated in chronological order to form a continuous sequence of user behaviors; Determine whether the timestamp of each user action in the user action sequence corresponds to the timestamp in the tender document, and mark it in the user action sequence.
3. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 1, characterized in that, Step S2 can be implemented in the following ways: S21, treat each scene segmentation dimension as a processing node, and verify the user behavior under each processing node relative to the historical behavior baseline of historical data; S22, determine the degree of deviation between user behavior and historical behavior baseline under each processing node, and set initial anomaly scores for behavior data sets corresponding to multiple degrees of deviation; S23, perform cross-validation on the behavioral data set corresponding to each degree of deviation, and if the validation results are inconsistent, trigger a consistency anomaly flag. S24: Using the initial anomaly score and consistency anomaly flag in the behavior data set, set a comprehensive risk score for user behavior, and output the abnormal behavior according to the value of the comprehensive risk score.
4. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 3, characterized in that, The implementation methods of step S21 include: S211, Divide the time interval according to the upload time of the user behavior, and use each time interval as a state interval to determine the operation type of the user behavior in each state interval; S212, in response to the operation type of user behavior, record the time point of user behavior change, and according to the operation object and operation quantity at the time of change, perform state identification on user behavior in continuous state interval, and determine the state transition path of user behavior under the corresponding operation type. S213, using the optimal state interval corresponding to each operation type in the state transition path, the frequency of executing that operation type in the optimal state interval is used as the historical behavior baseline for output.
5. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 1, characterized in that, Step S3 can be implemented in the following ways: S31, Receive the input abnormal behavior and abstract the abnormal behavior into graph nodes; S32, calculate the correlation value between any two graph nodes, and construct the abnormal behavior relationship graph using the correlation value between graph nodes as the edge; S33, Based on the abnormal behavior relationship graph, perform association mining on abnormal behaviors to obtain the association set corresponding to the abnormal behaviors; S34, extract the operation object for each abnormal behavior in the association set, and generate behavior association results about multiple operation objects.
6. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 5, characterized in that, The implementation methods of step S32 include: S321, For any two graph nodes, obtain the text similarity and behavior similarity between the graph nodes respectively; S322, when the current graph nodes are not from the same user, the similarity of the bidding text is used to represent the text similarity between the graph nodes; when the graph nodes are from the same user, the historical bidding text of the same user is extracted, and the similarity between the historical bidding text and the current bidding text is used to represent the text similarity between the graph nodes. S323: Based on the abnormal behavior corresponding to each graph node, the abnormal behavior is transformed into a behavior vector according to the label corresponding to the abnormal behavior, and the cosine similarity between any two graph nodes with respect to the behavior vector is used as the behavior similarity between the graph nodes. S324, Set association values based on text similarity and behavior similarity between graph nodes.
7. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 5, characterized in that, The implementation process of step S34 is further expressed as follows: S341, In response to the operation objects in the association set, set the feature pool for each operation object according to each association set; S342, Collect the frequency of occurrence of each element in the feature pool, sort the elements in the feature pool according to the frequency of occurrence, and form a behavior association list; use the behavior association list as the output behavior association result.
8. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 1, characterized in that, Step S4 can be implemented in the following ways: S41, traverse each operation object in the behavior association results, extract the monitoring points corresponding to each operation object, and sort all monitoring points in ascending order according to the global timestamp; S42, connect all monitoring points in the order in which they were identified to determine the associated paths belonging to the same operation object; S43. Based on the weight of the monitoring points on the associated path, determine the value range of each associated path, and sort the associated paths from largest to smallest according to the value range to output the associated paths under the corresponding scene segmentation dimensions.
9. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 8, characterized in that, The implementation of the value range of the associated path in step S43 also includes: For the value range of the associated path, if the value range of the associated path is positively correlated with the number of monitoring points on the associated path, the sorting position of the current associated path will be increased by one position; otherwise, the sorting position of the current associated path will be maintained.
10. The data security monitoring and analysis method for a bidding and tendering information platform according to claim 1, characterized in that, Step S5 can be implemented in the following ways: S51 treats a scene segmentation dimension as a subset, and the associated paths under all subsets are combined into a path list; S52, perform pairwise cross-validation on all subsets of associated path combinations, and calculate the intersection of associated paths for each pair of subsets; S53, verify the frequency of the intersection of associated paths. When the frequency of the intersection of associated paths exceeds the preset frequency threshold, merge the corresponding intersection of associated paths and input it into the path list as a new path. If the merged path exists in the path list, move the sorting position of the corresponding path up.
Citation Information
Patent Citations
Bidding and tendering system based on cloud service platform
CN114971796A
Electronic bidding and tendering transaction platform-based bidding and tendering information management method and system
CN119027223A
Data security tracing method, system and device based on artificial intelligence
CN118536093A
Abnormal transaction identification method and system based on correlation analysis
CN120509963A
Information management method and system based on electronic bidding transaction platform
CN120672472A
Cited By
Intelligent bid invitation risk control early warning method
CN121120222A