A method for monitoring and analyzing data security of an informationization platform for bidding
By using end-to-end data collection and scene segmentation, abnormal behaviors of bidding platforms are identified, an abnormal behavior relationship graph is constructed, and path priorities are dynamically adjusted. This solves the problem of difficulty in identifying abnormal behaviors in existing technologies and achieves efficient abnormal behavior identification and early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies ignore abnormal behavior in user uploads, making it difficult to identify anomalies in bidding data and respond quickly to complex abnormal behaviors.
By collecting data from the bidding platform across the entire process, user behavior is divided into multiple monitoring points based on the operation object, operation quantity, and operation type. Scene segmentation and abnormal behavior identification are performed, an abnormal behavior relationship graph is constructed, and the path sorting priority is dynamically adjusted to achieve efficient early warning.
It improves the accuracy of abnormal behavior identification and the timeliness of early warning, enabling more precise identification and response to abnormal behavior, and thus enhancing the accuracy of abnormal behavior detection.
Smart Images

Figure CN120880786B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of information monitoring, and in particular to a data security monitoring and analyzing method for a bidding information platform. BACKGROUND
[0002] Bidding refers to the act of a buyer issuing a procurement demand to the public, which includes the act of a production and manufacturing enterprise extending credit to a client with the legal person nature in credit management, that is, product credit sales. In the process of product credit sales, the credit grantor is usually a material supplier, a product manufacturer and a wholesaler, and the buyer is the beneficiary of product credit sales, which is various enterprise clients or agents.
[0003] For example, Chinese Patent Publication No. CN114971796A discloses a bidding system based on a cloud service platform. A transaction platform is used to receive a bidding transaction application of a mobile terminal user and to perform identity verification on the mobile terminal user. A public service platform is used to generate bidding transaction data after identity information of the mobile terminal user is audited and to perform bidding control. An administrative supervision platform is used to receive bidding data of the mobile terminal user and to generate bidding data by synchronously signing in the cloud.
[0004] For example, Chinese Patent Publication No. CN119027223A discloses a bidding information management method and system based on an electronic bidding transaction platform, which includes the following steps: S01, importing readable data from multiple data sources of a database; S02, based on project components stored in the database, managing related metadata of the project components in the imported data in the database to form a metadata repository matching the data volume of the database; S03, matching the metadata repository with a data import module and generating difference information between the metadata repository and multiple identifiable database entities located under the metadata repository. The application simplifies the filling process of the bidding template and intelligently processes the bidding template through the bidding book.
[0005] In the prior art, the transaction subject in the bidding scenario is explained, the bidding data is processed in an encrypted permission authentication manner, and the superposition processing of the bidding data is completed by checking the information under the bidding text. However, the prior art ignores the abnormal behavior of the user uploading behavior, so that the abnormality existing in the received data is difficult to identify, and the complex abnormal behavior cannot be quickly responded to. SUMMARY
[0006] In order to solve the above technical problems, the technical scheme adopted by the present application is: a bidding information platform data security monitoring and analysis method, comprising: S1, collecting bidding platform data in the whole link, according to the operation object, operation quantity and operation type of the bidding platform data in the continuous time period, the user behavior is divided into multiple monitoring points, and the user behavior sequence composed of each monitoring point is obtained.
[0007] S2, scene segmentation is carried out according to the user behavior implementation scene, the time distribution during the user behavior processing is taken as the main body, and the abnormal behavior relative to the historical behavior baseline in each scene segmentation dimension is determined.
[0008] S3, the correlation value between each abnormal behavior is obtained, and then the abnormal behavior is integrated into multiple correlation sets according to the correlation value, and the behavior correlation result of each correlation set and the user operation object is determined.
[0009] S4, data correlation tracing is carried out according to the behavior correlation result, the behavior correlation result is connected in series in the form of monitoring point distribution, and the correlation path under different scene segmentation dimensions is determined.
[0010] S5, the correlation path under all scene segmentation dimensions is verified circularly, the intersection during the circular verification is used to update the path list after the combination of each correlation path, and the updated path list is used as the output early warning target.
[0011] The beneficial effects of the present application are: first, the present application avoids the problem of user behavior recognition error in a single dimension by monitoring the bidding platform data according to the operation object, operation quantity and operation type, and deploying the user behavior sequence under multiple monitoring points; then, the user behavior is scene segmented, the abnormal behavior relative to the historical behavior baseline in the current user behavior sequence is identified in multiple scene segmentation dimensions, the adaptability of the historical behavior baseline is improved through state transition path modeling and cross validation, and more accurate abnormal behavior preliminary screening is realized by combining the comprehensive risk score.
[0012] Second, the present application abstracts the abnormal behavior as a graph node, calculates the correlation value through text similarity and behavior similarity, constructs an abnormal behavior relationship graph, extracts an operation object feature pool, realizes the correlation mining of abnormal behavior, and further improves the focusing of the correlation set through the feature pool sorting mechanism, and improves the correlation and accuracy of the isolated analysis of abnormal behavior.
[0013] Thirdly, the application sorts the monitoring points according to global timestamps, generates a correlation path, and dynamically adjusts the path sorting priority through the positive correlation check of the monitoring point number and the value range, uses the positive correlation check of the value range and the monitoring point number to realize the dynamic adjustment of the path priority, ensures the high concentration of the abnormal path priority response, finally calculates the path intersection through the subset cross-validation, combines the high-frequency intersection path according to the frequency threshold, and moves up the priority, completes the merging analysis under each scene segmentation dimension, realizes the dynamic update of the path list, and improves the timeliness and accuracy of the early warning. BRIEF DESCRIPTION OF DRAWINGS
[0014] The application will be further described below in combination with the drawings and embodiments.
[0015] Figure 1 It is a flowchart of a data security monitoring and analysis method of an informationization platform for bidding and tendering.
[0016] Figure 2 It is a flowchart of step S2 of a data security monitoring and analysis method of an informationization platform for bidding and tendering.
[0017] Figure 3 It is a flowchart of step S3 of a data security monitoring and analysis method of an informationization platform for bidding and tendering.
[0018] Figure 4 It is a flowchart of step S4 of a data security monitoring and analysis method of an informationization platform for bidding and tendering.
[0019] Figure 5 It is a flowchart of step S5 of a data security monitoring and analysis method of an informationization platform for bidding and tendering. DETAILED DESCRIPTION
[0020] The embodiments of the application will be described in detail below. The embodiments described below are exemplary and are only used to explain the application, and cannot be understood as a limitation of the application. If the specific technology or condition is not indicated in the embodiments, the technology or condition described in the literature in the art or according to the product instruction is used.
[0021] Reference Figure 1 A data security monitoring and analysis method of an informationization platform for bidding and tendering, comprising: S1, collecting data of a bidding and tendering platform in a whole link, dividing user behaviors into multiple monitoring points according to operation objects, operation quantities and operation types of the bidding and tendering platform data in a continuous time period, and obtaining user behavior sequences composed of each monitoring point.
[0022] S2, scene segmentation is performed according to user behavior implementation scenes, the time distribution during user behavior processing is used as a main body, and abnormal behaviors relative to historical behavior baselines under each scene segmentation dimension are determined.
[0023] S3, obtain the correlation values between the abnormal behaviors, and then integrate the abnormal behaviors into a plurality of correlation sets according to the correlation values, and determine the behavior correlation results of each correlation set and the user operation object.
[0024] S4, perform data correlation tracing according to the behavior correlation results, connect the behavior correlation results in series according to the distribution of the monitoring points, and determine the correlation paths under different scene segmentation dimensions.
[0025] S5, perform cyclic verification on the correlation paths under all scene segmentation dimensions, update the path list after combination of the correlation paths according to the intersection during the cyclic verification, and take the updated path list as the output pre-warning target.
[0026] When collecting the data of the bidding platform, the data chain process, upload time, and related data of the bidding text and file attributes during uploading of the bidding text are taken as the user behaviors during uploading of the user on the bidding platform, the operation type, operation object, operation time, operation quantity, and text data involved in the operation of the user are comprehensively analyzed to analyze the rationality of the user in uploading the bidding file, the specific situation of whether the received bidding text exists string bidding or bid rigging is monitored to determine whether the user exists abnormal behavior during uploading of the related bidding documents.
[0027] The data collected at this time include but are not limited to business operation logs, file metadata, and network flow metadata; the business operation logs indicate the username, timestamp, operation type (upload, modify, delete, etc.), operation object (such as file ID, file name, size, etc.), upload IP, used encryption certificate ID, and upload result.
[0028] The file metadata is the specific text of the uploaded bidding file; the text contained in the file, and the creator, last modifier, creation time, and last modification time are batch processed to determine whether the bidding files of a plurality of different users under the same bidding file exist abnormal correlation and the like.
[0029] The network flow data indicates the source IP, destination IP, port, protocol, transmission data volume, and timestamp, which represent the uploading situation under the business operation, and are used to identify whether there exists abnormal uploading or user behavior of large flow abnormal access. These data are combined into a complete user behavior to indicate the specific operation of the user during uploading.
[0030] The operation time represents the timestamp of the user behavior implementation, and the operation quantity represents the number of files involved in each operation. Further, the total number of files received by the system in a certain time period is analyzed to determine whether there is abnormal behavior and whether there is excessive similarity in the text. The user behavior is determined in the form of nodes to determine the change of user behavior in the time sequence of multiple time period combinations.
[0031] That is, the implementation of step S1 further comprises: traversing all user behaviors according to the time window corresponding to each monitoring point, determining the context information of user behavior in each time window, and the context information is represented as login success→upload file→download file→logout, and subsequent re-authentication→modify file→logout, etc. Context information is used to mark the overall process of uploading bid files by users, and mark the corresponding user, time and IP to assist subsequent abnormal identification of each user behavior data.
[0032] According to the context information of the user behavior, the discrete user behavior is aggregated in time sequence to form a continuous user behavior sequence; and whether the timestamp of each user behavior in the user behavior sequence corresponds to the timestamp of the bid file is judged and marked in the user behavior sequence.
[0033] After obtaining the data of these user behaviors, it is converted into a unified standardized format to facilitate subsequent viewing of inconsistent or abnormal behavior information for user upload and processing events.
[0034] As for the timestamp, it is determined whether the last modification and last storage timestamp of the bid file corresponds to the timestamp specified by the bid file, that is, it meets the time requirement or appears to be a timeout submission, modification, etc. The processing method of this part is equivalent to data preprocessing to check the relevant operations of the user when uploading the bid file in the obtained data of the bidding platform, and to check the timestamp to determine whether there is a part with inconsistent timestamp marking.
[0035] Preferably, the monitoring point represents the data index corresponding to each user behavior, and the number of settings is equal to the number of bid files currently processed, that is, the monitoring point represents the number of operation quantities in the continuous time period, and monitors the uploading, modifying and deleting of each bid file. The monitoring point is arranged at the operation object, operation quantity and operation type to indicate the corresponding bid file in the operation object, operation quantity and operation type, and the data form to be concerned by each bid file.
[0036] In one embodiment of the present invention, when performing scene segmentation, a comprehensive verification is performed based on the data chain flow and upload time of the user-uploaded bid documents. The data chain flow verification verifies the data of the bid documents under the creation time, modification records, version iterations, etc., to illustrate the process of the current file upload, modification, and deletion corresponding to the historical user behavior. The upload time emphasizes the time anomalies of the bid document submission, and the upload time of the bid documents in the platform is centrally analyzed. By comparing the time distribution of the bid documents uploaded in the past, abnormal concentrated submissions are identified, and it is checked whether the file upload distribution conforms to the normal operating rhythm. The user's upload behavior pattern is identified according to the form of its upload time distribution to identify abnormal behaviors relative to the historical behavior baseline, such as abnormal behaviors of bid rigging and collusion.
[0037] like Figure 2 As shown, the implementation of step S2 includes: S21, treating each scene segmentation dimension as a processing node, verifying the user behavior under each processing node relative to the historical behavior baseline of historical data, which will be set according to the time distribution of historically uploaded tender documents.
[0038] For example, analyzing the time distribution of historical uploaded bid documents helps determine the usual upload time periods by analyzing the time distribution of bid documents uploaded by a particular user or all users of the same company. Then, by merging the behavior of multiple users, it can be determined whether the bid documents are similar in time distribution, thus enabling anomaly detection in the time distribution dimension.
[0039] Preferably, the scenario segmentation dimension represents the segmentation based on the time period in which the user's behavior occurs, the user's geographical location, and the bidding type corresponding to the user's behavior; the time period in which the user's behavior occurs is segmented by quarter to illustrate the time distribution of bidding relative to each quarter; the geographical location represents the relative situation of users bidding in different regions and provinces; and the bidding type represents the classification of bidding documents into different types such as entrusted bidding documents, construction project bidding documents, procurement bidding documents, and information system bidding documents, in order to conduct security monitoring of the bidding platform.
[0040] In step S21, when verifying the historical behavior baseline of each user behavior under the processing node relative to the historical data, the implementation method includes: S211, dividing the time interval according to the upload time of the user behavior, taking each time interval as a state interval, determining the operation type of the user behavior in each state interval, dividing the time interval into equal sizes, taking one hour as a state interval, and counting the operation types of all users in the corresponding time period.
[0041] As for the statistical operation type, record all user tendencies to upload bid files in which time period, and the corresponding time period of bid file modification, deletion and other corresponding operations, so as to subsequently segment the relationship between multiple user bid files under the same bid file, and analyze the relevant content of a single user and multiple bid files on the bid file.
[0042] S212, in response to the operation type of the user behavior, record the time point of the user behavior change, and identify the state of the user behavior in the continuous state interval according to the operation object and the operation quantity at the time of change, and determine the state transition path of the user behavior under the corresponding operation type.
[0043] S213, with the optimal state interval corresponding to each operation type in the state transition path, the frequency of executing the operation type under the optimal state interval is taken as the output historical behavior baseline.
[0044] At this time, the user behavior sequence aggregated by multiple users is divided into multiple state intervals, and each state interval corresponds to multiple states, such as A1, A2, A3, A4, etc. These multiple states represent the number of files uploaded and processed by the user in this time period. These quantities are converted into the total frequency of processing files by the corresponding user in a time period, and the value of each frequency corresponds to a value range in A1, A2, A3, A4, etc. At this time, the highest frequency submitted in a single time in the historical data is taken as the upper limit value, and the average value of the minimum frequency submitted in the historical data is taken as the lower limit value. The interval between the upper and lower limit values is divided into four equal parts to form the value range of A1, A2, A3, A4, etc. At this time, the maximum value and the minimum value in each hour period within the whole day are set as the state, and then the probability value of adjacent two state intervals under the change of A1→A2 is solved. The probability value is represented as the frequency of adjacent two state intervals under the state transition of A1→A2 divided by the frequency of all state transitions. After that, the states marked in all state intervals are connected to obtain the state transition path under the corresponding operation type. If there is a part that exceeds the upper limit value and the lower limit value in a certain time interval, it is marked as A5, A6, etc. Formally described state to complete the labeling of relative states in all state intervals.
[0045] After continuously iterating the probability value of the adjacent two state intervals, until the data of continuous multiple days is counted, the probability value represented by every two state intervals reaches stability, that is, the average value of multiple days in the same time interval is stable, the corresponding probability average value is taken as the optimal state interval under the corresponding state interval, and the frequency under the optimal state interval is taken as the output historical behavior baseline. At this time, the optimal state interval output represents the optimal transition of the adjacent two state intervals under the state transition such as A1→A2, and the frequency of the historical data belonging to the two states such as A1→A2 is counted at this time. The weighted average value of the frequency is taken as the output historical behavior baseline, and the weight used at this time is based on the ratio of the frequency of each state interval under the optimal state interval to the sum of the frequencies in all optimal state intervals. If the data of multiple days is introduced, the probability value of the state interval still fluctuates, and the judgment standard is adjusted to: the average value of the probability of the same state transition path in continuous 7 days fluctuates ≤3%, at this time the corresponding state interval is the optimal state interval, to obtain the optimal state interval under each hour division state interval.
[0046] Finally, all state intervals are compared one by one to obtain the historical behavior baseline of uploading the bid file within 24 hours of the whole day. At this time, it is for the processing of uploading the bid file on the bidding platform. The processing method of the operation types of deletion and modification is the same as that of uploading. It needs to be explained that uploading is for the first submission of the bid file, and modification is for the modification of the bid file subsequently, to identify the operation of the bid file in the data chain process.
[0047] The optimal state interval obtaining process can be illustrated by an example. At this time, the number of bid file submissions within ten o'clock and the number of bid file submissions within eleven o'clock can be used to identify the optimal state interval from ten o'clock to eleven o'clock, to obtain the relatively stable historical behavior baseline originally belonging to ten o'clock. Then the optimal state interval from eleven o'clock to twelve o'clock is judged, and the processing of each divided state interval is completed.
[0048] S22, judge the deviation degree of user behavior under each processing node from the historical behavior baseline, and set an initial abnormal score for each behavior data set corresponding to multiple deviation degrees.
[0049] At this time, the deviation of the current user behavior from the historical behavior baseline is calculated based on the operations of multiple users in a time interval to indicate the deviation degree value of the user behavior under the current scene segmentation dimension relative to the historical behavior baseline, and multiple behavior data sets are combined according to the deviation degree value. Each behavior data set sets an initial abnormal score according to the deviation degree, which is set based on the deviation degree value, such as using Z-Score to set the standard score, that is, the frequency of the current user behavior in a single time interval is subtracted from the historical behavior baseline, and divided by the frequency standard deviation of the historical behavior baseline in the historical data, to obtain a standard initial abnormal score.
[0050] S23, cross-validation is performed on the behavior data set corresponding to each deviation degree, and if the verification is inconsistent, a consistency abnormality flag is triggered.
[0051] When step S23 performs cross-validation, it includes identity and operation consistency verification, file attribute association verification, and traffic and operation matching verification. Identity and operation consistency mainly verifies whether the operation IP is consistent with the authentication IP to exclude account theft. File attribute association verification is used to compare the creator, modification time, and MAC address of files uploaded by different users to identify shared devices or collaborative editing behavior and determine whether the editing and processing belong to the same person. Traffic and operation matching verification is used to check whether the upload traffic is consistent with the file size to prevent fragmented transmission or hidden data.
[0052] If there is inconsistency in the current bidding file submission process under the three verification methods, the consistency abnormality flag is triggered to indicate that there is a problem with the uploaded related files.
[0053] For example, identity and operation consistency verification mainly compares the geographical location and network attribute consistency of the operation IP and the authentication IP, sets a risk score based on the impossibility of IP address changes, and the farther the geographical distance and the shorter the time difference, the higher the risk.
[0054] In the same city, different IP segments may use dynamic IP or different network exits, such as switching WiFi / 4G, which is common in normal situations, and the current situation is considered as low risk and set to 10 points.
[0055] In the same province, different cities may share accounts or operate during business trips, at which time a medium risk is set and a score of 30 is used to indicate it. If the operation time interval is <1 hour, it is set to 50 to indicate its relative irrationality.
[0056] If there are different provinces, countries, high probability of account theft or malicious sharing, at this time will be set to high risk, and using 70 points of value to represent, while the operation time interval is very short, such as 10 minutes, directly set to 100 points, to illustrate the cross-validation when there is a serious inconsistency.
[0057] File attribute association verification compares the metadata of different bidding files, such as creator, last saver, last printer, company information, MAC address, editing duration, etc. According to the accuracy and uniqueness of metadata matching, set the risk score. The more unique and accurate the matching items are, the higher the risk is. At this time, file attribute association verification is mainly aimed at different users. For parts submitted by the same user, the current part will be split into multiple data sets and processed respectively to verify the association between text attributes.
[0058] After verifying the creator and last saver username, it is shown that the file comes from the same person. If multiple files match this item, it is directly determined as the highest risk, with a score of 100 points.
[0059] When verifying the computer name and MAC address, it is shown that the file comes from the same physical device. If the current user file matches other user files, it is considered high risk, with a score of 90 points.
[0060] When verifying company information, the company name embedded in the file metadata is the same, but the submitter account belongs to different companies, which is considered high risk, with a score of 80 points.
[0061] When verifying the editing duration, the total editing time of multiple files is completely consistent or highly similar, which may be a script batch generation. These data are considered medium risk, with a risk score of 60 points (value range 0-100).
[0062] Traffic and operation matching verification mainly compares the actual transmission data volume recorded by the network layer with the file size declared by the application layer. According to the significant degree and direction of traffic deviation, set the risk score. If the upload traffic is much larger than the file size, it is a high-risk signal.
[0063] If the actual upload traffic is significantly larger than the file size, it is extremely likely that other data is carried during the upload of the file, resulting in data leakage or communication. The current situation is high risk, with a score of 90 points as the explanation score.
[0064] If the actual upload traffic is less than the file size, it may be due to upload failure, split upload, or traffic monitoring anomaly, which is a technical problem. The current situation is low risk, with a score of 30 points.
[0065] If the same file is uploaded multiple times in small quantities, a data fragmentation transmission attempt may be performed to avoid triggering an alarm due to a large amount of data at one time, and the current situation is a relatively abnormal mode, which is considered as a medium risk, and is explained with a score of 70.
[0066] It should be noted that the above content is only used to express the consistency anomaly flag when it appears, and the content used to explain the abnormal situation. These scores will be identified according to the labels of each user behavior according to the logical rules set in the database in advance, and the corresponding abnormal situation will be set according to the labels. The score and label are used to complete the setting of the comprehensive risk score using the score of the consistency anomaly flag, and the abnormal behavior existing in the current user behavior is output.
[0067] S24, setting a comprehensive risk score for user behavior with the initial abnormal score and the consistency anomaly flag in the behavior data set, and outputting abnormal behavior according to the value of the comprehensive risk score.
[0068] At this time, the consistency anomaly indicates that the score will be set according to the inconsistent part of the cross-validation. If there is no inconsistent part, the score is set to 0; if there is, the score is weighted and summed with the initial abnormal score. The weights are set according to the identity and operation consistency verification, file attribute association verification, traffic and operation matching verification, and deviation degree setting. For example, the weights are 0.3, 0.2, 0.2 and 0.3, or the machine learning aggregation method is used to cluster the consistency anomaly flag and the deviation degree respectively, and the probability of the corresponding value is output. The probability is used as the score, and the sum of the scores is considered as the comprehensive risk score. This machine learning aggregation method clusters the parts with obvious abnormal behaviors in the current data and historical data to identify the corresponding probability value. At this time, the probability value can be based on the ratio of frequency to explain the comprehensive risk score of the abnormal behavior.
[0069] When outputting abnormal behavior according to the comprehensive risk score, the comprehensive risk score is divided into 0-30 points of low risk, which only needs to be marked with a log and does not enter the subsequent direct analysis process; 31-70 points of medium risk, which generates an alarm and is pushed to the subsequent processing process; and 71-100 points of high risk, which needs to be automatically intervened in the account and analyze whether there is a batch generation of related data text.
[0070] It should be noted that the comprehensive risk score represents a set of multiple labels indicating abnormal behavior, that is, according to the calculation value of the deviation degree and the part of the cross-validation, a label is added to each abnormal behavior to indicate the type of the current abnormal behavior. When each abnormal behavior is output, it not only covers the comprehensive risk score but also includes multiple labels used to calculate the comprehensive risk score to explain the comprehensive situation of the abnormal behavior.
[0071] In one embodiment of the present application, as shown in Figure 3 The implementation of step S3 includes: S31, receiving the input abnormal behavior, abstracting the abnormal behavior as a graph node, and the attributes of each graph node represent the data information including the specific user behavior such as the file ID, user ID, comprehensive risk score, and bid text corresponding to the abnormal behavior.
[0072] S32, calculating the correlation value between any two graph nodes, and constructing an abnormal behavior relationship graph with the correlation value between the graph nodes as the edge.
[0073] In the current step, the correlation value between any two graph nodes is determined based on the bid text similarity. That is, the implementation of step S32 includes: S321, for any two graph nodes, the text similarity and behavior similarity between the graph nodes are obtained respectively.
[0074] S322, when the current graph node is not the same user, the similarity of the bid text is used to represent the text similarity between the graph nodes; when the graph node is the same user, the historical bid text of the same user is extracted, and the similarity between the historical bid text and the current bid text is used to represent the text similarity between the graph nodes.
[0075] When representing the correlation value of the graph node, the comprehensive risk score of the graph node also needs to be obtained. The abnormal behavior represented by the comprehensive risk score is used to form a behavior vector with the cross-validation part and the comprehensive risk score. At this time, the data related to the abnormal behavior is standardized to calculate the similarity between multiple graph nodes.
[0076] S323, based on the abnormal behavior corresponding to each graph node, the abnormal behavior is converted into a behavior vector according to the label corresponding to the abnormal behavior, and the cosine similarity of any two graph nodes with respect to the behavior vector is used as the behavior similarity between the graph nodes.
[0077] S324, according to the text similarity and behavior similarity between the graph nodes, the correlation value is set.
[0078] It should be noted that the text similarity is calculated by NLP to calculate the similarity of the project implementation scheme, implementation plan, and protection measures in the bid text. The similarity can be based on converting the text into a word vector, and the cosine similarity calculated based on the word vector is used as the text similarity. The behavior similarity is calculated in the same way. The abnormal behavior in the current scene is selected until the correlation values between all abnormal behaviors are calculated.
[0079] Preferably, the correlation value is set based on a weighted sum of the text similarity and the behavior similarity when the correlation value is set according to the text similarity and the behavior similarity, at this time, the weight can be set in the form of 0.5, 0.5, or the weight is set in the form of a ratio of the text similarity and the behavior similarity to the average value of the text similarity and the behavior similarity of the historical data, if the weight is set in the form of a ratio, the correlation value needs to be normalized after the calculation of the correlation value to prevent the value from being too large and causing the value range to exceed the limit in the subsequent process.
[0080] After the calculation of the correlation value is completed, the graph nodes of the abnormal behaviors and the edges corresponding to the correlation values can be connected one by one to finally form an abnormal behavior relationship graph.
[0081] S33, based on the abnormal behavior relationship graph, the correlation of the abnormal behaviors is mined to obtain a correlation set corresponding to the abnormal behaviors.
[0082] When the correlation set is selected and the correlation of the abnormal behaviors is mined, the Louvain algorithm can be used for processing, the Louvain algorithm aims to maximize the modularity of the entire graph, and the modularity is an index for measuring the quality of community division. By dividing the community, the divided community is taken as the correlation set of the output abnormal behaviors.
[0083] At this time, the following steps will be performed: A1, each graph node is assigned to an independent community, and there are several communities as there are several graph nodes initially.
[0084] A2, each graph node is tried to be moved from its original community to the community in which a neighbor node of the graph node is located.
[0085] A3, after each movement, the change value AQ of the modularity Q of the entire graph is calculated, the modularity measures the degree to which the connections within a community are more dense than random connections, and the higher the Q value, the better the community division.
[0086] A4, the graph node is moved to the community that can increase AQ the most, if all movements cannot produce a positive AQ, the graph node remains in the original community.
[0087] A5, cycle A2-A4 until the movement of any graph node cannot improve the modularity Q any more, at this time, a local optimal solution is obtained.
[0088] A6, the communities of the local optimal solution are aggregated into a new super node to obtain a new graph, the weight of the edge between the super nodes in the new graph is the sum of the weights of all edges between the corresponding two communities in the original graph; the edge weight within the community is converted into the self-loop weight of the super node.
[0089] A7, repeat A2-A6 on this new graph until the modularity cannot be increased, at which point stop the calculation and take the final aggregation of a supernode as an output association set, and thus obtain multiple association sets related to abnormal behaviors.
[0090] The calculation method of modularity can be expressed as: ; wherein, represents the weight between graph node i and graph node j, i and j represent the index of the graph node within the community at this time; and represent the degree of graph node i and graph node j respectively, represents the sum of the weights of all edges connected to graph node i, represents the sum of the weights of all edges connected to graph node j, indicating the connection strength of the graph node, the higher the degree of a graph node, the more important it is; represents the sum of the weights of all edges, representing the sum of the weights of all edges between graph nodes in the entire abnormal behavior relationship graph; represents the Kronecker delta function, which is 1 if graph node i and graph node j belong to the same community, otherwise 0, and represent the community index of graph node i and graph node j respectively, to indicate the relative position under clustering division. Here, the modularity continuously searches for graph nodes within the community, and after two-way traversal is completed, the sum of the values is taken as the modularity to indicate the aggregation manner of the current abnormal behavior.
[0091] S34, each abnormal behavior in the association set is operated to extract the operation object, and the behavior association result about multiple operation objects is generated.
[0092] When the operation object is extracted in step S34, the bid files under multiple abnormal behaviors are described according to the file ID, file name, size and other contents corresponding to the corresponding operation object, to obtain the association result about different file descriptions.
[0093] That is, the implementation process of step S34 is further represented as: S341, in response to the operation object in the association set, the operation object is set according to the characteristic pool of each association set, and the operation object includes but is not limited to file identifier, metadata and source information, etc. The file identifier mainly embodies the file ID and the file name; the metadata mainly embodies the file size, the creation time and the modification time; the source information mainly embodies the upload IP address, the device fingerprint and the user identity; the data described here is used to point to the specific object of the abnormal behavior when the abnormal behavior occurs, and these objects will embody the corresponding problem set.
[0094] S342, collect the occurrence frequency of each element in the feature pool, sort the elements in the feature pool according to the occurrence frequency, and form a behavior association list; the behavior association list is taken as an output behavior association result.
[0095] The behavior association result is expressed in the form of an association set ID + a bid file content similarity + a corresponding bid file ID occurrence frequency + an IP, and a set of data results are combined to indicate the data source and the data set of the specific performance of the current abnormal behavior.
[0096] For example, in the abnormal behavior of the current association set aggregation, there is an IP address with the highest occurrence frequency, and this IP address is taken as the output behavior association result, and the data related to the IP in the current association set is recorded for subsequent analysis of the time distribution of the data.
[0097] In an embodiment of the application, when data association is performed, the behavior data set contained in the result of the abnormal behavior association analysis is mapped to the initially extracted monitoring points, and the monitoring points are connected to indicate the similar parts of the abnormal behavior in the entire processing flow, and the similar parts are combined according to different scene segmentation dimensions to determine an association path of the table detection point state change.
[0098] As shown in Figure 4 The implementation mode of step S4 includes: S41, traversing each operation object in the behavior association result, extracting the monitoring points corresponding to each operation object, and sorting all the monitoring points in ascending order according to the global timestamp; at this time, the operation object is traced back to find the monitoring points directly corresponding to each operation object in the behavior association result, and these monitoring points will represent the user behavior identified in the initial situation to indicate the specific distribution of the user behavior under the abnormal behavior analysis.
[0099] S42, connecting all the monitoring points in the order of the identified monitoring points to determine the association path belonging to the same operation object.
[0100] When all the monitoring points are connected, the result analyzed in step S3 is connected, the abnormal behaviors represented by multiple monitoring points are connected according to the contents aggregated by the association set, the multiple bid files with high correlation degree in the analysis of the association set are connected to indicate the association path combined by similar bid files in each scene segmentation dimension.
[0101] S43, determining the value range of each association path according to the weight of the monitoring points on the association path, and sorting the association paths according to the value range from large to small to output the association paths under different scene segmentation dimensions.
[0102] The weight of the monitoring point at this time is determined based on the correlation value calculated between each abnormal behavior during the correlation set processing. As for the value range of each correlation path, it represents the value range of the correlation values of all monitoring points on this correlation path. The smaller the value range, the more concentrated the correlation values on this path, and the more consistent the user behavior that appears to be problematic. The abnormal behavior will be highly correlated with the identified operation object. When the value range is larger, it means that the abnormal behavior on the aggregated correlation path is more dispersed, indicating that the abnormal behavior is related to multiple data on the current correlation path and is not an abnormal behavior that occurs for a single operation object.
[0103] It should be noted that the value range of the above correlation path represents the difference between the upper limit value and the lower limit value of the correlation value on the current correlation path.
[0104] The implementation of the value range of the correlation path in step S43 also includes: for the value range of the correlation path, when the value range of the correlation path and the number of monitoring points on the correlation path are positively correlated, the sorting position of the current correlation path is promoted by one; otherwise, the sorting position of the current correlation path is maintained.
[0105] At this time, the correlation between the value range of the correlation path and the number of monitoring points is analyzed using significance test. After normalizing the data of the correlation path and calculating the Pearson correlation coefficient with historical data, the t-statistic is used to observe the change in t-value, and the p-value obtained using two-sided test is used to indicate whether the value range of the current correlation path and the number of monitoring points are positively correlated.
[0106] That is, the data on the current correlation path and other correlation paths with the same number of monitoring points or historical data are sequentially calculated to obtain the Pearson correlation coefficient of each correlation path. At this time, the correlation value of each monitoring point is taken as the main body for calculating the Pearson correlation coefficient.
[0107] The t-statistic is then represented as: ; wherein, represents the Pearson correlation coefficient, which indicates the Pearson correlation coefficient of the current correlation path; represents the number of samples, and represents the number of monitoring points on the current correlation path. At this time, the distribution of t-value under different sample quantities is calculated to obtain the p-value. When the p-value is less than 0.05, it means that there is a significant relationship between the correlation value calculated and the sample quantity corresponding to the monitoring point. At this time, it is indicated that the value range of the corresponding correlation path and the number of monitoring points on the correlation path are positively correlated, and the monitoring priority of the corresponding data on the current correlation path needs to be improved to improve the sorting and attention to related problems.
[0108] The p value of significance is then expressed as: ; wherein, n-2; The cumulative distribution function of t distribution is usually obtained by numerical integration or statistical software.
[0109] The correlation path under different scene segmentation dimensions explained in step S4 will form a path list according to the value range of the correlation path, and each path list represents data under a scene segmentation dimension; the order of these data in the path list represents the priority of the correlation path during monitoring, and these contents will be the main target of subsequent early warning, which is convenient for background staff to verify the uploading of the bidding document and to process these data with related abnormalities in a timely manner.
[0110] In an embodiment of the present application, step S5 mainly finds out the intersection part under the cycle processing and whether the intersection part is an abnormal behavior by taking the current acquired bidding platform data in multiple subsets in the cycle analysis scene, and takes the abnormal behavior as the target data under the current security monitoring.
[0111] As shown in Figure 5 , the implementation of step S5 includes: S51, taking a scene segmentation dimension as a subset, and combining the correlation paths under all subsets into a path list; at this time, the formed subset is related to the specific content of the scene segmentation dimension, including the time period subset, the geographical location subset and the bidding type subset. At this time, the correlation paths under each subset are preliminarily combined into an integral path list, which is equivalent to a list without intersection cycle verification. After the intersection verification is completed, the priority of the part of data with intersection will be adjusted, and the ordering of this part of data will be adjusted.
[0112] S52, two-by-two intersection verification is performed on the combined correlation paths of all subsets, and the intersection of the correlation paths of each pair of subsets is calculated.
[0113] S53, the frequency of the intersection of the correlation paths is verified, and when the frequency of the intersection of the correlation paths exceeds the preset frequency threshold, the corresponding intersection of the correlation paths is merged and input into the path list as a new path; if the merged path exists in the path list, the ordering position of the corresponding path is moved up. The priority of the intersection path is adjusted by moving up, or an initial priority is configured at each ordering position, and the priority of the data corresponding to each correlation path in the path list is updated according to the number of times of moving up. The current path list is updated in turn, and then the part of the data with sorting adjustment and data integration is output as the early warning target.
[0114] At this time, the following examples can be illustrated: as the path list of time period subset Q1: [path A (entrusted bidding), path B (construction engineering bidding)]; the path list of geographic location subset Beijing: [path A (entrusted bidding), path C (procurement bidding)]; the combined initial path list: [path A, path B, path C] (not de-duplicated, only preliminary combination).
[0115] After the frequency verification, the number of times each intersection path appears in all subset combinations is counted, which represents the frequency of the corresponding intersection path; if the frequency exceeds the preset threshold, it is marked as an abnormal path and merged into the path list.
[0116] If the intersection path does not exist in the main list, it is added and set to an initial priority, such as the lowest priority; if the intersection path already exists, the priority is moved up according to the frequency, and each time it is moved up by one position; when outputting the early warning target, the path with high priority is output first to explain the main problem part in the current scenario.
[0117] As for the above-mentioned preset frequency threshold, the frequency distribution of the corresponding associated path intersection in the historical data is used, and the quartile is used as the preset frequency threshold at this time, that is, the upper limit value of the frequency in the first 75% of the data after sorting the associated path intersection frequency from small to large, as the preset frequency threshold at this time, to explain the part that is obviously abnormal relative to the normal situation, so as to facilitate the subsequent checking and processing of the staff.
[0118] Although the embodiments of the present application have been shown and described above, it can be understood that the above-mentioned embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-mentioned embodiments within the scope of the present application, which are still covered by the protection scope of the present application.
Claims
1. A method for monitoring and analyzing data security of an informationization platform for bidding and tendering, characterized in that, Comprise: S1, full link acquisition bidding platform data, according to the operation object, operation quantity and operation type of bidding platform data under continuous time period, the user behavior is divided into multiple monitoring points, and the user behavior sequence composed of each monitoring point is obtained; S2, according to the scene segmentation of user behavior implementation scene, taking the time distribution of user behavior processing as the main body, the abnormal behavior relative to the historical behavior baseline under each scene segmentation dimension is determined; S3, the correlation value between each abnormal behavior is obtained, and then the abnormal behavior is integrated into multiple correlation sets according to the correlation value, and the behavior correlation result of each correlation set and user operation object is determined; S4, data correlation tracing is carried out with the behavior correlation result, the behavior correlation result is connected in series in the form of monitoring point distribution, and the correlation path under different scene segmentation dimensions is determined; S5, the correlation path under all scene segmentation dimensions is verified circularly, the intersection during circular verification is used to update the path list after combination of each correlation path, and the updated path list is used as the output of the early warning target; The implementation mode of step S2 comprises: S21, each scene segmentation dimension is regarded as a processing node, and the historical behavior baseline of user behavior relative to historical data under each processing node is verified; S22, the deviation degree of user behavior under each processing node from the historical behavior baseline is judged, and the initial abnormal score of the behavior data set corresponding to multiple deviation degrees is set; S23, cross validation is carried out on the behavior data set corresponding to each deviation degree, and if the verification is inconsistent, the consistency abnormality flag is triggered; S24, the comprehensive risk score of user behavior is set according to the initial abnormal score and the consistency abnormality flag in the behavior data set, and the abnormal behavior is output according to the value of the comprehensive risk score; The implementation mode of step S21 comprises: S211, the time interval is divided according to the upload time of user behavior, each time interval is regarded as a state interval, and the operation type of user behavior in each state interval is determined; S212, in response to the operation type of user behavior, the time point of user behavior change is recorded, and the state recognition of user behavior in continuous state interval is carried out according to the operation object and operation quantity at the time of change, and the state transition path of user behavior under the corresponding operation type is determined; S213, the optimal state interval corresponding to each operation type in the state transition path is taken as the output of the historical behavior baseline, and the frequency of executing the operation type in the optimal state interval is taken as the output of the historical behavior baseline.
2. The method of claim 1, wherein the method further comprises: The implementation mode of step S1 further comprises: According to the time window corresponding to each monitoring point, all user behaviors are traversed, and the context information of user behavior in each time window is determined; According to the context information of user behavior, the discrete user behavior is aggregated in time sequence to form continuous user behavior sequence; It is judged whether the timestamp of each user behavior in the user behavior sequence corresponds to the timestamp of the bidding file, and is marked in the user behavior sequence.
3. The method of claim 1, wherein the method further comprises: The implementation mode of step S3 comprises: S31, the input abnormal behavior is received, and the abnormal behavior is abstracted as a graph node; S32, calculate the association value between any two graph nodes, and construct an abnormal behavior relationship graph by taking the association value between the graph nodes as an edge; S33, based on the abnormal behavior relationship graph, perform association mining on the abnormal behaviors to obtain an association set corresponding to the abnormal behaviors; S34, extract operation objects for each abnormal behavior in the association set to generate a behavior association result about multiple operation objects.
4. The method of claim 3, wherein the method further comprises: The implementation of step S32 includes: S321, for any two graph nodes, respectively obtain the text similarity and behavior similarity between the graph nodes; S322, when the current graph node is not the same user, use the similarity of the bid text to represent the text similarity between the graph nodes; when the graph node is the same user, extract the historical bid text of the same user, and use the similarity between the historical bid text and the current bid text to represent the text similarity between the graph nodes; S323, based on the abnormal behavior corresponding to each graph node, convert the abnormal behavior into a behavior vector according to the label corresponding to the abnormal behavior, and take the cosine similarity between any two graph nodes with respect to the behavior vector as the behavior similarity between the graph nodes; S324, set the association value according to the text similarity and behavior similarity between the graph nodes.
5. The method of claim 3, wherein the method further comprises: The implementation process of step S34 is further represented as: S341, in response to the operation objects in the association set, set a feature pool for the operation objects according to each association set; S342, collect the occurrence frequency of each element in the feature pool, sort the elements in the feature pool according to the occurrence frequency to form a behavior association list, and take the behavior association list as the output behavior association result.
6. The method of claim 1, wherein the method further comprises: The implementation of step S4 includes: S41, traverse each operation object in the behavior association result, extract the monitoring points corresponding to each operation object, and sort all the monitoring points in ascending order according to the global timestamp; S42, connect all the monitoring points in the order of identifying the monitoring points to determine the association paths belonging to the same operation object; S43, according to the weights of the monitoring points on the association paths, determine the value range of each association path respectively, and sort the association paths according to the value range from large to small to output as the association paths under different scene segmentation dimensions.
7. The method of claim 6, wherein the method further comprises: The implementation of the value range of the association path in step S43 further includes: For the value range of the association path, when the value range of the association path is positively correlated with the number of monitoring points on the association path, the sorting position of the current association path is promoted by one; otherwise, the sorting position of the current association path is maintained.
8. The method of claim 1, wherein the method further comprises: The implementation of step S5 includes: S51, take one scene segmentation dimension as a subset, and combine the association paths under all subsets into a path list; S52, perform pairwise cross-validation on the combination of association paths of all subsets, and calculate the intersection of association paths of each pair of subsets; S53, verify the frequency of the intersection of association paths, and when the frequency of the intersection of association paths exceeds a preset frequency threshold, merge the corresponding intersection of association paths as a new path and input it into the path list; if the merged path exists in the path list, the sorting position of the corresponding path is moved up.
Citation Information
Patent Citations
Bidding and tendering system based on cloud service platform
CN114971796A
Electronic bidding and tendering transaction platform-based bidding and tendering information management method and system
CN119027223A
Data security tracing method, system and device based on artificial intelligence
CN118536093A
Information management method and system based on electronic bidding transaction platform
CN120672472A