Data marking system and method based on database intelligent management

Through the data labeling system of multi-source data fusion and dynamic modeling of knowledge graphs, the problem of difficulty in fusion of multi-source heterogeneous data in traditional database management systems is solved, efficient data labeling and adaptability are achieved, and false alarm rates and manual maintenance costs are reduced.

CN120336389APending Publication Date: 2025-07-18天津云象科技发展有限公司
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510476474.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional data labeling systems based on database management have difficulty in multi-source heterogeneous data fusion labeling, insufficient real-time and adaptability of dynamic behavioral labeling, resulting in low marking coverage, high false alarm rate and dramatic increase in rule maintenance costs.

Method used

The database multi-source data acquisition module, log image feature extraction module, mark prediction module, scoring early warning module and visual feedback module are adopted to realize the semantic correlation analysis of cross-modal data through multi-source data fusion and knowledge graph dynamic modeling. The weighted scoring mechanism of user behavior embedding vectors and device abnormal probability is used to dynamically update the knowledge graph adjacency matrix and model parameters, and combine it with the visual feedback optimization system.

Benefits of technology

It improves the accuracy of data tags, reduces the false alarm rate, and reduces the cost of manual rule maintenance, achieving system adaptability and real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336389A_ABST
    Figure CN120336389A_ABST
Patent Text Reader

Abstract

The invention discloses a data marking system and method based on database intelligent management, and relates to the technical field of artificial intelligence, and the system comprises a database multi-source data collection module, a log image feature extraction module, a marking prediction module, a scoring early warning module and a visual feedback module. The data acquisition module acquires a user operation log and an equipment image, the feature extraction module extracts an operation frequency and a red channel pixel proportion, the mark prediction module outputs a user behavior embedding vector and an equipment abnormal probability, and the score early warning module calculates a comprehensive risk score, sets a mark logic and triggers early warning. And the visual feedback module carries out visual display and dynamically updates the node weight of the knowledge graph. Through multi-source heterogeneous data fusion and knowledge graph dynamic modeling, user-equipment behavior association analysis and double risk quantitative evaluation are realized, the marking accuracy and real-time performance are improved, and through visual feedback closed-loop optimization, the manual maintenance cost is reduced, and a self-adaptive intelligent marking system is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a data marking system and method based on intelligent database management. Background Technique

[0002] A data marking system based on database management is a management tool that uses database technology to classify, mark, and efficiently retrieve data. This system stores data in a structured manner, achieving high sharing, low redundancy, and ensuring data independence and security. Its core functions include data definition, manipulation, transaction management, and operation and maintenance, supporting enterprises to accurately analyze and apply data. This system is widely used in the retail, financial, and medical industries to help enterprises improve their data-driven decision-making capabilities.

[0003] In order to solve the problems of difficult fusion marking of multi-source heterogeneous data and insufficient real-time performance and adaptability of dynamic behavior marking in traditional data marking systems based on database management, the existing technology is to use a combination of a rule engine and a static threshold for processing, and simple marking is achieved through predefined operation frequency thresholds and fixed device status rules. However, there will still be cases of cross-modal data association breakage and missed detection of dynamic behavior patterns, which will further lead to problems such as low marking coverage, high false alarm rate, and sharp increase in rule maintenance costs. In order to solve the above problems, a data marking system and method based on intelligent database management are proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a data marking system and method based on intelligent database management to solve the problems raised in the above background technique.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is: A data marking system based on intelligent database management, including a database multi-source data acquisition module, a log image feature extraction module, a marking prediction module, a scoring and warning module, and a visualization feedback module; The database multi-source data acquisition module collects and preprocesses user operation log data and device operation status image data; The log image feature extraction module extracts the user operation frequency and the proportion of red channel pixels from the preprocessed user operation log data and device operation status image data; The marking prediction module constructs a knowledge graph marking model and a device anomaly prediction model based on the user operation frequency and the proportion of red channel pixels, and generates user behavior embedding vectors and device anomaly probabilities; The scoring and warning module weights and concatenates the user behavior embedding vectors and device anomaly probabilities into a comprehensive risk score, and based on the comprehensive risk score, sets data marking logic and conducts system-level warnings; The visualization feedback module draws a line chart of the user operation frequency over time, a heat map of the device status, and a histogram of the risk score distribution, and updates the node weights of the knowledge graph through the feedback of the manual review results.

[0006] A further improvement of the technical solution of the present invention is that in the multi-source data acquisition module of the database, the acquisition and preprocessing process of the user operation log data and the device operation status image data includes: Deploy a log acquisition agent on the database server. The acquisition agent listens to the database transaction log through the database interface, grabs the fields including the operation type, timestamp, and operation object number, and captures the user operation log data in real time; Deploy high-definition cameras at key positions of industrial equipment, equipped with ring fill lights. Deploy the image processing node on the edge computing gateway, which is directly connected to the high-definition camera through optical fiber. The camera captures the device operation status at a frequency of 2 frames per second, saves it as an RGB three-channel image, and the single-frame storage size ≤ 500KB, so as to obtain the device operation status image data; Among them, a data cleaning node is configured between the acquisition agent and the transmission pipeline, connected to the memory buffer. The acquisition frequency is set to a maximum of 1000 items per second. When the acquisition frequency threshold is exceeded, the cache queue is triggered to temporarily store. The fill light intensity is dynamically adjusted according to the ambient brightness, so that the image brightness value L satisfies: , where R, G, B are the average values of the image pixels; Obtain the hash value of the continuously acquired log entries. If the hash values of two adjacent log entries are equal, it is determined as duplicate data, and the latter is deleted. The timestamp is uniformly converted into the format of year-month-day - hour:minute:second, and the operation object number is filled to a 12-bit fixed-length string. The preprocessed user operation log data is encapsulated in the key-value pair format and distributed to the topic partition of the message middleware through the polling strategy. When the backlog data volume of the message middleware exceeds 10,000 items, the camera acquisition frequency is automatically reduced to 1 frame / second, and the non-critical log acquisition is paused; Adopt a median filter to sort the pixel values in the 3×3 neighborhood of each pixel point P(x, y), take the median to replace the original pixel value, obtain the interpolation weight coefficient from the pixel coordinate offset, and adopt the bilinear interpolation algorithm to scale the original device operation status image to 640×480 pixels in a 4:3 ratio. The compressed device operation status image is converted into a binary stream, and a timestamp header is appended according to the frame number. Adopt the breakpoint resumption mechanism, and resume transmission from the nearest complete frame when the transmission is interrupted. If the image preprocessing node outputs abnormal data continuously for 5 times, the device self-check process is triggered.

[0007] A further improvement of the technical solution of the present invention is that in the log image feature extraction module, the process of extracting the user operation frequency and the proportion of red channel pixels from the preprocessed user operation log data and device operation status image data includes: Based on the preprocessed user operation log data, with a fixed window length of 10 minutes, the window intervals are divided according to the log timestamps. For each user ID, the number of operations within its window is counted to obtain the user operation frequency. As the system clock advances by 1 second, the window intervals slide synchronously, and the operation count within the latest 600 seconds is updated in real time. For cross-window operations, they are assigned to the corresponding windows according to the timestamp attribution, and a user-frequency mapping table is generated. The user operation frequency is transmitted through the message middleware topic partition, and the partition is assigned according to the hash value of the user ID, so that the data of the same user is routed to a fixed partition. When the user operation frequency is greater than 1 time per second, it is marked as a feature to be verified and temporarily stored in the buffer queue; Based on the preprocessed device operation status image data, the image is decomposed into three independent channel matrices of red, green, and blue. The red channel matrix is extracted. The percentage of the total intensity value of the red channel to 255 times the total number of effective pixels is used as the normalized red channel intensity ratio, and a device-red ratio mapping table is generated. The red channel pixel ratio is appended with the device ID and timestamp and converted into a binary stream for transmission through an independent channel. When the red channel pixel ratio is greater than 95%, the image re-acquisition mechanism is triggered to request the edge gateway to re-obtain the current frame image. When the red channel pixel ratio exceeds the maximum theoretical red channel pixel ratio for 3 consecutive times, it is determined that the sensor is faulty, and the database multi-source data acquisition module is notified to switch to the backup camera.

[0008] A further improvement of the technical solution of the present invention lies in: in the marker prediction module, based on the user operation frequency, through a graph convolutional network architecture, the process of constructing a knowledge graph marker model to generate user behavior embedding vectors includes: Construct a knowledge graph marker model based on three types of nodes of users, devices, and operations and their association relationships. The association relationships include user-operation relationships and user-device relationships. Define a graph structure matrix. If there is a relationship between nodes i and j, then ; if there is no relationship between nodes i and j, then ; a weighted adjacency matrix is used to structurally represent the node connection status through the adjacency matrix; Using the operation frequency as the initial feature of the node, aggregate neighborhood information through multi-layer graph convolution. Each layer fuses the node's own feature and the weighted feature of the first-order neighbor, and after parameter matrix transformation and non-linear activation, finally outputs a 32-dimensional user behavior embedding vector; The edge weight of the user-operation relationship is calculated by normalizing the historical operation times, and the edge weight of the user-device relationship is determined by the frequency of the user operating the device. The adjacency matrix is iteratively refreshed every 10 minutes to make the structure of the knowledge graph marker model dynamically adapt to the changes in user behavior; Load user operation frequency data from the topic partitions of the message middleware, inject it into the node features of the knowledge graph tagging model, perform graph convolution calculations, and output the user behavior embedding vector to the scoring and warning module. When the norm of the user behavior embedding vector is abnormal, trigger the reconstruction of the knowledge graph tagging model.

[0009] A further improvement of the technical solution of the present invention is that in the tagging prediction module, based on the proportion of red channel pixels, an equipment anomaly prediction model is constructed through a long short-term memory network architecture. The process of generating the equipment anomaly probability includes: Divide the equipment red channel pixel proportion data into 10 time series points according to a 5-minute window to construct a continuous input sequence; Design a two-layer memory unit network architecture, regulate the information flow through the forget gate, input gate, and output gate, construct an equipment anomaly prediction model, capture the time series dependence relationship of the red channel pixel proportion, the bottom memory unit extracts local fluctuation features, the top memory unit aggregates long-range patterns, and output the equipment anomaly probability through a fully connected layer; The newly added equipment red channel pixel proportion data points are used to slide and update the input window in real time, and incremental training is performed regularly every day. Optimize the parameters of the equipment anomaly prediction model based on the latest 30-day data; Adopt temperature scaling technology to suppress the extreme values of the equipment anomaly probability output, and apply weighted average filtering to the equipment anomaly probability for three consecutive times for correction.

[0010] A further improvement of the technical solution of the present invention is that in the scoring and warning module, the process of weighted splicing the user behavior embedding vector and the equipment anomaly probability into a comprehensive risk score includes: The 32-dimensional user behavior embedding vector Is compressed into a 1-dimensional scalar through a fully connected layer , where Is the weight vector, Is the bias term; The 1-dimensional scalar And the equipment anomaly probability Are combined into a comprehensive risk score according to a preset weight , where the weight of the 1-dimensional scalar is 0.6, the weight of the equipment anomaly probability is 0.4, and the weight coefficient is determined by grid search of historical data.

[0011] A further improvement of the technical solution of the present invention is that in the scoring and warning module, the process of setting data marking logic and performing system-level warning based on the comprehensive risk score includes: Based on the comprehensive risk score, the data marking logic is set at different levels. The data marking logic includes: when s > 0.8, it is marked as a high-risk operation - equipment anomaly, triggering real-time alerts and locking the relevant accounts; when 0.6 < s ≤ 0.8, it is marked as a suspicious behavior and pushed to the manual review queue; when s ≤ 0.6, it is marked as a compliant operation and the data is normally stored in the database; When multiple devices are associated with the same user, the highest comprehensive risk score is taken as the final marking basis. If the manual review overrides the automatic marking result, then according to the formula: 、 Update the weight coefficient, where, is the learning rate; Taking 1 hour as a window, the mean and standard deviation of the user's comprehensive risk score are statistically calculated. When the mean plus 3 times the standard deviation exceeds 0.75, a system-level warning is triggered, notifying the administrator to start a full scan, and the warning threshold is dynamically calculated based on the distribution of the comprehensive risk score for the day where, and are the mean and standard deviation of the comprehensive risk score for the day. When the warning threshold exceeds the benchmark value of 0.75, the warning sensitivity is automatically increased.

[0012] A further improvement in the technical solution of the present invention lies in: in the visualization feedback module, the process of drawing the user operation frequency time series line chart, the device status heat map, and the risk score distribution histogram includes: Through the user-frequency mapping table, for the user operation frequency time series line chart, the horizontal axis is time and the vertical axis is the operation frequency value. It is grouped by user ID, and each user generates an independent line. The data points are connected in ascending order of time. After each 10-minute window slides, new data points are automatically appended and the view is refreshed; Through the device-red ratio mapping table, for the device status heat map, the device ID is used as the row and the time window is used as the column. The color intensity is generated by mapping the red channel pixel occupancy ratio to the red channel; The risk score distribution histogram divides the data into three intervals to calculate the proportion, calculates the proportion of the data volume in each level interval and generates a three-color stacked bar chart, marks the percentage value, and refreshes the view in real time to show the risk distribution trend.

[0013] A further improvement in the technical solution of the present invention lies in: in the visualization feedback module, the process of updating the node weights of the knowledge graph through the feedback of the manual review results includes: The manual review results are accessed through the message middleware. The manual review results include confirmation and false alarm. When confirmed, the user-device association weight is enhanced according to the comprehensive risk score. When it is a false alarm, the user-operation relationship weight is weakened proportionally. When the cumulative adjustment amount of the single-user weight exceeds 0.5, the neighbor node relationship is asynchronously reconstructed and the non-zero elements in the corresponding rows and columns of the adjacency matrix are updated.

[0014] Second aspect, a data marking method based on intelligent database management, is implemented based on the above-mentioned data marking system for intelligent database management, and includes the following steps: S1. Collect and preprocess user operation log data and device operation status image data; S2. Extract the user operation frequency and the proportion of red channel pixels from the preprocessed user operation log data and device operation status image data; S3. Based on the user operation frequency and the proportion of red channel pixels, construct a knowledge graph marking model and a device anomaly prediction model, and output user behavior embedding vectors and device anomaly probabilities; S4. Perform weighted splicing on the user behavior embedding vectors and device anomaly probabilities, output a comprehensive risk score, and based on the comprehensive risk score, set data marking logic and conduct system-level early warnings; S5. Draw a time series line chart of user operation frequency, a heat map of device status, and a histogram of risk score distribution, and update the node weights of the knowledge graph through feedback from manual review results.

[0015] Due to the adoption of the above technical solutions, the technical progress achieved by the present invention compared with the prior art is: The present invention provides a data marking system and method based on intelligent database management. Through multi-source data fusion and dynamic modeling of the knowledge graph, semantic association analysis of cross-modal data is realized, and the problems of low marking coverage rate and high misjudgment rate caused by data isolation in traditional methods are solved. Based on the weighted scoring mechanism of user behavior embedding vectors and device anomaly probabilities, the dual risk contributions are quantified, and the data marking accuracy is improved.

[0016] The present invention provides a data marking system and method based on intelligent database management. By adopting real-time sliding window statistics and incremental learning optimization technologies, the adjacency matrix of the knowledge graph and the parameters of the LSTM model are dynamically updated, so that the system can still maintain high-precision marking when the user behavior pattern changes suddenly and the device status drifts abnormally, reducing the false alarm rate compared with the static rule method and reducing the manual rule maintenance cost.

[0017] The present invention provides a data marking system and method based on intelligent database management. Through a visual feedback closed loop, human-machine collaborative optimization is realized. The administrator can quickly locate the source of anomalies through multi-dimensional views, and inversely update the node weights of the knowledge graph in combination with the results of manual review, forming a closed-loop control of risk discovery-artificial verification-model iteration, so that the system self-optimizes to a stable state and the frequency of manual intervention decreases. Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments described in the present invention. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a block diagram of the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0021] Embodiment 1, as Figure 1 shown, the present invention provides a data marking system based on database intelligent management, including a database multi-source data acquisition module, a log image feature extraction module, a marking prediction module, a scoring and early warning module, and a visualization feedback module; The database multi-source data acquisition module collects and preprocesses user operation log data and device operation status image data. A log collection agent is deployed on the database server. The database interface collection agent listens to the database transaction log, grabs fields including operation type, timestamp, and operation object number, and captures user operation log data in real time. High-definition cameras are deployed at key positions of industrial equipment, equipped with ring fill lights. The image processing node is deployed on the edge computing gateway and directly connected to the high-definition camera through optical fiber. The camera captures the device operation status at a frequency of 2 frames per second and saves it as an RGB three-channel image, with a single-frame storage size ≤ 500KB, thereby obtaining device operation status image data. Among them, a data cleaning node is configured between the collection agent and the transmission pipeline, connected to the memory buffer. The collection frequency is set to a maximum of 1000 items per second. When the collection frequency threshold is exceeded, the cache queue is triggered to temporarily store. The fill light intensity is dynamically adjusted according to the environmental brightness so that the image brightness value L satisfies: , where R, G, B are the average values of image pixels. Obtain the hash value of continuously collected log entries. If the hash values of two adjacent log entries are equal, it is determined as duplicate data, and the latter is deleted. The timestamp is uniformly converted into the format of year-month-day - hour:minute:second. The operation object number is supplemented to a 12-bit fixed-length string. The preprocessed user operation log data is encapsulated in the key-value pair format and distributed to the topic partitions of the message middleware through the polling strategy. When the amount of backlogged data in the message middleware exceeds 10,000 entries, the camera acquisition frequency is automatically reduced to 1 frame / second, and the collection of non-critical logs is suspended. The median filter is used to sort the pixel values within the 3×3 neighborhood of each pixel point P(x, y), and the median is taken to replace the original pixel value. The interpolation weight coefficient is obtained from the pixel coordinate offset, and the bilinear interpolation algorithm is used to scale the original device operation status image to 640×480 pixels in a 4:3 ratio. The compressed device operation status image is converted into a binary stream, and a timestamp header is appended according to the frame number. The breakpoint resumption mechanism is adopted, and when the transmission is interrupted, it is restarted from the nearest complete frame. If the image preprocessing node outputs abnormal data continuously for 5 times, the device self-check process is triggered; The log image feature extraction module extracts the user operation frequency and the proportion of red channel pixels from the preprocessed user operation log data and device operation status image data. Based on the preprocessed user operation log data, with a fixed window length of 10 minutes, the window interval is divided according to the log timestamp. For each user number, the number of operations within its window is counted to obtain the user operation frequency. Every time the system clock advances 1 second, the window interval slides synchronously, and the operation count within the latest 600 seconds is updated in real time. For cross-window operations, they are assigned to the corresponding window according to the timestamp attribution, and a user-frequency mapping table is generated. The user operation frequency is transmitted through the topic partitions of the message middleware and assigned to partitions according to the hash value of the user number, so that the data of the same user is routed to a fixed partition. When the user operation frequency is greater than 1 time / second, it is marked as a feature to be verified and temporarily stored in the buffer queue. Based on the preprocessed device operation status image data, the image is decomposed into three independent channel matrices of red, green, and blue, and the red channel matrix is extracted. The percentage of the total intensity value of the red channel to 255 times the total number of effective pixels is used as the normalized red channel intensity proportion, and a device-red proportion mapping table is generated. The proportion of red channel pixels is appended with the device number and timestamp and converted into a binary stream for transmission through an independent channel. When the proportion of red channel pixels is greater than 95%, the image re-acquisition mechanism is triggered to request the edge gateway to re-obtain the current frame image. When the proportion of red channel pixels exceeds the maximum theoretical value of the red channel pixel proportion continuously for 3 times, it is determined as a sensor failure, and the database multi-source data acquisition module is notified to switch to the backup camera; The marker prediction module constructs a knowledge graph marking model and a device anomaly prediction model based on the user operation frequency and the proportion of red channel pixels, generates a user behavior embedding vector and a device anomaly probability, constructs a knowledge graph marking model based on three types of nodes including users, devices, and operations and their association relationships. The association relationships include user-operation relationships and user-device relationships. Define a graph structure matrix. If there is a relationship between node i and j, then , if there is no relationship between node i and j, then . The weighted adjacency matrix structurally represents the node connection status through the adjacency matrix. Using the operation frequency as the initial feature of the node, aggregate neighborhood information through multi-layer graph convolution. Each layer fuses the node's own feature and the weighted feature of the first-order neighbors, and through parameter matrix transformation and non-linear activation, finally outputs a 32-dimensional user behavior embedding vector. The edge weight of the user-operation relationship is calculated by normalizing the historical operation times, and the edge weight of the user-device relationship is determined by the frequency of the user operating the device. Iteratively refresh the adjacency matrix every 10 minutes to make the structure of the knowledge graph marking model dynamically adapt to the changes in user behavior. Load the user operation frequency data from the topic partition of the message middleware, inject it into the node features of the knowledge graph marking model, and then perform graph convolution calculation, and output the user behavior embedding vector to the scoring and warning module. When the norm of the user behavior embedding vector is abnormal, trigger the reconstruction of the knowledge graph marking model. Divide the device red channel pixel proportion data into 10 time series points according to a 5-minute window, construct a continuous input sequence, design a two-layer memory unit network architecture, regulate the information flow through the forget gate, input gate, and output gate, construct a device anomaly prediction model, capture the time series dependence relationship of the red channel pixel proportion, the bottom memory unit extracts local fluctuation features, the top memory unit aggregates long-range patterns, and outputs the device anomaly probability through a fully connected layer. The newly added data points of the device red channel pixel proportion are slid and updated in real time for the input window, and incremental training is performed daily at a fixed time. Optimize the parameters of the device anomaly prediction model based on the latest 30 days of data, use temperature scaling technology to suppress extreme values in the output of the device anomaly probability, and apply weighted average filtering to the device anomaly probability for three consecutive times for correction; The scoring and warning module weights and splices the user behavior embedding vector and the device anomaly probability into a comprehensive risk score. Based on the comprehensive risk score, set data marking logic and conduct system-level warnings. The 32-dimensional user behavior embedding vector is compressed into a 1-dimensional scalar through a fully connected layer , where is the weight vector, is the bias term. The 1-dimensional scalar and the device anomaly probability are combined into a comprehensive risk score according to the preset weight , where the 1D scalar weight is 0.6, the device anomaly probability weight is 0.4, and the weight coefficient is determined by historical data grid search. Based on the comprehensive risk score, the data marking logic is set hierarchically. The data marking logic includes: when s > 0.8, it is marked as high-risk operation - device anomaly, triggering real-time alarm and locking related accounts; when 0.6 < s ≤ 0.8, it is marked as suspicious behavior and pushed to the manual review queue; when s ≤ 0.6, it is marked as compliant operation and the data is normally stored in the database. When multiple devices are associated with the same user, the highest comprehensive risk score is taken as the final marking basis. If the manual review overrides the automatic marking result, then according to the formula: 、 Update the weight coefficient, where is the learning rate. Taking 1 hour as the window, the mean and standard deviation of the user's comprehensive risk score are statistically calculated. When the mean plus 3 times the standard deviation exceeds 0.75, a system-level warning is triggered to notify the administrator to start a full scan, and the warning threshold is dynamically calculated based on the distribution of the comprehensive risk score on the same day , where and are the mean and standard deviation of the comprehensive risk score on the same day. When the warning threshold exceeds the benchmark value of 0.75, the alarm sensitivity is automatically increased; Visual feedback module, which draws a time-series line chart of user operation frequency, a heat map of device status, and a histogram of risk score distribution. The node weights of the knowledge graph are updated through the feedback of the manual review results. Through the user-frequency mapping table, the time-series line chart of user operation frequency has the time on the horizontal axis and the operation frequency value on the vertical axis, grouped by user number, and each user generates an independent line. The data points are connected in ascending order of time. After each 10-minute window slides, new data points are automatically appended and the view is refreshed. Through the device-red ratio mapping table, the heat map of device status has the device number as the row and the time window as the column, and the color intensity is generated by mapping the red channel pixel occupancy ratio to the red channel. The histogram of risk score distribution divides into three intervals to statistically calculate the proportion, calculates the proportion of the data volume in each interval and generates a three-color stacked bar chart, marking the percentage value, and refreshing the view in real time to show the risk distribution trend. The manual review results are accessed through the message middleware. The manual review results include confirmation and false alarm. When confirmed, the user-device association weight is enhanced according to the comprehensive risk score. When it is a false alarm, the user-operation relationship weight is weakened proportionally. When the cumulative adjustment amount of the single-user weight exceeds 0.5, the neighbor node relationship is asynchronously reconstructed and the non-zero elements in the corresponding rows and columns of the adjacency matrix are updated.

[0022] Embodiment 2, as Figure 1 shown, on the basis of Embodiment 1, the present invention provides a technical solution: a data marking method based on intelligent management of a database, implemented based on the above-mentioned data marking system based on intelligent management of a database, including the following steps: S1. Collect and preprocess user operation log data and device operation status image data; S2. Extract the user operation frequency and the proportion of red channel pixels from the preprocessed user operation log data and device operation status image data; S3. Based on the user operation frequency and the proportion of red channel pixels, construct a knowledge graph marking model and a device anomaly prediction model, and output user behavior embedding vectors and device anomaly probabilities; S4. Perform weighted splicing on the user behavior embedding vectors and device anomaly probabilities, output a comprehensive risk score, and based on the comprehensive risk score, set data marking logic and conduct system-level early warnings; S5. Draw a time series line chart of user operation frequency, a heat map of device status, and a histogram of risk score distribution, and update the node weights of the knowledge graph through feedback from manual review results.

[0023] As mentioned above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claimed rights.

Claims

1. A data marking system based on intelligent management of a database, characterized in that: It includes a database multi-source data acquisition module, a log image feature extraction module, a marking prediction module, a scoring and warning module, and a visualization feedback module; The database multi-source data acquisition module collects and preprocesses user operation log data and device operation status image data; The log image feature extraction module extracts the user operation frequency and the proportion of red channel pixels from the preprocessed user operation log data and device operation status image data; The marking prediction module constructs a knowledge graph marking model and a device anomaly prediction model based on the user operation frequency and the proportion of red channel pixels, and generates a user behavior embedding vector and a device anomaly probability; The scoring and warning module weights and concatenates the user behavior embedding vector and the device anomaly probability into a comprehensive risk score, and based on the comprehensive risk score, sets data marking logic and conducts system-level warnings; The visualization feedback module draws a time series line chart of user operation frequency, a heat map of device status, and a histogram of risk score distribution, and updates the node weights of the knowledge graph through the feedback of manual review results.

2. A data marking system based on database intelligent management according to claim 1, characterized in that: In the database multi-source data acquisition module, the process of collecting and preprocessing user operation log data and device operation status image data includes: Deploy a log collection agent on the database server. The collection agent listens to the database transaction log through the database interface, grabs the fields including the operation type, timestamp, and operation object number, and captures the user operation log data in real time; Deploy high-definition cameras at key positions of industrial equipment, equipped with ring-shaped fill lights. Deploy the image processing node on the edge computing gateway, directly connect to the high-definition camera through optical fiber. The camera captures the device operation status at a frequency of 2 frames per second, saves it as an RGB three-channel image, and the single-frame storage size ≤ 500KB, so as to obtain the device operation status image data; Among them, a data cleaning node is configured between the acquisition agent and the transmission pipeline, which is connected to the memory buffer. The acquisition frequency is set to a maximum of 1000 pieces per second. When the acquisition frequency threshold is exceeded, the cache queue is triggered for temporary storage, and the fill light intensity is dynamically adjusted according to the environmental brightness, so that the image brightness value L satisfies: , where R, G, and B are the average values of image pixels; Obtain the hash value of continuously collected log entries. If the hash values of two adjacent log entries are equal, it is determined as duplicate data, and the latter is deleted. The timestamp is uniformly converted into the format of year-month-day - hour:minute:second, and the operation object number is filled to a 12-bit fixed-length string. The preprocessed user operation log data is encapsulated in a key-value pair format and distributed to the topic partition of the message middleware through a polling strategy. When the backlog of data in the message middleware exceeds 10,000 entries, the camera acquisition frequency is automatically reduced to 1 frame / second, and the collection of non-critical logs is paused; Adopt a median filter to sort the pixel values within the 3×3 neighborhood of each pixel point P(x,y), replace the original pixel value with the median value, obtain the interpolation weight coefficient from the pixel coordinate offset, adopt a bilinear interpolation algorithm to scale the original device operation status image to 640×480 pixels in a 4:3 ratio. The compressed device operation status image is converted into a binary stream, and a timestamp header is appended according to the frame number. Adopt a breakpoint resumption mechanism to resume transmission from the nearest complete frame when the transmission is interrupted. If the image processing node outputs abnormal data continuously for 5 times, trigger the device self-check process.

3. A data marking system based on intelligent management of a database according to claim 2, characterized in that: In the log image feature extraction module, the process of extracting the user operation frequency and the proportion of red channel pixels from the preprocessed user operation log data and device operation status image data includes: Based on the preprocessed user operation log data, with a fixed window length of 10 minutes, the window intervals are divided according to the log timestamps. For each user ID, the number of operations within its window is counted to obtain the user operation frequency. Every time the system clock advances 1 second, the window intervals slide synchronously, and the operation counts within the latest 600 seconds are updated in real-time. For cross-window operations, they are assigned to the corresponding windows according to the timestamp attribution, generating a user-frequency mapping table. The user operation frequency is transmitted through the message middleware topic partition, and the partitions are assigned according to the hash value of the user ID, so that the data of the same user is routed to a fixed partition. When the user operation frequency is greater than 1 time per second, it is marked as a feature to be verified and temporarily stored in the buffer queue; Based on the preprocessed device running status image data, the image is decomposed into three independent channel matrices of red, green, and blue. The red channel matrix is extracted, and the percentage of the total intensity value of the red channel to 255 times the total number of effective pixels is used as the normalized red channel intensity ratio, generating a device-red ratio mapping table. The red channel pixel ratio is appended with the device ID and timestamp and converted into a binary stream for transmission through an independent channel. When the red channel pixel ratio is greater than 95%, the image re-acquisition mechanism is triggered to request the edge gateway to re-obtain the current frame image. When the red channel pixel ratio exceeds the maximum theoretical red channel pixel ratio three times in a row, it is determined that the sensor is faulty, and the database multi-source data acquisition module is notified to switch to the backup camera.

4. A data marking system based on intelligent database management according to claim 3, characterized in that: In the said label prediction module, based on the user operation frequency, through the graph convolutional network architecture, the process of constructing the knowledge graph label model and generating the user behavior embedding vector includes: Construct a knowledge graph marking model based on three types of nodes: users, devices, and operations, and their associated relationships. The associated relationships include user-operation relationships and user-device relationships. Define a graph structure matrix. If there is a relationship between node i and j, then ; if there is no relationship between node i and j, then , a weighted adjacency matrix, which structurally represents the node connection status through the adjacency matrix; Using the operation frequency as the initial node feature, aggregating neighborhood information through multi-layer graph convolution. Each layer fuses the node's own feature and the weighted feature of the first-order neighbors, and after parameter matrix transformation and non-linear activation, finally outputs a 32-dimensional user behavior embedding vector; The edge weight of the user-operation relationship is calculated by normalizing the historical operation times, and the edge weight of the user-device relationship is determined by the frequency of the user operating the device. The adjacency matrix is iteratively refreshed every 10 minutes to make the structure of the knowledge graph label model dynamically adapt to the changes in user behavior; Loading the user operation frequency data from the topic partition of the message middleware, injecting the node features into the knowledge graph label model and then performing graph convolution calculation, and outputting the user behavior embedding vector to the scoring and early warning module. When the norm of the user behavior embedding vector is abnormal, the knowledge graph label model is triggered to be reconstructed.

5. A data marking system based on intelligent management of a database according to claim 4, characterized in that: In the said label prediction module, based on the red channel pixel ratio, through the long short-term memory network architecture, the process of constructing the device anomaly prediction model and generating the device anomaly probability includes: Dividing the device red channel pixel ratio data into 10 time series points according to a 5-minute window to construct a continuous input sequence; Designing a two-layer memory unit network architecture, regulating the information flow through the forget gate, input gate, and output gate, constructing a device anomaly prediction model to capture the temporal dependence relationship of the red channel pixel ratio. The bottom memory unit extracts local fluctuation features, and the top memory unit aggregates long-range patterns, and outputs the device anomaly probability through a fully connected layer; The real-time sliding update input window for the pixel ratio data points of the newly added device red channel, perform incremental training at a fixed time every day, and optimize the parameters of the device anomaly prediction model based on the latest 30-day data; Adopt temperature scaling technology to suppress the extreme values of the device anomaly probability output, and apply weighted average filtering to the device anomaly probability for three consecutive times for correction.

6. The data marking system based on database intelligent management according to claim 5, characterized in that: In the scoring and warning module, the process of weighted splicing the user behavior embedding vector and the device anomaly probability into the comprehensive risk score includes: Embed the 32-dimensional user behavior vector Compress it into a 1-dimensional scalar through a fully connected layer , where is the weight vector is the bias term; Combine a 1D scalar with the device exception probability into a comprehensive risk score according to preset weights , where the weight of the 1D scalar is 0.6, the weight of the device exception probability is 0.4, and the weight coefficients are determined by historical data grid search.

7. The data marking system based on database intelligent management according to claim 6, characterized in that: In the scoring and warning module, the process of setting the data marking logic and performing system-level warning based on the comprehensive risk score includes: Based on the comprehensive risk score, the data marking logic is set at different levels. The data marking logic includes: when s>0.8, it is marked as high-risk operation - device anomaly, triggering real-time alarm and locking the relevant accounts; when 0.6<s≤0.8, it is marked as suspicious behavior and pushed to the manual review queue; when s≤0.6, it is marked as compliant operation and the data is normally stored in the database; When multiple devices are associated with the same user, the highest comprehensive risk score is taken as the final marking basis. If the automatic marking result is overturned by manual review, the weight coefficient is updated according to the formula: , where is the learning rate; Taking 1 hour as a window, statistically calculate the mean and standard deviation of the comprehensive risk scores of users. When the mean plus 3 times the standard deviation exceeds 0.75, trigger a system-level warning, notify the administrator to start a full scan, and dynamically calculate the warning threshold based on the distribution of the comprehensive risk scores on the same day , where and are the mean and standard deviation of the comprehensive risk scores on the same day. When the warning threshold exceeds the benchmark value of 0.75, automatically increase the warning sensitivity.

8. A data marking system based on intelligent management of a database according to claim 7, characterized in that: In the visualization feedback module, the process of drawing the time series line chart of user operation frequency, the heat map of device status, and the histogram of risk score distribution includes: Through the user-frequency mapping table, the time series line chart of user operation frequency takes the horizontal axis as time and the vertical axis as the operation frequency value, grouped by user number, and each user generates an independent line. The data points are connected in ascending order of time. After each 10-minute window slides, new data points are automatically appended and the view is refreshed; Through the device-red ratio mapping table, the heat map of device status takes the device number as the row and the time window as the column, and the color intensity is generated by mapping the red channel from the pixel ratio value of the red channel; The risk score distribution histogram divides into three intervals to count the proportion, calculates the proportion of the data volume in each interval and generates a three-color stacked bar chart, marks the percentage value, and refreshes the view in real time to show the risk distribution trend.

9. A data marking system based on intelligent management of a database according to claim 8, characterized in that: In the visualization feedback module, the process of updating the node weights of the knowledge graph through the feedback of the manual review result includes: The manual review result is accessed through the message middleware. The manual review result includes confirmation and false alarm. When confirmed, the user-device association weight is enhanced according to the comprehensive risk score. When there is a false alarm, the user-operation relationship weight is weakened proportionally. When the cumulative adjustment amount of the single-user weight exceeds 0.5, the neighbor node relationship is asynchronously reconstructed and the non-zero elements in the corresponding rows and columns of the adjacency matrix are updated.

10. A data marking method based on intelligent management of a database, implemented based on the data marking system for intelligent management of a database according to any one of the above claims 1-9, characterized in that: It includes the following steps: S1. Collect and preprocess the user operation log data and the device running status image data; S2. Extract the user operation frequency and the pixel ratio of the red channel from the preprocessed user operation log data and the device running status image data; S3. Based on the user operation frequency and the pixel ratio of the red channel, construct a knowledge graph marking model and a device anomaly prediction model, and output the user behavior embedding vector and the device anomaly probability; S4. Perform weighted splicing on the user behavior embedding vector and the device anomaly probability, output the comprehensive risk score, and based on the comprehensive risk score, set the data marking logic and perform system-level warning; S5. Draw the time series line chart of user operation frequency, the heat map of device status, and the histogram of risk score distribution, and update the node weights of the knowledge graph through the feedback of the manual review result.

Citation Information

Cited By

  • Smart factory equipment data acquisition system based on extended access

    CN121008519A

  • Dynamic layer increment updating method and system for large-scale three-dimensional scene

    CN121147421A

  • A method and system for dynamic layer incremental update of large-scale three-dimensional scenes

    CN121147421B

  • Examination data auditing method and system based on knowledge graph

    CN122266598A

  • A knowledge graph-based inspection data auditing method and system

    CN122266598B