Big data-based multi-modal data management system and method
Through a multimodal data governance system based on big data, integrating the processing layer, distributed storage layer, intelligent governance layer and multimodal analysis layer, the problems of data quality and processing speed are solved, and data quality is guaranteed and processing speed is improved.
Patent Information
- Application Number
- CN202510464019.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
Existing data governance technologies are difficult to ensure the data quality requirements of different industries, and the data processing speed is difficult to meet the data application standards.
A multimodal data governance system based on big data is adopted, including an integrated processing layer, a distributed storage layer, an intelligent governance layer, a multimodal analysis layer and a decision-making control layer. Through data reception strategies, a layered storage mechanism and a dynamic security governance mechanism, massive legal entity data are protected in full process.
It ensures the data quality of massive legal entity data in different industries and improves the data processing speed during data governance.
Smart Images

Figure CN120337146A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data governance, and in particular to a multimodal data governance system and method based on big data. Background Technique
[0002] With the accelerating advancement of digital and intelligent transformation, the demand for the integration and analysis of multimodal data in fields such as public security, justice, and administrative law enforcement is becoming increasingly urgent. The business of public security, procuratorate, and courts involves multimodal data such as images (such as surveillance videos), texts (case files), and audios (communication records). Traditional single-modal data management methods are difficult to achieve cross-modal correlation analysis. For example, anti-fraud early warning requires the combination of multi-source data such as communication records, fund flows, and geographical locations, and problems such as scattered data storage and heterogeneous formats lead to low analysis efficiency. Although large models (such as natural language processing models) have powerful capabilities in theory, their actual applications are often limited by insufficient data quality, poor timeliness, and lack of interpretability.
[0003] The public security, procuratorate, and court industries have extremely high requirements for data security and privacy protection. Therefore, encryption technology, permission management, and auditing mechanisms need to be embedded in data governance to ensure the security of the entire data life cycle. In addition, when sharing cross-departmental data, it is necessary to balance openness and confidentiality, and traditional systems are difficult to meet the requirements of dynamic permission control.
[0004] Patent No. CN201910957435.8 discloses a system and method for pollution control in college chemistry laboratories based on big data analysis. The steps include: obtaining pollution data of college chemistry laboratories; summarizing and processing the pollution data according to a preset method and then sending and storing it in a big data platform supporting green chemistry; in the big data platform, visualizing and displaying the summarized data in real time; using a multimodal data mining method and combining with a neural network method for the summarized data to extract feature information corresponding to the pollution data; according to the feature information, looking for feature information that is completely or partially opposite to the feature information in the big data platform supporting green chemistry according to the above method; according to the feature information, deducing the dynamic evolution law of the individual combined by the two; and obtaining a performance evaluation index of real-time data by combining the dynamic evolution law of the individual. The above invention provides a decision-making basis and technical support for the control of pollution in college chemistry laboratories, and improves the control efficiency and control quality.
[0005] Patent No. CN202311805638.8 discloses a method and system for information data management based on big data, including the following steps: Based on the original data set, a multi-modal data analysis algorithm and a long short-term memory network model are used to monitor and evaluate data quality, and a data quality evaluation report is generated. In the above invention, through an anomaly pattern recognition model that combines ensemble learning and incremental learning algorithms, abnormal types can be effectively identified and predicted, thereby improving data security and reliability. Through a data index structure optimized by reinforcement learning and cloud-native technology, the data access efficiency and storage optimization strategy are significantly improved. The application of the graph convolutional network model strengthens the in-depth analysis of data relationships and provides a richer perspective for data association. The combination of the predictive model and the intelligent data perception mechanism not only optimizes the data governance strategy but also provides a more accurate and dynamic adjustment for data management.
[0006] Although the above patent can solve the problem of information data management for massive data, the levels of legal affairs data governance vary, and the existing data governance technologies are difficult to guarantee the data quality requirements of different industries, and the data processing speed is difficult to meet the data application standards. Summary of the Invention
[0007] The purpose of the present invention is to provide a multi-modal data governance system and method based on big data, which can protect the entire process of receiving, storing, and securely governing massive legal affairs entity data through big data technology combined with data receiving strategies, hierarchical storage mechanisms, and dynamic security governance mechanisms, so as to ensure the data quality of massive legal affairs entity data in different industries and improve the data processing speed during the data governance process.
[0008] The present invention uses the following technical solutions: A multi-modal data governance system based on big data, including an integrated processing layer, a distributed storage layer, an intelligent governance layer, a multi-modal analysis layer, and a decision control layer; among them, The integrated processing layer is used to collect and normalize massive legal affairs entity data in different formats according to the data receiving strategy, and generate metadata tags; The distributed storage layer is used to distribute and store massive legal affairs entity data according to the metadata tags combined with the access frequency, and establish a multi-source data lake; The intelligent governance layer is used to automatically detect and mark abnormal legal affairs entity data in the multi-source data lake according to the dynamic security governance mechanism, and obtain a multi-modal encrypted data set; The multi-modal analysis layer is used to cross-modal associate different multi-modal encrypted data sets based on a hybrid computing strategy to obtain a multi-modal knowledge graph; The decision control layer is used to dynamically interact with the multi-modal knowledge graph and optimize the dynamic security governance mechanism according to business requirements, and then complete multi-modal data governance.
[0009] Preferably, the data receiving policy judges and parses the request message according to the Internet protocol address: if the Internet protocol address is within the preset private protocol segment, it is determined that the request message is a confidentiality request, and the request message is parsed to obtain the confidentiality level, request line, header line, blank line, and request body; if the Internet protocol address is outside the preset private protocol segment, it is determined that the request message is a public request, and the request message is parsed to obtain the request line, header line, blank line, and request body; the request line is decrypted and decomposed according to the confidentiality level to obtain the request method, resource locator, and protocol version; and the occupied space of the request body is allocated according to the request method according to the preset space configuration threshold to obtain the initial static space; at the same time, after performing a hash operation on the resource locator combined with an N-bit random number, a unique digital identifier is obtained; subsequently, the unique digital identifier is converted into a space code and a serial number, which are respectively assigned to the initial static space and the request body, and then a response message is sent.
[0010] Preferably, the data receiving policy counts the request messages within the preset receiving time interval to obtain the total number of requests, and compares it with the number of internal computing nodes: if the total number of requests is greater than or equal to the number of internal computing nodes and the quantity difference is calculated, edge nodes are enabled according to the quantity difference to cache and batch the request messages; if the total number of requests is less than the number of internal computing nodes, the header line is parsed to obtain the format types of all legal entity data in the request body, and at the same time, the boundaries of the legal entity data of different format types are marked to obtain a format separation sequence list; the initial static space is sorted according to the format separation sequence list to obtain a space link list.
[0011] Preferably, the integrated processing layer encrypts the transmission channel according to the data receiving policy in combination with the confidentiality level to obtain a confidentiality exclusive channel; in the confidentiality exclusive channel, the sliding window is used to decrypt and measure the confidentiality request to obtain the original requests and request quantities of each confidentiality level; the request quantities in each confidentiality exclusive channel are compared with the channel processing threshold: if the request quantity is greater than or equal to the channel processing threshold, the sliding window is dynamically corrected through the slow start algorithm or congestion avoidance algorithm to limit the receiving rate, and at the same time, the packet loss rate of the current confidentiality exclusive channel is detected until timeout retransmission; if the request quantity is less than the channel processing threshold, according to the format separation sequence list combined with the serial number of the unique digital identifier, the request bodies of each original request are normalized and mapped to a unified semantic space according to the format type in combination with the cross-modal embedding algorithm to generate metadata labels.
[0012] Preferably, the distributed storage layer divides the database into several storage units of different capacities according to the spatial link table combined with the spatial encoding of the unique digital identifier, and uses the initial static space for correction to obtain the original storage blocks. At the same time, the access frequencies of all legal entity data are statistically analyzed in combination with the hierarchical storage mechanism, and then the original storage blocks are classified in combination with the metadata tags. The clustering algorithm is used to cluster the legal entity data of each layer in the original storage block according to the data type to obtain the aggregated data stream, and the aggregated data stream is clustered in combination with the metadata tags according to the Internet protocol address to obtain the multi-source data lake.
[0013] Preferably, the hierarchical storage mechanism is as follows: If the access frequency is greater than or equal to the first access threshold, then in real time for the legal entity data of the current batch, the adaptive cache replacement algorithm is used to allocate the transmission and storage to the hot data layer of the original storage block, and at the same time the metadata tags are classified into the first category; If the access frequency is greater than or equal to the second access threshold and less than the first access threshold, then once a week for the legal entity data of the current batch, the time decay heat algorithm is used to allocate the transmission and storage to the warm data layer of the original storage block, and at the same time the metadata tags are classified into the second category; If the access frequency is greater than or equal to the third access threshold and less than the second access threshold, then once a month for the legal entity data of the current batch, the distributed storage algorithm is used to allocate the transmission and storage to the cold data layer of the original storage block, and at the same time the metadata tags are classified into the third category; If the access frequency is less than the third access threshold, then once a year for the legal entity data of the current batch, the distributed columnar sliding window is used to allocate the transmission and storage to the ice data layer of the original storage block, and at the same time the metadata tags are classified into the fourth category.
[0014] Preferably, according to the level of metadata tags, the intelligent governance layer maps the legal entity data of each layer to the feature spaces of different dimensions in combination with the dynamic security governance mechanism, and at the same time encrypts the feature spaces according to the confidentiality level to obtain an encrypted dimension space; uses the data association algorithm to perform association analysis on the legal entity data of different modalities in the encrypted dimension space, and combines the time series to obtain a spatio-temporal association data table; according to the preset anomaly determination threshold and the spatio-temporal association data table, uses the distributed multi-modal quality inspection algorithm to perform quality inspection on all legal entity data in the multi-source data lake, identifies and marks the abnormal legal entity data, and at the same time calls the corresponding data cleaning algorithm to perform fitting correction on the abnormal legal entity data; at the same time, uses the named entity recognition algorithm to measure the sensitivity of all legal entity data in the multi-source data lake, combines the preset sensitivity threshold to obtain the sensitivity level of each legal entity data, and then obtains the data sensitivity level table according to the metadata tags; issues a warning mark for the abnormal legal entity data according to the data sensitivity level table, and inserts the response message generated by the data reception policy; at the same time, according to the adaptive feedback loop and the data sensitivity level table, performs elastic encryption on the legal entity data to obtain a multi-modal encrypted data set.
[0015] Preferably, the operation process of the dynamic security governance mechanism is as follows: A: Automatically adapt and hierarchically manage the multi-source data lake according to the metadata tags according to the industry data standard to obtain several single-modal data processing sets; B: Classify each single-modal data processing set according to data sensitivity, business criticality, and data timeliness to obtain classification elements, and then fill in the metadata tags; C: According to the filled metadata tags, use the tensor decomposition and fusion algorithm to map the single-modal data processing sets of different layers to the full-dimensional feature space; D: Each multi-source data lake generates feature confusion information according to the full-dimensional feature space, and exchanges part of the feature confusion information using the oblivious transfer protocol; E: When each single-modal data processing set is not decrypted, calculate the single-modal feature similarity to obtain a feature similarity matrix; F: Compare the single-modal feature similarity with the similarity threshold: If the single-modal feature similarity is greater than or equal to the similarity threshold, construct a spatio-temporal association data table for the single-modal data processing set according to the space coding and time stamp, and obtain the number of associated entities; If the single-modal feature similarity is less than the similarity threshold, perform anomaly detection and dynamic cleaning on the single-modal data processing set, and then calculate the single-modal feature similarity again until the spatio-temporal association data table is constructed; G: Analyze the sensitivity of legal entity data based on the single-modal data processing set, and combine the occurrence times of legal entity data and the number of associated entities to calculate the sensitivity of each legal entity data, and then complete the elastic encryption of legal entity data.
[0016] Preferably, the multi-modal analysis layer generates a tensor product joint matrix according to the encryption dimension space and the data sensitivity level table; at the same time, according to the access frequency of all legal entity data, measure the global feature density of the multi-modal encrypted data set; then use the data classification algorithm to classify the multi-modal encrypted data set according to the data type and time series to obtain a multi-modal data tuple constructed by several single-modal tuples; according to the sensitivity of legal entity data and the confidentiality level, use the dynamic adaptive weight allocation algorithm to generate a spatio-temporal weight matrix for the multi-modal data tuple; then use the segment recurrence mechanism to perform temporal sequence modeling on the single-modal tuple and the tensor product joint matrix to extract local sensitive hash features; at the same time, bilinear pooling performs inter-modal feature interaction on several single-modal tuples in combination with the global feature density, and uses the generative adversarial network to simulate abnormal legal entity data in combination with the metadata label to obtain a cross-modal feature matrix; at the same time, according to the local sensitive hash features, the spatio-temporal weight matrix and the cross-modal feature matrix, map the multi-modal data tuple to a low-dimensional shared data space, and combine the cross-modal alignment loss function to transform it into a four-dimensional spatio-temporal tensor to obtain a dynamically evolving multi-modal knowledge graph.
[0017] A multi-modal data governance method based on big data, applied to a multi-modal data governance system, includes the following steps: S1: The integrated processing layer collects and normalizes massive legal entity data in different formats, and generates metadata labels according to the data reception strategy; S2: The distributed storage layer distributes and stores massive legal entity data, and at the same time completes the construction of a multi-source data lake according to the metadata label and the access frequency; S3: The intelligent governance layer automatically detects and marks abnormal legal entity data in the multi-source data lake, and obtains a multi-modal encrypted data set according to the dynamic security governance mechanism; S4: The multi-modal analysis layer performs cross-modal association on different multi-modal encrypted data sets based on the hybrid computing strategy to obtain a multi-modal knowledge graph; S5: The decision control layer dynamically interacts with the multi-modal knowledge graph and optimizes the dynamic security governance mechanism according to the business requirements, and then completes the multi-modal data governance.
[0018] The present invention uses big data technology combined with a data reception strategy, a hierarchical storage mechanism and a dynamic security governance mechanism to perform full-process protection on the reception, storage and security governance of massive legal entity data, ensuring the data quality of massive legal entity data in different industries and improving the data processing speed in the data governance process. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the related art, the following will briefly introduce the drawings required for the description of the embodiments or the related art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0020] Figure 1 It is a schematic diagram of a multi-modal data governance system; Figure 2 It is a flowchart of a dynamic security governance mechanism; Figure 3 It is a flowchart of a multi-modal data governance method. Detailed implementation manners
[0021] The following will describe the present invention in detail in conjunction with the drawings and embodiments: As Figures 1 to 3 shown, a multi-modal data governance system based on big data according to the present invention includes an integrated processing layer, a distributed storage layer, an intelligent governance layer, a multi-modal analysis layer, and a decision and control layer; wherein, The integrated processing layer is used to collect and normalize massive legal entity data in different formats according to the data reception strategy, and generate metadata tags; The distributed storage layer is used to distribute and store massive legal entity data according to the metadata tags combined with the access frequency, and establish a multi-source data lake; In this embodiment, the legal entity data includes legal documents (indictments, judgments), communication records (calls, text messages, social media chat records), transaction flows (bank transfers, virtual currency transactions), audio and video files (surveillance videos, audio evidence); Law enforcement records: law enforcement personnel information, law enforcement process records (such as on-site law enforcement videos, transcripts); Procedural data: key time nodes for case handling (such as detention time, court session date), approval processes (such as arrest warrant issuance, property seizure), execution results (sentence, fine amount); Associated network data: kinship, fund transfers, communication links, etc. among involved persons; Spatio-temporal trajectory data: geographical location information (such as base station location, IP address) and timestamp records of involved persons or devices; Government affairs data: public security household registration information, court public judgment documents, market supervision data, etc.; Industry feature data: such as suspicious transaction reports in the financial field, abnormal number libraries in the communication field.
[0022] The intelligent governance layer is used to automatically detect and mark abnormal legal entity data in the multi-source data lake according to the dynamic security governance mechanism, and obtain a multi-modal encrypted data set; The multi-modal analysis layer is used to cross-modally associate different multi-modal encrypted data sets based on a hybrid computing strategy to obtain a multi-modal knowledge graph; The decision-making and control layer is used to dynamically interact with the multi-modal knowledge graph and optimize the dynamic security governance mechanism according to business requirements, thereby completing multi-modal data governance.
[0023] In this embodiment, the business requirements include data quality improvement requirements, compliance and security control requirements, business process optimization requirements, decision-making support requirements, innovation-driven requirements, and unified standard construction requirements; The data quality improvement requirements include data accuracy and data integrity; the compliance and security control requirements include data classification and grading, audit traceability mechanism, and enhanced privacy protection; the business process optimization requirements include breaking data islands and improving process efficiency; the decision-making support requirements include providing trustworthy data and dynamic business insights; the innovation-driven requirements include data asset operation and business model innovation; the unified standard construction requirements include term and index specifications and metadata management systems.
[0024] In the present invention, the data reception strategy judges and parses the request message according to the Internet protocol address: if the Internet protocol address is within the preset private protocol segment, it is determined that the request message is a confidential request, and the request message is parsed to obtain the confidentiality level, request line, header line, blank line, and request body; if the Internet protocol address is outside the preset private protocol segment, it is determined that the request message is a public request, and the request message is parsed to obtain the request line, header line, blank line, and request body; the request line is decrypted and decomposed according to the confidentiality level to obtain the request method, resource locator, and protocol version; and the occupied space of the request body is allocated according to the request method according to the preset space configuration threshold to obtain the initial static space; at the same time, after performing a hash operation on the resource locator combined with an N-bit random number, a unique digital identifier is obtained; then the unique digital identifier is converted into a space code and a serial number, which are respectively assigned to the initial static space and the request body, and then a response message is sent; then, the request messages within the preset reception time interval are counted to obtain the total number of requests, and compared with the number of internal computing nodes: if the total number of requests is greater than or equal to the number of internal computing nodes and the quantity difference is calculated, edge nodes are enabled according to the quantity difference to cache and batch the request messages; if the total number of requests is less than the number of internal computing nodes, the header line is parsed to obtain the format types of all legal entity data in the request body, and at the same time, the boundaries of the legal entity data of different format types are marked to obtain a format separation sequence list; the initial static space is sorted according to the format separation sequence list to obtain a space link list.
[0025] In this embodiment, the confidentiality levels include top secret, secret, and confidential; the format types include JSON, XML, Protobuf, etc.; the Internet protocols include more than 20 protocols such as HTTP / HTTPS, MQTT, and FTP.
[0026] In the present invention, the integrated processing layer encrypts the transmission channel according to the data reception strategy and in combination with the confidentiality level to obtain a confidentiality-exclusive channel; in the confidentiality-exclusive channel, the confidentiality request is decrypted and measured by using a sliding window to obtain the original requests and the request quantities of each confidentiality level; the request quantities in each confidentiality-exclusive channel are compared with the channel processing threshold: if the request quantity is greater than or equal to the channel processing threshold, the sliding window is dynamically corrected by using a slow start algorithm or a congestion avoidance algorithm to limit the reception rate, and at the same time, the packet loss rate of the current confidentiality-exclusive channel is detected until timeout retransmission; if the request quantity is less than the channel processing threshold, according to the format separation sequence list and in combination with the serial number of the unique digital identifier, the request body of each original request is normalized and mapped to a unified semantic space according to the format type by using a cross-modal embedding algorithm to generate metadata tags.
[0027] In the present invention, the distributed storage layer divides the database into several storage units of different capacities according to the spatial link table and in combination with the spatial coding of the unique digital identifier, and corrects it by using the initial static space to obtain the original storage blocks; at the same time, the access frequencies of all legal entity data are statistically analyzed in combination with the hierarchical storage mechanism, and then the original storage blocks are classified in combination with the metadata tags; the legal entity data of each layer in the original storage blocks are clustered according to the data type by using a clustering algorithm to obtain an aggregated data stream; and the aggregated data stream is clustered in combination with the metadata tags according to the Internet protocol address to obtain a multi-source data lake.
[0028] In the present invention, the hierarchical storage mechanism is as follows: If the access frequency is greater than or equal to the first access threshold, the legal entity data of the current batch are allocated for transmission and storage to the hot data layer of the original storage block in real time by using an adaptive cache replacement algorithm, and at the same time, the metadata tags are classified into the first category. In this embodiment, the hot data layer is a set of currently active data that is frequently accessed or modified. If the access frequency is greater than or equal to the second access threshold and less than the first access threshold, the legal entity data of the current batch are allocated for transmission and storage to the warm data layer of the original storage block on a weekly basis by using a time decay heat algorithm, and at the same time, the metadata tags are classified into the second category. In this embodiment, the warm data layer is non-instantaneous status and behavior data. If the access frequency is greater than or equal to the third access threshold and less than the second access threshold, the legal entity data of the current batch is allocated for transmission and storage to the cold data layer of the original storage block on a monthly basis using a distributed storage algorithm, and at the same time, the metadata label is classified into the third level; In this embodiment, the cold data layer refers to data that is not frequently accessed offline and is used for disaster recovery backup or must be retained for a certain period of time due to legal requirements, such as enterprise backup data, business and operation log data, call detail records and statistical data; If the access frequency is less than the third access threshold, the legal entity data of the current batch is allocated for transmission and storage to the ice data layer of the original storage block on an annual basis using a distributed columnar sliding window, and at the same time, the metadata label is classified into the fourth level.
[0029] In this embodiment, the ice data layer is data that is rarely accessed (such as a few times a year) and is used for compliance archiving.
[0030] In the present invention, the intelligent governance layer maps the legal entity data of each layer to different dimensional feature spaces according to the level of the metadata label, and combines the dynamic security governance mechanism to encrypt the feature space according to the confidentiality level to obtain an encrypted dimensional space; uses the data association algorithm to perform association analysis on the legal entity data of different modalities in the encrypted dimensional space, and combines the time series to obtain a spatio-temporal association data table; according to the preset anomaly determination threshold and the spatio-temporal association data table, uses the distributed multi-modal quality inspection algorithm to perform quality inspection on all the legal entity data in the multi-source data lake, identify and mark the abnormal legal entity data (as shown in Table 1), and at the same time call the corresponding data cleaning algorithm to perform fitting correction on the abnormal legal entity data (as shown in Table 2); at the same time, uses the named entity recognition algorithm to measure the sensitivity of all the legal entity data in the multi-source data lake, combines the preset sensitivity threshold to obtain the sensitivity level of each legal entity data, and then obtains the data sensitivity level table according to the metadata label; issues a warning mark for the abnormal legal entity data according to the data sensitivity level table and inserts the response message generated by the data reception policy; at the same time, according to the adaptive feedback loop and the data sensitivity level table, performs elastic encryption on the legal entity data to obtain a multi-modal encrypted data set.
[0031] In this embodiment, Table 1 Abnormal Legal Entity Data Type Table Model type Detection object Algorithm implementation Statistical model Abnormality of numerical fields Improved 3σ criterion (dynamic baseline adjustment) Graph neural network Abnormality of association relationship GraphSAGE + Attention mechanism Multimodal autoencoder Abnormality of cross-modal consistency VAE-GAN hybrid architecture Table 2 Mapping Table of Abnormal Legal Entity Data Types and Data Cleaning Algorithms: Abnormal code Feature description Cleaning algorithm Repair verification mechanism E001 Outliers of numerical fields Local interpolation method based on KNN Residual analysis + confidence interval test E002 Cross-modal semantic conflict Knowledge graph-guided semantic alignment Ontology reasoning consistency verification E003 Spatiotemporal association break Trajectory completion algorithm (such as ST-Matching) Spatiotemporal topological continuity test In this embodiment, the data processing process of the intelligent governance layer is also adjusted through a load balancing algorithm: , where is the utility function of the th node, is the resource allocation vector, is the equilibrium coefficient.
[0032] In the present invention, the operation process of the dynamic security governance mechanism is as follows: A: Automatically adapt and hierarchically govern the multi-source data lake according to the metadata tags according to the industry data standards to obtain several layers of single-modal data processing sets; In this embodiment, the multi-source data lake includes several types of heterogeneous data lakes (such as government affairs, finance, medical care, IoT, social media); The industry data standard is a unified rule system formulated for data definition, storage, exchange, and application within a specific domain. Its core lies in enhancing the availability, security, and collaboration efficiency of data through standardization and structuring.
[0033] B: Classify each layer of single-modal data processing sets according to data sensitivity, business criticality, and data timeliness to obtain classification elements, and then fill in the metadata tags; In this embodiment, data sensitivity is the degree to which the privacy information or confidential information contained in the data needs to be protected, reflecting the legal, financial, or reputational risk level that may be caused by data leakage.
[0034] Business criticality is the support intensity of the data for the continuity of the organization's core business, revenue creation, or strategic decision-making, measuring the impact of data loss / damage on the enterprise.
[0035] Data timeliness is the effective value of the data within a specific time window, reflecting the rate at which it decays over time and its scenario adaptability.
[0036] C: According to the filled metadata tags, use the tensor decomposition fusion algorithm to map different layers of single-modal data processing sets to the full-dimensional feature space : , where is the sum of multiple vectors, is the total number of modalities, is the original feature tensor of the th class of modality, is the Hadamard operation, is the modality adaptation matrix; In this embodiment, the working principle of the tensor decomposition fusion algorithm is: Unify the heterogeneous data from different fields (such as images, texts, sensors) into a high-order tensor structure to form a cross-modal joint representation; Decompose the N-dimensional tensor into a chain product of rank-R factor matrices, which is applicable to sparse data processing. By multiplying the core tensor and the factor matrices, the main features are retained, making it suitable for high-dimensional data dimensionality reduction; The decomposed low-dimensional factor matrices reveal cross-modal shared features (such as visual-semantic associations), and the stability of the decomposition is ensured by optimizing the constraint with the F-norm; Based on the shared factor matrices of the decomposition results, construct cross-modal conversion functions (such as text description → image feature generation).
[0037] D: Each multi-source data lake generates feature obfuscation information according to the full-dimensional feature space, and exchanges part of the feature obfuscation information using the oblivious transfer protocol; E: Calculate the single-modal feature similarity when the single-modal data processing set at each layer is not decrypted : , represents the similarity function, represents the th layer of the single-modal data processing set, and then obtain the feature similarity matrix; F: Compare the single-modal feature similarity with the similarity threshold: If the single-modal feature similarity is greater than or equal to the similarity threshold, construct a spatio-temporal association data table for the single-modal data processing set according to the spatial coding combined with the timestamp, and obtain the number of associated entities; If the single-modal feature similarity is less than the similarity threshold, perform anomaly detection and dynamic cleaning on the single-modal data processing set, and then calculate the single-modal feature similarity again until the spatio-temporal association data table is constructed; G: Analyze the sensitivity of the legal entity data according to the single-modal data processing set, and at the same time combine the number of occurrences of the legal entity data and the number of associated entities , calculate the sensitivity of each legal entity data : , is the first adjustment factor, is the weight coefficient, is the second adjustment factor, and then complete the elastic encryption of the legal entity data.
[0038] In the present invention, the multi-modal analysis layer generates a tensor product joint matrix according to the encrypted dimensional space in combination with the data sensitivity level table; at the same time, according to the access frequencies of all legal entity data, the global feature density of the multi-modal encrypted data set is measured; then, the multi-modal encrypted data set is classified according to the data type in combination with the time series by using a data classification algorithm to obtain a multi-modal data tuple constructed by a number of single-modal tuples; according to the sensitivity of the legal entity data in combination with the confidentiality level, a spatio-temporal weight matrix is generated for the multi-modal data tuple by using a dynamic adaptive weight allocation algorithm; then, a fragment recursion mechanism is used to perform time series modeling on the single-modal tuple in combination with the tensor product joint matrix to extract local sensitive hash features; at the same time, bilinear pooling performs inter-modal feature interaction on a number of single-modal tuples in combination with the global feature density, and an adversarial generation network is used to simulate abnormal legal entity data in combination with the metadata label to obtain a cross-modal feature matrix; at the same time, according to the local sensitive hash features and the spatio-temporal weight matrix in combination with the cross-modal feature matrix, the multi-modal data tuple is mapped to a low-dimensional shared data space, and is transformed into a four-dimensional spatio-temporal tensor in combination with a cross-modal alignment loss function to obtain a dynamically evolving multi-modal knowledge graph.
[0039] The present invention further includes a multi-modal data governance method based on big data, which is applied to a multi-modal data governance system and includes the following steps: S1: The integrated processing layer collects and normalizes massive legal entity data in different formats, and generates metadata labels according to the data reception strategy; S2: The distributed storage layer stores the massive legal entity data in a distributed manner, and at the same time completes the construction of a multi-source data lake according to the metadata labels in combination with the access frequencies; S3: The intelligent governance layer automatically detects and marks abnormal legal entity data in the multi-source data lake, and obtains a multi-modal encrypted data set according to the dynamic security governance mechanism; S4: The multi-modal analysis layer performs cross-modal association on different multi-modal encrypted data sets based on a hybrid computing strategy to obtain a multi-modal knowledge graph; S5: The decision control layer dynamically interacts with the multi-modal knowledge graph and optimizes the dynamic security governance mechanism according to the business requirements, so as to complete the multi-modal data governance.
[0040] Embodiment: The integrated processing layer judges and parses the request message according to the Internet protocol address: if the Internet protocol address is within the preset private protocol segment, it judges that the request message is a confidential request, parses the request message, and obtains the confidentiality level, request line, header line, blank line, and request body; if the Internet protocol address is outside the preset private protocol segment, it judges that the request message is a public request, parses the request message, and obtains the request line, header line, blank line, and request body; decrypts and decomposes the request line according to the confidentiality level to obtain the request method, resource locator, and protocol version; and allocates the occupied space of the request body according to the request method according to the preset space configuration threshold to obtain the initial static space; at the same time, after performing a hash operation on the resource locator combined with an N-bit random number, obtains a unique digital identifier; then converts the unique digital identifier into a space code and a serial number, assigns them to the initial static space and the request body respectively, and then sends a response message; Then, the request messages within the preset reception time interval are counted to obtain the total number of requests, and compared with the number of internal operation nodes: if the total number of requests is greater than or equal to the number of internal operation nodes and the quantity difference is calculated, edge nodes are enabled according to the quantity difference to cache and batch the request messages; if the total number of requests is less than the number of internal operation nodes, the header line is parsed to obtain the format types of all legal entity data in the request body, and at the same time, the boundaries of the legal entity data of different format types are marked to obtain a format separation sequence list; the initial static space is sorted according to the format separation sequence list to obtain a space link list; And encrypts the transmission channel in combination with the confidentiality level to obtain a confidential exclusive channel; in the confidential exclusive channel, uses a sliding window to decrypt and measure the confidential request to obtain the original requests and request quantities of each confidentiality level; compares the request quantity in each confidential exclusive channel with the channel processing threshold: if the request quantity is greater than or equal to the channel processing threshold, dynamically corrects the sliding window through a slow start algorithm or a congestion avoidance algorithm to limit the reception rate, and at the same time detects the packet loss rate of the current confidential exclusive channel until timeout retransmission; if the request quantity is less than the channel processing threshold, according to the format separation sequence list combined with the serial number of the unique digital identifier, the request body of each original request is normalized and mapped to a unified semantic space according to the format type in combination with a cross-modal embedding algorithm to generate a metadata label.
[0041] The distributed storage layer divides the database into several storage units of different capacities according to the space link list combined with the space code of the unique digital identifier, and uses the initial static space for correction to obtain the original storage block; at the same time, statistically analyzes the access frequencies of all legal entity data in combination with a hierarchical storage mechanism: if the access frequency is greater than or equal to the first access threshold, in real time, for the current batch of legal entity data, uses an adaptive cache replacement algorithm to allocate and transmit and store the hot data layer of the original storage block, and at the same time classifies the metadata label into the first category; If the access frequency is greater than or equal to the second access threshold and less than the first access threshold, the legal entity data of the current batch is distributed, transmitted, and stored weekly in the warm data layer of the original storage block using the time decay heat algorithm. Meanwhile, the metadata label is classified into the second category. If the access frequency is greater than or equal to the third access threshold and less than the second access threshold, the legal entity data of the current batch is distributed, transmitted, and stored monthly in the cold data layer of the original storage block using the distributed storage algorithm. Meanwhile, the metadata label is classified into the third category. If the access frequency is less than the third access threshold, the legal entity data of the current batch is distributed, transmitted, and stored annually in the ice data layer of the original storage block using the distributed columnar sliding window. Meanwhile, the metadata label is classified into the fourth category. The original storage blocks are classified in combination with the metadata labels. The legal entity data of each layer in the original storage block is clustered using the clustering algorithm according to the data type to obtain the aggregated data stream. And the aggregated data stream is clustered in combination with the metadata labels according to the Internet protocol address to obtain the multi-source data lake.
[0042] The intelligent governance layer maps the legal entity data of each layer to the feature space of different dimensions in combination with the dynamic security governance mechanism according to the category of the metadata label, and encrypts the feature space in combination with the confidentiality level to obtain the encrypted dimension space. The data association algorithm is used to perform association analysis on the legal entity data of different modalities in the encrypted dimension space, and the spatio-temporal association data table is obtained in combination with the time series. According to the preset anomaly determination threshold and the spatio-temporal association data table, the distributed multi-modal quality inspection algorithm is used to perform quality inspection on all the legal entity data in the multi-source data lake, identify and mark the abnormal legal entity data, and simultaneously call the corresponding data cleaning algorithm to perform fitting correction on the abnormal legal entity data. Meanwhile, the named entity recognition algorithm is used to measure the sensitivity of all the legal entity data in the multi-source data lake, and in combination with the preset sensitivity threshold, the sensitivity level of each legal entity data is obtained, and then the data sensitivity level table is obtained according to the metadata label. The abnormal legal entity data is warned and marked according to the data sensitivity level table, and the response message generated by the data reception policy is inserted. Meanwhile, according to the adaptive feedback loop and in combination with the data sensitivity level table, the legal entity data is elastically encrypted to obtain the multi-modal encrypted data set.
[0043] The multi-modal analysis layer generates a tensor product joint matrix according to the encrypted dimension space and the data sensitivity level table; at the same time, according to the access frequencies of all legal entity data, it measures the global feature density of the multi-modal encrypted data set; then it classifies the multi-modal encrypted data set according to the data type and time series by using a data classification algorithm, and obtains multi-modal data tuples constructed by several single-modal tuples; according to the sensitivity of the legal entity data and the security level, it uses a dynamic adaptive weight allocation algorithm to generate a spatio-temporal weight matrix for the multi-modal data tuples; then it uses a fragment recursive mechanism to perform temporal sequence modeling on the single-modal tuples combined with the tensor product joint matrix to extract local sensitive hash features; at the same time, bilinear pooling performs inter-modal feature interaction on several single-modal tuples combined with the global feature density, and uses a generative adversarial network combined with metadata labels to simulate abnormal legal entity data to obtain a cross-modal feature matrix; at the same time, according to the local sensitive hash features, the spatio-temporal weight matrix and the cross-modal feature matrix, it maps the multi-modal data tuples to a low-dimensional shared data space, and combines the cross-modal alignment loss function to transform them into a four-dimensional spatio-temporal tensor to obtain a dynamically evolving multi-modal knowledge graph.
[0044] Finally, the decision and control layer dynamically interacts with the multi-modal knowledge graph and optimizes the dynamic security governance mechanism, thereby completing multi-modal data governance.
Claims
1. A multi-modal data governance system based on big data, characterized in that: including an integrated processing layer, which is used to collect and normalize a large amount of legal entity data in different formats according to a data reception policy, and generate metadata tags; a distributed storage layer, which is used to distribute and store a large amount of legal entity data according to metadata tags combined with access frequencies, and establish a multi-source data lake; an intelligent governance layer, which is used to automatically detect and mark abnormal legal entity data in the multi-source data lake according to a dynamic security governance mechanism, and obtain a multi-modal encrypted data set; a multi-modal analysis layer, which is used to cross-modally associate different multi-modal encrypted data sets based on a hybrid computing strategy, and obtain a multi-modal knowledge graph; a decision control layer, which is used to dynamically interact with the multi-modal knowledge graph and optimize the dynamic security governance mechanism according to business requirements, so as to complete multi-modal data governance.
2. The multimodal data governance system based on big data according to claim 1, characterized in that: The data reception policy judges and analyzes the request message according to the Internet protocol address: if the Internet protocol address is within a preset private protocol segment, it is judged that the request message is a confidential request, and the request message is analyzed to obtain the confidentiality level, request line, header line, blank line, and request body; if the Internet protocol address is outside the preset private protocol segment, it is judged that the request message is a public request, and the request message is analyzed to obtain the request line, header line, blank line, and request body; the request line is decrypted and decomposed according to the confidentiality level to obtain the request method, resource locator, and protocol version; and the occupied space of the request body is allocated according to the request method according to a preset space configuration threshold to obtain an initial static space; at the same time, after performing a hash operation on the resource locator combined with an N-bit random number, a unique digital identifier is obtained; then the unique digital identifier is converted into a space code and a serial number, which are respectively assigned to the initial static space and the request body, and then a response message is sent.
3. The multimodal data governance system based on big data according to claim 1, characterized in that: The data reception policy counts the request messages within a preset reception time interval to obtain the total number of requests, and compares it with the number of internal computing nodes: if the total number of requests is greater than or equal to the number of internal computing nodes and the quantity difference is calculated, edge nodes are enabled according to the quantity difference to cache and batch the request messages; if the total number of requests is less than the number of internal computing nodes, the header line is analyzed to obtain the format types of all legal entity data in the request body, and at the same time, the boundaries of legal entity data of different format types are marked to obtain a format separation sequence list; the initial static space is sorted according to the format separation sequence list to obtain a space link list.
4. The multimodal data governance system based on big data according to claim 1, characterized in that: The integrated processing layer encrypts the transmission channel according to the data reception policy combined with the confidentiality level to obtain a confidential exclusive channel; decrypts and measures the confidential request in the confidential exclusive channel to obtain the original requests and request quantities of each confidentiality level; compares the request quantity in each confidential exclusive channel with the channel processing threshold: if the request quantity is greater than or equal to the channel processing threshold, the sliding window is dynamically corrected to limit the reception rate, and at the same time, the packet loss rate of the current confidential exclusive channel is detected until timeout retransmission; if the request quantity is less than the channel processing threshold, according to the format separation sequence list combined with the serial number of the unique digital identifier, the request body of each original request is normalized and mapped to a unified semantic space according to the format type to generate metadata tags.
5. The multimodal data governance system based on big data according to claim 1, wherein: The distributed storage layer divides the database into several storage units with different capacities according to the spatial link table and the spatial encoding with unique digital identifiers, and uses the initial static space for correction to obtain the original storage blocks. At the same time, the access frequencies of all legal entity data are statistically analyzed in combination with the hierarchical storage mechanism, and then the original storage blocks are classified in combination with metadata tags. Cluster the legal entity data of each layer in the original storage block according to the data type to obtain an aggregated data stream. And cluster the aggregated data stream in combination with metadata tags according to the Internet protocol address to obtain a multi-source data lake.
6. The multimodal data governance system based on big data according to claim 1, characterized in that: The hierarchical storage mechanism is as follows: If the access frequency is greater than or equal to the first access threshold, the legal entity data of the current batch is allocated, transmitted and stored in real time to the hot data layer of the original storage block, and the metadata tag is classified into the first category at the same time. If the access frequency is greater than or equal to the second access threshold and less than the first access threshold, the legal entity data of the current batch is allocated, transmitted and stored to the warm data layer of the original storage block on a weekly basis, and the metadata tag is classified into the second category at the same time. If the access frequency is greater than or equal to the third access threshold and less than the second access threshold, the legal entity data of the current batch is allocated, transmitted and stored to the cold data layer of the original storage block on a monthly basis, and the metadata tag is classified into the third category at the same time. If the access frequency is less than the third access threshold, the legal entity data of the current batch is allocated, transmitted and stored to the ice data layer of the original storage block on an annual basis, and the metadata tag is classified into the fourth category at the same time.
7. The multimodal data governance system based on big data according to claim 1, wherein: The intelligent governance layer maps the legal entity data of each layer to different-dimensional feature spaces according to the levels of metadata tags in combination with the dynamic security governance mechanism, and encrypts the feature spaces in combination with the confidentiality level to obtain an encrypted dimensional space. In the encrypted dimensional space, correlation analysis is performed on legal entity data of different modalities, and a spatio-temporal correlation data table is obtained in combination with the time series. According to the preset anomaly determination threshold and the spatio-temporal correlation data table, quality detection is performed on all legal entity data in the multi-source data lake, abnormal legal entity data is identified and marked, and the abnormal legal entity data is fitted and corrected at the same time. At the same time, sensitivity measurement is performed on all legal entity data in the multi-source data lake, and in combination with the preset sensitivity threshold, the sensitivity level of each legal entity data is obtained, and then according to the metadata tag, a data sensitivity level table is obtained. Warn and mark the abnormal legal entity data according to the data sensitivity level table, and insert the response message generated by the data reception policy. At the same time, in combination with the data sensitivity level table, elastic encryption is performed on the legal entity data to obtain a multi-modal encrypted data set.
8. The multimodal data governance system based on big data according to claim 1, characterized in that: The operation process of the dynamic security governance mechanism is as follows: A: Automatically adapt and hierarchically manage the multi-source data lake according to the metadata tag according to the industry data standard to obtain several single-modal data processing sets at each layer. B: Classify each layer of single-modal data processing sets according to data sensitivity, business criticality and data timeliness to obtain classification elements, and then fill in the metadata tags. C: According to the filled metadata tags, use the tensor decomposition and fusion algorithm to map different layers of single-modal data processing sets to the full-dimensional feature space. D: Each multi-source data lake generates feature obfuscation information based on the full-dimensional feature space and exchanges some of the feature obfuscation information using the oblivious transfer protocol; E: When each layer of the single-modal data processing set is not decrypted, calculate the single-modal feature similarity, and then obtain the feature similarity matrix; F: Compare the single-modal feature similarity with the similarity threshold: If the single-modal feature similarity is greater than or equal to the similarity threshold, then construct a spatio-temporal association data table for the single-modal data processing set according to the space encoding combined with the time stamp, and obtain the number of associated entities; If the single-modal feature similarity is less than the similarity threshold, perform anomaly detection and dynamic cleaning on the single-modal data processing set, and then calculate the single-modal feature similarity again until the construction of the spatio-temporal association data table is completed; G: Analyze the sensitivity of the legal entity data according to the single-modal data processing set, and at the same time combine the number of occurrences of the legal entity data and the number of associated entities to calculate the sensitivity of each legal entity data, and then complete the elastic encryption of the legal entity data.
9. The multimodal data governance system based on big data according to claim 1, wherein: The multi-modal analysis layer generates a tensor product joint matrix according to the encrypted dimension space combined with the data sensitivity level table; at the same time, according to the access frequencies of all legal entity data, measure the global feature density of the multi-modal encrypted data set; Then use the data classification algorithm to classify the multi-modal encrypted data set according to the data type combined with the time series to obtain a multi-modal data tuple constructed by several single-mode tuples; generate a spatio-temporal weight matrix for the multi-modal data tuple according to the sensitivity of the legal entity data combined with the security level; then perform time series modeling on the single-mode tuple combined with the tensor product joint matrix to extract local sensitive hash features; at the same time, for several single-mode tuples, perform inter-modal feature interaction combined with the global feature density, and simulate abnormal legal entity data combined with the metadata label to obtain a cross-modal feature matrix; at the same time, according to the local sensitive hash features and the spatio-temporal weight matrix combined with the cross-modal feature matrix, map the multi-modal data tuple to a low-dimensional shared data space and transform it into a four-dimensional spatio-temporal tensor to obtain a dynamically evolving multi-modal knowledge graph.
10. A multi-modal data governance method based on big data, applied to the multi-modal data governance system described in any one of claims 1 to 9, characterized in that: Including the following steps: S1: The integrated processing layer collects and normalizes massive legal entity data in different formats and generates metadata labels according to the data reception strategy; S2: The distributed storage layer distributes and stores massive legal entity data, and at the same time completes the construction of the multi-source data lake according to the metadata label combined with the access frequency; S3: The intelligent governance layer automatically detects and marks the abnormal legal entity data in the multi-source data lake and obtains the multi-modal encrypted data set according to the dynamic security governance mechanism; S4: The multi-modal analysis layer performs cross-modal association on different multi-modal encrypted data sets based on the hybrid computing strategy to obtain the multi-modal knowledge graph; S5: The decision control layer dynamically interacts with the multi-modal knowledge graph and optimizes the dynamic security governance mechanism according to the business requirements, and then completes the multi-modal data governance.
Citation Information
Patent Citations
College chemical laboratory pollution treatment system and method based on big data analysis
CN110619011A
Information data management method and system based on big data
CN117785858A