A data governance method, device and computer equipment based on big data

By performing real-time preprocessing and multi-parallel subsystem processing on the received data stream, combined with dynamic caching and encryption shielding of the data warehouse, the problems of slow response, poor compatibility and insufficient security in big data processing are solved, and fast, stable and secure data governance is achieved.

CN115705321BActive Publication Date: 2026-02-13JINZHONG HLDG GRP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110927693.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-10
Publication Date
2026-02-13
Estimated Expiration
2041-08-10

AI Technical Summary

Technical Problem

Existing technologies suffer from slow response and low timeliness in big data processing, fragmented information flow, poor data consistency, weak compatibility, and lack of encryption protection in database storage.

Method used

By receiving data streams in real time, segmenting them into character segments for preprocessing, and utilizing multiple parallel partitioning subsystems for processing and transformation, the system indexes and retrieves historical records and parses evaluation information. Combined with the dynamic caching and encryption shielding units of the data warehouse, it achieves rapid data governance.

Benefits of technology

It achieves rapid response, high timeliness, excellent compatibility, and data privacy protection, improving the stability and security of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115705321B_ABST
    Figure CN115705321B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of data governance method, device and computer equipment based on big data, the data governance method is applied to the data stream associated with the exchange application layer, it includes the following steps: obtaining data stream from application layer, data stream is segmented / translated into specified data length character segment signal;Character segment signal is preprocessed;Multiple partition parallel subsystems process and transform character segment signal, obtain its parameter attribute, according to parameter attribute, index call history record and compare, mine and extract information, and obtain evaluation information by parsing;Update storage character segment signal, dynamically cache parameter attribute and evaluation information, and feedback evaluation information to application layer.The present application is fast in response, time effectiveness is strong, and data compatibility, consistency is good.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a data management method and device based on big data and a computer device. BACKGROUND

[0002] With the development of Internet technology and scale, more and more data systems are connected to each other, which supports user activities through data processing technology. Data processing is the processing and arrangement of collected data to form a suitable data analysis style, which is an essential stage before data analysis. The basic purpose of data processing is to extract and deduce data that is valuable and meaningful to solve problems from a large amount of chaotic data.

[0003] The application layer communication client related information, the processing equipment collects, cleans and arranges the data, analyzes and feeds back useful information to solve specific problems, and reliably reasons the current conditions and future events. In the prior art, such as patent application No. "201910985525.8" discloses a "big data processing flow execution plan generation method", which ensures the unity of the big data processing process and provides the contrast correlation in the big data processing process. The application scale is considerable, but its response is slow, the timeliness is low, the information flow is discrete, and the data consistency is poor. For example, patent application No. "201811617520.1" discloses a "metadata management method, device and computer readable medium", which manages metadata including business system metadata, data warehouse metadata and data application metadata in the process links from data generation to data processing to data result application, meets the daily use needs of the audience group for metadata, and ensures that data application data is correctly, efficiently and conveniently reused, but its compatibility is weak, the carrying load is concentrated, the database storage lacks encryption protection, and the reliability is poor. SUMMARY

[0004] To solve the problems in the background art, the present application provides a data management method and device based on big data and a computer device, which has fast response, strong timeliness, good data compatibility and consistency.

[0005] The present application provides the following technical solutions:

[0006] In a first aspect, the present application provides a data management method based on big data, characterized in that the data management method is applied to the data flow associated with the application layer, and the method comprises the following steps:

[0007] Step a: Real-time receiving data flow from the application layer, and dividing / interpreting the data flow into character segment signals of a specified data length;

[0008] Step b: pre-processing operation on the character segment signal;

[0009] Step c: multiple partitioned parallel subsystems process and transform the character segment signal to obtain the parameter attribute of at least one word segment, index and call historical records according to the parameter attribute, compare, mine and extract information, and parse to obtain evaluation information;

[0010] Step d: update and store the character segment signal, dynamically cache the parameter attribute and the evaluation information, and feed back the evaluation information to the application layer.

[0011] Preferably, the step b comprises the following steps:

[0012] Screening the repeated character segment signal, the missing character segment signal and the character segment signal containing sensitive information in the character segment signal and performing multi-path thread classification processing;

[0013] Filtering and removing the repeated character segment signal, correcting and supplementing the missing character segment signal, and performing deformation marking processing on the character segment signal containing sensitive information, and outputting the classified character segment signal to the subsystem.

[0014] Preferably, the step c comprises the following steps:

[0015] Grouping and transforming the parallel partitioned character segment signal according to the information transmission chain direction in the subsystem;

[0016] Displaying different interval subsets in the partitioned character segment signal, arranging the same interval subsets in the character segment signal to the same group for processing, and obtaining the parameter attribute of at least one word segment from the same interval subset;

[0017] According to the parameter attribute, the static data of the historical data information is indexed and called, compared and statistically analyzed, and the evaluation information is parsed.

[0018] Preferably, the data governance method independently collects and samples the classification processing results of the multi-path thread or the associated information of the subsystem at regular intervals.

[0019] In the second aspect, the present application provides a data governance device based on big data, comprising:

[0020] The application layer transmits the data stream related to the user data in real time and receives the evaluation information;

[0021] A server for interfacing with a data stream associated with an application layer, the server comprising a receiving module, a sending module and a preprocessing system, the receiving module configured to receive the data stream from the application layer, the sending module configured to send evaluation information to the application layer, and the preprocessing system configured to preprocess the data stream into character segment signals of a specified length;

[0022] A processor configured to process the character segment signals to obtain parameter attributes of at least one word segment, the processor comprising a plurality of sub-systems in parallel, and a merge sorting module connected to an output end of an information transmission chain of the plurality of sub-systems, the merge sorting module configured to sort the character segment signals;

[0023] A data warehouse configured to store the processed data of the processor, the data warehouse comprising a storage module, the storage module comprising a dynamic storage end, a cache unit and a historical storage end, the dynamic storage end configured to update and write the character segment signals and the parameter attributes, and the historical storage end configured to prepare an index according to the parameter attributes to call static data of historical data information in the data warehouse.

[0024] Preferably, the preprocessing system comprises:

[0025] A screening unit configured to screen the character segment signals for repetition, deletion and sensitive information;

[0026] A filtering unit configured to filter the repeated character segment signals;

[0027] A correction unit configured to correct the deleted character segment signals;

[0028] A marking unit configured to mark the character segment signals containing sensitive information;

[0029] An output unit configured to output the processed character segment signals to the sub-systems in parallel in the processor.

[0030] Preferably, each of the sub-systems comprises a grouping module, a conversion processing module and a data mining module, the grouping module disposed at an initial end of an information transmission chain of the sub-system;

[0031] The grouping module configured to arrange different interval subsets of the character segment signals, and to arrange the same interval subsets of the character segment signals to the same group for processing and inputting to the conversion processing module;

[0032] The conversion processing module configured to obtain the parameter attributes of at least one word segment from the same interval subsets, to index the static data of the historical data information in the data warehouse, and to make a comparison and statistics;

[0033] The data mining module mines and extracts information based on static data of historical data information, and obtains evaluation information through analysis.

[0034] Preferably, the data governance device further comprises a self-checking metric system configured with a sampling unit for periodically and independently sampling the associated information of the multi-path thread or subsystem in the preprocessing system.

[0035] Preferably, the storage module is provided with a data encryption shielding unit, which provides visual information of corresponding permissions according to the reading level of the application layer, and the encryption shielding unit is connected with the application layer.

[0036] In a third aspect, the present application provides a computer device storing a software program for running the above method or the above device.

[0037] The present application has the following advantages:

[0038] (1) It is applied to the data stream associated with the application layer, and the data information is processed and stored through the cooperation of the application layer, the server, the processor and the data warehouse. The multiple subsystems in the processor process and transform the character segment signals to obtain the parameter attributes of at least one word segmentation. According to the parameter attributes, the historical records are called and compared, the information is mined and extracted, and the evaluation information is obtained through analysis. Based on the historical records, the information is quickly compared and mined, the response is fast, the evaluation information related changes are comprehensively considered, the timeliness is strong, and the application layer can obtain the evaluation information feedback by the device as a reference basis after the data governance optimization of the data stream;

[0039] (2) The preprocessing system in the server preliminarily classifies and processes the repeated character segment signals, filters and removes the missing character segment signals, corrects and supplements the character segment signals containing sensitive information, and processes the character segment signals containing sensitive information. The load carrying pressure of the processor is relieved, the partitioned parallel subsystem cooperates with the grouping module configured therein to process different interval subsets in the character segment signals, and the running is stable and the compatibility is excellent;

[0040] (3) The data encryption shielding unit provided in the storage module can enhance the privacy storage and reading ability of the data warehouse, and protect the authority of the data warehouse. BRIEF DESCRIPTION OF DRAWINGS

[0041] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, which together with the embodiments of the present application, serve to explain the present application, and do not constitute a limitation of the present application.

[0042] Figure 1 The system architecture diagram provided for the embodiments of the present application is shown in the figure;

[0043] Figure 2 The self-checking measurement system and the connection diagram of the subsystems are provided for the embodiments of the present application. DETAILED DESCRIPTION

[0044] The embodiments of the present application will be described in detail below with the accompanying drawings of the embodiments of the present application.

[0045] Please refer to Figure 1 , which shows the system architecture diagram provided by the embodiments of the present application, a data governance method based on big data, the application layer updates the data stream of real-time communication, and the data information is summarized through the network interface, and the communication includes the data stream associated with the HTTP, FTP, SMTP, DNS, and TCP transmission protocol, the data governance method is applied to the data stream associated with the application layer, and the method comprises the following steps:

[0046] Step a: real-time receiving and acquiring the data stream from the application layer, and dividing / interpreting the data stream into character segment signals of a specified data length;

[0047] Step b: pre-processing the character segment signals;

[0048] Specifically, step b comprises the following steps:

[0049] Screening the repeated character segment signals, the missing character segment signals, and the character segment signals containing sensitive information in the character segment signals and performing multi-path thread classification processing;

[0050] Filtering and removing the repeated character segment signals, correcting and supplementing the missing character segment signals, performing deformation marking processing on the character segment signals containing sensitive information, and outputting the classified character segment signals to the subsystem.

[0051] Step c: the multiple partition parallel subsystems process and transform the character segment signals to obtain the parameter attributes of at least one word segmentation, index and call the historical records according to the parameter attributes, compare, mine and extract information, and parse to obtain evaluation information;

[0052] Please refer to Figure 2 the connection diagram of the subsystems, step c specifically comprises the following steps:

[0053] Grouping and transforming the character segment signals in parallel partition according to the information transmission chain direction in the subsystem;

[0054] Displaying different interval subsets in the partition character segment signals, arranging the same interval subsets in the character segment signals to the same group for processing, and obtaining the parameter attributes of at least one word segmentation from the same interval subset;

[0055] According to the parameter attributes, index and call the static data of the historical data information, compare and statistically analyze, and parse to obtain evaluation information.

[0056] Step d: updating the storage character segment signal, dynamic cache parameter attribute and evaluation information, and feeding back the evaluation information to the application layer.

[0057] In one embodiment, the data governance method can independently collect samples of the classification processing results of the multi-path thread or the associated information of the subsystem at a fixed time.

[0058] In one embodiment, the present application provides a big data-based data governance device, comprising an application layer, a server, a processor, and a data warehouse,

[0059] The application layer transmits data streams related to user data in real time and receives evaluation information.

[0060] The server interfaces the data streams associated with the application layer, divides / interprets the data streams into character segment signals of a specified data length, and comprises a receiving module, a sending module, and a preprocessing system. The receiving module obtains the data streams input by the application layer, the sending module outputs the evaluation information to the application layer, and the preprocessing system performs preprocessing operations on the character segment signals.

[0061] The processor processes and transforms the character segment signals to obtain the parameter attributes of at least one segmented word. The processor is provided with a plurality of partitioned and parallel subsystems. The information transmission chain output ends of the plurality of subsystems are connected with a merge sorting module.

[0062] The data warehouse stores the processed data of the processor. The data warehouse is provided with a storage module. The storage module is provided with a dynamic storage end, a cache unit, and a historical storage end. The dynamic storage end updates and writes the character segment signals and the parameter attributes. The historical storage end prepares an index according to the parameter attributes to call the static data of the historical data information in the data warehouse.

[0063] In one embodiment, the preprocessing system comprises:

[0064] The screening unit is used to screen the character segment signals that are repeated, missing, and contain sensitive information.

[0065] The filtering unit filters the repeated character segment signals.

[0066] The correction unit corrects and supplements the missing character segment signals.

[0067] The marking unit performs deformation marking processing on the character segment signals that contain sensitive information. The deformation marking processing can be one or more of the following processing methods: adding prefixes / suffixes to the character segment signals, loading insertion / nesting, function transformation, semantic equivalence, etc.

[0068] An output unit outputs the processed character segment signal to a subsystem in the processor partitioned in parallel.

[0069] In one embodiment, each subsystem is configured with a grouping module, a conversion processing module, and a data mining module.

[0070] The grouping module is used to display different interval subsets in the partitioned character segment signal, and to arrange the same interval subsets in the character segment signal to the same group for processing and introduction into the conversion processing module.

[0071] The conversion processing module obtains at least one parameter attribute of a word in the same interval subset, indexes static data of historical data information in the data warehouse through a historical storage end, and performs comparative statistics.

[0072] The data mining module mines and extracts information based on the static data of the historical data information, and obtains evaluation information through analysis.

[0073] Based on the historical record quick comparison, the response is fast, the data mining module refers to the difference characteristics and correlation characteristics of the parameter attributes of the word, comprehensively considers the related change trend of the evaluation information, and has strong timeliness, which can timely and reliably infer the current conditions and future events.

[0074] Specifically, the subsystem outputs the character segment signal, the parameter attribute of the word, and the evaluation information to the data warehouse, and updates the transmission to the dynamic storage end according to the sequence mode specified by the merge sorting module. The dynamic storage end dynamically transmits the character segment signal, the parameter attribute of the word, and the evaluation information to the cache module. The cache module stores the character segment signal and the parameter attribute of the word in different storage locations in the storage module, and dynamically executes the integration of the parsed evaluation information into the sending module of the server through an independent path, so that the application layer interfaces the evaluation information.

[0075] Through the preliminary screening and classification processing of the pre-processing system in the server, the repeated character segment signals are filtered and removed, the missing character segment signals are corrected and supplemented, and the character segment signals containing sensitive information are marked for deformation processing, which relieves the load carrying pressure of the processor. The partitioned and parallel subsystem cooperates with the grouping module configured therein to process different interval subsets in the character segment signal, which runs stably and has excellent compatibility.

[0076] The receiving module and the sending module are provided with a buffer area for data transmission, which is used to allocate and adjust the load carrying pressure in the device, limit the high load operation, and further improve the compatibility of the device.

[0077] In one embodiment, the data governance device further comprises a self-checking measurement system, which is configured with a sampling unit for regularly sampling the associated information of the multi-path threads or subsystems in the preprocessing system, and for the preprocessing system part, the sampling unit targets the threads connected to the screening unit, the filtering unit, the correction unit, the marking unit and the output unit for regular sampling, and for the part of each subsystem of the server, the sampling unit targets the channel threads in each subsystem for regular sampling, so as to determine the associated attribute of the character segment signal transmitted by the relevant path, check the error in time, and promote the reliable operation of the device.

[0078] In one embodiment, the storage module is provided with a data encryption shielding unit, which is connected to the application layer, and the data encryption shielding unit is used to enhance the private storage and reading ability of the data warehouse, protect the authority of the data warehouse, and access the data warehouse through the network communication equipment of the user, and the data encryption shielding unit sends the visual table information based on the corresponding authority according to the reading level provided by the application layer, and feeds back the processed evaluation information to the user.

[0079] In one embodiment, the present application provides a computer device storing a software program for running the above method or the above device.

[0080] The preferred embodiments of the present application are described in detail above, but the present application is not limited to the above embodiments, and various changes and improvements can be made within the knowledge of those skilled in the art without departing from the purpose of the present application.

Claims

1. A data governance method based on big data, characterized in that, The data governance method is applied to data flows associated with the application layer during communication, and the method includes the following steps: Step a: Receive and acquire data streams from the application layer in real time, and segment / translate the data streams into character segment signals of a specified data length; Step b: Perform preprocessing operations on the character segment signal; Step c: Multiple partitioned parallel subsystems process and transform the character segment signal to obtain at least one word segmentation parameter attribute. Based on the parameter attribute, the index calls the historical records and compares them to mine and extract information, and then parses the information to obtain evaluation information. Step d: Update the stored character segment signal, dynamically cache the parameter attributes and the evaluation information, and feed back the evaluation information to the application layer; Step b includes the following steps: The signal segments containing repeated characters, missing characters, and sensitive information are filtered out and classified using multi-path threads. Repeated character segment signals are filtered out, missing character segment signals are corrected and supplemented, character segment signals containing sensitive information are deformed and marked, and the classified character segment signals are output to the subsystem. Step c includes the following steps: Following the direction of the information transmission chain within the subsystem, the character segment signal is divided into parallel partitions and grouped for transformation. Different interval subsets in the segmented character signal are displayed, and the same interval subsets in the segmented character signal are grouped together for processing. At least one segmentation parameter attribute is obtained from the same interval subset. Static data from historical data is retrieved based on parameter attribute indexes, compared and statistically analyzed, and then the evaluation information is obtained.

2. The method according to claim 1, characterized in that: It performs timed independent sampling of the classification and processing results of multi-path threads or the associated information of subsystems.

3. A data governance device based on big data, based on the method of claim 1 or 2, characterized in that, include: At the application layer, real-time data streams related to user data are sent and evaluation information is received. The server interfaces with the data stream associated with the application layer, and segments / translates the data stream into character segment signals of a specified data length. The server includes a receiving module, a sending module, and a preprocessing system. The receiving module acquires the data stream input from the application layer, the sending module outputs evaluation information to the application layer, and the preprocessing system performs preprocessing operations on the character segment signals. The processor processes and transforms the character segment signal to obtain at least one word segmentation parameter attribute. It has multiple partitioned parallel subsystems. The output end of the information transmission chain of the multiple subsystems is connected to a merge sorting module. A data warehouse stores the processed data of a processor. It is equipped with a storage module, which includes a dynamic storage terminal, a cache unit, and a historical storage terminal. The dynamic storage terminal updates and writes character segment signals and parameter attributes. The historical storage terminal prepares an index based on the parameter attributes to retrieve static data of historical data information in the data warehouse.

4. The apparatus according to claim 3, characterized in that: The preprocessing system includes: The filtering unit is used to filter out repeated, missing, and sensitive information-containing character segment signals within the character segment signals. The filtering unit filters out repetitive character segment signals; The correction unit corrects and supplements the missing character segment signal; The marking unit performs deformation marking processing on character segment signals containing sensitive information; The output unit outputs the processed character segment signal to the partitioned parallel subsystem within the processor.

5. The apparatus according to claim 3, characterized in that: Each of the subsystems is equipped with a grouping module, a transformation and processing module, and a data mining module. The grouping module is located at the initial end of the information transmission chain of the subsystem. The grouping module is used to display different interval subsets in the partitioned character segment signal, organize the same interval subsets in the character segment signal into the same group for processing, and import them into the conversion and processing module. The conversion and processing module obtains at least one word segmentation parameter attribute from the same interval subset, indexes and calls the static data of historical data information in the data warehouse, and performs comparative statistics. The data mining module extracts information from static historical data and analyzes it to obtain evaluation information.

6. The apparatus according to any one of claims 4 or 5, characterized in that: It also includes a self-testing measurement system, which is equipped with a sampling unit, which is used to periodically and independently collect and sample the associated information of multi-path threads or subsystems within the preprocessing system.

7. The apparatus according to claim 3, characterized in that: The storage module is equipped with a data encryption and shielding unit. The data encryption and shielding unit provides corresponding permission-based visual display information according to the read level of the application layer. The encryption and shielding unit is connected to the application layer.

8. A computer device, characterized in that: A software program storing the method as described in any one of claims 1-2.

Citation Information

Patent Citations

  • A metadata management method, device, and computer-readable medium

    CN109739893B

  • A method for generating execution plans for big data processing workflows

    CN110765163B

  • Data quality-based data management system

    CN107748775A

  • Data retrieval method and device, data sorting method and device, terminal and storage medium

    CN109657044A