Intelligent Data Analysis and Storage Device, Medium and Electronic Device Based on Large Model
By using intelligent data analysis and storage devices based on large models in data storage, building large models and combining source security allocation storage lakes, the existing data storage methods are solved and more efficient data secure storage is achieved.
Patent Information
- Application Number
- CN202510345544.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing data storage methods are prone to security threats such as hacker attacks and virus infections during the storage process, resulting in data leakage and tampering, reducing the security of data storage.
Using intelligent data analysis and storage devices based on large models, data processing requirements are obtained through the requirements analysis module, the data analysis module builds a large model and analyzes data, and the storage allocation module allocates storage lakes based on the analysis results and source security to realize secure storage of data.
Improve the security of data storage, and combine the source security and processing path of the data source to achieve security analysis and encryption of data to prevent data leakage and tampering.
Smart Images

Figure CN119885241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis and storage, and particularly to an intelligent data analysis and storage device, medium, and electronic device based on a large model. Background Art
[0002] In the current digital wave, large models, like a bright new star, have risen strongly and quickly become the focus of the technology field. From the initial theoretical exploration to the current wide application in various industries, large models are reshaping the current life and work patterns at an astonishing speed. Using large models for data analysis to determine data anomalies is a common method, and data analysis is closely related to data storage.
[0003] A common data storage method is to directly transfer some data to a specified location (such as different disks) for storage. However, during the storage process, the data to be stored may be subject to security threats such as hacker attacks and virus infections, resulting in data leakage and tampering. For example, if the data itself and the storage device do not have good security protection, it is easy for hackers to use vulnerabilities to steal sensitive information, etc., resulting in a reduction in the security of data storage.
[0004] Therefore, the present invention proposes an intelligent data analysis and storage device, medium, and electronic device based on a large model. Summary of the Invention
[0005] The present invention provides an intelligent data analysis and storage device, medium, and electronic device based on a large model, which is used to build a large model according to data processing requirements, analyze the data to determine existing abnormal situations, and then combine the source security and processing path to realize the security analysis of the data, so as to improve the security of data storage.
[0006] The present invention provides an intelligent data analysis and storage device based on a large model, including:
[0007] A requirements analysis module, which is used to obtain data processing requirements, extract the requirements, obtain a number of requirement items, and then obtain the requirement analysis objectives and objective weights of each requirement item;
[0008] A data analysis module, which is used to build a large model according to the requirement analysis objectives and objective weights, and input the data to be processed into the large model to obtain the analysis results of sub-data under different data sources;
[0009] A storage allocation module, which is used to allocate a storage lake to the corresponding sub-data according to the processing path corresponding to the analysis results and the data anomaly situation, and combine the source security of the data source for storage.
[0010] Preferably, the requirements analysis module includes:
[0011] A description extraction unit for separately extracting the data processing requirements based on set predefined metrics to obtain requirement descriptions under each predefined metric;
[0012] A parameter extraction unit for extracting requirement parameters related to the corresponding predefined metric from the requirement description to obtain a parameter expression, which is regarded as a requirement entry;
[0013] A matching unit for matching the requirement entry with an entry - target comparison table to obtain the corresponding requirement analysis target and target weight.
[0014] Preferably, the data analysis module includes:
[0015] An initial unit for obtaining an initial model consistent with all set predefined metrics from a metric - model database and initializing the relevant parameters of each set predefined metric in the initial model;
[0016] An allocation unit for respectively allocating to the corresponding initial layer based on the requirement analysis target and target weight under each set predefined metric to construct a large model.
[0017] Preferably, the initial model contains initial layers with the same number as the set predefined metrics, and each initial layer corresponds to a set predefined metric.
[0018] Preferably, the storage allocation module includes:
[0019] A path drawing unit for processing the nodes in sequence in the corresponding large model for the sub - data under each data source captured by the capture tool in real - time and drawing them into a processing path;
[0020] A variation calculation unit for determining the processing variation coefficient of the processing path in combination with the processing logs of each processing node;
[0021] A situation determination unit for determining the data anomaly situation of the corresponding sub - data according to the analysis result, where the data anomaly situation includes: the data self - anomaly coefficient and the factor set affecting the generation of data anomalies;
[0022] ;
[0023] Wherein, represents the data self - anomaly coefficient of the corresponding sub - data; represents the normalized processing value of the j2 - th parameter under the corresponding sub - data; represents the standard normalized value of the j2 - th parameter under the corresponding sub - data; represents the total number of parameters under the corresponding sub - data; j2 represents the j2 - th parameter under the corresponding sub - data;
[0024] A reliability determination unit, configured to determine the reliability of corresponding sub-data according to the processed coefficient of variation, in combination with data anomaly conditions and the source security of the corresponding data source;
[0025] A lake determination unit, configured to determine a storage lake from a type-storage comparison table according to the data type of the corresponding sub-data;
[0026] A storage analysis unit, configured to directly store the corresponding sub-data into the corresponding storage lake if the reliability is greater than or equal to a preset value;
[0027] Otherwise, determine a security level to be adjusted according to the reliability and the preset value;
[0028] A security encryption unit, configured to retrieve a specified number of encryption algorithms from an encryption database according to the security level to be adjusted and in combination with the reliability, perform algorithm combination, and perform security encryption on the corresponding sub-data;
[0029] A security upgrade unit, configured to perform a security upgrade on a corresponding storage space according to the space security of a specified storage space of the corresponding sub-data in the corresponding storage lake;
[0030] A storage unit, configured to store the security-encrypted sub-data in the security-upgraded storage space, where the storage space is a storage unit in the corresponding storage lake, and the storage lake contains a plurality of storage units.
[0031] Preferably, the variation calculation unit is configured to:
[0032] ;
[0033] Wherein, represents the processed coefficient of variation of the processing path; n1 represents the number of processing nodes existing on the processing path; represents the historical actual working times of the i1-th processing node; represents the historical set working times of the i1-th processing node; represents the repeated working times of the i1-th processing node on the processing path; represents the processing log of the i1-th processing node and the set log of the similarity function; represents the data after the corresponding sub-data is processed by the i1-th processing node and the data before the i1-th processing node processes the corresponding sub-processing of the similarity function; i1 represents the i1-th processing node.
[0034] Preferably, the reliability determination unit is configured to:
[0035] ;
[0036] ;
[0037] Among them, represents the reliability of the corresponding sub-data; represents the factor set under the corresponding sub-data the reliable influence coefficient of the data; represents the source security of the data source of the corresponding sub-data; represents the corresponding factor set the number of influencing factors involved; represents the average value of the data influence length of the j1th influencing factor on all historical sub-data; represents that from the data influence lengths of all historical sub-data corresponding to the j1th historical factor, according to randomly screened to obtain the variance of the corresponding number of lengths; represents the data influence length of the j1th influencing factor on the corresponding sub-data; represents the data length of the corresponding sub-data, and rand represents the random function symbol.
[0038] Preferably, the security encryption unit includes:
[0039] a number calculation sub-unit for calculating the specified number Zn of the encryption algorithm;
[0040] ;
[0041] Among them, represents the presupposition of the corresponding sub-data; represents the security attenuation unit quantity based on reliability; represents the floor function; represents the ceiling function;
[0042] an algorithm locking sub-unit for sequentially locking the first algorithm whose security is adjacent to the presupposition from the encryption database according to the specified number Zn;
[0043] a random sorting unit for randomly sorting and combining all the obtained first algorithms, and performing security encryption on the corresponding sub-data in sequence according to the first algorithms in the combination result.
[0044] The present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute any one of the intelligent data processing methods of the intelligent data processing device based on a large model.
[0045] The present invention provides an electronic device, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the intelligent data processing method of any intelligent data processing device based on a large model.
[0046] Compared with the prior art, the beneficial effects of the present application are as follows:
[0047] Build a large model according to data processing requirements and analyze the data to determine existing abnormal situations. Then, subsequently, combine source security and processing paths to implement security analysis of the data, so as to improve the security of data storage. Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 It is a structural diagram of an intelligent data analysis and storage device based on a large model provided by an embodiment of the present invention;
[0050] Figure 2 It is a structural diagram of a large model provided by an embodiment of the present invention;
[0051] Figure 3 It is a flowchart of the intelligent data processing method of an intelligent data processing device based on a large model provided by an embodiment of the present invention. Detailed Embodiments
[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0053] The present invention provides an intelligent data analysis and storage device based on a large model, as Figure 1 shown, including:
[0054] A requirements analysis module, configured to obtain data processing requirements and perform requirement extraction to obtain a number of requirement entries, and then obtain the requirement analysis objectives and objective weights of each requirement entry;
[0055] A data analysis module, which is used to build a large model according to the demand analysis target and target weight, and input the data to be processed into the large model to obtain the analysis results of sub-data under different data sources;
[0056] A storage allocation module, which is used to allocate a storage lake to the corresponding sub-data according to the processing path corresponding to the analysis result and the data anomaly situation, and in combination with the source security of the data source, and perform storage.
[0057] In this embodiment, the data processing requirement refers to a processing description of what kind of processing needs to be performed on the data. For example, it is necessary to identify typos in the data, identify whether the typos are individual errors or consecutive errors, and the typo recognition accuracy is 90%. Or, it is necessary to identify missing information in the data, and the missing information recognition accuracy is 95%, etc.
[0058] In this embodiment, the large model framework is based on a neural network model. Through the processing of resource data with different data processing requirements and the extraction requirements under different predefined indicators as samples, the neural network model is trained. Therefore, there is an initial framework. However, due to different requirements, it is necessary to adjust the parameters of different layers in the model to ensure that the model can strictly meet the requirements and ensure the reliability of data analysis.
[0059] In this embodiment, the data to be processed refers to network resource data. For example, resource data A1 is obtained from platform R1, and resource data A2 is obtained from platform R2, etc. And the data to be processed refers to the data that needs to be analyzed for anomalies and security. For example, it is the resource data required by users for AI model construction. Since the data is distributed on different platforms, and there may be situations such as data tampering due to insecure factors during the data acquisition process from the platform, it is necessary to perform a preliminary analysis on the data to be processed based on the large model to determine anomalies and the possible security impacts of the data on the processing path during the analysis process, so as to further enhance the security protection of the data and provide anticipatory protection.
[0060] In this embodiment, the source security is obtained by matching the source type - security comparison table of the data source, and the source security under different source types is different, with a value range of 0 to 1, which is set before the construction of the relevant platform of the corresponding data source and belongs to known content.
[0061] In this embodiment, the allocated storage lake is for the secure storage of sub-data, and the storage lake refers to storage media such as storage disks that can store data.
[0062] The beneficial effects of the above technical solution are as follows: A large model is constructed according to data processing requirements, and the data is analyzed to determine existing abnormal situations. Subsequently, in combination with source security and processing paths, security analysis of the data is realized to improve the security of data storage.
[0063] The present invention provides an intelligent data analysis and storage device based on a large model. The requirements analysis module includes:
[0064] A description extraction unit, configured to separately extract the data processing requirements based on set predefined metrics, and obtain requirement descriptions under each predefined metric;
[0065] A parameter extraction unit, configured to extract requirement parameters related to corresponding predefined metrics from the requirement descriptions, obtain parameter expressions, and regard them as requirement entries;
[0066] A matching unit, configured to match the requirement entries with an entry-target comparison table to obtain corresponding requirement analysis targets and target weights.
[0067] In this embodiment, the set predefined metrics are typo recognition description metrics and missing information description metrics. Then, corresponding descriptions can be obtained by extracting from the requirements according to the relevant metrics, that is: it is necessary to recognize typos in the data, recognize single typos, consecutive two typos, consecutive multiple typos, etc., and the typo recognition accuracy is 90%. It is necessary to recognize missing information in the data, recognize single-character missing and multi-character missing, and the missing recognition accuracy is 95%.
[0068] The parameter expression of the requirement parameters under the typo recognition description metric is: @1@2@3 + 90%. The requirement parameters are single typo @1, consecutive two typos @2, consecutive multiple typos @3, and typo recognition accuracy 90%.
[0069] The parameter expression of the requirement parameters under the missing information description metric is: #1#3 + 95%. The requirement parameters are single missing character #1, multiple missing characters #3, and missing recognition accuracy 95%.
[0070] In this embodiment, the entry-target comparison table includes a number of pre-stored requirement entries, requirement analysis targets and metric weights matching the entries. The requirement analysis targets include: recognizing typos in different quantities, etc. Among them, the sum of the target weights of all set predefined metrics is 1, and the target weight = the metric weight of the corresponding set predefined metric / the sum of the metric weights of all set predefined metrics involved.
[0071] The beneficial effects of the above technical solution are as follows: extracting requirements according to indicators to obtain requirement descriptions, and then obtaining requirement analysis objectives and objective weights through parameter extraction and reference table matching, providing a basis for subsequent construction of a large model.
[0072] The present invention provides an intelligent data analysis and storage device based on a large model. The data analysis module includes:
[0073] An initial unit, configured to obtain an initial model consistent with all set predefined indicators from an indicator-model database, and initialize the relevant parameters of each set predefined indicator in the initial model;
[0074] An allocation unit, configured to allocate the requirement analysis objectives and objective weights under each set predefined indicator to corresponding initial layers respectively, and construct a large model.
[0075] Preferably, the initial model includes initial layers with the same number as the set predefined indicators, and each initial layer corresponds to a set predefined indicator.
[0076] In this embodiment, the indicator-model database includes different combinations of set predefined indicators and the initial models matching the combinations. The purpose of initialization is to facilitate reallocation of objectives and weights according to indicators to ensure that the model fits the requirements as much as possible. It should be noted that the initial model is pre-planned, and only the relevant parameters need to be initialized and reset. For example, the weights of the layers corresponding to the indicators are set to 0 and replaced with the newly obtained objective weights.
[0077] For example, as Figure 2 shown, it is the structure diagram of the large model, including layer H1, layer H2, etc., and layer H1 corresponds to the set predefined indicator a1, layer H2 corresponds to the set predefined indicator a2, etc.
[0078] In this embodiment, the number of predefined indicators is at least 3.
[0079] For example, first identify typos, then identify missing information, and finally analyze data anomalies.
[0080] The beneficial effects of the above technical solution are as follows: obtaining an initial model under a set combination of predefined indicators from the database, and then resetting the objectives and weights for the initial layers to obtain a large model, ensuring the rationality of data analysis.
[0081] The present invention provides an intelligent data analysis and storage device based on a large model. The storage allocation module includes:
[0082] A path drawing unit for processing nodes successively in a corresponding large model according to sub - data under each data source captured in real - time by a capture tool and drawing them into a processing path;
[0083] A mutation calculation unit for determining a processing mutation coefficient of the processing path in combination with the processing logs of each processing node;
[0084] A situation determination unit for determining a data anomaly situation corresponding to the sub - data according to the analysis result, where the data anomaly situation includes: a data self - anomaly coefficient and a set of factors affecting the generation of data anomalies;
[0085] ;
[0086] Among them, represents the data self - anomaly coefficient of the corresponding sub - data; represents the normalized processing value of the j2 - th parameter under the corresponding sub - data; represents the standard normalized value of the j2 - th parameter under the corresponding sub - data; represents the total number of parameters under the corresponding sub - data; j2 represents the j2 - th parameter under the corresponding sub - data;
[0087] A reliability determination unit for determining the reliability of the corresponding sub - data according to the processing mutation coefficient, in combination with the data anomaly situation and the source security of the corresponding data source;
[0088] A lake determination unit for determining a storage lake from a type - storage comparison table according to the data type of the corresponding sub - data;
[0089] A storage analysis unit for directly storing the corresponding sub - data into the corresponding storage lake if the reliability is greater than or equal to a preset value;
[0090] Otherwise, determine a security level to be adjusted according to the reliability and the preset value;
[0091] A security encryption unit for retrieving a specified number of encryption algorithms from an encryption database according to the security level to be adjusted and combining the algorithms to perform security encryption on the corresponding sub - data;
[0092] A security upgrade unit for upgrading the security of the specified storage space in the corresponding storage lake according to the space security of the corresponding sub - data in the specified storage space of the corresponding storage lake;
[0093] A storage unit for storing the securely encrypted sub - data in the securely upgraded storage space, where the storage space is a storage unit in the corresponding storage lake, and the storage lake contains several storage units.
[0094] Preferably, the mutation calculation unit is used for:
[0095] ;
[0096] wherein, represents the processing coefficient of variation of the processing path; n1 represents the number of processing nodes existing on the processing path; represents the historical actual working times of the i1-th processing node; represents the historical set working times of the i1-th processing node; represents the repeated working times of the i1-th processing node on the processing path; represents the processing log of the i1-th processing node and the set log of the similarity function; represents the data after the corresponding sub-data is processed by the i1-th processing node and the data before the corresponding sub-processing is not processed by the i1-th processing node of the similarity function; i1 represents the i1-th processing node.
[0097] Preferably, the reliability determination unit is used for:
[0098] ;
[0099] ;
[0100] wherein, represents the reliability of the corresponding sub-data; represents the factor set under the corresponding sub-data of the reliable influence coefficient on the data; represents the source security of the data source of the corresponding sub-data; represents the corresponding factor set involving the number of influencing factors; represents the average value of the data influence lengths of the j1-th influencing factor on all historical sub-data; represents the variance of the corresponding number of lengths randomly selected from the data influence lengths of all historical sub-data corresponding to the j1-th historical factor according to ; represents the data influence length of the j1-th influencing factor on the corresponding sub-data; represents the data length of the corresponding sub-data, and rand represents the random function symbol.
[0101] In this embodiment, the processing nodes correspond one-to-one with the initial layers. For example, the typo recognition layer corresponds to one node, the missing information recognition layer corresponds to one node, and the data anomaly analysis layer corresponds to one node. For example, the typo recognition corresponding to the typo recognition layer is 3 times. At this time, the missing recognition corresponding to the missing information recognition layer is 2 times, and the anomaly analysis corresponding to the data anomaly analysis layer is 1 time. At this time, the processing path is: typo recognition layer - typo recognition layer - typo recognition layer - missing information recognition layer - missing information recognition layer - data anomaly analysis layer.
[0102] In this embodiment, the processing log is the background data automatically generated by the corresponding node during the processing, and can be directly captured. For example, the situation of typo recognition, such as which word is a typo and which position it is in the sub-data.
[0103] In this embodiment, for example, the word at position e1 should be "I", but the recognized word is "oh", and so on.
[0104] In this embodiment, the set of factors affecting data anomalies is determined based on the analysis results. That is, there are U1 set influencing factors. Each factor is used to conduct a security test on the sub-data to determine whether there is a change under the corresponding influencing factor. The factors with changes are summarized to obtain the factor set. That is, the number in the factor set is less than U1. For example, a Trojan virus of model w1 will change some characters in the sub-data.
[0105] In this embodiment, the normalization process refers to normalizing the values of relevant parameters to between 0 and 1 to facilitate subsequent calculations, and the standard normalization values of the values of each parameter are preset.
[0106] The data type refers to the type of resource obtained. For example, the sub-data obtained from the search results of "semiconductor" can be regarded as the device type.
[0107] In this embodiment, the type-storage correspondence table contains different data types and the storage lakes for this type, which is convenient for direct storage allocation.
[0108] In this embodiment, the presupposition is preset and the value is 0.8.
[0109] In this embodiment, the specified storage space refers to a certain storage address in the storage lake (disk), such as the address 0011-ff01.
[0110] In this embodiment, the space security is obtained by matching from the space-security correspondence table, and this table contains different address segments and the security of the space for this address segment, which are all preset in advance. That is, it has been factory-tested before the storage lake is used for storage.
[0111] In this embodiment, security upgrade means that if the space security is greater than or equal to the preset security, no upgrade is required; otherwise, the storage space needs to be securely upgraded according to the preset security, and an upgrade package can be used for the upgrade. The upgrade package contains content related to security encryption algorithms, which belongs to the prior art.
[0112] In this embodiment, the set log is pre-planned and belongs to known content.
[0113] In this embodiment, the value of M1 is greater than 5, and corresponding lengths are obtained by random screening. It should be noted that the number of random screening is a positive number.
[0114] For example, if the length of the sub-data is 100 and the length of the influence of factor a1 on the sub-data is 10, then the calculation result of factor a1 is: 10 / 100.
[0115] The beneficial effects of the above technical solution are: starting from the processing path, processing log, and analysis results to determine the processing coefficient of variation and data anomaly conditions, and then determining the security level to be adjusted for the sub-data and combining with reliability to obtain a specified number of encryption algorithms, ensuring the security of data encryption, and further combining with the security situation of the space to ensure the security of data storage from two aspects.
[0116] The present invention provides an intelligent data analysis and storage device based on a large model. The security encryption unit includes:
[0117] A number calculation sub-unit for calculating the specified number Zn of the encryption algorithm;
[0118] ;
[0119] Among them, represents the presupposition corresponding to the sub-data; represents the security attenuation unit quantity based on reliability; represents the floor function; represents the ceiling function;
[0120] An algorithm locking sub-unit for sequentially locking the first algorithm whose security is adjacent to the presupposition from the encryption database according to the specified number Zn;
[0121] A random sorting unit for randomly sorting and combining all the obtained first algorithms, and securely encrypting the corresponding sub-data in sequence according to the first algorithm in the combination result.
[0122] In this embodiment, is the security level to be adjusted, and its value is an integer.
[0123] In this embodiment, each encryption algorithm has its corresponding security, which is pre-tested. Therefore, the security of all encryption algorithms can be sorted in order to obtain an algorithm close to the preset security. For example, the security of the encryption algorithms sorted in order are: o1, o2, o3, o4, o5, o6. At this time, the preset security is between o4 and o5. At this time, Zn is 2, and the encryption algorithm corresponding to o4 and o5 is used as the first algorithm. If the preset security is after o6, Zn is 2, and the encryption algorithm corresponding to o5 and o6 is used as the first algorithm, and so on. The principle is to sort the algorithms closest to the preset security.
[0124] In this embodiment, the random ordering combination is to encrypt after disrupting the order to ensure encryption security.
[0125] In this embodiment, the reliability-based safety attenuation unit is generally 0.2.
[0126] The beneficial effect of the above technical solution is: based on the preset and reliability, the specified number is calculated, and then the first algorithm is obtained and randomly combined to achieve secure encryption of data and ensure the security of data.
[0127] The present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes any one of the intelligent data processing methods of the intelligent data processing device based on a large model, such as Figure 3 As shown:
[0128] Step 1: Obtain data processing requirements and extract requirements to obtain several requirement items, and then obtain the requirement analysis target and target weight of each requirement item;
[0129] Step 2: According to the demand analysis target and target weight, a large model is constructed, and the data to be processed is input into the large model to obtain the analysis results of the sub-data under different data sources;
[0130] Step 3: According to the processing path and data anomalies corresponding to the analysis results, and in combination with the source security of the data source, a storage lake is allocated to the corresponding sub-data and stored.
[0131] The present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any one of the intelligent data processing methods of the intelligent data processing device based on a large model:
[0132] Step 1: Obtain data processing requirements and perform requirement extraction to obtain several requirement entries, and then obtain the requirement analysis objectives and objective weights for each requirement entry;
[0133] Step 2: Construct a large model according to the requirement analysis objectives and objective weights, and input the data to be processed into the large model to obtain the analysis results of sub-data under different data sources;
[0134] Step 3: According to the processing path corresponding to the analysis results and data anomaly situations, and in combination with the source security of the data source, allocate storage lakes to the corresponding sub-data and perform storage.
[0135] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent data analysis and storage device based on a large model, characterized in that: include: The demand analysis module is used to obtain data processing requirements and extract requirements to obtain a number of requirement items, and then obtain the demand analysis target and target weight of each requirement item; A data analysis module is used to build a large model according to the demand analysis target and target weight, and input the data to be processed into the large model to obtain the analysis results of the sub-data under different data sources; A storage allocation module, configured to allocate a storage lake to corresponding sub-data and perform storage according to the processing path and data anomaly corresponding to the analysis result and in combination with the source security of the data source; Wherein, the storage allocation module includes: A path drawing unit is used to process nodes in the corresponding large model in sequence according to the sub-data under each data source captured in real time by the capture tool, and draw a processing path; a variation calculation unit, for determining a processing variation coefficient of the processing path in combination with a processing log of each processing node; A situation determination unit, used to determine the data anomaly situation of the corresponding sub-data according to the analysis result, wherein the data anomaly situation includes: an anomaly coefficient of the data itself and a set of factors that affect the generation of data anomaly; ; in, Indicates the data anomaly coefficient of the corresponding sub-data; Indicates the normalized value of the j2th parameter under the corresponding sub-data; represents the standard normalized value of the j2th parameter under the corresponding sub-data; represents the total number of parameters under the corresponding sub-data; j2 represents the j2th parameter under the corresponding sub-data; A reliability determination unit, configured to determine the reliability of the corresponding sub-data according to the processing variation coefficient and in combination with data anomalies and source security of the corresponding data source; A lake determination unit, configured to determine a storage lake from a type-storage comparison table according to a data type of corresponding sub-data; A storage analysis unit, configured to directly store the corresponding sub-data in a corresponding storage lake if the reliability is greater than or equal to a preset reliability; The variation calculation unit is used to calculate the processing variation coefficient of the processing path: ; in, represents the processing variation coefficient of the processing path; n1 represents the number of processing nodes existing on the processing path; Indicates the historical actual working times of the i1th processing node; represents the number of historical set jobs of the i1th processing node; represents the number of times the i1th processing node repeats work on the processing path; Indicates the processing log of the i1th processing node and setting log Similarity function of ; Represents the data after the corresponding sub-data is processed by the i1th processing node The data before the corresponding subprocessing of the i1th processing node Similarity function of ; i1 represents the i1th processing node.
2. The large model-based intelligent data analysis and storage device according to claim 1, characterized in that: The demand analysis module includes: A description extraction unit, used to extract the data processing requirements respectively based on the set predefined indicators, and obtain the requirement description under each predefined indicator; A parameter extraction unit, used to extract the requirement parameters related to the corresponding predefined indicators from the requirement description, obtain parameter expressions, and regard them as requirement items; The matching unit is used to match the demand items with the item-target comparison table to obtain corresponding demand analysis targets and target weights.
3. The large model-based intelligent data analysis and storage device according to claim 1, characterized in that: The data analysis module comprises: An initialization unit, used for acquiring an initial model consistent with all predefined indicators from an indicator-model database, and initializing relevant parameters of each predefined indicator in the initial model; The allocation unit is used to allocate the demand analysis objectives and target weights under each set predefined indicator to the corresponding initial layers to construct a large model.
4. The large model-based intelligent data analysis and storage device according to claim 3, characterized in that: The initial model includes initial layers whose number is consistent with the set predefined indicators, and each initial layer corresponds to a set predefined indicator.
5. The large model-based intelligent data analysis and storage device according to claim 1, characterized in that: The storage allocation module further includes: A storage analysis unit, configured to determine a security level to be adjusted according to the reliability and the preset reliability if the reliability is less than the preset reliability; A security encryption unit, used to retrieve a specified number of encryption algorithms from an encryption database according to the security level to be adjusted and in combination with reliability, and to combine the algorithms to securely encrypt the corresponding sub-data; A security upgrade unit, used to perform security upgrade on the corresponding storage space according to the space security of the designated storage space of the corresponding sub-data in the corresponding storage lake; The storage unit is used to store the securely encrypted sub-data in the security-upgraded storage space, wherein the storage space is a storage unit in the corresponding storage lake, and the storage lake includes several storage units.
6. The large model-based intelligent data analysis and storage device according to claim 5, characterized in that: The reliability determination unit is used to: ; ; in, Indicates the reliability of the corresponding sub-data; Represents the factor set under the corresponding sub-data Reliable influence coefficient on data; Indicates the source security of the data source of the corresponding sub-data; Represents the corresponding factor set The number of influencing factors involved; It represents the average value of the data impact length of the j1th influencing factor on all historical sub-data; Indicates that the data impact length of all historical sub-data corresponding to the j1th historical factor is calculated according to Random screening obtains the variance of the corresponding number of lengths; Indicates the length of the impact of the j1th influencing factor on the corresponding sub-data; Indicates the data length of the corresponding sub-data, and rand indicates the random function symbol.
7. The large model-based intelligent data analysis and storage device according to claim 6, characterized in that: The security encryption unit comprises: A number calculation subunit, used to calculate the specified number Zn of the encryption algorithm; ; in, Indicates the preset nature of the corresponding sub-data; It represents the unit quantity of safety degradation based on reliability; represents the floor function; represents the ceiling function; An algorithm locking subunit, used to sequentially lock the first algorithm with a predetermined security level from the encryption database according to the specified number Zn; The random sorting unit is used to randomly sort and combine all the acquired first algorithms, and securely encrypt the corresponding sub-data in sequence according to the first algorithms in the combination result.
8. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes the intelligent data processing method of the large model-based intelligent data processing device according to any one of claims 1-7.
9. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the intelligent data processing method of the large model-based intelligent data processing device according to any one of claims 1-7.
Citation Information
Patent Citations
Node detection method and device, computer, storage medium and program product
CN118101427A
Abnormal data positioning and tracing method
CN118426997A