Big data-based method for collecting full data of higher vocational colleges
By adopting a big data-based full-volume data collection method, and employing unique identifier link comparison, status threshold discrimination, and partitioned storage mechanisms, the problems of data redundancy and omission in the data collection process of higher vocational colleges have been solved, and efficient and compliant data storage and transmission have been achieved.
Patent Information
- Application Number
- CN202511339956.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-19
AI Technical Summary
In existing technologies, data redundancy and omissions exist in the data collection process of higher vocational colleges. The format conversion and classification summary lack real-time verification, resulting in low data utilization efficiency and uneven storage, making it difficult to meet the proportional control requirements of large-scale scientific research data.
By employing a full-scale data aggregation method based on big data, and using unique identifier link comparison, status threshold discrimination, citation frequency detection, and partitioned storage mechanisms, we can achieve accurate data differentiation, real-time verification, and differentiated storage, ensuring the accuracy and integrity of data push and optimizing the storage structure.
It improves the retrieval and storage efficiency of high-value data, ensures the stability and compliance of data transmission, achieves efficient transmission and reasonable partitioned storage, and solves the problems of low data utilization efficiency and unbalanced storage.
Smart Images

Figure CN120821735B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data aggregation, in particular to a full data aggregation method for higher vocational colleges based on big data. BACKGROUND
[0002] The technical field of data aggregation involves collecting, sorting, converting and integrating raw data scattered in different sources, formats and systems to achieve unified management and efficient use of data, including diversification of data collection methods, establishment of data cleaning rules, data standardization processing, and centralized storage and management of multi-source data, which is completed by centralized databases. Among them, the full data aggregation method for higher vocational colleges refers to the collection and integration of multi-dimensional business data of students, teachers, courses, scientific research and logistics in higher vocational colleges. Usually, it relies on the periodic provision of electronic spreadsheets and business management system export files by each department through artificial batch import, and uses rule scripts to convert and classify each item, and then uses a relational database to realize unified storage.
[0003] In the existing data aggregation process of colleges and universities, batch import and periodic file sorting are mostly relied on. When collecting multi-source data across departments, data redundancy and omission are prone to occur. The lack of real-time verification means in the format conversion and classification summary process makes it difficult to ensure consistency in the transmission link. In the multi-dimensional business data processing scenario, frequent manual intervention increases the processing delay. The lack of dynamic monitoring of reference records cannot effectively distinguish high-value scientific research achievements. The centralized storage mode of the relational database cannot meet the proportion control requirements when dealing with large-scale scientific research data distribution, resulting in reduced data utilization efficiency and uneven storage problems. SUMMARY
[0004] The purpose of the present application is to solve the shortcomings in the prior art and to provide a full data aggregation method for higher vocational colleges based on big data.
[0005] In order to achieve the above purpose, the present application adopts the following technical scheme: a full data aggregation method for higher vocational colleges based on big data, comprising the following steps:
[0006] S1: Collect the source system code and time sequence of the higher vocational college, splice to form a unique identification string and perform repetitive verification item by item to obtain a unique link comparison sequence;
[0007] S2: According to the unique link comparison sequence, when the state satisfies the data quality threshold condition, execute the link activation action to push the data to the scientific research achievement storage area, and when the state does not satisfy the condition, keep the link standby, to obtain a link activation sequence;
[0008] S3: According to the link activation sequence, the reference record is detected and the frequency of each achievement is calculated, the high-performance storage partition write action is performed when the high-cited threshold is exceeded, and the ordinary storage area aggregation write action is performed when the high-cited threshold is less than or equal to, to obtain a scientific research achievement partition list;
[0009] S4: Based on the scientific research achievement partition list, the consistency of the check state is compared and checked, the retransmission mechanism is performed when the check state appears to be different, and the confirmation flag is recorded when the state is consistent, to obtain a partition link integrity record;
[0010] S5: According to the partition link integrity record, it is judged whether the data distribution of each storage area meets the data minimization principle, if there is deviation, the cache reconstruction is performed, if there is no deviation, the current state is maintained, and the high-vocational college big data full collection record is obtained.
[0011] As a further scheme of the application, the unique link comparison sequence includes identification uniqueness information, cache update state, and repeated check result, the link activation sequence includes data quality state, transmission trigger flag, and push storage instruction, the scientific research achievement partition list includes high-performance storage index, ordinary storage aggregation record, and frequency classification result, the partition link integrity record includes state consistency identification, retransmission confirmation record, and log comparison result, and the high-vocational college big data full collection record includes storage area distribution proportion data, cache reconstruction information, and link storage state.
[0012] As a further scheme of the application, the unique link comparison sequence specifically acquires the steps of:
[0013] S111: Based on the high-vocational college source system code and time sequence received by the acquisition gateway, the code and time sequence are spliced to form a unique identification string, and the unique identification string is compared with the identification sequence in the cache list one by one to obtain a difference comparison value;
[0014] S112: According to the difference comparison value, the identification sequence of the cache list is compared, when the comparison result is inconsistent, the received data is transmitted, and in the process, the cache list order is updated according to the comparison difference and the earliest identification is deleted, to obtain a cache update sequence;
[0015] S113: According to the cache update sequence, the joint judgment is performed in combination with the difference comparison value, when the comparison result is completely consistent, the current data is discarded, and a unique link comparison sequence is generated.
[0016] As a further scheme of the application, the link activation sequence specifically acquires the steps of:
[0017] S211: Based on the unique link comparison sequence read link transmission parameters, extract the transmission rate, delay time and error rate parameters in each link sequence, and store the integrity information of each link with the number of received data packets, integrate the parameters of all links one by one, and obtain the link parameter set;
[0018] S212: Based on the link parameter set, judge each dimension of integrity, timeliness and accuracy compared with the data quality threshold, record the qualified or unqualified state after comparison, and obtain the link state judgment result;
[0019] S213: According to the link state judgment result, when the state mark is qualified, execute the link activation action to push the data to the scientific research achievement storage area, when the state mark is unqualified, keep the link standby and do not push the data, carry out joint evaluation and arrange the qualified link items in order, and generate the link activation sequence.
[0020] As a further scheme of the application, the scientific research achievement partition list specifically obtains the steps of:
[0021] S311: Based on the scientific research achievement data in the link activation sequence, detect the reference record of each achievement, read the achievement number one by one and associate the reference log item, store the reference frequency and mark the corresponding achievement number, and obtain the reference record frequency data;
[0022] S312: According to the reference record frequency data, compare the frequency value of each achievement with the high cited threshold one by one, if the frequency is greater than the high cited threshold, mark it as high cited, otherwise mark it as ordinary, and obtain the frequency status record;
[0023] S313: According to the frequency status record, execute the high performance storage partition write action on the high cited achievement and record the index information, execute the ordinary storage area aggregation write action on the ordinary achievement, and establish the scientific research achievement partition list.
[0024] As a further scheme of the application, the partition link integrity record specifically obtains the steps of:
[0025] S411: Based on the scientific research achievement partition list, obtain the partition number and corresponding storage location data of each achievement, read the transmission state record and time node information of each number in the transmission log, combine the storage record of the corresponding number in the storage record and the write time, compare the time and state fields of the same number in the two data sources one by one, arrange each number comparison result into a state comparison item, and obtain the state consistency comparison record;
[0026] S412: According to the state consistency comparison record, it is judged whether all achievement number corresponding state fields are consistent or not, if consistent, 1 is assigned, if inconsistent, 0 is assigned, and the consistency deviation influence value is obtained by operation.
[0027] S413: Based on the consistency deviation influence value, the number whose difference is greater than the determination threshold is marked as an abnormal number, and the corresponding number is retriggered to form a retransmission list, the remaining number is marked as a confirmation item and a confirmation flag bit is generated, all number corresponding flag bits generate a chain confirmation index queue, and a partition link integrity record is established.
[0028] As a further scheme of the application, the high vocational college big data full collection record specifically obtains the steps of:
[0029] S511: Based on the partition link integrity record, the achievement number set corresponding to each storage area is extracted, and the belonging label and the belonging partition identifier of each number in the link data are called, the relationship between the number of data numbers in each storage area and the total number of corresponding storage area is counted, and the partition data belonging proportion is obtained;
[0030] S512: According to the partition data belonging proportion, a belonging proportion determination reference value is set, the belonging proportion corresponding to each storage area is compared with the reference value, and each partition is identified according to the comparison result, the partition whose proportion is higher than the reference value is marked as irregular, and the partition whose proportion is lower than or equal to the reference value is marked as regular, and a data distribution compliance identification sequence is generated;
[0031] S513: According to the data distribution compliance identification sequence, the information content is sorted for the data partition marked as irregular, and the storage structure is combed according to the record state of the number in the cache structure, the update operation of the storage area number distribution is completed, and the high vocational college big data full collection record is obtained.
[0032] Compared with the prior art, the application has the advantages and positive effects that:
[0033] In the application, the source data is accurately distinguished in the receiving link through link comparison of unique identification, the repeated information occupies the resources, the data quality in the transmission process is judged in real time through the setting of the state threshold, the accuracy and integrity of data pushing are ensured, the differentiated partition storage of scientific research achievements is realized through the dynamic detection of the frequency, the retrieval and storage efficiency of high value data is effectively improved, the transmission state integrity is ensured through comparison and inspection, and the retransmission mechanism is triggered under the difference, the link stability is improved, the data minimization principle is implemented through the proportion control of the storage area distribution, the data meets the compliance requirements and the storage structure is optimized, the overall effect of efficient transmission, reasonable partition and compliance storage is realized. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 is a main step flowchart of the present application;
[0035] Figure 2 is a unique link comparison sequence acquisition flowchart of the present application;
[0036] Figure 3 is a link activation sequence acquisition flowchart of the present application;
[0037] Figure 4 is a scientific research achievement subarea list acquisition flowchart of the present application;
[0038] Figure 5 is a subarea link complete record acquisition flowchart of the present application;
[0039] Figure 6 is a higher vocational college big data full amount collection record acquisition flowchart of the present application. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0041] In the description of the present application, it should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, in the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.
[0042] Please refer to Figure 1 , the full amount data collection method of higher vocational colleges based on big data, comprising the following steps:
[0043] S1: obtaining the higher vocational college source system coding and time sequence received by the collection gateway, splicing the two to form a unique identification string, and performing repetitive verification with the identification sequence in the cache list, when the comparison result is inconsistent, performing data transmission and updating the cache sorting to delete the earliest identification, when the comparison result is completely consistent, discarding the current data, and obtaining the unique link comparison sequence;
[0044] S2: According to the unique link comparison sequence, read the link transmission parameters, judge whether the sequence state triggers data sending under the condition of data quality threshold (covering integrity (≥99%), timeliness (delay ≤1s), accuracy (error rate ≤0.1%) and other dimensions), when the state meets the threshold condition, execute the link activation sequence to push the data to the scientific research achievement storage area, when the state does not meet the condition, keep the link standby, get the link activation sequence;
[0045] S3: According to the scientific research data in the link activation sequence, detect the reference record and calculate the frequency of each achievement, compare the calculation result with the high cited threshold (set dynamic threshold according to the cited frequency of the top 1% papers), when the frequency exceeds the high cited threshold, execute the high performance storage partition write action and record the index information, when the frequency is lower than or equal to the high cited threshold, execute the ordinary storage area aggregation write action, get the scientific research partition list;
[0046] S4: Based on the scientific research partition list, combine the transmission log and storage record to check the consistency of the state, when the state difference appears, execute the retransmission mechanism, when the state is consistent, record the confirmation mark, get the partition link complete record;
[0047] S5: According to the partition link complete record, judge whether the data distribution of each storage area meets the data minimization principle (meets the storage proportion limit of general data protection regulations), if there is deviation in the distribution, execute the cache reconstruction protocol, if there is no deviation in the distribution, maintain the current link and storage state, get the high vocational college big data full collection record.
[0048] The unique link comparison sequence includes identification uniqueness information, cache update state and repeated check result, the link activation sequence includes data quality state, transmission trigger mark and push storage instruction, the scientific research partition list includes high performance storage index, ordinary storage aggregation record and frequency classification result, the partition link complete record includes state consistency mark, retransmission confirmation record and log comparison result, the high vocational college big data full collection record includes storage area distribution proportion data, cache reconstruction information and link storage state.
[0049] Please refer to Figure 2 , the specific steps of S1 are:
[0050] S111: Based on the high vocational college source system code and time sequence received by the collection gateway, splice the code and time sequence to form a unique identification string, and compare the unique identification string with the identification sequence in the cache list one by one to get the difference comparison value;
[0051] Based on the high vocational college source system code and time sequence received by the collection gateway, the code data is extracted in the form of string, and the time sequence is formatted to form a time stamp that increases by milliseconds. In the refinement process, the source system code such as A101, A102 is spliced with the time stamp such as 1722300000, 1722300001 in turn to form unique identification strings such as A101-1722300000, A102-1722300001, and then these unique identification strings are compared with the existing identification in the cache list one by one. In the comparison process, it is judged whether there is a difference by scanning the characters bit by bit. If there is a difference, the difference index is recorded and the difference content is stored. For example, when the cache list contains A101-1722299999, the difference index is the last position 1, and the difference value is 1. Then, the difference indexes are processed. According to the absolute value of the difference index, it is divided into three intervals: less than 2, between 2 and 4, and greater than 4. The difference level can be determined by the numerical value. For example, A101-1722300000 and A101-1722299998 have a difference value of 2, which falls into the middle interval. See the example comparison data shown in Table 1. Finally, all the difference indexes and corresponding difference levels are recorded by the above comparison operation to obtain the difference comparison value.
[0052] Table 1: Difference comparison data table
[0053]
[0054] As shown in Table 1, the difference index is obtained by bit-by-bit scanning, and the difference level is directly determined by interval division.
[0055] S112: According to the difference comparison value, compare the identification sequence of the cache list. When the comparison result is inconsistent, transmit the received data, and in the process, update the cache list order according to the comparison difference and delete the earliest identification to obtain the cache update sequence.
[0056] According to the difference comparison value, all the identification sequences in the cache list are compared. In the refinement process, first, the entries marked as greater than 4 interval in the difference comparison value are selected, and the received data corresponding to these entries is directly transmitted to the data receiving end, for example, A102-1722300001 in Table 1 meets the condition, and after the transmission is completed, the cache list is updated immediately, the latest received identification is inserted into the head of the cache list, and the oldest identification such as A102-1722299997 is deleted in the original cache sequence order. In the middle interval difference entry such as A103-1722300003, the cache is updated in sequence but the old identification is retained for subsequent comparison processing, and the interval less than 2 is only retained without updating the existing cache. In this process, the cache sequence index table is regenerated every time the update is performed, such as [0: A102-1722300001, 1: A101-1722300000], and the new sequence index table is sorted and checked. If the index order is out of order, it is reordered and adjusted to ensure that the final cache list is continuous and complete, and the cache update sequence is finally obtained.
[0057] S113: According to the cache update sequence, combined with the difference comparison value, when the comparison result is completely consistent, the current data is discarded, and a unique link comparison sequence is generated;
[0058] According to the cache update sequence, when the comparison result is completely consistent, the current data is discarded, and in the refinement process, first, the identification with a difference comparison value less than 2 interval in the cache update sequence is screened. These identifications are marked as discarded items because they are completely consistent with the cache identification. Then, the middle interval and large interval parts of the difference comparison value are associated with the index of the cache update sequence to construct a joint judgment table. The update state of each identification and the comparison difference are judged one by one, for example, A101-1722300000 is less than 2 interval and the state is not updated, so it is discarded directly, and A103-1722300003 is in the middle interval and the state has been updated, so it is retained. After marking the processing results of each item in the joint judgment table, all the identifications marked as retained are rearranged into a link data chain according to the index order, and the final unique sequence set is formed by removing all duplicates and discarded items. This set is the unique link comparison sequence.
[0059] Please refer to Figure 3 , the specific steps of S2 are:
[0060] S211: Based on the unique link comparison sequence, read the link transmission parameters, extract the transmission rate, delay time and error rate parameters in each link sequence, and associate and store the integrity information of each link with the number of received data packets. Integrate the parameters of all links one by one to obtain the link parameter set;
[0061] Based on the unique link alignment sequence read link transmission parameters, first, the transmission rate, delay time and error rate corresponding to each link in the sequence are collected and recorded in the tabular structure, then the integrity is quantitatively processed according to the ratio of received packets to the number of packets, for example, a link receives 990 packets and should receive 1000 packets, then the integrity is 99%, then the delay time is extracted through the transmission record, for example, the delay time of link 1 is 0.8s, the delay time of link 2 is 1.2s, and the error rate is calculated by the ratio of the number of error packets to the total number of packets, for example, the error rate of link 1 is 0.05%, and the error rate of link 2 is 0.12%, by recording the integrity, delay time and error rate, a link parameter set is constructed and stored in the system parameter table, ensuring that the parameters of each link are arranged according to the integrity, delay and error rate dimensions, and through actual application examples, a scientific research network link records the integrity of 99.2%, the delay of 0.7s and the error rate of 0.06% when collecting data, the link parameters are recorded in the system parameter set at the same time as other links in the same batch, see the example parameter record shown in Table 2, and finally a link parameter set is generated.
[0062] Table 2: Link parameter record table
[0063]
[0064] As shown in Table 2, the parameters are obtained by collection and calculation and stored according to the link number.
[0065] S212: Based on the link parameter set, the integrity, timeliness and accuracy of each dimension are compared with the data quality threshold, and the qualified or unqualified state is recorded after comparison, and the link state judgment result is obtained;
[0066] According to the link parameter set, the integrity, timeliness and accuracy of three dimensions are compared with the preset data quality threshold, and in actual execution, the integrity threshold is set to 99%, the timeliness threshold is set to 1s, and the accuracy threshold is set to 0.1%, then each link parameter is read and compared one by one, if the integrity ≥ 99%, the delay ≤ 1s, and the error rate ≤ 0.1%, it is marked as qualified, if any condition is not met, it is marked as unqualified, for example, in Table 2, L1 integrity 99.2% delay 0.7s error rate 0.06% meets all conditions and is marked as qualified, L2 integrity 98.9% does not meet the threshold and is marked as unqualified, L3 meets the conditions and is marked as qualified, the state array [qualified, unqualified, qualified] is arranged through the comparison result, the link state record is completed, and the link state judgment result is obtained.
[0067] S213: According to the link state judgment result, when the state mark is qualified, the link activation operation is performed to push data to the scientific research achievement storage area, when the state mark is unqualified, the link standby state is maintained without pushing data, the joint evaluation is carried out, the qualified link items are sorted and arranged in order, and the link activation sequence is generated;
[0068] According to the link state judgment result, the link marked as qualified is executed to push the corresponding data to the scientific research achievement storage area, and the link marked as unqualified is kept in standby state without pushing. In the operation process, the link state judgment value is associated with the link parameter set value to form a state parameter mapping table. Through the mapping table, the link numbers that meet the conditions are screened out and arranged in order. For example, under the state array [qualified, unqualified, qualified], only L1 and L3 links enter the activation list, and L2 is excluded. Finally, the activation link set arranged in order by number is sorted out, and the link activation sequence is finally generated.
[0069] Please refer to Figure 4 The specific steps of S3 are as follows:
[0070] S311: Based on the scientific research achievement data in the link activation sequence, the reference record of each achievement is detected, the achievement number is read one by one, the reference log item is associated, the reference times are stored and marked with the corresponding achievement number, and the reference record frequency data is obtained;
[0071] Based on the scientific research achievement data in the link activation sequence, first, the order of each achievement number is detected, all reference log items related to the number are extracted, the reference source and time node contained in the log are recorded one by one, and the multiple references under different time nodes are integrated in the reference collection corresponding to the single achievement number. For example, the achievement with number R101 is recorded three times in the source S1, S2 on January 5, January 9 and January 12, and the three records are arranged in ascending order of time and stored in the collection. There is only one reference source S3 on January 6 for the achievement with number R102, and a single item is recorded. Then, the frequency of each achievement is counted according to the length of the collection, the frequency value is bound with the achievement number, and stored in the record table. Through such read and sorting operation one by one, the reference data set is formed, for example, the reference times of number R101 is 3, the reference times of number R102 is 1, and the reference times of number R103 is 5. All frequency records are in table 3 for subsequent operation calling, see the example shown in table 3. Finally, the reference record frequency data is obtained.
[0072] Table 3: Frequency count table
[0073]
[0074] As shown in table 3, the reference times are summarized according to the number of log records.
[0075] S312: According to the cited record frequency data, compare the frequency value of each achievement with the high-cited threshold value one by one, if the frequency is greater than the high-cited threshold value, mark it as high-cited, otherwise mark it as ordinary, and get the frequency status record;
[0076] Call the cited record frequency data, and compare the frequency of each scientific research achievement with the high-cited threshold value one by one. In this process, first, according to the scientific research achievement frequency data collected by the experiment, the frequency of the top 1% of papers is determined to be 50 to 72 times, and the high-cited threshold value 50 times is set as the comparison standard. Then read the frequency value of each achievement in Table 3 one by one, and compare each frequency value with the threshold value. When the number of citations is greater than 50, mark it as high-cited, otherwise mark it as ordinary. For example, the frequency of R101 is 3, which is less than 50, and is marked as ordinary. The frequency of R103 is also less than 50, and is marked as ordinary. To verify the distinction, introduce R201 with a frequency of 60, which is greater than 50 and is marked as high-cited. Write the marking results of all numbers into the state array, such as [ordinary, ordinary, high-cited], and add the index corresponding bit in the state array. The index 0 corresponds to R101, the index 1 corresponds to R103, and the index 2 corresponds to R201. The index information will be stored with the state result to facilitate subsequent operations. Finally, arrange the number and state into a table form for state registration. See Table 4 for details. Get the frequency status record.
[0077] Table 4: Frequency status table
[0078]
[0079] As shown in Table 4, the state is directly determined by comparing the number of citations with the threshold value 50.
[0080] S313: According to the frequency status record, perform high-performance storage partition write action on high-cited achievements and record index information, and perform ordinary storage area aggregation write action on ordinary achievements, and establish a scientific research achievement partition list;
[0081] According to the record of the frequency of citations, the high-cited research results are extracted to perform the high-performance storage partition write operation, and the index information of each high-cited research result is recorded for storage and retrieval. During the operation, the state table 4 is read, the R201 with the state of high-cited is screened into the high-performance storage write list, and the R101 and R103 with the state of ordinary are classified into the ordinary storage write list. In the write operation, the high-performance storage list numbers are arranged in ascending order of index to form a sequence [R201], and the ordinary storage list numbers are arranged in ascending order of index to form a sequence [R101, R103]. Then, two list storage instruction sets are generated, each of which contains a number, a state, and a write partition position information such as P1 representing a high-performance storage area and P2 representing an ordinary storage area. The example storage instructions are {R201, high-cited, P1}, {R101, ordinary, P2}, and {R103, ordinary, P2}. Finally, all instruction sets are integrated into a research result partition list, which records the complete mapping of each research result number corresponding to the storage area and index information, and establishes the research result partition list.
[0082] Please refer to Figure 5 The specific steps of S4 are as follows:
[0083] S411: Based on the research result partition list, the data of each research result number and the corresponding storage position in the partition are obtained, the transmission state record and time node information of each number in the transmission log are read, the storage record of the corresponding number is combined with the storage record of the corresponding number, and the time and state fields of the same number in the two data sources are compared one by one. Each number comparison result is arranged into a state comparison item to obtain a state consistency comparison record.
[0084] Based on the research achievement partition list, the research achievement number and the corresponding storage partition number information are extracted in turn, the transmission state value and the end time recorded in the transmission log of each number are called, and the transmission state and the transmission completion time are marked respectively, and then the storage state and the write completion time identified in the storage record of the number are called, the number-transmission log and number-storage record double table structure is constructed, the number is aligned as the primary key, whether there is consistency difference between the two states is analyzed by cross comparison of fields, according to the state code rule, the transmission state "complete" is assigned as 1, and the "failure" is assigned as 0, the storage state rule is consistent, taking sample number R301 as an example, its transmission state is "complete" (1), and the storage state is "failure" (0), the states are different, which constitutes a logical inconsistency item, and sample number R302 is "complete" (1) in transmission and storage state, which indicates that the states are consistent, in order to further verify the accuracy of comparison, the state consistency logical judgment and the time deviation are combined to review, the difference value between the transmission end time and the write completion time is extracted as the delay index, if the time interval exceeds the tolerance interval, it is also recorded as an abnormality, the maximum allowed delay threshold is set as 60 milliseconds, the transmission delay of number R301 is 120 milliseconds, the write delay is 130 milliseconds, and the delay difference is 10 milliseconds, which is less than the threshold, and the states are inconsistent; The delay difference of number R302 is 1 millisecond, and the states are consistent; Number R303 fails to transmit (0) and succeeds to store (1), and the states are inconsistent; Number R304 has a delay of 8 milliseconds, and the states are consistent; Number R305 fails to transmit and store, and the states are consistent.
[0085] Table 5: Partition consistency test example table
[0086]
[0087] As shown in Table 5, sample state information and corresponding time delay values are completely collected and listed, after the field extraction operation and the construction of field relationship and flag bit, the state consistency comparison record is generated.
[0088] S412: According to the state consistency comparison record, whether the state fields corresponding to all achievement numbers are consistent is judged, if consistent, assign 1, if inconsistent, assign 0, use the formula:
[0089] ;
[0090] The consistency deviation influence value is obtained by operation, wherein, represents the number the state flag in the transmission log, the numerical value is 1 or 0, indicating that the state is complete and failed, represents the number the state flag in the storage record, the rule is consistent, is the number The network transmission latency, measured in milliseconds, is obtained from the time difference between the start and end of the transmission. Number The write latency, measured in milliseconds, is obtained from the difference between the request and the actual write completion time. The value represents the impact of state consistency deviation, expressing the proportion of time influence under the overall state comparison. The larger the value, the more obvious the consistency deviation.
[0091] Based on the state consistency comparison records, a combined analysis is performed on the state differences and time delay values for each sample number. The index of the number is... The transmission status is Storage status is The state difference is Further call the transmission delay of the sample number With write latency Construct the time perturbation adjustment factor structure The sample consistency deviation contribution value is obtained by combining it with the state difference value, and finally the average value is used to obtain the consistency deviation impact value. The specific calculations are as follows:
[0092] Sample R301: , , ;
[0093] ;
[0094] ;
[0095] ;
[0096] Contribution value ;
[0097] Sample R302: Contribution value = 0;
[0098] Sample R303: , , ;
[0099] ;
[0100] ;
[0101] ;
[0102] Contribution value ;
[0103] Samples R304 and R305 have the same status, and their contribution value is 0.
[0104] Finally:
[0105] ;
[0106] The result shows that the consistency deviation influence value of the current sample group is 0.3823, which is used for subsequent judgment whether there is a need to retransmit the numbering state.
[0107] The consistency deviation influence value is a comprehensive index for measuring the coupling degree of state logical consistency and physical delay fluctuation of scientific research achievements in transmission and storage process. The value combines the consistency difference between transmission state and storage state, and the delay amplitude and symmetry of each numbering in network transmission and partition writing link, reflects the data link reliability deviation degree caused by state inconsistency or time delay in the system. When the value is closer to 1, it means that there are a large number of state inconsistencies and high delay intensity in the numbering sample, indicating that the system data consistency risk is larger. When the value is close to 0, it means that the transmission and writing state consistency is good and the delay distribution is stable, and the overall link structure is reliable, so the value can be used as a criterion for dynamically determining whether to trigger the retransmission mechanism, whether to need link compensation control and other key control processes. It is a compound state index that integrates logical deviation and physical interference.
[0108] The operation logic of the formula aims to comprehensively measure the coupling influence of the numbering sample in state logical consistency and transmission delay difference, where represents the state difference of the numbering in the transmission and storage two stages, which is only 1 when the state is inconsistent, otherwise it is 0, which is used to determine whether the numbering is an abnormal state basis, and the right part of the product is the delay disturbance compensation factor, the numerator characterizes the overall delay intensity of the sample in network transmission and storage writing two stages, and adopts the square root of the sum of squares to reflect the nonlinear superposition effect of the two time items, emphasizing the promotion effect of delay absolute intensity on error contribution, and the denominator adds characterizes the delay asymmetry degree between stages, and this additive structure is used to simulate the lowering effect of delay fluctuation and structure separation on consistency comparison, so the fractional form is the delay coupling compensation factor, and the value closer to 1 represents that the transmission and writing delay is concentrated, large and equal, and smaller represents that the time difference is large and the structure is unstable. By multiplying the state difference value, a logical deviation item with time physical influence weight is finally constructed, and the average of all numbering calculation results is used to measure the amplification effect of state inconsistency in the physical network fluctuation in the whole data link.
[0109] S413: Based on the consistency deviation influence value, screening the number marked as an abnormal number where the difference is greater than the determination threshold, and retriggering the transmission command to form a retransmission list for the corresponding number, marking the remaining number as an acknowledgment item and generating an acknowledgment flag bit, generating a chain confirmation index queue for all number corresponding flag bits, and establishing a partition link integrity record;
[0110] Based on the consistency deviation influence value, the number of samples with values exceeding the set reference threshold is screened for retransmission operation. The threshold value is 0.5, which is set based on the average state deviation value and the delay coupling ratio in the statistical distribution median of the typical stable link operation stage. Specifically, the value comes from the concentration interval formed by the proportion of delay fluctuations corresponding to the state consistent samples in the total number of samples. In samples where the transmission state and storage state are consistent, the delay coupling ratio is generally concentrated between 0.2 and 0.4, and rarely exceeds 0.5. Therefore, 0.5, which is slightly higher than the upper limit, is selected as the determination threshold value. This value can be dynamically adjusted according to the fluctuations of link transmission rate or changes in physical partition write load. The determination basis is whether the value is greater than the threshold value. A direct comparison method is used. For example, the index value of sample R301 is 0.9465 and the index value of R303 is 0.9650, both of which are greater than the set threshold value of 0.5, and are determined as abnormal numbers, written into the retransmission list and assigned a retransmission flag bit of 1. The contribution values of R302, R304, and R305 are 0, which are lower than the threshold value, and are marked as confirmation numbers, assigned a confirmation flag bit of 0. The retransmission numbers form a chain task scheduling item, forming a structure like {R301, 1}, {R303, 1}, {R302, 0}, {R304, 0}, {R305, 0}, which is used for task scheduling controller scheduling logic execution in subsequent data back mechanism, and finally completes the registration of data flag index queue under control structure, establishing a partition link integrity record.
[0111] Please refer to Figure 6 , the specific steps of S5 are:
[0112] S511: Based on the partition link integrity record, extract the achievement number set corresponding to each storage area, and call the ownership label and partition identification of each number in the link data. Statistically analyze the relationship between the number of data numbers in each storage area and the total number of corresponding storage area, and obtain the partition data ownership ratio;
[0113] Based on the storage area structure information registered in the partition link integrity record, the achievement number set corresponding to each storage area is extracted, the ownership tag field of the number is read in the set, and the field is marked as "personal data" by judging whether it is "personal data". The number marked as "personal data" is recorded as a subset, and the proportion of the number of subsets to the total number of numbers in the storage area is counted. In actual operation, if the total number of Z1 area is 120, and the number of personal data is 48, the ownership ratio is 48 divided by 120, which is 0.40. The ownership ratio of Z2 area is 41 divided by 105, which is 0.39. The ownership ratio of Z3 area is 62 divided by 98, which is 0.63. The ownership ratio of Z4 area is 59 divided by 143, which is 0.41. The ownership ratio of Z5 area is 38 divided by 112, which is 0.34. The proportion values of all storage areas are uniformly listed to construct the following structure:
[0114] Table 6: Data ownership ratio table of each storage area:
[0115]
[0116] As shown in Table 6, the ownership ratio of each partition has been calculated, and the ownership ratios of Z3 and Z4 are both greater than 0.40, which provides basic data for subsequent compliance judgment operation, and obtains the partition data ownership ratio.
[0117] S512: According to the partition data ownership ratio, set the ownership ratio judgment reference value, compare the ownership ratio corresponding to each storage area with the reference value, and identify each partition according to the comparison result. The partition with a ratio higher than the reference value is marked as non-compliant, and the partition with a ratio lower than or equal to the reference value is marked as compliant, and a data distribution compliance identification sequence is generated;
[0118] The partition data ownership ratio is called, and the proportion reference value is set to 0.40 according to the personal data storage proportion limit specified in the general data protection regulation. The ownership ratio in Table 6 is compared with the reference value one by one. If the proportion of a certain storage area is higher than 0.40, it is marked as "non-compliant", otherwise it is marked as "compliant". For example, Z1 area is 0.40, which is considered as compliant. Z2 area is 0.39, which is less than the reference value, and is considered as compliant. Z3 area is 0.63, which is higher than 0.40, and is considered as non-compliant. Z4 area is 0.41, which is considered as non-compliant. Z5 area is 0.34, which is still considered as compliant. The result structure after comparison is as follows:
[0119] Table 7: Storage area compliance judgment result table:
[0120]
[0121] Referring to Table 7, Z3 and Z4 are marked as non-compliant areas, and the remaining storage areas meet the reference value condition. The above judgment identification is integrated to generate a data distribution compliance identification sequence.
[0122] S513: According to the data distribution compliance identification sequence, the information content is reorganized for the data partition marked as non-compliant, and the storage structure is processed according to the record state in the cache structure, the update operation of the storage area number distribution is completed, and the high vocational college big data full collection record is obtained.
[0123] According to the data distribution compliance identification sequence, the data structure extraction processing is performed on the storage areas Z3 and Z4 marked as "non-compliant", and in the Z3 area, all achievement number corresponding cache occupancy rate fields, number ownership target fields, number transmission frequency fields and generation time fields are obtained, sample numbers R304, R309 and R310 are extracted in numbers R301 to R312 as processing objects, corresponding cache occupancy rate values are 82%, 91% and 87%, daily transmission frequency is 3 times, 2 times and 4 times, and generation time is earlier than the current date by more than 30 days, which meets the reconstruction condition, and in the Z4 area, numbers R411 to R425 are extracted, numbers R416, R421 and R423 with cache occupancy rate higher than 75% are marked, their transmission frequency is 3 times, 2 times and 3 times, and the generation time is earlier than the beginning of the month, which is recorded as the object of arrangement, the above numbers meeting the condition are added to the arrangement index structure and the state flag bit is registered, and the record content is established as follows:
[0124] Table 8 Non-compliant area reconstruction number registration table:
[0125]
[0126] As shown in Table 8, the above numbers meet the comprehensive conditions of high cache occupancy, high frequency and early generation time, and have been locked as reconstruction targets and generated index table, and the final arrangement structure is used for archiving operation to obtain the high vocational college big data full collection record.
[0127] The above is only a preferred embodiment of the present application, and does not limit the form of the present application, any skilled person in the art can use the above disclosed technical content to make changes or modifications as equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made according to the technical essence of the present application to the above embodiments without departing from the technical solution content of the present application, still belongs to the protection scope of the technical solution of the present application.
Claims
1. A method for collecting full data from higher vocational colleges based on big data, characterized in that: Includes the following steps: S1: Collect source system codes and time series from higher vocational colleges, concatenate them to form a unique identifier string, and perform duplicate checks on each item to obtain a unique link comparison sequence; S2: Based on the unique link comparison sequence, when the state meets the data quality threshold condition, the link activation action is executed to push the data to the scientific research results storage area; when the state does not meet the condition, the link is kept in standby, and the link activation sequence is obtained. S3: Based on the link activation sequence, detect citation records and calculate the citation frequency of each result. When the citation frequency exceeds the high citation threshold, execute the high-performance storage partition write operation. When the citation frequency is lower than or equal to the high citation threshold, execute the ordinary storage area aggregation write operation to obtain a list of scientific research results partitions. S4: Based on the research results partition list, compare and check the consistency of the status. When the status is different, execute the retransmission mechanism. When the status is consistent, record the confirmation flag to obtain the complete record of the partition link. S5: Based on the complete record of the partition link, determine whether the data distribution of each storage area conforms to the data minimization principle. If there is a deviation in the distribution, perform cache reconstruction. If there is no deviation, maintain the current state and obtain the full collection record of big data from higher vocational colleges. The unique link comparison sequence includes unique identifier information, cache update status, and duplicate verification results. The link activation sequence includes data quality status, transmission trigger flag, and push storage instruction. The scientific research achievement partition list includes high-performance storage index, ordinary storage aggregation record, and citation frequency classification results. The partition link complete record includes status consistency identifier, retransmission confirmation record, and log comparison results. The vocational college big data full collection record includes storage area distribution ratio data, cache reconstruction information, and link storage status.
2. The method for collecting full data from higher vocational colleges based on big data according to claim 1, characterized in that, The specific steps for obtaining the unique link alignment sequence are as follows: S111: Based on the source system code and time series of higher vocational colleges received by the acquisition gateway, the code and time series are concatenated to form a unique identifier string, and the unique identifier string is compared with the identifier sequence in the cache list item by item to obtain the difference comparison value; S112: Based on the difference comparison value, compare the identifier sequence of the cache list. When the comparison results are inconsistent, transmit the received data. During this process, update the cache list order according to the comparison difference and delete the earliest identifier to obtain the cache update sequence. S113: Based on the cache update sequence and the difference comparison value, a joint judgment is made. When the comparison results are completely consistent, the current data is discarded and a unique link comparison sequence is generated.
3. The method for collecting full data from higher vocational colleges based on big data as described in claim 1, characterized in that, The specific steps for obtaining the link activation sequence are as follows: S211: Based on the unique link comparison sequence, read the link transmission parameters, extract the transmission rate, delay time and error rate parameters in each link sequence, and associate the integrity information of each link with the number of received data packets for storage. Integrate the parameters of all links one by one to obtain the link parameter set. S212: Based on the set of link parameters, each dimension of integrity, timeliness, and accuracy is judged against the data quality threshold, and the qualified or unqualified status is recorded after comparison to obtain the link status judgment result. S213: Based on the link status judgment result, when the status is marked as qualified, the link activation action is executed to push the data to the scientific research results storage area; when the status is marked as unqualified, the link is kept in standby and no data is pushed. A joint evaluation is performed, qualified link items are sorted and arranged in order to generate a link activation sequence.
4. The method for collecting full data from higher vocational colleges based on big data as described in claim 1, characterized in that, The specific steps for obtaining the research achievement zoning list are as follows: S311: Based on the scientific research results data in the link activation sequence, detect the citation record of each result, read the result number one by one and associate it with the citation log entry, store the citation count and label the corresponding result number to obtain the citation record frequency data; S312: Based on the citation record frequency data, compare the frequency value of each result with the high citation threshold one by one. If the frequency is greater than the high citation threshold, mark it as highly cited; otherwise, mark it as ordinary, and obtain the citation frequency status record. S313: Based on the citation frequency status record, perform high-performance storage partition writing for highly cited results and record index information, perform ordinary storage area aggregation writing for ordinary results, and establish a list of scientific research results partitions.
5. The method for collecting full data from higher vocational colleges based on big data according to claim 1, characterized in that, The specific steps for obtaining the complete record of the partitioned link are as follows: S411: Based on the scientific research achievement partition list, obtain the achievement number and corresponding storage location data within the partition, read the transmission status record and time node information of each number in the transmission log, combine the entry flag and writing time of the corresponding number in the storage record, compare the time and status fields of the same number in the two data sources one by one, organize the comparison results of each number into a status comparison item, and obtain the status consistency comparison record. S412: Based on the status consistency comparison record, determine whether the status fields corresponding to all result numbers are consistent. If they are consistent, assign a value of 1; if they are inconsistent, assign a value of 0. Calculate and obtain the consistency deviation impact value. S413: Based on the consistency deviation impact value, select the numbers whose differences are greater than the judgment threshold and mark them as abnormal numbers. Re-trigger the transmission command for the corresponding numbers to form a retransmission list. Mark the remaining numbers as confirmation items and generate confirmation flag bits. Generate a chained confirmation index queue for all number corresponding flag bits and establish a complete record of the partition link.
6. The method for collecting full data from higher vocational colleges based on big data according to claim 1, characterized in that, The specific steps for obtaining the full collection and record of big data from higher vocational colleges are as follows: S511: Based on the complete record of the partition link, extract the set of result numbers corresponding to each storage area, and call the ownership tag and partition identifier of each number in the link data, calculate the relationship between the number of data numbers in each storage area and the total number of numbers in the corresponding storage area, and obtain the partition data ownership ratio; S512: Based on the partition data ownership ratio, set an ownership ratio judgment benchmark value, compare the ownership ratio corresponding to each storage area with the benchmark value, and mark each partition according to the comparison result. Partitions with a ratio higher than the benchmark value are marked as non-compliant, and partitions with a ratio lower than or equal to the benchmark value are marked as compliant, generating a data distribution compliance identifier sequence. S513: Based on the data distribution compliance identifier sequence, for the data partitions marked as non-compliant, the information content is organized, and the storage structure is sorted out according to the record status of the number in the cache structure, completing the update operation of the storage area number distribution, and obtaining the full collection record of big data of higher vocational colleges.
Citation Information
Patent Citations
Dual-park data verification method and device
CN113391956A
Carbon emission data management method and device based on block chain
CN120492543A