High vocational college total data collection method based on big data
By using a big data-based full-volume data aggregation method, the problems of data redundancy and omission in the data aggregation of higher vocational colleges were solved, and the accurate differentiation and real-time verification of data were achieved, thereby improving data utilization efficiency and the rationality of storage structure.
Patent Information
- Application Number
- CN202511339956.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-19
AI Technical Summary
In existing technologies, data redundancy and omissions exist in the data collection process of higher vocational colleges, format conversion lacks real-time verification, making it difficult to guarantee the consistency of transmission links, requiring frequent manual intervention, resulting in low data utilization efficiency, and relational databases are unable to meet the storage requirements of large-scale scientific research data.
By adopting a full-scale data collection method based on big data, and through unique identifier link comparison, status threshold judgment, reference frequency detection and partitioned storage, we can achieve accurate data differentiation, real-time verification and dynamic monitoring, ensuring the accuracy and integrity of data push. We also improve link stability and optimize storage structure through comparison and retransmission mechanisms.
It improves the retrieval and storage efficiency of high-value data, ensures the accuracy and integrity of data transmission, achieves reasonable partitioning and compliant storage, optimizes data utilization efficiency, and avoids storage imbalance problems.
Smart Images

Figure CN120821735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data collection technology, and in particular to a method for collecting full data of higher vocational colleges based on big data. Background Art
[0002] The field of data collection technology involves the collection, organization, conversion, and integration of raw data scattered across different sources, formats, and systems to achieve unified management and efficient use of data. This includes the diversification of data collection methods, the formulation of data cleaning rules, data standardization, and the centralized storage and management of multi-source data, using centralized databases to complete the collection and storage. Among them, the full-volume data collection method for higher vocational colleges refers to the collection and integration of multi-dimensional business data such as students, teachers, courses, scientific research, and logistics within higher vocational colleges. This is usually done through manual batch imports, relying on spreadsheets and business management system export files periodically provided by various departments, and using rule scripts to perform format conversion and classification aggregation item by item, and then using relational databases for unified storage.
[0003] In the existing data collection process of colleges and universities, batch import and periodic file organization are relied upon in most cases. Data redundancy and omissions are prone to occur when collecting multi-source data across departments. The format conversion and classification summary process lacks real-time verification methods, making it difficult to ensure consistency in the transmission link. In multi-dimensional business data processing scenarios, manual intervention frequently increases processing delays. The lack of dynamic monitoring of citation records makes it impossible to effectively distinguish high-value scientific research results. The centralized storage model of relational databases is difficult to meet the proportion control requirements when dealing with large-scale scientific research data distribution, resulting in reduced data utilization efficiency and prone to storage imbalance problems. Summary of the Invention
[0004] The purpose of the present invention is to solve the shortcomings of the existing technology and propose a full data collection method for higher vocational colleges based on big data.
[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solution: a method for collecting full data of higher vocational colleges based on big data, comprising the following steps: S1: Collect the source system codes and time series of higher vocational colleges, splice them into a unique identification string, and perform duplication verification item by item to obtain a unique link comparison sequence; S2: Based on the unique link comparison sequence, when the status meets the data quality threshold condition, execute the link activation action to push the data to the scientific research results storage area; when the status does not meet the condition, keep the link standby, and obtain the link activation sequence; S3: Based on the link activation sequence, citation records are detected and the citation frequency of each achievement is calculated. When the citation frequency exceeds the high citation threshold, a high-performance storage partition write action is executed. When the citation frequency is lower than or equal to the high citation threshold, a normal storage area aggregate write action is executed to obtain a scientific research achievement partition list; S4: Based on the scientific research results partition list, check the status consistency, execute the retransmission mechanism when there is a difference in the check status, record the confirmation mark when the status is consistent, and obtain a complete record of the partition link; S5: Based on the complete record of the partition link, determine whether the data distribution of each storage area complies with the principle of data minimization. If there is a deviation in the distribution, perform cache reconstruction. If there is no deviation, maintain the current state and obtain the full collection record of big data of higher vocational colleges.
[0006] As a further solution of the present invention, the unique link comparison sequence includes identification uniqueness information, cache update status, and duplicate verification results; the link activation sequence includes data quality status, transmission trigger flag, and push storage instructions; the scientific research results partition list includes high-performance storage index, ordinary storage aggregation record, and citation frequency classification results; the partition link complete record includes status consistency identification, retransmission confirmation record, and log comparison result; the full collection record of big data of higher vocational colleges includes storage area distribution ratio data, cache reconstruction information, and link storage status.
[0007] As a further solution of the present invention, the specific steps of obtaining the unique link comparison sequence are: S111: Based on the source system code and time sequence of the higher vocational college received by the acquisition gateway, the code and time sequence are concatenated to form a unique identification string, and the unique identification string is compared with the identification sequence in the cache list item by item to obtain a difference comparison value; S112: comparing the identifier sequence of the cache list according to the difference comparison value, transmitting the received data when the comparison result is inconsistent, and updating the cache list order according to the comparison difference and deleting the earliest identifier in the process to obtain a cache update sequence; S113: Perform a joint judgment based on the cache update sequence and the difference comparison value, discard the current data when the comparison results are completely consistent, and generate a unique link comparison sequence.
[0008] As a further solution of the present invention, the link activation sequence is specifically obtained by: S211: Reading link transmission parameters based on the unique link comparison sequence, extracting the transmission rate, delay time, and error rate parameters in each link sequence, and associating and storing the integrity information of each link with the number of received data packets, integrating the parameters of all links one by one to obtain a link parameter set; S212: Based on the link parameter set, each dimension of integrity, timeliness, and accuracy is judged against the data quality threshold, and after comparison, a qualified or unqualified status is recorded to obtain a link status judgment result; S213: Based on the link status judgment result, when the status is marked as qualified, the link activation action is executed to push the data to the scientific research results storage area; when the status is marked as unqualified, the link is kept in standby mode and no data is pushed; a joint evaluation is performed and qualified link entries are sorted and arranged in order to generate a link activation sequence.
[0009] As a further solution of the present invention, the specific steps for obtaining the scientific research achievement partition list are: S311: Based on the scientific research achievement data in the link activation sequence, detect the citation record of each achievement, read the achievement number one by one and associate it with the citation log entry, count and store the citation times and mark the corresponding achievement number to obtain citation record frequency data; S312: Based on the citation record frequency data, the frequency value of each achievement is compared with the high citation threshold one by one. If the frequency is greater than the high citation threshold, it is marked as highly cited; otherwise, it is marked as normal, thereby obtaining a citation frequency status record; S313: Based on the citation frequency status record, a high-performance storage partition write action is performed on highly cited achievements and index information is recorded; an ordinary storage area aggregation write action is performed on ordinary achievements, and a partition list of scientific research achievements is established.
[0010] As a further solution of the present invention, the specific steps for obtaining the complete record of the partition link are: S411: Based on the scientific research achievement partition list, obtain the achievement number and corresponding storage location data in the partition, read the transmission status record and time node information of each number in the transmission log, combine the storage flag and write time of the corresponding number in the storage record, compare the time and status fields of the same number in the two data sources one by one, organize each number comparison result into a status comparison item, and obtain a status consistency comparison record; S412: According to the status consistency comparison record, determine whether the status fields corresponding to all achievement numbers are consistent or not. If they are consistent, assign a value of 1; if they are inconsistent, assign a value of 0; and calculate to obtain the consistency deviation impact value.
[0011] S413: Based on the consistency deviation impact value, the numbers whose differences are greater than the judgment threshold are screened and marked as abnormal numbers, and the transmission command of the corresponding numbers is re-triggered to form a retransmission list, and the remaining numbers are marked as confirmation items and confirmation flags are generated. A chain confirmation index queue is generated for the flags corresponding to all numbers to establish a complete record of the partition link.
[0012] As a further solution of the present invention, the specific steps for obtaining the full collection of big data records of higher vocational colleges are as follows: S511: Based on the complete record of the partition link, extract the set of achievement numbers corresponding to each storage area, call the attribution label and the partition identifier of each number in the link data, calculate the relationship between the number of data numbers in each storage area and the total number of numbers in the corresponding storage area, and obtain the partition data attribution ratio; S512: Based on the partition data attribution ratio, a reference value for determining the attribution ratio is set, and the attribution ratio corresponding to each storage area is compared with the reference value. Each partition is identified according to the comparison result, and partitions with a ratio higher than the reference value are marked as non-compliant, and partitions with a ratio lower than or equal to the reference value are marked as compliant, thereby generating a data distribution compliance identification sequence; S513: According to the data distribution compliance identification sequence, for the data partitions marked as non-compliant, the information content is sorted out, and the storage structure is sorted out according to the record status of the number in the cache structure, completing the update operation of the storage area number distribution and obtaining the full collection record of the big data of higher vocational colleges.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, accurate distinction of source data in the receiving link is achieved through link comparison of unique identification, which avoids duplicate information from occupying resources. Real-time judgment of data quality during transmission is achieved through the setting of status thresholds, which ensures the accuracy and completeness of data push. Differentiated partition storage of scientific research results is achieved through dynamic detection of citation frequency, which effectively improves the retrieval and storage efficiency of high-value data. The integrity of the transmission status is ensured through comparison and inspection, and the retransmission mechanism is triggered in the case of differences, which improves the stability of the link. The principle of data minimization is implemented through proportional control of the storage area distribution, which ensures that the data meets the compliance requirements and optimizes the storage structure, achieving the overall effect of efficient transmission, reasonable partitioning, and compliant storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flow chart of the main steps of the present invention; Figure 2 The flowchart of obtaining the unique link comparison sequence of the present invention; Figure 3 Obtaining a flow chart for the link activation sequence of the present invention; Figure 4 Obtain a flow chart for the partition list of the scientific research results of the present invention; Figure 5 A flowchart for obtaining a complete record of a partition link of the present invention; Figure 6 This is a flowchart for obtaining the full collection and record of big data from higher vocational colleges in the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0016] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0017] See also Figure 1 The method for collecting full data of higher vocational colleges based on big data includes the following steps: S1: Obtain the source system code and time series of the higher vocational college received by the collection gateway, concatenate the two to form a unique identification string, and perform a duplication check against the identification sequence in the cache list item by item. If the comparison results are inconsistent, perform data transmission and update the cache sort to delete the earliest identification. If the comparison results are completely consistent, discard the current data to obtain a unique link comparison sequence. S2: Read the link transmission parameters based on the unique link comparison sequence and determine whether the sequence status triggers data transmission under the data quality threshold (covering dimensions such as integrity (≥99%), timeliness (delay ≤1s), and accuracy (error rate ≤0.1%)). When the status meets the threshold conditions, the link activation action is executed to push the data to the scientific research results storage area. When the status does not meet the conditions, the link remains in standby mode to obtain the link activation sequence. S3: Based on the scientific research results data in the link activation sequence, the citation records are detected and the citation frequency of each result is calculated. The calculated result is compared with the high citation threshold (a dynamic threshold is set based on the citation frequency of the top 1% of papers). When the frequency exceeds the high citation threshold, a high-performance storage partition write action is executed and the index information is recorded. When the frequency is lower than or equal to the high citation threshold, an ordinary storage area aggregation write action is executed to obtain a scientific research results partition list. S4: Based on the scientific research results partition list, the transmission log and storage records are compared to check the status consistency. When the check status is different, the retransmission mechanism is executed. When the status is consistent, the confirmation mark is recorded to obtain a complete record of the partition link; S5: Based on the complete record of the partition link, determine whether the data distribution of each storage area complies with the principle of data minimization (complying with the storage ratio limit of the General Data Protection Regulation). If there is a deviation in the distribution, execute the cache reconstruction protocol. If there is no deviation in the distribution, maintain the current link and storage status and obtain the full collection record of the big data of higher vocational colleges.
[0018] The unique link comparison sequence includes identification uniqueness information, cache update status, and duplicate verification results. The link activation sequence includes data quality status, transmission trigger flag, and push storage instructions. The scientific research results partition list includes high-performance storage indexes, ordinary storage aggregation records, and citation frequency classification results. The partition link complete record includes status consistency identification, retransmission confirmation records, and log comparison results. The full collection record of big data of higher vocational colleges includes storage area distribution ratio data, cache reconstruction information, and link storage status.
[0019] See also Figure 2 , the specific steps of S1 are: S111: Based on the source system code and time sequence of the higher vocational college received by the acquisition gateway, the code and time sequence are concatenated to form a unique identification string, and the unique identification string is compared with the identification sequence in the cache list item by item to obtain a difference comparison value; Based on the source system code and time series of higher vocational colleges received by the acquisition gateway, the coded data is extracted in the form of a string, and the time series is formatted to form a timestamp that increases by milliseconds. In the refinement process, the source system codes such as A101, A102, etc. and the timestamps such as 1722300000, 1722300001 are first concatenated into unique identification strings A101-1722300000, A102-1722300001, etc., and then these unique identification strings are compared one by one with the existing identifications in the cache list. During the comparison process, the characters are scanned bit by bit to determine whether there are any differences. If there are any differences, If there is a difference, the difference index is recorded and the difference content is stored. For example, when the cache list contains A101-1722299999, the difference index is the last position 1, and the difference value stored is 1. Then, each difference index is differentiated and divided into three intervals according to the absolute value of the difference index: less than 2, between 2 and 4, and greater than 4. The difference level can be divided by numerical judgment. For example, if the difference value between A101-1722300000 and A101-1722299998 is 2, it falls into the middle interval. See the example comparison data shown in Table 1. Finally, all difference indexes and corresponding difference levels are recorded through the above comparison operation to obtain the difference comparison value.
[0020] Table 1 Difference comparison data table:
[0021] As shown in Table 1, the difference index is obtained by bit-by-bit scanning, and the difference level is directly determined by interval division.
[0022] S112: Compare the identifier sequences of the cache list based on the difference comparison value. If the comparison results are inconsistent, transmit the received data. In this process, update the cache list order based on the comparison difference and delete the earliest identifier to obtain a cache update sequence. According to the difference comparison value, all the identification sequences in the cache list are compared. In the refinement process, the entries marked as greater than 4 intervals in the difference comparison value are first selected, and the received data corresponding to these entries are directly transmitted to the data receiving end. For example, A102-1722300001 in Table 1 meets the conditions. After the transmission is completed, the cache list is updated immediately, and the latest received identification is inserted into the cache list head. At the same time, the earliest identification such as A102-1722299997 is deleted according to the original cache sequence order. The difference entries in the middle interval are For example, if A103-1722300003 is updated in sequence, but the old identifier is retained for subsequent comparison processing. If the interval entries are less than 2, the existing cache is retained without updating. During this process, the cache sequence index table is regenerated for each update, such as [0: A102-1722300001, 1: A101-1722300000], and then the new sequence index table is sorted and checked. If the index order is disordered, it is re-sorted and adjusted to ensure that the final cache list order is continuous and complete, and finally the cache update sequence is obtained.
[0023] S113: Perform a joint judgment based on the cache update sequence and the difference comparison value. When the comparison results are completely consistent, the current data is discarded to generate a unique link comparison sequence; According to the cache update sequence, when the comparison results are completely consistent, the current data is directly discarded. In the refinement process, the identifiers with difference comparison values less than 2 intervals in the cache update sequence are first screened. These identifiers are marked as discarded items because they are completely consistent with the cache identifiers. Then, the middle and large interval parts of the difference comparison values are associated with the index of the cache update sequence, and a joint judgment table is constructed. The update status and comparison difference of each identifier are judged one by one. For example, if A101-1722300000 is less than 2 intervals and the status is not updated, it is directly discarded. If A103-1722300003 is in the middle interval and the status is updated, it is retained. After marking each processing result one by one in the joint judgment table, all identifiers marked as retained are reorganized into a link data chain in index order. By removing all duplicates and discarded items, the final unique sequence set is formed. This set is the unique link comparison sequence.
[0024] See also Figure 3 , the specific steps of S2 are: S211: Reading link transmission parameters based on the unique link comparison sequence, extracting the transmission rate, delay time, and error rate parameters in each link sequence, and associating and storing the integrity information of each link with the number of received data packets. The parameters of all links are integrated one by one to obtain a link parameter set. Based on the unique link comparison sequence, the link transmission parameters are read. First, the transmission rate, delay time, and error rate corresponding to each link in the sequence are collected and recorded in a tabular structure. Then, the integrity is quantified according to the ratio of the number of received packets to the number of receivable packets. For example, if a link receives 990 packets and should receive 1000 packets, the integrity is 99%. The delay time is then extracted from the transmission record. For example, the delay time of link 1 is 0.8s and the delay time of link 2 is 1.2s. At the same time, the error rate is calculated by the ratio of the number of error packets to the total number of packets. For example, the error rate of link 1 is 0.05% and the error rate of link 2 is 0.12%. Based on the recorded integrity, delay time, and error rate, a link parameter set is constructed and stored in the system parameter table, ensuring that the parameters of each link are arranged according to the integrity, delay, and error rate dimensions by number. Through actual application examples, it can be seen that a scientific research network link records 99.2% integrity, 0.7s delay, and 0.06% error rate when collecting data. The link parameters are entered into the system parameter set at the same time as other links in the same batch. See the example parameter record shown in Table 2, and finally a link parameter set is generated.
[0025] Table 2 Link parameter record table:
[0026] As shown in Table 2, the parameters are obtained through collection and calculation and stored according to the link number.
[0027] S212: Based on the link parameter set, each dimension of integrity, timeliness, and accuracy is judged against the data quality threshold, and after comparison, a qualified or unqualified status is recorded to obtain a link status judgment result; Based on the link parameter set, the three dimensions of integrity, timeliness, and accuracy are judged item by item against the preset data quality threshold. In actual implementation, the integrity threshold is set to 99%, the timeliness threshold is set to 1s, and the accuracy threshold is set to 0.1%. Then, each link parameter is read and compared one by one. During the judgment process, if the integrity is ≥99%, the delay is ≤1s, and the error rate is ≤0.1%, the link is marked as qualified. If any condition is not met, the link is marked as unqualified. For example, in Table 2, L1 has an integrity of 99.2%, a delay of 0.7s, and an error rate of 0.06%, which meets all conditions and is marked as qualified. L2 has an integrity of 98.9% but does not meet the threshold and is marked as unqualified. L3 meets the conditions and is marked as qualified. The comparison results are sorted into a status array [qualified, unqualified, qualified] to complete the link status record and obtain the link status judgment result.
[0028] S213: Based on the link status judgment result, when the status is marked as qualified, the link activation action is executed to push the data to the scientific research results storage area; when the status is marked as unqualified, the link is kept in standby mode and no data is pushed. A joint evaluation is performed and qualified link entries are sorted and arranged in order to generate a link activation sequence; According to the link status judgment result, the activation action is performed on the links marked as qualified and the corresponding data is pushed to the scientific research results storage area. The links marked as unqualified remain in standby state and are not pushed. During the operation, the link status judgment value and the link parameter set value are associated with each other to form a status parameter mapping table. Through this mapping table, the link numbers that meet the conditions are screened out and arranged in sequence. For example, under the status array [qualified, unqualified, qualified], only L1 and L3 links enter the activation list, and L2 is eliminated. Finally, a set of activated links arranged in order of number is sorted out, and a link activation sequence is finally generated.
[0029] See also Figure 4 , the specific steps of S3 are: S311: Based on the scientific research achievement data in the link activation sequence, detect the citation record of each achievement, read the achievement number one by one and associate it with the citation log entry, count and store the citation times and mark the corresponding achievement number to obtain the citation record frequency data; Based on the scientific research results data in the link activation sequence, each result number is first sequentially detected, and all citation log entries related to the number are extracted. The citation sources and time nodes contained in the log are recorded item by item, and multiple citations at different time nodes are integrated into the citation set corresponding to the single result number. For example, the result number R101 is recorded three times in sources S1 and S2 on January 5, January 9, and January 12, respectively. The three records are sorted in ascending time order and stored in the set. For the result number R102, there is only one citation from source S3 on January 6, so a single entry is recorded. Then, the citation frequency of each result is counted according to the set length, and the frequency value is bound to the result number and stored in the record table. Through such item-by-item reading and sorting operations, a citation data set is formed. For example, the number of citations for number R101 is 3, the number of citations for number R102 is 1, and the number of citations for number R103 is 5. All frequencies are recorded in Table 3 for subsequent operation calls. See the example shown in Table 3, and finally the citation record frequency data is obtained.
[0030] Table 3 Citation frequency statistics:
[0031] As shown in Table 3, the number of citations is summarized by the number of log records.
[0032] S312: Based on the citation record frequency data, the frequency value of each achievement is compared with the high citation threshold one by one. If the frequency is greater than the high citation threshold, it is marked as highly cited; otherwise, it is marked as normal, and a citation frequency status record is obtained; The citation record frequency data is called, and the frequency of each scientific research result is compared with the high citation threshold in turn. In this process, the citation frequency range of the top 1% papers in this batch is first determined to be 50 to 72 times based on the citation frequency data of scientific research results collected in the experiment. The high citation threshold of 50 times is dynamically set as the comparison standard. Then, the citation frequency value of each result in Table 3 is read one by one, and each frequency value is directly compared with the threshold. When the number of citations is greater than 50, it is marked as highly cited, otherwise it is marked as ordinary. For example, the comparison result of the frequency 3 of No. R101 with 50 is less than, Marked as normal, number R103 frequency 5 is also less than, marked as normal, to verify the distinction, number R201 frequency 60 is introduced, then greater than 50 is marked as highly cited, all numbered marking results are written into the status array such as [normal, normal, highly cited], and at the same time, the index corresponding bit is added to the status array, status array index 0 corresponds to R101, index 1 corresponds to R103, index 2 corresponds to R201, the index information will be stored with the status result for subsequent operations, and finally the number and status are organized into a table for status registration, as shown in Table 4, to obtain the citation frequency status record.
[0033] Table 4 Citation frequency status table:
[0034] As shown in Table 4, the status is directly determined by comparing the number of citations with the threshold of 50.
[0035] S313: Based on the citation frequency status record, a high-performance storage partition write operation is performed for highly cited achievements and index information is recorded. For ordinary achievements, an ordinary storage area aggregate write operation is performed to establish a scientific research achievement partition list. According to the citation frequency status record, the results number marked as highly cited is extracted to perform the high-performance storage partition write action. At the same time, the index information of each highly cited result number is recorded for storage and retrieval. During the operation, the status table 4 is read, and the R201 with a highly cited status is filtered into the high-performance storage write list, and the R101 and R103 with a normal status are classified into the normal storage write list. In the write action, the high-performance storage list numbers are arranged in ascending order by index to form a sequence [R201], and the normal storage list numbers are arranged in ascending order by index to form a sequence [R201]. Sequence [R101, R103], then generate a storage instruction set from the two lists. Each instruction set contains the number, status, and write partition location information. For example, P1 represents the high-performance storage area and P2 represents the ordinary storage area. Example storage instructions are {R201, highly cited, P1}, {R101, ordinary, P2}, {R103, ordinary, P2}. Finally, all instruction sets are integrated into a scientific research achievement partition list. The scientific research achievement partition list records the complete mapping of each achievement number to the storage area and index information, and establishes a scientific research achievement partition list.
[0036] See also Figure 5 , the specific steps of S4 are: S411: Based on the scientific research achievement partition list, obtain the achievement number and corresponding storage location data in the partition, read the transmission status record and time node information of each number in the transmission log, combine the storage mark and write time of the corresponding number in the storage record, compare the time and status fields of the same number in the two data sources one by one, organize the comparison results of each number into a status comparison item, and obtain a status consistency comparison record; Based on the scientific research results partition list, each result number and its corresponding storage partition number information are extracted in turn, and the transmission status value and end time recorded in the transmission log of each number are retrieved, and marked as transmission status and transmission completion time respectively. Then the storage status and write completion time identified by the number in the storage record are retrieved, and a double table structure of number-transmission log and number-storage record is constructed. The number is aligned as the primary key, and the field cross-comparison is used to analyze whether there is a consistency difference between the two states. According to the status coding rule, the transmission status "completed" is assigned to 1 and "failed" is assigned to 0. The storage status rules are consistent. Taking sample number R301 as an example, its transmission status is "completed", that is, 1, and the storage status is "failed", that is, 0. The states are different, which constitutes a logical inconsistency item, and sample number R3 02 When both the transmission and storage status are "completed" or 1, it means the status is consistent. To further verify the accuracy of the comparison, the status consistency logic judgment and time deviation need to be jointly reviewed, and the difference between the transmission end time and the write completion time is extracted as the delay indicator. If the time interval exceeds the tolerance range, it is also recorded as an abnormality. The maximum allowable delay threshold is set to 60 milliseconds. The transmission delay of No. R301 is 120 milliseconds, the write delay is 130 milliseconds, and the delay difference is 10 milliseconds, which is less than the threshold and the status is inconsistent; the delay difference of No. R302 is 1 millisecond and the status is consistent; No. R303 transmission failed (0), storage completed (1), the status is inconsistent; No. R304 delay is 8 milliseconds, the status is consistent; No. R305 transmission failed, storage failed, the status is consistent.
[0037] Table 5 Partition consistency test example table:
[0038] As shown in Table 5, the sample status information and the corresponding time delay values are completely collected and listed. After the above-mentioned field extraction operation and the construction of field relationships and flag bits, a status consistency comparison record is generated.
[0039] S412: Based on the status consistency comparison record, determine whether the status fields corresponding to all achievement numbers are consistent. If they are consistent, assign a value of 1; if they are inconsistent, assign a value of 0. Use the formula: ; The consistency deviation impact value is obtained by operation, where Representative No. The status flag in the transmission log is a numeric value of 1 or 0, indicating completion or failure. Representative No. The status flags in the storage records follow the same rules. Number The network transmission delay time, in milliseconds, is obtained by the difference between the start and end time of the transmission. Number The write latency, in milliseconds, is obtained by the difference between the request and the actual write completion time. is the state consistency deviation impact value, which expresses the proportion of time impact under the overall state comparison. The larger the value, the more obvious the consistency deviation. According to the state consistency comparison record, the state difference and time delay value of each sample number are combined and analyzed. The number index is , the transmission status is , the storage status is , the state difference is , further call the transmission delay of this sample number and write latency , construct the time disturbance adjustment factor structure , and then combine it with the state difference to get the consistency deviation contribution value of the sample, and finally average it to get the consistency deviation impact value , the specific calculation is as follows: Sample R301: , , ; ; ; ; Contribution value ; Sample R302: , contribution value = 0; Sample R303: , , ; ; ; ; Contribution value ; Samples R304 and R305 have the same status, and the contribution value = 0; final: ; The result shows that the consistency deviation impact value in the current sample group is 0.3823, which is used to subsequently determine whether there is a numbering state that requires retransmission.
[0040] The consistency deviation impact value is a comprehensive indicator used to measure the degree of coupling between the logical consistency of the state and the physical delay fluctuations during the transmission and storage process of scientific research results. This value combines the consistency differences between the transmission state and the storage state, as well as the delay amplitude and symmetry of each number in the network transmission and partition writing links. It reflects the degree of data link reliability deviation in the system due to inconsistent state or uneven delay. When the value is closer to 1, it means that there are a large number of inconsistent states and high delay intensity in the number sample, indicating that the system data consistency risk is greater. When the value is close to 0, it means that the transmission and writing states are consistent, the delay distribution is stable, and the overall link structure is reliable. Therefore, this value can be used as a criterion for dynamically judging whether to trigger the retransmission mechanism, whether link compensation control is required, and other key control processes. It is a composite state indicator that integrates logical deviation and physical interference.
[0041] The operation logic of the formula is to comprehensively measure the coupling effect of the numbered samples in terms of state logic consistency and transmission delay difference, where Indicates the difference in the state of the number in the transmission and storage stages. It takes the value 1 only when the state is inconsistent, otherwise it is 0. It is used to determine whether the number is in an abnormal state. The right part of the product is the delay disturbance compensation factor, the numerator Characterizes the overall delay intensity of the sample in the two stages of network transmission and storage writing. The square root of the sum of squares is used to reflect the nonlinear superposition effect of the two time terms, emphasizing the improvement effect of the absolute delay intensity on the error contribution. The denominator term is added while retaining the overall delay intensity. Characterizing the degree of delay asymmetry between stages, this additive structure is used to simulate the depressing effect of delay fluctuations and structural separation on consistency comparison. Therefore, the overall fractional form is constructed as a delay coupling compensation factor. The closer the value is to 1, the more concentrated, large, and equal the transmission and write delays are. The smaller the value, the larger the time difference and the unstable structure. By multiplying the state difference value, a logical deviation term with a time-physical influence weight is finally constructed. All numbered calculation results are then averaged to measure the amplification effect of state inconsistency in the entire data link under physical network fluctuations.
[0042] S413: Based on the consistency deviation impact value, the numbers with differences greater than the judgment threshold are screened and marked as abnormal numbers. The transmission command is re-triggered for the corresponding numbers to form a retransmission list. The remaining numbers are marked as confirmed items and confirmation flags are generated. A chain confirmation index queue is generated for the corresponding flags of all numbers to establish a complete record of the partition link. Based on the consistency deviation impact value, the numbered samples whose value exceeds the set reference threshold are screened for retransmission operations. The threshold is set to 0.5 this time. This threshold is set based on the median of the statistical distribution of the average state deviation value and the delay coupling ratio corresponding to the sample number in the typical stable link operation stage. Specifically, this value comes from the concentrated interval composed of the delay fluctuation proportion items corresponding to the samples with consistent states in the total sample volume. In samples where the transmission state and storage state are consistent, their delay coupling ratios are generally concentrated between 0.2 and 0.4, and rarely exceed 0.5. Therefore, 0.5, which is slightly above the upper boundary, is selected as the judgment threshold value. This value can be dynamically adjusted with the fluctuation of the link transmission rate or the change of the physical partition write load. The judgment is based on According to whether the value is greater than this value, a direct comparison method is used. For example, the index value obtained for sample R301 is 0.9465, and that for R303 is 0.9650, both of which are greater than the set threshold of 0.5. They are judged as abnormal numbers, written into the retransmission list, and assigned the retransmission flag bit 1. The contribution values of R302, R304, and R305 are 0, which is lower than the threshold. They are marked as confirmation numbers and assigned the confirmation flag bit 0. The retransmission numbers are constructed into chain task scheduling items, forming a structure such as {R301, 1}, {R303, 1}, {R302, 0}, {R304, 0}, and {R305, 0}, which are used for the scheduling logic execution of the task scheduling controller in the subsequent data return mechanism. Finally, the registration of the data flag index queue under the control structure is completed, and a complete record of the partition link is established.
[0043] See also Figure 6 , the specific steps of S5 are: S511: Based on the complete record of the partition link, extract the result number set corresponding to each storage area, and call the attribution label and partition identifier of each number in the link data, calculate the relationship between the number of data numbers in each storage area and the total number of numbers in the corresponding storage area, and obtain the partition data attribution ratio; Based on the storage area structure information registered in the complete record of the partition link, the achievement number set corresponding to each storage area is extracted. The attribution tag field of the number in the set is read in sequence. It is marked by judging whether the field is "personal data". The numbers marked as "personal data" are recorded as a subset, and the ratio of the number of numbers in the subset to the total number of numbers in the storage area is calculated. In actual operation, if the total number of numbers in Z1 area is 120, of which the personal data number is 48, the attribution ratio is 48 divided by 120, which is 0.40; for Z2 area, it is 41 divided by 105, which is 0.39; for Z3 area, it is 62 divided by 98, which is 0.63; for Z4 area, it is 59 divided by 143, which is 0.41; and for Z5 area, it is 38 divided by 112, which is 0.34. All storage areas are listed uniformly according to the proportion value to construct the following structure: Table 6 Data ownership ratio of each storage area:
[0044] As shown in Table 6, the attribution ratio of each partition has been calculated. The attribution ratios of Z3 and Z4 are both greater than 0.40, providing basic data for subsequent compliance judgment operations and obtaining the attribution ratio of partition data.
[0045] S512: Based on the partition data ownership ratio, a benchmark value for determining the ownership ratio is set. The ownership ratio corresponding to each storage area is compared with the benchmark value. Each partition is identified based on the comparison result. Partitions with a ratio higher than the benchmark value are marked as non-compliant, and partitions with a ratio lower than or equal to the benchmark value are marked as compliant, thereby generating a data distribution compliance identification sequence. The partition data attribution ratio is called. According to the personal data storage ratio limit stipulated in the General Data Protection Regulation, the ratio benchmark value is set to 0.40. The attribution ratios in Table 6 are compared with the benchmark value item by item. If the ratio of a storage area is higher than 0.40, it is marked as "non-compliant", otherwise it is marked as "compliant". For example, the Z1 area is 0.40, which is considered compliant. The Z2 area is 0.39, which is lower than the benchmark value and is compliant. The Z3 area is 0.63, which is higher than 0.40 and is non-compliant. The Z4 area is 0.41, which is non-compliant. The Z5 area is 0.34, which is still compliant. The result structure after comparison is as follows: Table 7 Storage Area Compliance Determination Results:
[0046] As shown in Table 7, Z3 and Z4 are marked as non-compliant areas, and the remaining storage areas meet the benchmark value conditions. The above judgment identifiers are integrated to generate a data distribution compliance identifier sequence.
[0047] S513: Based on the data distribution compliance identification sequence, for the data partitions marked as non-compliant, the information content is sorted, and the storage structure is sorted according to the record status of the numbers in the cache structure, completing the update operation of the storage area number distribution and obtaining the full collection record of the vocational college big data; According to the data distribution compliance identification sequence, the data structure extraction processing is performed on the storage areas Z3 and Z4 marked as "non-compliant". The cache occupancy rate field, number attribution target field, number transmission frequency field and generation time field corresponding to all achievement numbers in the Z3 area are obtained. Sample numbers R304, R309 and R310 are extracted from numbers R301 to R312 as processing objects. The corresponding cache occupancy values are 82%, 91% and 87%, respectively. The average daily transmission frequency is 3 times, 2 times and 4 times. The generation time is more than 30 days earlier than the current date, respectively, which meets the reconstruction conditions. Numbers R411 to R425 are extracted from the Z4 area. Numbers R416, R421 and R423 with cache occupancy rates higher than 75% are marked. Their transmission frequencies are 3 times, 2 times and 3 times, and their generation time is earlier than the beginning of the month. They are recorded as sorting objects. The above-mentioned qualified numbers are added to the sorting index structure and the status flag is registered. The record content is established as follows: Table 8 Registration form for non-compliant areas requiring reconstruction numbers:
[0048] As shown in Table 8, the above numbers meet the comprehensive conditions of high cache occupancy, high frequency, and early generation time. They have been locked as reconstruction target numbers and an index table has been generated. The final structure is used for archiving operations to obtain the full collection records of big data from higher vocational colleges.
[0049] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for collecting full data of higher vocational colleges based on big data, characterized by: The following steps are involved: S1: Collect the source system codes and time series of higher vocational colleges, splice them into a unique identification string, and perform duplication verification item by item to obtain a unique link comparison sequence; S2: Based on the unique link comparison sequence, when the status meets the data quality threshold condition, execute the link activation action to push the data to the scientific research results storage area; when the status does not meet the condition, keep the link standby, and obtain the link activation sequence; S3: Based on the link activation sequence, citation records are detected and the citation frequency of each achievement is calculated. When the citation frequency exceeds the high citation threshold, a high-performance storage partition write action is executed. When the citation frequency is lower than or equal to the high citation threshold, a normal storage area aggregate write action is executed to obtain a scientific research achievement partition list; S4: Based on the scientific research results partition list, check the status consistency, execute the retransmission mechanism when there is a difference in the check status, record the confirmation mark when the status is consistent, and obtain a complete record of the partition link; S5: Based on the complete record of the partition link, determine whether the data distribution of each storage area complies with the principle of data minimization. If there is a deviation in the distribution, perform cache reconstruction. If there is no deviation, maintain the current state and obtain the full collection record of big data of higher vocational colleges.
2. The method for collecting full data of higher vocational colleges based on big data according to claim 1 is characterized in that: The unique link comparison sequence includes identification uniqueness information, cache update status, and duplicate verification results; the link activation sequence includes data quality status, transmission trigger flag, and push storage instructions; the scientific research results partition list includes high-performance storage index, ordinary storage aggregation record, and citation frequency classification results; the partition link complete record includes status consistency identification, retransmission confirmation record, and log comparison results; the full collection record of big data of higher vocational colleges includes storage area distribution ratio data, cache reconstruction information, and link storage status.
3. The method for collecting full data of higher vocational colleges based on big data according to claim 1 is characterized in that: The specific steps for obtaining the unique link comparison sequence are: S111: Based on the source system code and time sequence of the higher vocational college received by the acquisition gateway, the code and time sequence are concatenated to form a unique identification string, and the unique identification string is compared with the identification sequence in the cache list item by item to obtain a difference comparison value; S112: comparing the identifier sequence of the cache list according to the difference comparison value, transmitting the received data when the comparison result is inconsistent, and updating the cache list order according to the comparison difference and deleting the earliest identifier in the process to obtain a cache update sequence; S113: Perform a joint judgment based on the cache update sequence and the difference comparison value, discard the current data when the comparison results are completely consistent, and generate a unique link comparison sequence.
4. The method for collecting full data of higher vocational colleges based on big data according to claim 1 is characterized in that: The specific steps for obtaining the link activation sequence are: S211: Reading link transmission parameters based on the unique link comparison sequence, extracting the transmission rate, delay time, and error rate parameters in each link sequence, and associating and storing the integrity information of each link with the number of received data packets, integrating the parameters of all links one by one to obtain a link parameter set; S212: Based on the link parameter set, each dimension of integrity, timeliness, and accuracy is judged against the data quality threshold, and after comparison, a qualified or unqualified status is recorded to obtain a link status judgment result; S213: Based on the link status judgment result, when the status is marked as qualified, the link activation action is executed to push the data to the scientific research results storage area; when the status is marked as unqualified, the link is kept in standby mode and no data is pushed; a joint evaluation is performed and qualified link entries are sorted and arranged in order to generate a link activation sequence.
5. The method for collecting full data of higher vocational colleges based on big data according to claim 1 is characterized in that: The specific steps for obtaining the zoning list of scientific research achievements are as follows: S311: Based on the scientific research achievement data in the link activation sequence, detect the citation record of each achievement, read the achievement number one by one and associate it with the citation log entry, count and store the citation times and mark the corresponding achievement number to obtain citation record frequency data; S312: Based on the citation record frequency data, the frequency value of each achievement is compared with the high citation threshold one by one. If the frequency is greater than the high citation threshold, it is marked as highly cited; otherwise, it is marked as normal, thereby obtaining a citation frequency status record; S313: Based on the citation frequency status record, a high-performance storage partition write action is performed on highly cited achievements and index information is recorded; an ordinary storage area aggregation write action is performed on ordinary achievements, and a partition list of scientific research achievements is established.
6. The method for collecting full data of higher vocational colleges based on big data according to claim 1 is characterized in that: The specific steps for obtaining the complete record of the partition link are: S411: Based on the scientific research achievement partition list, obtain the achievement number and corresponding storage location data in the partition, read the transmission status record and time node information of each number in the transmission log, combine the storage flag and write time of the corresponding number in the storage record, compare the time and status fields of the same number in the two data sources one by one, organize each number comparison result into a status comparison item, and obtain a status consistency comparison record; S412: Based on the status consistency comparison record, determine whether the status fields corresponding to all achievement numbers are consistent. If they are consistent, assign a value of 1; if they are inconsistent, assign a value of 0. Calculate and obtain the consistency deviation impact value. S413: Based on the consistency deviation impact value, the numbers whose differences are greater than the judgment threshold are screened and marked as abnormal numbers, and the transmission command of the corresponding numbers is re-triggered to form a retransmission list, and the remaining numbers are marked as confirmation items and confirmation flags are generated. A chain confirmation index queue is generated for the flags corresponding to all numbers to establish a complete record of the partition link.
7. The method for collecting full data of higher vocational colleges based on big data according to claim 1 is characterized in that: The specific steps for obtaining the full collection records of big data of higher vocational colleges are as follows: S511: Based on the complete record of the partition link, extract the set of achievement numbers corresponding to each storage area, call the attribution label and the partition identifier of each number in the link data, calculate the relationship between the number of data numbers in each storage area and the total number of numbers in the corresponding storage area, and obtain the partition data attribution ratio; S512: Based on the partition data attribution ratio, a reference value for determining the attribution ratio is set, and the attribution ratio corresponding to each storage area is compared with the reference value. Each partition is identified according to the comparison result, and partitions with a ratio higher than the reference value are marked as non-compliant, and partitions with a ratio lower than or equal to the reference value are marked as compliant, thereby generating a data distribution compliance identification sequence; S513: According to the data distribution compliance identification sequence, for the data partitions marked as non-compliant, the information content is sorted out, and the storage structure is sorted out according to the record status of the number in the cache structure, completing the update operation of the storage area number distribution and obtaining the full collection record of the big data of higher vocational colleges.
Citation Information
Patent Citations
Dual-park data verification method and device
CN113391956A
Oracle to openGauss full-amount and incremental data verification method
CN116910025A
Cross-platform office equipment resource dynamic scheduling optimization method and system
CN120301899A
Multi-scene application data acquisition method based on atomization design
CN120407813A
Safety production informatization data management method and system
CN120429283A