Digital archive information data processing method

By performing fine-grained splitting and feature construction on archive information, the problem of not being able to identify and record the modification location of archive content in the prior art is solved, a safe and unique feature sequence is achieved, and the transmission efficiency of archive information data processing is improved.

CN119989407AInactive Publication Date: 2025-05-13SUZHOU HENGSHUNTONG INFORMATION SERVICE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510062134.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The digital signature constructed based on the hash algorithm in the prior art cannot identify and record the specific modification location of the archive content, and cannot meet certain needs that focus on archival change data.

Method used

By splitting the archive information into a single-page archive and splitting the single-page archive into multiple content areas, a content feature sequence is constructed based on the feature data of the content area, and combining text feature sequences and unit feature sequences, the feature description of the archive content and the positioning of the modified location are realized.

Benefits of technology

It is realized that under the premise of streamlining the recording of multiple features of the archive content, a relatively complex and multi-angle feature description combination is constructed, and a secure and unique feature sequence is established for archival information at different levels, which can identify and record tampering and modifications during the archive transmission process, which is conducive to improving the transmission efficiency in the process of archival information data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989407A_ABST
    Figure CN119989407A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a digital archive information data processing method, which comprises the following steps of: 1, splitting archive information into a plurality of single-page archives, calculating regional characteristic values by combining capacity values and angle characteristic values of each content region, and arranging the regional characteristic values in sequence to form a content characteristic sequence; 2, constructing a text feature sequence in combination with the three feature values of the content area; 3, splitting the text content into a plurality of unit areas, and combining the two feature sequences of each unit area to obtain a unit feature sequence; and 4, when the archive information is sent, obtaining the initial binding value of the archive information, when the archive information is received, initially and sequentially checking the initial binding value, and determining an abnormal single page based on a checking result and performing unit checking. The method has confidentiality and safety, can identify and record tampering and modification in the archive transmission process, and is beneficial to improving the transmission efficiency in the archive information data processing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to, in particular to, a digitalized archival information data processing method. Background Art

[0002] Digital signature is an encryption technology based on public key infrastructure (PKI). It is usually used to confirm the source, identity, timestamp and other information of archives or files to prevent illegal modification or forgery during the circulation of archives. It ensures the authenticity, integrity and legitimacy of the source of the file by encrypting and signing the archive content. The digital signature process includes the following steps: generating hash values, private key encryption, appending signatures, and verifying signatures. It mainly uses the uniqueness between the value generated by the hash algorithm and the original data.

[0003] The digital signature built based on the hash algorithm can effectively ensure the integrity of the archive content and provide privacy protection during the archive transmission process. However, the irreversibility of the hash algorithm also makes it impossible to determine the specific modification location of the archive content based on the change of the hash value. This leads to the fact that in some cases where the focus is on archive change data, the data processing method in the prior art cannot meet the needs of identifying and recording the modification location of the archive content. Summary of the invention

[0004] In view of the above-mentioned shortcomings of the prior art, the present invention provides a digital archival information data processing method, which can effectively solve the problem that the digital signature constructed based on the hash algorithm in the prior art cannot meet the needs of identifying and recording the modification location of the archival content.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0006] The present invention provides a digital archival information data processing method, which at least comprises:

[0007] Step 1: Split the archive information into multiple single-page archives, split the content in the single-page archives into multiple content areas, and sort each content area based on the center position of each content area;

[0008] A capacity value is assigned to each content area according to the amount of data stored in the content area, a characteristic graph is constructed based on the position of the center of the content area, and the angles of each internal angle of the characteristic graph are analyzed to obtain the angular characteristic value;

[0009] Combining the capacity value and the angular eigenvalue of each content area, the regional eigenvalue is calculated and arranged in order to form a content feature sequence;

[0010] Step 2: Perform independent analysis on the content area, calculate the text feature value, numerical feature value and sentence feature value based on the text features in the content area, and construct a text feature sequence by combining the three feature values ​​of the content area;

[0011] Step 3: Obtain the text content in the content area, split the text content into multiple unit areas, preset a keyword set and a symbol sequence, analyze the keyword set and the symbol sequence to obtain the keyword feature sequence and the symbol feature sequence of each unit area, and combine the two feature sequences of each unit area to obtain the unit feature sequence;

[0012] Step 4: When the archive information is sent, the content feature series corresponding to each single-page archive and the text feature series corresponding to each content area are obtained and used as the initial binding value. When the archive information is received, the archive information is initially verified by single page verification and content verification in sequence. Based on the verification results, the abnormal single page is determined and the process goes to step 5.

[0013] Step 5: Get the unit feature sequence initially bound to the unit area and perform unit verification. If the unit verification fails, generate an area modification signal and highlight the unit area.

[0014] Further, the file format of the single-page file is obtained, and when the single-page file is in a text table format, each cell is recorded as an independent content area;

[0015] When the single-page file is in text format, convert the single-page file in text format to text table format.

[0016] Furthermore, the process of converting a single-page file in text format to a text table format is as follows:

[0017] Split a single-page file into multiple separate paragraphs, obtain the paragraph line spacing between two adjacent paragraphs to obtain multiple paragraph intervals, extract the maximum value of the paragraph intervals as the division interval, and divide the single-page file into multiple primary areas using the multiple paragraph intervals corresponding to the division interval as the area interval;

[0018] Get the text formats of multiple paragraphs in the first-level area. The text formats include font size and font. Divide the same first-level area into multiple second-level areas according to different text formats. The text formats of paragraphs in each second-level area are consistent. Define a cell for each second-level area, and the edge of the cell coincides with the boundary of the text content in the second-level area.

[0019] Furthermore, the process of constructing the content feature array is as follows:

[0020] S1: Establish a plane rectangular coordinate system with the center of one of the content areas as the origin and two straight lines parallel to the edge of the content area as the coordinate axes. Obtain the center coordinate of any content area as the position coordinate of the content area (x1, y1). Sort the content areas in the order of the size of the horizontal coordinate x1 in the position coordinate and record them as Cell. i , when the horizontal coordinates are the same, they are sorted in the order of the size of the vertical coordinate y1, where i is the serial number of the content area, i = 1, 2, 3, ..., j, j is the total number of content areas;

[0021] S2: Get each content area Cell i The storage capacity of Chinese text data is recorded as cell capacity i ′, get the cell capacity Cell i The minimum value among them is recorded as the first reference capacity [Cell i ′] min , obtain the maximum value of the cell capacity and record it as the second reference capacity [Cell i ′] max , construct the capacity range [[Cell i ′] min ,[Cell i ′] max ] and split it into ten unit capacity intervals of equal length, and assign a capacity value Cell to each unit capacity interval in turn i RL , the capacity value is any integer from 0 to 9;

[0022] S3: Select z content areas and connect the center points in sequence to construct a closed figure, where z satisfies 3≤z≤j, calculate the area of ​​each closed figure and select the closed figure with the largest area as the feature figure, record the internal angle of the polygon with an angle greater than 180° as a concave angle, and record the internal angle of the polygon with an angle greater than 0° and less than 180° as a convex angle, count the number of acute angles and the number of obtuse angles in the internal angles of the feature figure and calculate the difference to obtain the angle characteristic value R;

[0023] S4: Get capacity value and the angular eigenvalue R and substitute into the formula Multiple regional characteristic values ​​are obtained by calculation in Indicates taking The last digit of the calculation result is the value of multiple regional characteristic values The content feature sequence is obtained by sorting and combining according to the sequence number of the subscript i.

[0024] Furthermore, the calculation process of text feature value is as follows:

[0025] The image corresponding to the content area is recorded as the content image. The total area of ​​the content image and the total area of ​​the text content in the content image are obtained. The total area of ​​the text content is equal to the total number of text content pixels. The total area of ​​the text content is divided by the total area of ​​the content image to obtain the text feature value TZ. txt .

[0026] Furthermore, the digital eigenvalue calculation process is as follows:

[0027] Get the text content in the content area, filter out the Arabic numerals in the text content and calculate the total storage occupied by all Arabic numerals as the digital capacity, get the total storage occupied by the text content as the text capacity, and substitute it into the formula The digital eigenvalue TZ is calculated in num ,in:

[0028] UTF num Indicates digital capacity;

[0029] UTF all Indicates the text capacity;

[0030] Indicates that the value is not less than The smallest integer.

[0031] Furthermore, the sentence feature value calculation process is as follows:

[0032] Obtain the total storage usage of different types of languages ​​and characters in the text content, select the language and character with the largest total storage usage as the target language, filter out the target language in the text content and combine and arrange them into a row to obtain the target sentence;

[0033] The image of the target sentence is obtained and recorded as the target image, a line segment l1 crossing the target image is constructed, the width wd and height hd of the line segment l1 are both preset values, and the number of pixels of the overlapped part of the target image and the line segment l1 is obtained and recorded as the overlap length L;

[0034] A secret semantic text is preset, and the secret semantic text is expressed in the target language to obtain a secret sentence. The image of the secret sentence is obtained and a line segment l2 with a width of wd and a height of hd that crosses the secret sentence image is constructed. The number of pixels of the overlapping part of the line segment l2 and the secret sentence image is calculated to obtain the sentence basic value Language ave ;

[0035] Get the total storage usage of the target statement and record it as statement capacity Language all , get the basic value of the sentence corresponding to the target language Language ave , substitute into the formula Calculate in and get the sentence feature value TZ la .

[0036] Furthermore, the construction process of the unit characteristic sequence is as follows:

[0037] There are multiple keyword sets composed of different keywords. n , where n is the sequence number of the keyword set, n = 1, 2, 3, ..., m, m is the total number of keyword sets, and each keyword set corresponds to a keyword coefficient GJ n ′, filter out all keywords in the unit area and record them as target keywords, and count the number of target keywords corresponding to each keyword set in the unit area and record them as Get The last digit of the value is sorted according to the sequence number of the keyword set to form a keyword feature sequence;

[0038] A symbol sequence consisting of m symbols is preset, the number of different symbols in the unit area is obtained and recorded as the symbol number, the last digit value of each symbol number is obtained and arranged according to the order of the symbols in the symbol sequence to form a symbol feature sequence;

[0039] The elements in the keyword feature sequence and the symbol feature sequence are obtained in turn, and are cross-sorted and combined to obtain a unit feature sequence.

[0040] Furthermore, the construction process of the unit characteristic sequence is as follows:

[0041] Get the preset width value wd, height value hd of line segment l1, get the secret semantic text, and get the keyword set GJ n And the symbol sequence together constitute the secret code set.

[0042] Furthermore, the process of step 4 is as follows:

[0043] Part 1: Obtain the current content feature sequence of the single-page file to be compared and record it as the content comparison sequence. Compare the content comparison sequence with the initially bound text feature sequence. When the two sequences are the same, generate a comparison pass signal and perform content verification.

[0044] When the lengths of the two number sequences are different, the single-page file is marked as an abnormal single page;

[0045] When two sequences have the same length but different elements, the number of different elements is recorded as the number of abnormalities. When the number of abnormalities is greater than the preset abnormal threshold, the single page file is marked as an abnormal single page.

[0046] When the number of anomalies is less than or equal to the preset anomaly threshold, the content areas corresponding to different elements are obtained and recorded as abnormal areas, the abnormal areas are highlighted, and Part 2 is performed on the abnormal areas;

[0047] Part 2: Content verification is performed on text feature sequences, specifically;

[0048] The content area to be verified is recorded as the content verification area, the preset width value wd, height value hd and secret code semantic text in the secret code set are obtained, the content verification area is processed in step 2 to obtain the current content feature sequence of the content verification area and recorded as the text comparison sequence, and the text comparison sequence is compared with the text feature sequence initially bound to the content verification area;

[0049] Part 3: When the two number sequences are the same, a comparison pass signal is generated and the content is verified;

[0050] When the two number sequences are the same, a comparison pass signal is generated, and when the two number sequences are different, step five is performed;

[0051] Part 4: When both Part 1 and Part 3 generate comparison pass signals, a preliminary verification completion signal is generated.

[0052] Compared with the known prior art, the technical solution provided by the present invention has the following beneficial effects:

[0053] The present invention generates a specific feature sequence of archive content, and can construct a relatively complex and multi-angle feature description combination as much as possible under the premise of concisely recording multiple features of the archive content, so as to establish a safe and unique feature sequence for archive information at different levels. Compared with the unique representation of file content composed of hash values, the feature sequence is characterized in that the numbers in the feature sequence have characteristics, and the characteristics are obtained based on the original text of the archive. The position of the modified content can be located according to the change of the characteristics, so that the staff can make further operations such as content verification, content abandonment, content repair, etc. based on the modified position.

[0054] 2. The content in the position area without feature changes obtained by comparison in the present invention is still original, which is more applicable to certain scenarios that rely on original data. Although the feature sequence has characteristics, the use of the feature is abstract and also irreversible. Even if the feature sequence is leaked, it will not cause the leakage of the archive content. Therefore, the present invention has confidentiality and security, and can identify and record tampering and modification in the process of archive transmission, which is beneficial to improving the transmission efficiency in the process of archive information data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0056] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] The present invention will be further described below in conjunction with the embodiments.

[0059] See also Figure 1 A digital archival information data processing method comprises at least the following steps:

[0060] Step 1: Split the archive information into multiple single-page archives, where a single-page archive is a single page in the archive information. Split the content in the single-page archive into multiple independent content areas, and construct a content feature sequence for the feature data of each content area in the single-page archive, where:

[0061] Get the file format of a single-page file, and use different methods to split the single-page file into regions according to different file formats. When the single-page file is in a text table format, the splitting and numbering process is as follows:

[0062] Remember each cell as a separate content area;

[0063] A plane rectangular coordinate system is established with the center of one of the content areas as the origin and two straight lines parallel to the edge of the content area as the coordinate axes. The center coordinates of any content area are obtained as the position coordinates of the content area (x1, y1). The content areas are sorted in the order of the size of the horizontal coordinate x1 in the position coordinate and recorded as Cell. i , when the horizontal coordinates are the same, they are sorted in the order of the size of the vertical coordinate y1, where i is the serial number of the content area, i = 1, 2, 3, ..., j, j is the total number of content areas; (the content areas are sorted based on the center position of each content area)

[0064] Since the content areas are independent of each other, there will not be two content areas with the same sequence number, so that i Any content area can be distinguished and identified.

[0065] Get each content area Cell i The storage capacity of Chinese text data (in bytes) is recorded as cell capacity i ′, get the cell capacity Cell i The minimum value of ′ is recorded as the first reference capacity [Cell i ′] min , obtain the maximum value of the cell capacity and record it as the second reference capacity [Cell i ′] max , construct the capacity range [[Cell i ′] min ,[Cell i ′] max ] and split it into ten unit capacity intervals of equal length, and assign a capacity value Cell to each unit capacity interval in turn i RL , the capacity value is any integer from 0 to 9; (a capacity value is assigned to each content area according to the size of the data in the content area)

[0066] Select z content areas and connect the center points in sequence to construct a closed figure, where z satisfies 3≤z≤j. Calculate the area of ​​each closed figure and select the closed figure with the largest area as the feature figure. The feature figure is a polygon. The inner angles of the polygon with an internal angle greater than 180° are recorded as concave angles, and the inner angles of the polygon with an internal angle greater than 0° and less than 180° are recorded as convex angles. Count the number of acute angles and obtuse angles in the inner angles of the feature figure and calculate the difference to obtain the angle eigenvalue R (when the centers of each cell are on the same straight line, the angle eigenvalue is 0); (Construct the feature figure based on the position of the center of the content area, and analyze the degrees of each internal angle of the feature figure to obtain the angle eigenvalue)

[0067] Get capacity value and the angular eigenvalue R and substitute into the formula Multiple regional characteristic values ​​are obtained by calculation in Indicates taking The last digit of the calculation result, for example, when the capacity value When the angular eigenvalues ​​R are 5 and 12 respectively, the corresponding regional eigenvalues Multiple regional feature values Sort and combine according to the sequence number of subscript i to obtain a content feature sequence with a length of j;

[0068] When the single-page file is in text format, convert it to a text table format, where:

[0069] Split a single-page file into multiple separate paragraphs, obtain the paragraph line spacing between two adjacent paragraphs to obtain multiple paragraph intervals, extract the maximum value of the paragraph intervals as the division interval, and divide the single-page file into multiple primary areas using the multiple paragraph intervals corresponding to the division interval as the area interval;

[0070] The text formats of multiple paragraphs in the first-level area are obtained, and the text formats include font size and font. The same first-level area is divided into multiple second-level areas according to different text formats. The text formats of the paragraphs in each second-level area are consistent. A cell is defined for each second-level area, and the edge of the cell coincides with the boundary line of the text content in the second-level area (that is, the cell is the smallest square that can cover all the text content in the second-level area). By converting a single-page file in a text format into a single-page file in a text table format composed of multiple cells, the single-page file in the text format can be also applied to the generation of the above-mentioned content feature series.

[0071] It should be noted that text content refers to a text sentence composed of multiple characters used to convey information content, such as a paragraph of Chinese characters or a paragraph of English, which serves as the carrier of the archival information content.

[0072] The text data in the archive information in the text format is distinguished by a simple format (such as paragraphs, font size, font); the text data in the archive information in the text table format is recorded in different cells of the simple table. Both file formats do not contain pictures, charts and other contents. In other words, the difference between the two lies in "whether the different contents of the text data in the archive information are distinguished by paragraphs or cells". When the format of the archive information is a combination of the two file formats, the above two methods can be combined for splitting and numbering, which will not be elaborated here.

[0073] Step 2: Perform independent analysis on each content area, and construct a text feature sequence in the content area based on the text features in the content area. Each content area corresponds to a text feature sequence, where:

[0074] The image corresponding to the content area is recorded as the content image. The total area of ​​the content image and the total area of ​​the text content in the content image are obtained. The total area of ​​the text content is equal to the total number of text content pixels. The total area of ​​the text content is divided by the total area of ​​the content image to obtain the text feature value TZ. txt ;

[0075] Get the text content in the content area, filter out the Arabic numerals in the text content and calculate the total storage occupied by all Arabic numerals as the digital capacity, get the total storage occupied by the text content as the text capacity, and substitute it into the formula The digital eigenvalue TZ is calculated in num ,in:

[0076] UTF num Indicates digital capacity;

[0077] UTF all Indicates the text capacity;

[0078] Indicates that the value is not less than The smallest integer of ;

[0079] It should be noted that both storage capacity and storage occupancy refer to the storage occupancy of the information in the computer, measured in bytes. For example, in UTF-8 encoding, a Chinese character usually occupies 3 bytes, an Arabic letter usually occupies 2 bytes, and a Latin letter (including uppercase and lowercase letters) usually occupies 1 byte.

[0080] Get the total storage usage of different types of languages ​​and characters in the text content, select the language and character with the largest total storage usage as the target language, filter out the target language in the text content and arrange them in a row to get the target sentence. The target sentence only contains the same type of languages ​​and the spacing between two adjacent languages ​​and characters is consistent. Get the image of the target sentence and record it as the target image. Construct a line segment l1 that crosses the target image. The width wd and height hd of the line segment l1 are both preset values. Get the number of pixels of the overlapping part of the target image and the line segment l1 and record it as the overlapping length L. Get the total storage usage of the target sentence and record it as the sentence capacity Language all , get the basic value of the sentence corresponding to the target language Language ave , substitute into the formula Calculate in and get the sentence feature value TZ la ;

[0081] The process of obtaining the basic value of the corresponding sentence in the target language is as follows:

[0082] A secret command semantic text is preset, and the secret command semantic text is expressed in the target language to obtain a secret command sentence (that is, the secret command semantic text is translated into the target language to obtain a text sentence composed of the target language), the image of the secret command sentence is obtained, and a line segment l2 with a width of wd and a height of hd and crossing the secret command sentence image is constructed (the width and height of the line segment l2 are consistent with those of the line segment l1), and the number of pixels of the overlapped part of the line segment l2 and the secret command sentence image is calculated to obtain the sentence basic value Language ave .

[0083] It should be noted that in order to ensure that there will be no conversion errors when converting the secret command semantic text into the secret command sentence expressed in the target language, the secret command semantic text should use sentence text with simple semantics and not easy to cause ambiguity, and the conversion program (usually translation software) used in the feature generation and subsequent verification process should be consistent.

[0084] The text feature value, numerical feature value and sentence feature value of each content area are obtained, and the text feature sequence of each content area is obtained by arranging and combining them in sequence. The text feature sequence consists of multiple digits, and the first two digits are the text feature value and the numerical feature value respectively. Since the text feature value and the numerical feature value are both represented by single digits, it is possible to quickly distinguish whether there is a change, and then quickly lock the abnormal features (i.e., the total area of ​​the text and the numerical changes in the text). However, the disadvantage is that single digits are relatively easy to crack. By introducing the sentence feature value, a numerical value with an uncertain length, to supplement the feature sequence, the length of the text feature sequence is uncertain, and thus the regularity of the text feature sequence is difficult to crack.

[0085] Step 3: Get the text content in the content area, split the text content into multiple unit areas using the period in the text content as the segmentation point, analyze each unit area, and construct the unit feature series corresponding to each unit area. The construction process of the unit feature series is as follows:

[0086] There are multiple keyword sets composed of different keywords. n , where n is the sequence number of the keyword set, n = 1, 2, 3, ..., m, m is the total number of keyword sets, and each keyword set corresponds to a keyword coefficient GJ n ′, the elements in the keyword set are preset keywords, different keyword sets do not intersect with each other, all keywords in the unit area are screened out and recorded as target keywords, and the number of target keywords corresponding to each keyword set in the unit area is recorded as Get The last digit of the value is sorted according to the sequence number of the keyword set to form a keyword feature sequence;

[0087] A symbol sequence consisting of m symbols is preset, the number of different symbols in the unit area is obtained and recorded as the symbol number, the last digit value of each symbol number is obtained and arranged according to the order of the symbols in the symbol sequence to form a symbol feature sequence (when the symbol number is 0, its last digit value is 0);

[0088] Get the elements in the keyword feature sequence and the symbol feature sequence in turn, and cross-sort and combine them to get the unit feature sequence. Let the keyword feature sequence be The symbolic characteristic sequence is Then the unit characteristic sequence is

[0089] Further, the preset width value wd and height value hd of the line segment l1 are obtained, and the secret semantic text is obtained to obtain the keyword set GJ n and a sequence of symbols together form a secret code set, which is sent separately from the archive information data as a separate secret code and is used for feature verification when received.

[0090] Step 4: When the archive information is sent, the content feature series corresponding to each single-page archive and the text feature series corresponding to each content area are obtained and used as the initial binding value. When the archive information is received, the archive information is initially verified by single page verification and content verification in sequence, where:

[0091] Single page verification is performed on the content feature series, specifically:

[0092] The single-page file to be compared is processed in step 1 to obtain the current content feature sequence of the single-page file, which is recorded as the content comparison sequence. The content comparison sequence is compared with the text feature sequence initially bound to the single-page file. When the two sequences are the same, a comparison pass signal is generated and content verification is performed;

[0093] When the lengths of the two number sequences are different, the single-page file is marked as an abnormal single page;

[0094] When two sequences have the same length but different elements, the number of different elements (different elements refer to elements with the same sequence number but different values) is obtained and recorded as the number of abnormalities. When the number of abnormalities is greater than the preset abnormal threshold, the single page file is marked as an abnormal single page.

[0095] When the number of anomalies is less than or equal to the preset anomaly threshold, the content areas corresponding to different elements are obtained and recorded as abnormal areas, the abnormal areas are highlighted, and the content of the abnormal areas is verified;

[0096] Content verification is performed on text feature sequences, specifically:

[0097] The content area to be verified is recorded as the content verification area, the preset width value wd, height value hd and secret code semantic text in the secret code set are obtained, the content verification area is processed in step 2 to obtain the current content feature sequence of the content verification area and recorded as the text comparison sequence, the text comparison sequence is compared with the text feature sequence initially bound to the content verification area, when the two sequences are the same, a comparison pass signal is generated and content verification is performed, when the two sequences are the same, a comparison pass signal is generated, when the two sequences are different, step 5 is performed;

[0098] When the single-page verification and content verification are passed, a preliminary verification completion signal is generated. At this time, the staff can choose whether to conduct further verification in step 5 based on the security and importance of the archive information content. When the staff does not choose further verification, the verification is terminated;

[0099] Step 5: When receiving the file information, the unit shall verify the file information, including:

[0100] The file information is split into multiple unit areas (refer to steps 1 to 3 for the splitting process), the unit area to be checked is selected as the unit check area, the keyword set and symbol sequence in the secret code set are obtained, and the unit feature sequence corresponding to the unit check area is constructed based on the keyword set and the symbol sequence, which is recorded as the unit comparison sequence. The unit comparison sequence is compared with the unit feature sequence initially bound to the unit check area. When the two sequences have different elements, an area modification signal is generated to highlight the text content in the unit area.

[0101] It should be noted that, by generating a specific file content feature sequence, it is possible to construct a relatively complex and multi-angle feature description combination as much as possible under the premise of streamlining and recording multiple features of the file content, and establish a safe and unique feature sequence for different levels of file information. Compared with the unique representation of the file content composed of the hash value, the feature sequence is characterized in that the numbers in the feature sequence are characterized, and the feature is obtained based on the original text of the file. According to the change of the feature, the position of the modified content can be located, so that the staff can make further operations such as content verification, content abandonment, and content repair based on the modified position, and the content of the position area without feature change is still original (i.e., it is consistent with the content in the original text), which is more applicable to some scenes that rely on original data. In addition, although the feature sequence has features, the adoption of the feature is abstract and also has irreversibility, that is, even if the feature sequence is leaked, it will not cause the leakage of the file content. In summary, the present invention has both confidentiality and security, and can identify and record the tampering and modification in the file transmission process, which is conducive to improving the transmission efficiency in the file information data processing process.

[0102] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A digital archival information data processing method, characterized in that: The following steps are involved: Step 1: Split the archive information into multiple single-page archives, split the content in the single-page archives into multiple content areas, and sort each content area based on the center position of each content area; A capacity value is assigned to each content area according to the amount of data stored in the content area, a characteristic graph is constructed based on the position of the center of the content area, and the angles of each internal angle of the characteristic graph are analyzed to obtain the angular characteristic value; Combining the capacity value and the angular eigenvalue of each content area, the regional eigenvalue is calculated and arranged in order to form a content feature sequence; Step 2: Perform independent analysis on the content area, calculate the text feature value, numerical feature value and sentence feature value based on the text features in the content area, and construct a text feature sequence by combining the three feature values ​​of the content area; Step 3: Obtain the text content in the content area, split the text content into multiple unit areas, preset a keyword set and a symbol sequence, analyze the keyword set and the symbol sequence to obtain the keyword feature sequence and the symbol feature sequence of each unit area, and combine the two feature sequences of each unit area to obtain the unit feature sequence; Step 4: When the archive information is sent, the content feature series corresponding to each single-page archive and the text feature series corresponding to each content area are obtained and used as the initial binding value. When the archive information is received, the archive information is initially verified by single page verification and content verification in sequence. Based on the verification results, the abnormal single page is determined and the process goes to step 5. Step 5: Get the unit feature sequence initially bound to the unit area and perform unit verification. If the unit verification fails, generate an area modification signal and highlight the unit area.

2. A digital archival information data processing method according to claim 1, characterized in that: Get the file format of a single-page file. When the single-page file is in a text table format, each cell is recorded as an independent content area. When the single-page file is in text format, convert the single-page file in text format to text table format.

3. A digital archival information data processing method according to claim 2, characterized in that: The process of converting a single-page file in text format to a text table format is as follows: Split a single-page file into multiple separate paragraphs, obtain the paragraph line spacing between two adjacent paragraphs to obtain multiple paragraph intervals, extract the maximum value of the paragraph intervals as the division interval, and divide the single-page file into multiple primary areas using the multiple paragraph intervals corresponding to the division interval as the area interval; Get the text formats of multiple paragraphs in the first-level area. The text formats include font size and font. Divide the same first-level area into multiple second-level areas according to different text formats. The text formats of paragraphs in each second-level area are consistent. Define a cell for each second-level area, and the edge of the cell coincides with the boundary of the text content in the second-level area.

4. A digital archival information data processing method according to claim 1, characterized in that: The process of constructing the content feature array is as follows: S1: Establish a plane rectangular coordinate system with the center of one of the content areas as the origin and two straight lines parallel to the edge of the content area as the coordinate axes. Obtain the center coordinate of any content area as the position coordinate of the content area (x1, y1). Sort the content areas in the order of the size of the horizontal coordinate x1 in the position coordinate and record them as Cell. i , when the horizontal coordinates are the same, they are sorted in the order of the size of the vertical coordinate y1, where i is the serial number of the content area, i = 1, 2, 3, ..., j, j is the total number of content areas; S2: Get each content area Cell i The storage capacity of Chinese text data is recorded as cell capacity i ′, get the cell capacity Cell i The minimum value among them is recorded as the first reference capacity [Cell i ′] min , obtain the maximum value of the cell capacity and record it as the second reference capacity [Cell i ′] max , construct the capacity range [[Cell i ′] min ,[Cell i ′] max ] and split it into ten unit capacity intervals of equal length, and assign a capacity value to each unit capacity interval in turn The capacity value is any integer from 0 to 9; S3: Select z content areas and connect the center points in sequence to construct a closed figure, where z satisfies 3≤z≤j, calculate the area of ​​each closed figure and select the closed figure with the largest area as the feature figure, record the internal angle of the polygon with an angle greater than 180° as a concave angle, and record the internal angle of the polygon with an angle greater than 0° and less than 180° as a convex angle, count the number of acute angles and the number of obtuse angles in the internal angles of the feature figure and calculate the difference to obtain the angle characteristic value R; S4: Get capacity value and the angular eigenvalue R and substitute into the formula Multiple regional characteristic values ​​are obtained by calculation in Indicates taking The last digit of the calculation result is the value of multiple regional characteristic values The content feature sequence is obtained by sorting and combining according to the sequence number of the subscript i.

5. A digital archival information data processing method according to claim 1, characterized in that: The calculation process of text feature value is as follows: The image corresponding to the content area is recorded as the content image. The total area of ​​the content image and the total area of ​​the text content in the content image are obtained. The total area of ​​the text content is equal to the total number of text content pixels. The total area of ​​the text content is divided by the total area of ​​the content image to obtain the text feature value TZ. txt .

6. A digital archival information data processing method according to claim 4, characterized in that: The digital eigenvalue calculation process is as follows: Get the text content in the content area, filter out the Arabic numerals in the text content and calculate the total storage occupied by all Arabic numerals as the digital capacity, get the total storage occupied by the text content as the text capacity, and substitute it into the formula The digital eigenvalue TZ is calculated in num ,in: UTF num Indicates digital capacity; UTF all Indicates the text capacity; Indicates that the value is not less than The smallest integer.

7. A digital archival information data processing method according to claim 6, characterized in that: The calculation process of sentence feature value is as follows: Obtain the total storage usage of different types of languages ​​and characters in the text content, select the language and character with the largest total storage usage as the target language, filter out the target language in the text content and combine and arrange them into a row to obtain the target sentence; The image of the target sentence is obtained and recorded as the target image, a line segment l1 crossing the target image is constructed, the width wd and height hd of the line segment l1 are both preset values, and the number of pixels of the overlapped part of the target image and the line segment l1 is obtained and recorded as the overlap length L; A secret semantic text is preset, and the secret semantic text is expressed in the target language to obtain a secret sentence. The image of the secret sentence is obtained and a line segment l2 with a width of wd and a height of hd that crosses the secret sentence image is constructed. The number of pixels of the overlapping part of the line segment l2 and the secret sentence image is calculated to obtain the sentence basic value Language ave ; Get the total storage usage of the target statement and record it as statement capacity Language all , get the basic value of the sentence corresponding to the target language Language ave , substitute into the formula Calculate in and get the sentence feature value TZ la .

8. A digital archival information data processing method according to claim 7, characterized in that: The construction process of the unit characteristic series is as follows: There are multiple keyword sets composed of different keywords. n , where n is the sequence number of the keyword set, n = 1, 2, 3, ..., m, m is the total number of keyword sets, and each keyword set corresponds to a keyword coefficient GJ n ′, filter out all keywords in the unit area and record them as target keywords, and count the number of target keywords corresponding to each keyword set in the unit area and record them as Get The last digit of the value is sorted according to the sequence number of the keyword set to form a keyword feature sequence; A symbol sequence consisting of m symbols is preset, the number of different symbols in the unit area is obtained and recorded as the symbol number, the last digit value of each symbol number is obtained and arranged according to the order of the symbols in the symbol sequence to form a symbol feature sequence; The elements in the keyword feature sequence and the symbol feature sequence are obtained in turn, and are cross-sorted and combined to obtain a unit feature sequence.

9. A digital archival information data processing method according to claim 8, characterized in that: The construction process of the unit characteristic series is as follows: Get the preset width value wd, height value hd of line segment l1, get the secret semantic text, and get the keyword set GJ n And the symbol sequence together constitute the secret code set.

10. A digital archival information data processing method according to claim 9, characterized in that: The specific process of step 4 is as follows: Part 1: Obtain the current content feature sequence of the single-page file to be compared and record it as the content comparison sequence. Compare the content comparison sequence with the initially bound text feature sequence. When the two sequences are the same, generate a comparison pass signal and perform content verification. When the lengths of the two number sequences are different, the single-page file is marked as an abnormal single page; When two sequences have the same length but different elements, the number of different elements is recorded as the number of abnormalities. When the number of abnormalities is greater than the preset abnormal threshold, the single page file is marked as an abnormal single page. When the number of anomalies is less than or equal to the preset anomaly threshold, the content areas corresponding to different elements are obtained and recorded as abnormal areas, the abnormal areas are highlighted, and Part 2 is performed on the abnormal areas; Part 2: Content verification is performed on text feature sequences, specifically; The content area to be verified is recorded as the content verification area, the preset width value wd, height value hd and secret code semantic text in the secret code set are obtained, the content verification area is processed in step 2 to obtain the current content feature sequence of the content verification area and recorded as the text comparison sequence, and the text comparison sequence is compared with the text feature sequence initially bound to the content verification area; Part 3: When the two number sequences are the same, a comparison pass signal is generated and the content is verified; When the two number sequences are the same, a comparison pass signal is generated, and when the two number sequences are different, step five is performed; Part 4: When both Part 1 and Part 3 generate comparison pass signals, a preliminary verification completion signal is generated.

Citation Information

Cited By

  • Digital management system and method for science and technology project archives

    CN120723960A