Intelligent Archival Management Method and System Based on Large Language Model

Through the intelligent archive management method based on the large language model, efficient verification and precise positioning of archives are achieved, the problems of archive authenticity and integrity are solved, and the efficiency and security of archive management are improved.

CN119670723BActive Publication Date: 2025-07-22HANGZHOU AOCHAO TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510186046.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-22
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

In the existing archive management system, the authenticity and integrity of the archives are difficult to guarantee, and it is difficult to accurately determine the tampered person after tampering, which affects the authority and credibility of the archives.

Method used

An intelligent archive management method based on a large language model is adopted, and the content verification group is formed by randomly selecting original archive files, and the integrity verification plug-in is used to perform integrity verification, determine the document status, and when tampering is detected, the tampering content and the adjacent operation end are identified, and the initial archive comparison layer is established for display.

Benefits of technology

It improves the efficiency of archive verification, ensures the originality and authenticity of archives, can quickly and accurately identify tampered content and operation ends, provides responsibility traceability support, and improves the accuracy and convenience of archive management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670723B_ABST
    Figure CN119670723B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent file management method and system based on a large language model. Among them, the method includes: in response to any current operation terminal performing a current scan on the file tag corresponding to the current file, retrieving the original file and randomly extracting to determine a content verification group; triggering a file verification plugin to perform file verification on the current file based on the content verification group, and determining the document status of the current file based on the verification result; in response to the current file being in a tampered state, determining the tampered content of the current file based on the large language model, and determining the adjacent operation terminal for performing an adjacent scan on the current file based on the retrieved operation record table corresponding to the current file; establishing an initial file comparison layer, filling the tampered content and the adjacent operation terminal into the content indication area and the terminal indication area in the initial file comparison layer respectively, and sending the obtained current file comparison layer to the management terminal for display. The present invention at least improves the file management efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data processing technologies, and in particular to an intelligent file management method and system based on a large language model. Background Art

[0002] Currently, in current file management practices, ensuring the authenticity and integrity of files is a crucial issue. However, due to various reasons, there is a risk that files may be tampered with, and once tampering occurs, it is often difficult to accurately identify the tampering personnel. This situation not only seriously threatens the authority and credibility of files but also poses great challenges to file management.

[0003] Therefore, how to develop an intelligent file management method and system that can make full use of large language models to achieve efficient and intelligent file management has become an urgent problem to be solved. Summary of the Invention

[0004] Based on the above problems, the present invention is proposed to provide an intelligent file management method and system based on a large language model that overcomes the above problems or at least partially solves the above problems.

[0005] According to one aspect of the present invention, there is provided an intelligent file management method based on a large language model, including the following steps:

[0006] In response to any current operation terminal performing a current scan on a file tag corresponding to a current file, retrieving the original file corresponding to the current file, randomly extracting each original file included in the original file based on a preset extraction strategy, and determining each extracted original file as a content verification group;

[0007] Triggering a file verification plugin to perform integrity-based file verification on the current file based on the content verification group, and determining a document status corresponding to the current file based on the verification result, where the document status includes a complete status and a tampered status;

[0008] In response to the current file being in a tampered status, determining the tampered content corresponding to the current file, and determining an adjacent operation terminal for performing an adjacent scan on the current file based on the retrieved operation record table corresponding to the current file;

[0009] Based on the current file, establishing an initial file comparison layer, filling the tampered content and the adjacent operation terminal into a content indication area and a terminal indication area located in the initial file comparison layer respectively, and sending the obtained current file comparison layer to a management terminal for display.

[0010] Optionally, in the method according to the present invention, randomly extracting each original document included in the original file based on a preset extraction strategy, and determining the extracted original documents as a content verification group, including:

[0011] Determining each original document with a pre-verification mark in the original file as a necessary verification group, and determining all other original documents included in the original file except the necessary verification group as a random verification group;

[0012] Evaluating the content of each original document in the random verification group, and calculating the sum of the extraction priority values corresponding to each original document to obtain an extraction total value;

[0013] Calculating the ratio of each extraction priority value to the extraction total value respectively to obtain the extraction probabilities corresponding to each original document;

[0014] Determining the number of files corresponding to each original document in the random verification group, and calculating the product based on the number of files and a preset extraction ratio retrieved to obtain the current extraction number;

[0015] Creating a random extraction turntable, and dividing the random extraction turntable into regions based on the extraction probabilities corresponding to each original document to obtain sector regions corresponding to each original document;

[0016] Controlling the random extraction turntable to rotate and extract according to the current extraction number, and determining the extracted original documents and all original documents in the necessary verification group as a content verification group based on the extraction result.

[0017] Optionally, in the method according to the present invention, evaluating the content of each original document in the random verification group, including:

[0018] Determining each original document in the random verification group based on file attributes, and determining each original document with a sensitive attribute as a sensitive document group and each original document with a normal attribute as a normal document group;

[0019] Retrieving a first attribute evaluation value to configure scores for each original document in the sensitive document group respectively, and retrieving a second attribute evaluation value to configure scores for each original document in the normal document group respectively, where the first attribute evaluation value is greater than the second attribute evaluation value;

[0020] Identifying pixel points of each original document in the random verification group respectively to obtain each file pixel point in the same original document, and obtaining the pixel quantity corresponding to each file pixel point to obtain the pixel quantities corresponding to each original document;

[0021] Compare the number of pixels with the retrieved first quantity threshold and second quantity threshold respectively, where the first quantity threshold is less than the second quantity threshold;

[0022] In response to the number of pixels corresponding to any one of the original files being less than the first quantity threshold, retrieve the first quantity evaluation value to configure a score for the original file based on the quantity dimension;

[0023] In response to the number of pixels corresponding to any one of the original files being greater than or equal to the first quantity threshold and less than the second quantity threshold, retrieve the second quantity evaluation value to configure a score for the original file based on the quantity dimension;

[0024] In response to the number of pixels corresponding to any one of the original files being greater than or equal to the second quantity threshold, retrieve the third quantity evaluation value to configure a score for the original file based on the quantity dimension;

[0025] Sum up the attribute evaluation value and the quantity evaluation value corresponding to the same original file to obtain the respective extraction priority values corresponding to each original file.

[0026] Optionally, in the method according to the present invention, determining each original file in the random verification group based on the file attribute, and determining each original file with a sensitive attribute as a sensitive file group and each original file with a common attribute as a common file group includes:

[0027] Obtain the respective file page numbers corresponding to each original file, and in the original file archive, determine the original file with the same page number as the retrieved preset directory page number as the directory file;

[0028] Extract the directory from the directory file to obtain each directory title and the respective page number ranges corresponding to each directory title;

[0029] Classify each directory title hierarchically to obtain a hierarchical structure tree including different hierarchical nodes, where each hierarchical node corresponds to a different directory title;

[0030] Retrieve a preset sensitive record table, where the preset sensitive record table includes different sensitive characters;

[0031] Based on the hierarchical order of the hierarchical structure tree, determine each directory title corresponding to each hierarchical node. In response to at least a part of the directory title corresponding to any one hierarchical node being located in the preset sensitive record table, determine that this hierarchical node and all hierarchical nodes connected downward to this hierarchical node have sensitive attributes;

[0032] Based on each page number interval corresponding to each hierarchical node with sensitive attributes respectively, in the random verification group, all the original documents whose corresponding file page numbers are within the page number interval are respectively determined to have sensitive attributes, and all the remaining original documents in the random verification group are determined to have ordinary attributes.

[0033] Optionally, in the method according to the present invention, obtain each file line drop arranged based on the vertical order from top to bottom in the directory file, and determine all the file line drops with leading characters as each directory title;

[0034] Perform character recognition on each directory title based on a large language model to obtain each file character in the same directory title;

[0035] Perform category grouping on each file character in the same directory title to obtain a text grouping including each file character corresponding to the text category and a number grouping including each file character corresponding to the number category;

[0036] Determine the title page number corresponding to the directory title based on each file character in the number grouping to obtain each title page number corresponding to each directory title respectively;

[0037] Determine the size of any file character in the text grouping respectively to obtain each title size corresponding to each text grouping respectively;

[0038] Arrange each directory title in a hierarchical order from largest to smallest according to the corresponding title size to obtain each hierarchical level corresponding to each directory title respectively;

[0039] Based on the vertical order, determine two directory titles in adjacent positions and corresponding to the same hierarchical level as an adjacent title group respectively, and determine the directory title corresponding to the first position in the same adjacent title group as the first title and the directory title corresponding to the last position as the last title;

[0040] Determine the title page number corresponding to the first title and each title page number corresponding to all the directory titles between the first title and the last title as an interval determination group respectively, and determine the maximum page number corresponding to the maximum value and the minimum page number corresponding to the minimum value based on all the title page numbers in the interval determination group;

[0041] Determine the page number interval composed of the minimum page number and the maximum page number as corresponding to the first title to obtain each page number interval corresponding to each directory title respectively.

[0042] Optionally, in the method according to the present invention, the method further includes:

[0043] Performing handwritten recognition on each original document located in the random verification group based on a handwritten recognition model, and in response to determining that there is a handwritten area in any original document based on the recognition result, determining that the original document has a handwritten attribute;

[0044] Determining a handwritten extension direction corresponding to the handwritten area, and comparing the handwritten extension direction with the horizontal direction corresponding to the original document to obtain a handwritten angle between the handwritten extension direction and the horizontal direction;

[0045] In response to the handwritten angle being within a preset angle range, determining a first handwritten weight for the corresponding direction dimension based on the handwritten angle;

[0046] Determining an area ratio of the handwritten area corresponding to the original document, and determining a second handwritten weight for the corresponding ratio dimension based on the area ratio;

[0047] Performing a summation calculation on the first handwritten weight and the second handwritten weight to obtain a handwritten reduction weight for the original document with the handwritten attribute;

[0048] Updating the extraction priority value corresponding to the original document by reducing it based on the handwritten reduction weight to obtain an updated extraction priority value.

[0049] Optionally, in the method according to the present invention, determining a handwritten extension direction corresponding to the handwritten area includes:

[0050] Establishing a file coordinate system corresponding to the original document, wherein the X-axis of the file coordinate system is parallel to the horizontal direction corresponding to the original document;

[0051] Performing pixelization processing on the handwritten area to obtain each handwritten pixel point constituting the handwritten area;

[0052] Obtaining each handwritten coordinate point corresponding to each handwritten pixel point based on the file coordinate system, and respectively determining the handwritten coordinate points corresponding to the maximum horizontal coordinate value, the minimum horizontal coordinate value, the maximum vertical coordinate value, and the minimum vertical coordinate value as the maximum horizontal coordinate point, the minimum horizontal coordinate point, the maximum vertical coordinate point, and the minimum vertical coordinate point;

[0053] Respectively using the maximum horizontal coordinate point and the minimum horizontal coordinate point as starting points to construct a first vertical line segment and a second vertical line segment extending in the vertical direction, and respectively using the maximum vertical coordinate point and the minimum vertical coordinate point as starting points to construct a first horizontal line segment and a second horizontal line segment extending in the horizontal direction;

[0054] Control the first longitudinal line segment, the second longitudinal line segment, the first transverse line segment, and the second transverse line segment to extend respectively to form an initial rectangle, and obtain the corner points corresponding to the initial rectangle;

[0055] Determine the center point of the rectangle corresponding to the initial rectangle, and obtain the connecting lines between each corner point and the center point of the rectangle;

[0056] Respectively take each corner point as a starting point, establish corner extension lines perpendicular to each connecting line, and move each corner extension line along the corresponding connecting line towards the center point of the rectangle;

[0057] In response to each corner extension line moving to contact different handwritten pixels respectively, control each corner extension line to stop moving respectively to form the current rectangle;

[0058] Obtain the rectangle frame lines forming the current rectangle, and determine the handwritten extension direction corresponding to the handwritten area based on the extension directions of the frame lines corresponding to the rectangle frame lines respectively.

[0059] Optionally, in the method according to the present invention, the method further includes:

[0060] Obtain the respective historical verification information corresponding to the respective original documents located in the random verification group, wherein each historical verification information includes the respective verification operations arranged in positive order based on the verification time, and each verification operation has corresponding respective verification results;

[0061] Perform reverse extraction on each historical verification information respectively based on the positive order, and determine the continuous verification operations obtained and corresponding to a preset number as a continuous verification group;

[0062] In response to all verification results corresponding to the same continuous verification group being verified, call a preset weight reduction to reduce and update the extraction priority value corresponding to the original document to obtain an updated extraction priority value.

[0063] Optionally, in the method according to the present invention, trigger the file verification plugin to perform integrity-based file verification on the current file based on the content verification group, and determine the document status corresponding to the current file based on the verification result, wherein the document status includes a complete status and a tampered status, including:

[0064] Retrieve from the current file, based on the file verification plug-in, each current file that has a traceability relationship with each of the original files located in the content verification group, and perform binarization processing on each original file and each current file respectively to obtain each original image and each current image, where any one of the original images includes each first image pixel point corresponding to a first pixel value, and any one of the current images includes each second image pixel point corresponding to a second pixel value;

[0065] Create a transparent fusion layer, and stack the original file and the current file with the same traceability relationship on the transparent fusion layer in sequence;

[0066] In response to any first image pixel point and / or any second image pixel point being included in the transparent fusion layer, determine the file status corresponding to the current file as the tampered status.

[0067] Optionally, in the method according to the present invention, based on the current file, establish an initial file comparison layer, fill the tampered content and the adjacent operation end into the content indication area and the terminal indication area located in the initial file comparison layer respectively, and send the obtained current file comparison layer to the management end for display, including:

[0068] Establish an initial file comparison layer based on the current file, where the initial file comparison layer includes a vertically arranged content indication area and a terminal indication area for filling the adjacent operation end;

[0069] Determine the current file in the current file where the tampered content exists, obtain each file line segment arranged in a vertical order from top to bottom in the current file, and determine all file line segments with blank groups composed of two adjacent blank characters as the paragraph start lines;

[0070] Based on the vertical order, respectively determine two adjacent paragraph start lines in adjacent positions as an adjacent paragraph group, and determine the paragraph start line corresponding to the first position in the same adjacent paragraph group as the first paragraph and the paragraph start line corresponding to the last position as the last paragraph;

[0071] Determine the first paragraph and all file line segments between the first paragraph and the last paragraph as the same file paragraph to obtain each file paragraph in the current file;

[0072] Determine the file paragraph including the tampered content as the tampered paragraph, and determine the original paragraph corresponding to the tampered paragraph in the original file having the same traceability relationship as the current file;

[0073] Generate a first bounding box that encloses the tampered paragraph in the current file, and generate a second bounding box that encloses the original paragraph in the original file;

[0074] Generate a first indicator area and a second indicator area arranged horizontally in the content indicator area, and fill the current file with the first bounding box into the first indicator area and fill the original file with the second bounding box into the second indicator area;

[0075] Fill the proximity operation end into the terminal indicator area, and send the obtained current file comparison layer to the management end for display.

[0076] According to another aspect of the present invention, there is provided an intelligent file management system based on a large language model, including:

[0077] An extraction module, configured to respond to any current operation end to perform a current scan on the file tag corresponding to the current file, retrieve the original file corresponding to the current file, randomly extract each original file included in the original file based on a preset extraction strategy, and determine each extracted original file as a content verification group;

[0078] A verification module, configured to trigger an archive verification plugin to perform file verification based on integrity on the current file based on the content verification group, and determine the document status corresponding to the current file based on the verification result, where the document status includes a complete status and a tampered status;

[0079] A determination module, configured to respond to the current file being in a tampered state, determine the tampered content corresponding to the current file, and determine the proximity operation end that performs a proximity scan on the current file based on the retrieved operation record table corresponding to the current file;

[0080] A creation module, configured to create an initial file comparison layer based on the current file, fill the tampered content and the proximity operation end into the content indicator area and the terminal indicator area located in the initial file comparison layer respectively, and send the obtained current file comparison layer to the management end for display.

[0081] According to the solution of the present invention, the present invention can randomly select original files to form a content verification group in response to the scanning of file tags by the operation terminal, significantly improving the efficiency of file verification. Subsequently, the present invention can efficiently verify files based on integrity through a file verification plugin, quickly and accurately determining the document status of the file (complete status or tampered status), thereby effectively ensuring the originality and authenticity of the file. Moreover, once the present invention detects that the file is in a tampered state, it can accurately identify and locate the tampered content, and at the same time quickly lock the adjacent operation terminal that performed the adjacent scan on the current file based on the operation record table. This dual confirmation mechanism not only helps to promptly discover and correct false information in the file, but also provides strong support for subsequent responsibility tracing. Finally, the present invention can present key information such as tampered content and adjacent operation terminals in a graphical manner by establishing an initial file comparison layer, enabling the management terminal to intuitively understand the true status and potential problems of the file. This visual management method not only facilitates the management personnel to quickly master the file verification results, but also greatly improves the convenience and transparency of file management. The present invention integrates various technical means such as random selection, integrity verification, large language model analysis, and graphical display, realizing efficient verification, accurate positioning, intuitive display, and intelligent management of files, which not only helps to improve the accuracy and efficiency of file management, but also lays a solid foundation for building a more secure and reliable file management system. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 FIG. shows a flowchart of an intelligent file management method based on a large language model according to an embodiment of the present invention;

[0083] Figure 2 FIG. shows a schematic diagram of the random selection turntable in this embodiment;

[0084] Figure 3 FIG. shows a structural block diagram of an intelligent file management system based on a large language model according to another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0085] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0086] To solve the problems existing in the background art, the inventors propose the solution of the present invention. An embodiment of the present invention provides an intelligent file management method based on a large language model, which can be executed in a computing device, where the computing device can be understood as a terminal with data processing capabilities.

[0087] Figure 1 FIG. 4 shows a flowchart of an intelligent file management method based on a large language model according to an embodiment of the present invention, which is suitable for execution in a computing device.

[0088] As Figure 1 shown, the intelligent file management method based on a large language model proposed in this embodiment starts from step S102, and in step S102, the following contents are included:

[0089] In response to any current operation terminal performing a current scan on the file tag corresponding to the current file, retrieve the original file corresponding to the current file, randomly extract each original file included in the original file based on a preset extraction strategy, and determine each extracted original file as a content verification group.

[0090] For example, in this embodiment, the current operation terminal can be understood as an operation terminal corresponding to the operator who performs corresponding operations on the current file. A corresponding file tag is set on each current file. When any operator expects to perform a corresponding viewing operation on the current file, the operator can perform a current scan on the file tag corresponding to the current file through the corresponding current operation terminal, so that the server will retrieve the original file corresponding to the current file, and then randomly extract each original file included in the original file according to the preset extraction strategy, thereby determining each extracted original file as a content verification group to perform the integrity comparison for the current file in the following content; here, the original file can be understood as the original version corresponding to the current file, that is, the file that has never been subjected to corresponding operations. The original file can exist in electronic form in the server, while the corresponding current file can exist in paper form and / or electronic form.

[0091] Furthermore, the above "randomly extract each original file included in the original file based on a preset extraction strategy, and determine each extracted original file as a content verification group" further includes the following steps:

[0092] Determine each original file with a pre-verification mark in the original file as a necessary verification group, and determine all other original files included in the original file except the necessary verification group as a random verification group;

[0093] Evaluate the content of each original document in the random verification group, and sum up the extraction priority values corresponding to each original document to obtain the total extraction value;

[0094] Calculate the ratio of each extraction priority value to the total extraction value to obtain the extraction probabilities corresponding to each original document;

[0095] Determine the number of files corresponding to each original document in the random verification group, and calculate the product of the number of files and the preset extraction ratio retrieved to obtain the current extraction quantity;

[0096] Create a random extraction turntable, and divide the area of the random extraction turntable based on the extraction probabilities corresponding to each original document to obtain the sector areas corresponding to each original document;

[0097] Control the random extraction turntable to rotate and extract the corresponding number of times of the current extraction quantity, and determine the extracted original documents and all the original documents in the necessary verification group as the content verification group based on the extraction results.

[0098] For example, in this embodiment, the server will first determine each original document with a pre-verification mark in the original file as the necessary verification group, and then determine all other original documents included in the original file except the necessary verification group as the random verification group; it can be stated that each original document with a pre-verification mark can be understood as an original document that needs to be compulsorily verified.

[0099] Then, the server will evaluate the content of each original document in the random verification group to obtain the extraction priority values corresponding to each original document, and then sum up the extraction priority values to obtain the total extraction value.

[0100] Next, the server will calculate the ratio of each extraction priority value to the total extraction value respectively to obtain the extraction probabilities corresponding to each original document, then determine the number of files of each original document in the content verification group, and calculate the product of the number of files and the preset extraction ratio retrieved to obtain the current extraction quantity.

[0101] Then, the server will create a random extraction turntable and divide the area of the random extraction turntable according to the extraction probabilities corresponding to each original document to obtain the sector areas corresponding to each original document, as Figure 2 shown in the random extraction turntable. If there are three original documents A, B, and C with corresponding extraction probabilities of 30%, 50%, and 20% respectively, then three sector areas with corresponding area ratios of 30%, 50%, and 20% will be generated in the random extraction turntable. Figure 2The black arrow in it is the corresponding extraction rotation needle, that is, when the extraction rotation needle rotates to any sector area, it means that the original file corresponding to that area is extracted.

[0102] Finally, the server will control the random extraction turntable to perform rotational extraction corresponding to the current extraction quantity. For example, if the current extraction quantity is 2, the server will control the random extraction turntable to perform two rotational extractions, so as to determine all the original files extracted according to the extraction results and all the original files located in the necessary verification group as the content verification group. Among them, each original file located in the content verification group is the file that needs to be verified accordingly.

[0103] Furthermore, the above "conduct content evaluation on each original file located in the random verification group" further includes the following steps:

[0104] Determine each original file located in the random verification group based on file attributes, and determine each original file with sensitive attributes as the sensitive file group and each original file with ordinary attributes as the ordinary file group;

[0105] Retrieve the first attribute evaluation value to configure scores for each original file located in the sensitive file group respectively, and retrieve the second attribute evaluation value to configure scores for each original file located in the ordinary file group respectively, where the first attribute evaluation value is greater than the second attribute evaluation value;

[0106] Perform pixel point recognition on each original file located in the random verification group respectively, obtain each file pixel point in the same original file, and obtain the pixel quantity corresponding to each file pixel point to obtain the respective pixel quantities corresponding to each original file;

[0107] Compare the pixel quantity with the retrieved first quantity threshold and second quantity threshold respectively, where the first quantity threshold is less than the second quantity threshold;

[0108] In response to the pixel quantity of any original file being less than the first quantity threshold, retrieve the first quantity evaluation value to configure a score for this original file based on the quantity dimension;

[0109] In response to the pixel quantity of any original file being greater than or equal to the first quantity threshold and less than the second quantity threshold, retrieve the second quantity evaluation value to configure a score for this original file based on the quantity dimension;

[0110] In response to the pixel quantity of any original file being greater than or equal to the second quantity threshold, retrieve the third quantity evaluation value to configure a score for this original file based on the quantity dimension;

[0111] Sum the attribute evaluation values and quantity evaluation values corresponding to the same original file to obtain the extraction priority values corresponding to each original file respectively.

[0112] For example, in this embodiment, since some original files may contain more important information, that is, have sensitive attributes, while some original files may contain more conventional information, that is, have ordinary attributes, the server will determine each original file in the random verification group based on the file attributes, so as to determine each original file with sensitive attributes as the sensitive file group and each original file with ordinary attributes as the ordinary file group.

[0113] Then, the server will retrieve the first attribute evaluation value to configure scores for each original file in the sensitive file group, and then retrieve the second attribute evaluation value to configure scores for each original file in the ordinary file group. Since the original files with sensitive attributes are more important, the first attribute evaluation value is greater than the second attribute evaluation value.

[0114] Next, the server will perform pixel point recognition on each original file in the random verification group respectively to obtain each file pixel point in the same original file, and then obtain the pixel quantity corresponding to each file pixel point in the same original file, that is, obtain the pixel quantity corresponding to each original file.

[0115] Different pixel quantities represent the number of words in the original files corresponding to different pixel quantities. The server will compare the pixel quantity with the retrieved first quantity threshold and second quantity threshold respectively. Here, the first quantity threshold is less than the second quantity threshold.

[0116] When the pixel quantity of any original file is less than the first quantity threshold, the server will retrieve the first quantity evaluation value to configure the score of this original file based on the quantity dimension; when the pixel quantity of any original file is greater than or equal to the first quantity threshold and less than the second quantity threshold, the server will retrieve the second quantity evaluation value to configure the score of this original file based on the quantity dimension; when the pixel quantity of any original file is greater than or equal to the second quantity threshold, the server will retrieve the third quantity evaluation value to configure the score of this original file based on the quantity dimension.

[0117] Then, the server will sum the attribute evaluation value and quantity evaluation value corresponding to the same original file to obtain the extraction priority value corresponding to each original file respectively.

[0118] Further, the above-mentioned "determining each original document in the random verification group based on document attributes, and determining each original document with sensitive attributes as a sensitive document group and each original document with normal attributes as a normal document group" further includes the following steps:

[0119] Obtain the page numbers of each original document respectively, and determine the original document with the same page number as the retrieved preset directory page number in the original file as the directory file;

[0120] Extract the directory from the directory file to obtain each directory title and each page number range corresponding to each directory title respectively;

[0121] Classify each directory title hierarchically to obtain a hierarchical structure tree including different hierarchical nodes, where each hierarchical node corresponds to a different directory title respectively;

[0122] Retrieve a preset sensitive record table, where the preset sensitive record table includes different sensitive characters;

[0123] Based on the hierarchical order of the hierarchical structure tree, determine each directory title corresponding to each hierarchical node respectively. In response to at least a part of the directory title corresponding to any hierarchical node being located in the preset sensitive record table, determine this hierarchical node and all hierarchical nodes connected downward to this hierarchical node as having sensitive attributes;

[0124] Based on each page number range corresponding to each hierarchical node with sensitive attributes respectively, determine all original documents in the random verification group whose corresponding file page numbers are located in the page number range as having sensitive attributes, and determine the remaining all original documents in the random verification group as having normal attributes.

[0125] For example, in this embodiment, the server will first obtain the page numbers of each original document, so as to determine the original document with the same page number as the retrieved preset directory page number in the original file as the directory file. Since usually, the page number where the directory is located is mostly the second page, that is, the directory page number corresponding to the directory file is 2, so the preset directory page number preset in the server can be 2, and thus the server can determine the corresponding directory file according to the preset directory page number.

[0126] After that, the server will perform directory extraction on the directory file to obtain each directory title and the corresponding page number intervals for each directory title. Since in a directory file, there may be directory titles corresponding to different levels, such as first-level directory titles, second-level directory titles, etc. Therefore, the server will classify the levels of each directory title to obtain a hierarchical structure tree including various different-level nodes, and each level node in the hierarchical structure tree corresponds to a different directory title.

[0127] Next, the server will retrieve a preset sensitive record table, which includes various different sensitive characters. For example, the sensitive characters can be disciplinary records, etc.

[0128] At this time, the server will determine the directory titles corresponding to each level node according to the hierarchical order of the hierarchical structure tree. When at least part of the content in the directory title corresponding to any level node is located in the preset sensitive record table, the server will determine that this level node and all level nodes connected downward to this level node have sensitive attributes.

[0129] After that, the server will determine that all the original files in the random verification group whose corresponding file page numbers are within the page number intervals corresponding to each level node with sensitive attributes have sensitive attributes, and then determine that all the remaining original files in the random verification group have ordinary attributes.

[0130] It can be explained that this embodiment can quickly determine the file attributes of each original file based on the comparison between the hierarchical structure tree and the preset sensitive record table, improving the corresponding data processing efficiency.

[0131] Furthermore, the above "performing directory extraction on the directory file to obtain each directory title and the corresponding page number intervals for each directory title" further includes the following steps:

[0132] Obtain each file line segment arranged in a vertical order from top to bottom in the directory file, and determine all file line segments with leading characters as each directory title;

[0133] Perform character recognition on each directory title based on a large language model to obtain each file character in the same directory title;

[0134] Group the file characters in the same directory title by category to obtain a text group including file characters corresponding to text categories and a number group including file characters corresponding to number categories;

[0135] Determine the title page numbers corresponding to the directory titles based on each file character in the digital grouping, and obtain the respective title page numbers corresponding to each directory title;

[0136] Determine the size of any file character in the text grouping respectively, and obtain the respective title sizes corresponding to each text grouping;

[0137] Arrange the directory titles in a hierarchical order from largest to smallest according to the corresponding title sizes, and obtain the respective hierarchical levels corresponding to each directory title;

[0138] Based on the vertical order, respectively determine two directory titles at adjacent positions and corresponding to the same hierarchical level as an adjacent title group, and determine the directory title corresponding to the first position in the same adjacent title group as the first title and the directory title corresponding to the last position as the last title;

[0139] Determine the title page number corresponding to the first title and the respective title page numbers corresponding to all directory titles between the first title and the last title as an interval determination group, and determine the maximum page number corresponding to the largest value and the minimum page number corresponding to the smallest value based on all the title page numbers in the interval determination group;

[0140] Determine the page number interval composed of the minimum page number and the maximum page number as corresponding to the first title, and obtain the respective page number intervals corresponding to each directory title.

[0141] For example, in this embodiment, the server will obtain each file line arranged in a vertical order from top to bottom in the directory file. Since in general, there is a leading symbol connecting the directory title and the corresponding directory page number, the server can determine all file lines with leading symbols as each directory title.

[0142] Then, the server will perform character recognition on each directory title based on the large language model respectively, so as to obtain each file character in the same directory title, and then group the file characters in the same directory title by category, so as to obtain a text grouping including file characters corresponding to text categories and a digital grouping including file characters corresponding to digital categories.

[0143] Next, the server will determine the title page numbers corresponding to the respective directory titles based on each file character located in the digital grouping, thereby obtaining the respective title page numbers corresponding to each directory title. Since, under normal circumstances, the sizes of each file character located in the same directory title will be consistent, and the title sizes corresponding to different hierarchical levels of directory titles are also different, therefore, the respective hierarchical levels corresponding to each directory title can be determined based on recognizing the corresponding title sizes. Here, generally speaking, the larger the title size, the higher the corresponding hierarchical level. For example, when there are three different directory titles corresponding to the third chapter, the first section of the third chapter, and the second section of the third chapter, the title size of the directory title corresponding to the third chapter should be the largest, while the titles corresponding to the first section of the third chapter and the second section of the third chapter should be smaller and the same. Further, in order to reduce a certain amount of data processing by the server, the server will respectively determine the size of any file character located in the text grouping, thereby obtaining the respective title sizes corresponding to each text grouping.

[0144] Then, the server will arrange the various directory titles in a hierarchical order from largest to smallest according to the title size, thereby obtaining the respective hierarchical levels corresponding to each directory title. Since in a general directory file, the various directory titles are arranged in a vertical order, and the directory titles corresponding to relatively lower hierarchical levels are generally located behind the directory titles corresponding to relatively higher hierarchical levels, therefore, based on this characteristic, two directory titles that are adjacent in position and correspond to the same hierarchical level can be respectively determined as an adjacent title group, and the directory title corresponding to the first position in the same adjacent title group can be determined as the first title and the directory title corresponding to the last position can be determined as the last title.

[0145] After that, the server will determine the title page number corresponding to the first title and all the title page numbers corresponding to the respective directory titles located between the first title and the last title as an interval determination group, and determine the maximum page number corresponding to the largest value and the minimum page number corresponding to the smallest value based on this interval determination group.

[0146] Finally, the server will determine the page number interval composed of the minimum page number and the maximum page number as corresponding to the first title, thereby obtaining the respective page number intervals corresponding to each directory title.

[0147] It can be explained that in this embodiment, the page number intervals corresponding to each directory title can be quickly obtained based on the layout characteristics of the various directory titles located in the directory file, thereby improving the corresponding acquisition accuracy and data processing efficiency.

[0148] Here, the aforementioned large language model (LLM) refers to a deep learning model trained using a large amount of text data. This model can understand and generate natural language text. Since the large language model is prior art, this embodiment will not elaborate on it specifically.

[0149] Furthermore, the above method further includes the following steps:

[0150] Perform handwriting recognition on each original document in the random verification group based on a handwriting recognition model, and in response to determining that there is a handwriting area in any original document based on the recognition result, determine that the original document has a handwriting attribute;

[0151] Determine the handwriting extension direction corresponding to the handwriting area, compare the handwriting extension direction with the horizontal direction corresponding to the original document, and determine the first handwriting weight in the corresponding direction dimension based on the obtained handwriting extension direction;

[0152] Determine the area ratio of the handwriting area corresponding to the original document, and determine the second handwriting weight in the corresponding ratio dimension based on the area ratio;

[0153] Perform a summation calculation on the first handwriting weight and the second handwriting weight to obtain the handwriting increase weight of the original document with a handwriting attribute;

[0154] Increase and update the extraction priority value corresponding to the original document based on the handwriting increase weight to obtain the updated extraction priority value.

[0155] For example, in an actual application scenario, the original document generally includes various document characters corresponding to machine printing. In some cases, for managers with management authority over the original document, in order to be able to record important content in real time based on the original file, they may add corresponding handwriting areas to at least one original document of the original file. The handwriting area may include file characters formed by handwriting, etc.; based on the occurrence of this situation, the extraction priority value of the original document with a handwriting area should be increased accordingly to increase the extraction probability of the corresponding original document. Therefore, it can be achieved based on the following method steps:

[0156] First, in order to determine whether there is a corresponding handwriting area in each original document in the random verification group, a pre-trained handwriting recognition model can be used to perform corresponding handwriting recognition on each original document. When it is determined that there is a handwriting area in any original document, the original document is determined to have a handwriting attribute; here, the handwriting recognition model can be established based on a neural network learning model or a machine learning model. Since its establishment process and training process are relatively common, this embodiment will not elaborate on it.

[0157] Next, it can be explained that when the angle between the handwriting extension direction corresponding to the handwriting area and the horizontal direction is smaller, it indicates that the corresponding management personnel may be in a more focused and concentrated state when writing. Therefore, the importance of the text characters corresponding to the handwriting area may be higher. Thus, by comparing the handwriting extension direction of the handwriting area with the horizontal direction of the corresponding original document and obtaining the corresponding handwriting angle based on the comparison result, here, it can be known that when the handwriting angle is larger, the corresponding importance is smaller, and when the corresponding handwriting angle is smaller, the corresponding importance is larger. Therefore, the first handwriting weight corresponding to the direction dimension can be obtained based on the following formula:

[0158]

[0159] wherein, is the first handwriting weight, is the handwriting angle, is the preset direction weight retrieved;

[0160] Then, when the area ratio between the handwriting area and the corresponding original document is larger, it indicates that there are more text characters in the corresponding handwriting area. Therefore, the importance of the text characters corresponding to the handwriting area may be higher. Thus, the second handwriting weight corresponding to the ratio dimension can be determined by the area ratio between the handwriting area and the corresponding original document. Here, it can be known that when the area ratio is larger, the corresponding importance is larger, and when the corresponding area ratio is smaller, the corresponding importance is smaller. Therefore, the second handwriting weight corresponding to the ratio dimension can be obtained based on the following formula:

[0161]

[0162] wherein, is the second handwriting weight, is the area ratio, is the preset ratio weight retrieved;

[0163] Finally, after obtaining the corresponding first handwriting weight and second handwriting weight, by performing a summation calculation on the two, the extraction priority value corresponding to the original document can be increased and updated (i.e., increased accordingly) based on the obtained handwriting increase weight, and the updated extraction priority value can be obtained.

[0164] Furthermore, the above-mentioned "determining the handwriting extension direction corresponding to the handwriting area" further includes the following steps:

[0165] Establish a file coordinate system corresponding to the original document, wherein the X-axis of the file coordinate system is parallel to the horizontal direction of the corresponding original document;

[0166] Pixelize the handwritten area to obtain each handwritten pixel point that makes up the handwritten area;

[0167] Based on the document coordinate system, obtain each handwritten coordinate point corresponding to each handwritten pixel point respectively, and determine the handwritten coordinate points corresponding to the maximum horizontal coordinate value, the minimum horizontal coordinate value, the maximum vertical coordinate value, and the minimum vertical coordinate value as the maximum horizontal coordinate point, the minimum horizontal coordinate point, the maximum vertical coordinate point, and the minimum vertical coordinate point respectively;

[0168] Taking the maximum horizontal coordinate point and the minimum horizontal coordinate point as starting points respectively, construct a first vertical line segment and a second vertical line segment extending in the vertical direction, and taking the maximum vertical coordinate point and the minimum vertical coordinate point as starting points respectively, construct a first horizontal line segment and a second horizontal line segment extending in the horizontal direction;

[0169] Control the first vertical line segment, the second vertical line segment, the first horizontal line segment, and the second horizontal line segment to extend respectively to form an initial rectangle, and obtain each corner point position corresponding to the initial rectangle;

[0170] Determine the rectangle center point corresponding to the initial rectangle, and obtain each position connection line between each corner point position and the rectangle center point;

[0171] Taking each corner point position as a starting point respectively, establish each corner extension line perpendicular to each position connection line, and move each corner extension line along the corresponding position connection line towards the rectangle center point;

[0172] In response to each corner extension line moving to contact different handwritten pixel points respectively, control each corner extension line to stop moving respectively to form the current rectangle;

[0173] Obtain each rectangle frame line that makes up the current rectangle, and determine the handwritten extension direction corresponding to the handwritten area based on each frame line extension direction corresponding to each rectangle frame line.

[0174] For example, in this embodiment, the handwritten extension direction corresponding to the handwritten area can be obtained in the following manner:

[0175] First, a document coordinate system corresponding to the original document can be established, and at the same time, the handwritten area in the original document is pixelized to obtain each handwritten pixel point that makes up the handwritten area. Here, the X-axis of the document coordinate system is parallel to the horizontal direction of the corresponding original document;

[0176] After that, the server can obtain respective handwritten coordinate points corresponding to each handwritten pixel point based on the file coordinate system, and then determine the handwritten coordinate points corresponding to the maximum horizontal coordinate value, the minimum horizontal coordinate value, the maximum vertical coordinate value, and the minimum vertical coordinate value as the maximum horizontal coordinate point, the minimum horizontal coordinate point, the maximum vertical coordinate point, and the minimum vertical coordinate point, respectively.

[0177] Next, the server constructs a first vertical line segment and a second vertical line segment extending along the vertical direction with the maximum horizontal coordinate point and the minimum horizontal coordinate point as starting points, respectively, and constructs a first horizontal line segment and a second horizontal line segment extending along the horizontal direction with the maximum vertical coordinate point and the minimum vertical coordinate point as starting points, respectively.

[0178] Then, the server controls the first vertical line segment, the second vertical line segment, the first horizontal line segment, and the second horizontal line segment to extend respectively to form an initial rectangle, and obtains respective corner points corresponding to the initial rectangle area;

[0179] Then, the server first determines the rectangle center point corresponding to the initial rectangle area, and then obtains respective connecting lines between each corner point and the rectangle center point. Thus, respective corner extension lines perpendicular to each connecting line are established with each corner point as a starting point;

[0180] Then, the server moves each corner extension line forward along the corresponding connecting line towards the rectangle center point. When any one corner extension line contacts any handwritten pixel point, the movement of the corner extension line is controlled to stop.

[0181] Finally, when each corner extension line moves to contact different handwritten pixel points respectively, the movement of each corner extension line can be controlled to stop respectively, thereby forming a corresponding current rectangle. By obtaining each rectangle frame line constituting the current rectangle, and determining the extension direction of the longest rectangle frame line in each rectangle frame line as the handwritten extension direction corresponding to the handwritten area.

[0182] Furthermore, for example, in this embodiment, the above method steps include updating the corresponding extraction priority value based on the original file with handwritten attributes. And in some cases, in order to further improve the security of the corresponding file management, when any original file is drawn continuously for multiple times, the extraction priority value of the original file can be updated by reducing it accordingly, reducing the extraction probability corresponding to the original file, so as to ensure that other original files can have a higher probability of being drawn and verified, in order to increase the comprehensiveness of extraction. And the corresponding process can be implemented based on the following method steps:

[0183] Obtain respective historical verification information corresponding to each original document located in the random verification group, where each of the historical verification information includes respective verification operations arranged in a forward order based on the verification time, and each verification operation has corresponding respective verification results;

[0184] Perform reverse extraction on each historical verification information respectively based on the reverse order of the forward arrangement, and determine the continuously arranged respective verification operations corresponding to a preset quantity as a continuous verification group;

[0185] In response to all verification results corresponding to the same continuous verification group being verified, retrieve a preset weight reduction to perform a reduction update on the extraction priority value corresponding to the original document, and obtain the updated extraction priority value.

[0186] For example, in this embodiment, during random extraction, respective historical verification information corresponding to each original document located in the random verification group can be obtained. Among them, each historical verification information includes respective verification operations arranged in a forward order based on the verification time, and each verification operation has corresponding respective verification results. That is, when the original file has been verified three times, the corresponding verification time includes the corresponding three time nodes, and each original document has a verification operation corresponding to each verification time. The verification operation can include verified and unverified. Verified can be understood as the original document being extracted and verified at the corresponding verification time, and unverified can be understood as the original document not being extracted and verified at the corresponding verification time. By performing reverse extraction on each historical verification information respectively based on the forward arrangement, and determining the continuously arranged respective verification operations corresponding to a preset quantity as a continuous verification group. When all verification results located in the same continuous verification group are verified, it indicates that the original document has been extracted and verified in all the corresponding preset quantity of verification operations recently. At this time, the corresponding preset weight can be retrieved from the server to perform a corresponding reduction update on the extraction priority value corresponding to the original document, so as to reduce the extraction probability of the original document in subsequent verification operations and increase the extraction comprehensiveness.

[0187] In step S104, the following contents are included:

[0188] Trigger the file verification plugin to perform integrity-based file verification on the current file based on the content verification group, and determine the document status corresponding to the current file based on the verification result, where the document status includes a complete status and a tampered status.

[0189] For example, in this embodiment, after determining the content verification group, the server triggers a pre-set file verification plugin to perform corresponding verification operations on the integrity of the current document based on each original file in the content verification group, so as to determine the document status corresponding to the current file according to the verification result. Among them, when the current file has not been tampered with, the corresponding document status is the complete status, and when the current file has been tampered with, the corresponding document status is the tampered status.

[0190] It should be noted that based on the foregoing content, the current file can exist in paper form and / or electronic form. In this embodiment, generally, the current file exists in paper form. However, when any current operating terminal performs a corresponding viewing operation on the current file, the current operating terminal needs to perform an integrity scan on the current file and store the corresponding electronic form of the current file in the server for corresponding archiving, so that when the next current operating terminal needs to perform a viewing operation on the current file, it can first retrieve the current file stored in the server to compare its integrity with the original document to determine whether the current file has been tampered with by the previous current operating terminal, thereby improving the security of corresponding file management.

[0191] Further, the above-mentioned "trigger the file verification plugin to perform integrity-based file verification on the current file based on the content verification group, and determine the document status corresponding to the current file according to the verification result, where the document status includes the complete status and the tampered status" further includes the following steps:

[0192] Based on the file verification plugin, retrieve each current file from the current file that has a traceability relationship with each original file in the content verification group, and perform binarization processing on each original file and each current file respectively to obtain each original image and each current image. Among them, any one of the original images includes each first image pixel corresponding to a first pixel value, and any one of the current images includes each second image pixel corresponding to a second pixel value;

[0193] Create a transparent fusion layer, and stack the original file and the current file with the same traceability relationship on the transparent fusion layer in sequence;

[0194] In response to any first image pixel and / or any second image pixel being included in the transparent fusion layer, determine the file status corresponding to the current file as the tampered status.

[0195] For example, in this embodiment, since the current file can be stored in the server in electronic form, the file verification plugin can retrieve from the current file each current file that has a traceability relationship with each original file located in the content verification group, and determine whether the current file has been tampered with based on the one-to-one file comparison between each current file and each original file. Here, the corresponding file comparison process can be implemented based on the following method steps:

[0196] First, each original file and each current file can be binarized respectively to obtain the corresponding original images and current images. Among them, any original image includes each first image pixel point corresponding to the first pixel value, and any current image includes each second image pixel point corresponding to the second pixel value;

[0197] Next, a transparent fusion layer can be created, and the original file and the current file with the same traceability relationship can be sequentially superimposed on the transparent fusion layer at the same time;

[0198] Finally, it can be stated that if the current file has not been tampered with, only the fusion pixel value corresponding to the first pixel value and the second pixel value should appear in the corresponding transparent fusion layer. For example, when the first pixel value corresponds to yellow and the second pixel value is blue, then the corresponding fusion pixel value corresponds to green; and if the current file has been tampered with, any first image pixel point and / or any second image pixel point will appear in the transparent fusion layer. Therefore, the file status of the current file can be determined as the tampered status.

[0199] In step S106, the following content is included:

[0200] In response to the current file being in the tampered status, determine the tampered content corresponding to the current file, and based on the retrieved operation record table corresponding to the current file, determine the adjacent operation terminal for performing an adjacent scan on the current file.

[0201] For example, in this embodiment, when the document status corresponding to the current file is the tampered status, for example, in this embodiment, based on the foregoing content, it can be known that after the corresponding current document is viewed by the previous operation terminal, it is necessary to perform an integrity scan on the current document for archiving in the server. Therefore, when the document status corresponding to the current file is the tampered status, the server will further determine the tampered content corresponding to the current file, and in order to ensure the accuracy of subsequent responsibility tracing, the server will also determine the adjacent operation terminal for performing an adjacent scan (i.e., a view operation) on the current file according to the retrieved operation record table corresponding to the current file.

[0202] In step S108, the following content is included:

[0203] An initial file comparison layer is established based on the current file, the tampered content and the adjacent operation terminal are respectively filled into the content indication area and the terminal indication area located in the initial file comparison layer, and the obtained current file comparison layer is sent to the management terminal for display.

[0204] For example, in this embodiment, in order to record relevant tampered content and the corresponding temporary operation terminal and inform the management terminal, the server will establish an initial file comparison layer according to the current file, so as to fill the tampered content and the adjacent operation terminal into the content indication area and the terminal indication area located in the initial file comparison layer respectively, and send the obtained current file comparison layer to the management terminal for display.

[0205] It should be noted that the management terminal can be understood as the terminal device used by the corresponding management personnel.

[0206] Furthermore, the above-mentioned "establish an initial file comparison layer based on the current file, fill the tampered content and the adjacent operation terminal into the content indication area and the terminal indication area located in the initial file comparison layer respectively, and send the obtained current file comparison layer to the management terminal for display" further includes the following steps:

[0207] An initial file comparison layer is established based on the current file, wherein the initial file comparison layer includes a vertically arranged content indication area and a terminal indication area for filling the adjacent operation terminal;

[0208] Determine the current file with the tampered content in the current file, obtain each file line drop arranged in a vertical order from top to bottom in the current file, and determine all file line drops with blank groups composed of two adjacent blank characters as the paragraph starting lines;

[0209] Based on the vertical order, respectively determine two adjacent paragraph starting lines in adjacent positions as an adjacent paragraph group, and determine the corresponding first paragraph starting line in the same adjacent paragraph group as the first paragraph and the corresponding last paragraph starting line as the last paragraph;

[0210] Determine the first paragraph and all file line drops between the first paragraph and the last paragraph as the same file paragraph, and obtain each file paragraph in the current file;

[0211] Determine the file paragraph including the tampered content as the tampered paragraph, and determine the original paragraph corresponding to the tampered paragraph in the original file having the same traceability relationship as the current file;

[0212] Generate a first bounding box surrounding the tampered paragraph in the current file, and generate a second bounding box surrounding the original paragraph in the original file;

[0213] Generate a first indication area and a second indication area arranged horizontally in the content indication area, and fill the current file with the first bounding box into the first indication area and fill the original file with the second bounding box into the second indication area;

[0214] Fill the adjacent operation end into the terminal indication area, and send the obtained current file comparison layer to the management end for display.

[0215] For example, in this embodiment, when it is determined that the current file is in a tampered state, the server can establish a corresponding initial file comparison layer based on the current file. Here, the initial file comparison layer can include a vertically arranged content indication area and a terminal indication area; since any current file is generally composed of different file paragraphs, when there is corresponding tampered content in any current file, the file paragraph where the tampered content is located can be determined based on the tampered content, and further the corresponding file paragraph can be determined as the tampered paragraph, and the original paragraph corresponding to the tampered paragraph can be obtained in the original file; after determining the corresponding tampered paragraph and the original paragraph, a first bounding box surrounding the tampered paragraph can be generated in the current file, and a second bounding box surrounding the original paragraph can be generated in the original file; after completing the corresponding generation of the first bounding box and the second bounding box, a first indication area and a second indication area arranged horizontally can be further generated in the content indication area located in the initial file comparison layer, and the current file with the first bounding box can be filled into the first indication area, and the original file with the second bounding box can be synchronously filled into the second indication area; finally, the adjacent operation end can be correspondingly filled into the terminal indication area to complete the update of the initial file comparison layer, and then the obtained current file comparison layer can be sent to the management end for display, facilitating the management end to simultaneously obtain the corresponding tampered content and the corresponding temporary operation end, and improving the corresponding information traceability.

[0216] It should be noted that, in order to quickly obtain each file paragraph in the current file, the following method steps can be implemented:

[0217] First, it can be known that there are file lines arranged in a vertical order from top to bottom in the current file. Based on common writing habits, generally, the beginning of each file paragraph is generally composed of two adjacent blank characters. Therefore, all file lines with blank groups composed of two adjacent blank characters can be determined as paragraph start lines;

[0218] Next, based on the corresponding vertical order, two paragraph starting lines at adjacent positions can be determined as an adjacent paragraph group, and further, the paragraph starting line corresponding to the first position in the same adjacent paragraph group can be determined as the first paragraph, and the paragraph starting line corresponding to the last position can be determined as the last paragraph;

[0219] Finally, the first paragraph and all file lines between the first paragraph and the last paragraph can be determined as the same file paragraph, and then the acquisition of each file paragraph in the current file can be quickly completed.

[0220] According to the solution of this embodiment, this embodiment can respond to the scanning of the file label by the operation terminal, randomly extract the original files to form a content verification group, significantly improving the efficiency of file verification; subsequently, this embodiment can implement efficient file verification of the current file based on integrity through the file verification plug-in, quickly and accurately determine the document status of the file (complete status or tampered status), thereby effectively ensuring the originality and authenticity of the file; and once this embodiment detects that the file is in the tampered status, it can accurately identify and locate the tampered content, and at the same time quickly lock the adjacent operation terminal that performs adjacent scanning on the current file based on the operation record table; this dual confirmation mechanism not only helps to promptly discover and correct the false information in the file, but also provides strong support for subsequent responsibility tracing; finally, this embodiment can present key information such as the tampered content and the adjacent operation terminal in a graphical manner by establishing an initial file comparison layer, enabling the management terminal to intuitively understand the true status and potential problems of the file; this visual management method not only facilitates the management personnel to quickly master the file verification results, but also greatly improves the convenience and transparency of file management; this embodiment integrates various technical means such as random extraction, integrity verification, large language model analysis, and graphical display, realizing efficient verification, accurate positioning, intuitive display, and intelligent management of files, which not only helps to improve the accuracy and efficiency of file management, but also lays a solid foundation for building a more secure and reliable file management system.

[0221] Another embodiment of the present invention provides an intelligent file management system based on a large language model, Figure 3 For its corresponding system block diagram, the system includes:

[0222] An extraction module, configured to respond to the current scanning of the file label corresponding to the current file by any current operation terminal, retrieve the original file corresponding to the current file, randomly extract each original file included in the original file based on a preset extraction strategy, and determine each extracted original file as a content verification group;

[0223] A verification module, configured to trigger an archive verification plugin to perform integrity-based document verification on the current archive based on the content verification group, and determine a document status corresponding to the current archive based on the verification result, where the document status includes a complete status and a tampered status;

[0224] A determination module, configured to, in response to the current archive being in a tampered status, determine tampered content corresponding to the current archive, and determine an adjacent operation end for performing adjacent scanning on the current archive based on an operation record table retrieved corresponding to the current archive;

[0225] A creation module, configured to create an initial archive comparison layer based on the current archive, fill the tampered content and the adjacent operation end into a content indication area and a terminal indication area located in the initial archive comparison layer respectively, and send the obtained current archive comparison layer to a management end for display.

[0226] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the examples of the present invention. The structure required to construct such a system will be apparent from the above description. Additionally, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the descriptions of specific languages above are for disclosing the preferred embodiments of the present invention.

[0227] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0228] Similarly, it should be understood that, in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof.

[0229] Those skilled in the art should understand that the modules, units, or components of the devices in the examples disclosed herein may be arranged in the devices as described in this embodiment, or alternatively may be located in one or more devices different from the devices in this example. The modules in the foregoing examples may be combined into one module or further divided into multiple sub-modules.

[0230] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components.

[0231] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments.

[0232] In addition, some of the embodiments herein are described as a combination of methods or method elements that can be implemented by a processor of a computer system or by other devices performing the functions. Therefore, a processor having the necessary instructions for implementing the method or method elements forms a device for implementing the method or method elements. In addition, the elements described herein in the device embodiments are examples of the following devices: the device is used to implement the functions performed by the elements for the purpose of implementing the present invention.

[0233] As used herein, unless otherwise specified, the use of ordinal numbers "first", "second", "third", etc. to describe ordinary objects only represents different instances of similar objects, and does not intend to imply that the objects so described must have a given order in terms of time, space, sorting, or in any other way.

[0234] Although the present invention is described based on a limited number of embodiments, those skilled in the art of this technology understand that other embodiments can be conceived within the scope of the present invention described herein. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes, rather than for the purpose of interpreting or limiting the subject matter of the present invention.

Claims

1. An intelligent file management method based on a large language model, characterized in that, Including the following steps: In response to any current operation terminal scanning the file label corresponding to the current file, retrieving the original file corresponding to the current file, determining each original document marked for pre-verification in the original file as a necessary verification group, and determining all other original documents included in the original file except the necessary verification group as a random verification group; Evaluating the content of each original document in the random verification group, and summing up the extraction priority values respectively corresponding to each original document to obtain a total extraction value; Calculating the ratio of each extraction priority value to the total extraction value respectively to obtain the extraction probabilities respectively corresponding to each original document; Determining the number of files corresponding to each original document in the random verification group, and calculating the product based on the number of files and the preset extraction ratio retrieved to obtain the current extraction number; Creating a random extraction turntable, and dividing the random extraction turntable into regions based on the extraction probabilities respectively corresponding to each original document to obtain the sector regions respectively corresponding to each original document; Controlling the random extraction turntable to rotate and extract according to the current extraction number, and determining the extracted original documents and all original documents in the necessary verification group as a content verification group based on the extraction result; Triggering the file verification plugin to perform integrity-based file verification on the current file based on the content verification group, and determining the document status corresponding to the current file based on the verification result, where the document status includes a complete status and a tampered status; In response to the current file being in a tampered status, determining the tampered content corresponding to the current file, and determining the adjacent operation terminal for performing adjacent scanning on the current file based on the operation record table retrieved corresponding to the current file; Establishing an initial file comparison layer based on the current file, filling the tampered content and the adjacent operation terminal into the content indication area and the terminal indication area respectively located in the initial file comparison layer, and sending the obtained current file comparison layer to the management end for display.

2. The intelligent file management method based on a large language model according to claim 1, wherein Evaluating the content of each original document in the random verification group includes: Determining each original document in the random verification group based on file attributes, and determining each original document with sensitive attributes as a sensitive document group and each original document with common attributes as a common document group; Retrieving the first attribute evaluation value to configure scores for each original document in the sensitive document group respectively, and retrieving the second attribute evaluation value to configure scores for each original document in the common document group respectively, where the first attribute evaluation value is greater than the second attribute evaluation value; Performing pixel point recognition on each original document in the random verification group respectively to obtain each file pixel point in the same original document, and obtaining the pixel quantity corresponding to each file pixel point to obtain the pixel quantities respectively corresponding to each original document; Compare the number of pixels with the retrieved first quantity threshold and second quantity threshold respectively, where the first quantity threshold is less than the second quantity threshold; In response to the number of pixels corresponding to any one of the original files being less than the first quantity threshold, retrieve the first quantity evaluation value to configure the score of the original file based on the quantity dimension; In response to the number of pixels corresponding to any one of the original files being greater than or equal to the first quantity threshold and less than the second quantity threshold, retrieve the second quantity evaluation value to configure the score of the original file based on the quantity dimension; In response to the number of pixels corresponding to any one of the original files being greater than or equal to the second quantity threshold, retrieve the third quantity evaluation value to configure the score of the original file based on the quantity dimension; Sum up the attribute evaluation value and the quantity evaluation value corresponding to the same original file to obtain the respective extraction priority values corresponding to each original file.

3. The intelligent file management method based on a large language model according to claim 2, wherein Determine each original file in the random verification group based on file attributes, and determine each original file with sensitive attributes as a sensitive file group and each original file with ordinary attributes as an ordinary file group, including: Obtain the respective file page numbers corresponding to each original file, and determine the original file with the same page number as the retrieved preset directory page number as the directory file in the original file; Extract the directory from the directory file to obtain each directory title and each page number interval corresponding to each directory title; Classify the levels of each directory title to obtain a hierarchical structure tree including different hierarchical nodes, where each hierarchical node corresponds to a different directory title; Retrieve a preset sensitive record table, where the preset sensitive record table includes different sensitive characters; Determine each directory title corresponding to each hierarchical node based on the hierarchical order of the hierarchical structure tree. In response to at least a part of the directory title corresponding to any hierarchical node being located in the preset sensitive record table, determine the hierarchical node and all hierarchical nodes connected downward to the hierarchical node as having sensitive attributes; Based on each page number interval corresponding to each hierarchical node with sensitive attributes, determine all the original files in the random verification group whose corresponding file page numbers are located in the page number interval as having sensitive attributes, and determine the remaining all original files in the random verification group as having ordinary attributes.

4. The intelligent file management method based on a large language model according to claim 3, wherein Extract the directory from the directory file to obtain each directory title and each page number interval corresponding to each directory title, including: Obtain each file line drop arranged in a vertical order from top to bottom in the directory file, and determine all file line drops with leading characters as each directory title; Perform character recognition on each directory title based on a large language model to obtain each file character in the same directory title; Group the file characters in the same directory title by category to obtain a text grouping including file characters corresponding to respective text categories and a number grouping including file characters corresponding to respective number categories; Determine a title page number corresponding to the directory title based on each file character in the number grouping to obtain respective title page numbers corresponding to the respective directory titles; Determine the size of any file character in the text grouping respectively to obtain respective title sizes corresponding to the respective text groupings; Arrange the directory titles in a hierarchical order from largest to smallest according to the corresponding title sizes to obtain respective hierarchical levels corresponding to the respective directory titles; Based on the vertical order, respectively determine two directory titles at adjacent positions and corresponding to the same hierarchical level as an adjacent title group, and determine the directory title corresponding to the first position in the same adjacent title group as the first title and the directory title corresponding to the last position as the last title; Determine the title page numbers corresponding to the first title and all directory titles between the first title and the last title as an interval determination group, and determine the maximum page number corresponding to the largest value and the minimum page number corresponding to the smallest value based on all the title page numbers in the interval determination group; Determine the page number interval composed of the minimum page number and the maximum page number as corresponding to the first title to obtain respective page number intervals corresponding to the respective directory titles.

5. The intelligent file management method based on a large language model according to claim 2, wherein, The method further includes: Perform handwriting recognition on each original file in the random verification group based on a handwriting recognition model, and in response to determining that there is a handwriting area in any original file based on the recognition result, determine that the original file has a handwriting attribute; Determine a handwriting extension direction corresponding to the handwriting area, compare the handwriting extension direction with the horizontal direction corresponding to the original file, and determine a first handwriting weight corresponding to the direction dimension based on the obtained handwriting extension direction; Determine the area ratio of the handwriting area corresponding to the original file, and determine a second handwriting weight corresponding to the ratio dimension based on the area ratio; Perform a summation calculation on the first handwriting weight and the second handwriting weight to obtain a handwriting increase weight corresponding to the original file with the handwriting attribute; Increase and update the extraction priority value corresponding to the original file based on the handwriting increase weight to obtain an updated extraction priority value.

6. The intelligent file management method based on a large language model according to claim 2, wherein, The method further includes: Obtain respective historical verification information corresponding to each original file in the random verification group, wherein each historical verification information includes respective verification operations arranged in a forward order based on the verification time, and each verification operation has respective corresponding verification results; Perform reverse extraction on each historical verification information respectively based on the forward arrangement, and determine the consecutive verification operations obtained and corresponding to a preset number as a consecutive verification group; In response to all verification results corresponding to the same continuous verification group being verified, retrieve the preset weight reduction to reduce and update the extraction priority value corresponding to the original document, and obtain the updated extraction priority value.

7. The intelligent file management method based on a large language model according to claim 1, wherein trigger the file verification plugin to perform integrity-based file verification on the current file based on the content verification group, and determine the document status corresponding to the current file based on the verification results, where the document status includes a complete status and a tampered status, including: Based on the file verification plugin, retrieve each current file having a traceability relationship with each original file located in the content verification group from the current file, and perform binarization processing on each original file and each current file respectively to obtain each original image and each current image, where any one of the original images includes each first image pixel point corresponding to a first pixel value, and any one of the current images includes each second image pixel point corresponding to a second pixel value; Create a transparent fusion layer, and stack the original file and the current file having the same traceability relationship on the transparent fusion layer in sequence; In response to any first image pixel point and / or any second image pixel point being included in the transparent fusion layer, determine the file status corresponding to the current file as the tampered status.

8. The intelligent file management method based on a large language model according to claim 7, wherein Based on the current file, establish an initial file comparison layer, fill the tampered content and the adjacent operation end into the content indication area and the terminal indication area located in the initial file comparison layer respectively, and send the obtained current file comparison layer to the management end for display, including: Based on the current file, establish an initial file comparison layer, where the initial file comparison layer includes a vertically arranged content indication area and a terminal indication area for filling the adjacent operation end; Determine the current file in the current file where the tampered content exists, obtain each file line drop arranged in a vertical order from top to bottom in the current file, and determine all file line drops having a blank group composed of two adjacent blank characters as the paragraph start line; Based on the vertical order, respectively determine two adjacent paragraph start lines in adjacent positions as an adjacent paragraph group, and determine the paragraph start line corresponding to the first position in the same adjacent paragraph group as the first paragraph and the paragraph start line corresponding to the last position as the last paragraph; Determine the first paragraph and all file line drops between the first paragraph and the last paragraph as the same file paragraph, and obtain each file paragraph in the current file; Determine the file paragraph including the tampered content as the tampered paragraph, and determine the original paragraph corresponding to the tampered paragraph in the original file having the same traceability relationship as the current file; Generate a first bounding box surrounding the tampered paragraph in the current file, and generate a second bounding box surrounding the original paragraph in the original file; Generate a first indication area and a second indication area arranged horizontally in the content indication area, and fill the current file with a first bounding box into the first indication area and the original file with a second bounding box into the second indication area; Fill the proximity operation end into the terminal indication area, and send the obtained current file comparison layer to the management end for display.

9. An intelligent file management system based on a large language model, characterized in that, Comprising: An extraction module, configured to respond to any current operation end to perform a current scan on the file label corresponding to the current file, retrieve the original file corresponding to the current file, determine each original file with a pre-verification mark in the original file as a necessary verification group, and determine all other original files included in the original file except the necessary verification group as a random verification group; Perform content evaluation on each original file in the random verification group, and perform a summation calculation on each extraction priority value obtained corresponding to each original file to obtain an extraction total value; Perform a ratio calculation on each extraction priority value and the extraction total value respectively to obtain each extraction probability corresponding to each original file; Determine the number of files corresponding to each original file in the random verification group, and perform a product calculation based on the number of files and the preset extraction ratio retrieved to obtain the current extraction quantity; Create a random extraction turntable, and perform area division on the random extraction turntable based on each extraction probability corresponding to each original file to obtain each sector area corresponding to each original file; Control the random extraction turntable to perform rotational extraction corresponding to the current extraction quantity, and determine the extracted original files and all original files in the necessary verification group as a content verification group based on the extraction result; A verification module, configured to trigger an archive verification plugin to perform file verification based on integrity on the current file based on the content verification group, and determine the document status corresponding to the current file based on the verification result, where the document status includes a complete status and a tampered status; A determination module, configured to respond to the current file being in a tampered status, determine the tampered content corresponding to the current file, and determine the proximity operation end for performing a proximity scan on the current file based on the retrieved operation record table corresponding to the current file; A creation module, configured to create an initial file comparison layer based on the current file, fill the tampered content and the proximity operation end into the content indication area and the terminal indication area located in the initial file comparison layer respectively, and send the obtained current file comparison layer to the management end for display.

Citation Information

Patent Citations

  • Special merchant archive modification record display method and device, medium and equipment

    CN113704556A

  • Remodification, identification and alarm system and method for sensitive archives

    CN119046933A