An artificial intelligence-based archive management system

By segmenting, identifying, repairing, and classifying files uploaded by user terminals, the problem of inaccurate file classification in existing file management systems has been solved, and the compatibility between files and storage space and the search efficiency have been improved.

CN115328854BActive Publication Date: 2026-01-02SHANGHAI XINYINGJIE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211036277.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-28
Publication Date
2026-01-02
Estimated Expiration
2042-08-28

AI Technical Summary

Technical Problem

Existing document management systems are unable to accurately save documents to the appropriate storage space based on the document's built-in classification information, resulting in reduced accuracy and reliability in retrieval.

Method used

Artificial intelligence technology is used to segment files uploaded by user terminals, identify and repair text and image file blocks, and reassemble and classify them for storage in a multi-dimensional manner to ensure compatibility with the classification system of archive storage space.

Benefits of technology

It improves the efficiency and accuracy of finding files from archive storage space and ensures that the file and storage space classification systems are compatible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115328854B_ABST
    Figure CN115328854B_ABST
Patent Text Reader

Abstract

The application provides an artificial intelligence-based archive management system, which performs segmentation processing on files uploaded by a user terminal, obtains a plurality of text file blocks and a plurality of picture file blocks, performs inspection processing on each text file block and each picture file block, obtains corresponding text content and picture content, and performs repair processing on the text file blocks and the picture file blocks; all the text file blocks and all the picture file blocks are recombined to restore corresponding files, and the files are subjected to multi-dimensional classification and storage according to the text content and the picture content; the above method can identify and analyze each uploaded file in terms of text and pictures, thereby realizing reclassification of the files, ensuring that the classified files are compatible with the classification system of the archive storage space, and improving the efficiency and accuracy of subsequent file searching from the archive storage space.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of file management, and in particular to an artificial intelligence-based file management system. BACKGROUND

[0002] The existing file management system directly saves files in corresponding storage spaces according to file classification information carried by the uploaded files, and does not reclassify the files. The file classification information carried by the files is usually obtained by rough analysis of the files, and cannot accurately reflect the data content of the files. In addition, the file classification information carried by the files is not necessarily compatible with the classification system of the storage spaces, so that the files cannot be accurately saved in the corresponding storage spaces, and the accuracy and reliability of subsequent file searching are reduced. SUMMARY

[0003] In view of the defects of the prior art, the application provides an artificial intelligence-based file management system, which segments the files uploaded by a user terminal to obtain a plurality of text file blocks and a plurality of picture file blocks, checks each text file block and each picture file block to obtain corresponding text content and picture content, and repairs the text file blocks and the picture file blocks. All text file blocks and all picture file blocks are recombined to restore the corresponding files, and the files are classified and stored in multiple dimensions according to the text content and the picture content. The above method can identify and analyze the text and the picture of each uploaded file, so as to reclassify the files, ensure that the classified files are compatible with the classification system of the file storage space, and improve the efficiency and accuracy of subsequent file searching from the file storage space.

[0004] The application provides an artificial intelligence-based file management system, which comprises:

[0005] A file sending source terminal identification module is configured to analyze and process a file uploading request from a user terminal, and determine whether the user terminal has file uploading permission.

[0006] A file receiving module is configured to connect with the user terminal in a predetermined file uploading mode according to the determination result of the file uploading permission, so as to receive the files uploaded by the user terminal.

[0007] A file segmentation module is configured to segment the received files into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the files, and determine the original data positions of each text file block and each picture file block in the files.

[0008] a first file block processing module, configured to perform text content inspection processing on each text file block, and perform text repair processing and text content marking processing on the corresponding text file block according to the result of the text content inspection processing;

[0009] a second file block processing module, configured to perform picture inspection processing on each picture file block, and perform picture repair processing and picture content marking processing on the corresponding picture file block according to the result of the picture inspection processing;

[0010] a file block integration module, configured to recombine all the text file blocks and picture file blocks that have completed repair processing according to the original data positions, so as to restore the corresponding files;

[0011] a file classification and storage module, configured to perform multi-dimensional classification on the restored files according to the results of the text content marking processing and the picture content marking processing, and save the files to corresponding archive storage spaces.

[0012] Further, a file sending source terminal identification module is configured to analyze and process a file upload request from a user terminal, and determine whether the user terminal has file upload permission, specifically including:

[0013] The file sending source terminal identification module extracts terminal identity information of the user terminal from the file upload request from the user terminal, compares the terminal identity information with a preset terminal identity information library, and if the terminal identity information exists in the preset terminal identity information library, determines that the user terminal has file upload permission; otherwise, determines that the user terminal does not have file upload permission.

[0014] Further, a file receiving module is connected with the user terminal in a predetermined file upload mode according to the determination result of the file upload permission, so as to receive a file uploaded by the user terminal, specifically including:

[0015] When the user terminal does not have file upload permission, the file receiving module is not connected with the user terminal;

[0016] When the user terminal has file upload permission, the file receiving module determines an average file upload duration of the user terminal in a historical file upload process according to historical file upload log information of the user terminal, takes the average file upload duration as a connection life cycle between the user terminal and the file receiving module, so that the user terminal accesses the file receiving module, and determines half of the minimum of a maximum file upload rate of the user terminal and a maximum file receiving rate of the file receiving module as an average file upload rate of the user terminal.

[0017] Further, the file splitting module splits the received file into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the file, and determines the original data position of each text file block and each picture file block in the file, specifically comprising:

[0018] When the file splitting module confirms that the user terminal completes uploading a complete file to the file receiving module, the file splitting module performs data content recognition on the file, determines the position of the start text code and the position of the end text code of each text file block in the file, and the position of the first pixel and the position of the last pixel of each picture file block;

[0019] According to the position of the start text code and the position of the end text code, all text file blocks contained in the file are extracted from the file, and the original data position of each text file block in the file is determined;

[0020] According to the position of the first pixel and the position of the last pixel, all picture file blocks contained in the file are extracted from the file, and the original data position of each picture file block in the file is determined.

[0021] Further, the file splitting module splits the received file into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the file, and determines the original data position of each text file block and each picture file block in the file, specifically comprising:

[0022] After splitting a plurality of text file blocks and a plurality of picture file blocks, the file splitting module first determines whether each file block has a splitting-to-text situation and a splitting-to-complete-picture situation according to the position of the start text code and the position of the end text code of each text file block, and the position of the first pixel and the position of the last pixel of each picture file block, records the situation as a splitting abnormal situation, and if the splitting abnormal situation exists, locates the position points of the splitting edge head and tail ends of the splitting abnormal situation, then finds the remaining file blocks that are spliced with the file block abnormal edge of the current splitting abnormal situation according to the position points of the splitting edge head and tail ends of the splitting abnormal situation, and performs re-splicing and splitting on the re-spliced file blocks, and then performs the detection of the above steps again, until the plurality of text file blocks and the plurality of picture file blocks are ensured not to be split to text and not to be split to complete pictures, and the process is as follows:

[0023] Step S1, using the following formula (1), according to the position of the start text code and the position of the end text code of each text file block, and the position of the first pixel and the position of the last pixel of each picture file block, to determine whether each file block exists the case of being divided into text and complete picture,

[0024]

[0025] In the above formula (1), W(a) represents the determination value of whether the a-th file block exists the case of being divided into text and complete picture; ∨{} represents that if there is one or more formulae in the brackets, the overall result value is 1, otherwise the overall result value is 0; [X0(a), Y0(a)] represents the position point of the start text code or the position point of the first pixel of the a-th file block; [X(a), Y(a)] represents the position point of the end text code or the position point of the last pixel of the a-th file block; G{→} represents that if there is a pixel point between the left position point of the arrow in the brackets and the right position point of the arrow, the overall result value is 1, otherwise the overall result value is 0;

[0026] If W(a) = 0, it means that the a-th file block does not exist the case of being divided into text and complete picture;

[0027] If W(a) = 1, it means that the a-th file block exists the case of being divided into text and complete picture;

[0028] Step S2, using the following formula (2), according to the position points of the head and tail of the division edge of the division abnormal case, and the position points of the four vertices of the original file before division, to determine whether the division edge of the division abnormal case coincides with the four edges of the original file, so as to avoid the case that the original file exists text being divided,

[0029]

[0030] In the above formula (2), F b (i) represents the determination value of whether the i-th edge of the b-th file block existing the division abnormal case is divided into text and complete picture coincides with the four edges of the original file; [x b (i_1), y b (i_1)] represents the position point of the head of the i-th edge of the b-th file block existing the division abnormal case; [x b (i_2), y b (i_2)] represents the position point of the tail of the i-th edge of the b-th file block existing the division abnormal case; (Xk Y k represents the kth vertex position point of the original file; represents the value of k is taken from 1 to 4 into the formula, if one or more than one formula in the square brackets is true, the overall value is 1, otherwise the overall value is 0;

[0031] If F b (i) = 0, it means that the i-th split to text and the split to the edge of the complete picture in the b-th file block with split abnormality does not coincide with the four edges of the original file;

[0032] If F b (i) = 1, it means that the i-th split to text and the split to the edge of the complete picture in the b-th file block with split abnormality coincides with the four edges of the original file, then the b-th file block with split abnormality is included in the file block without split abnormality;

[0033] Step S3, if the split edge of the split abnormality does not coincide with the four edges of the original file, then according to the position points of the head and tail of the split edge of the split abnormality, the remaining file blocks that are spliced with the abnormal edge of the file block of the current split abnormality are obtained by using the following formula (3),

[0034]

[0035] In the above formula (3), P(a) represents the control value of the a-th file block spliced with the i-th split to text and the split to the edge of the complete picture in the b-th file block with split abnormality;

[0036] If P(a) = 0, the a-th file block is controlled not to be spliced with the i-th split to text and the split to the edge of the complete picture in the b-th file block with split abnormality in any form;

[0037] If P(a) = 1, the a-th file block is controlled to be spliced with the i-th split to text and the split to the edge of the complete picture in the b-th file block with split abnormality according to the corresponding overlapping coordinate points.

[0038] Further, the first file block processing module performs text content inspection processing on each text file block, and according to the result of the text content inspection processing, performs text repair processing and text content marking processing on the corresponding text file block, which specifically includes:

[0039] The first file block processing module performs text syntax checking and text error checking on each text file block to determine text syntax error areas and errors in each text file block, and corrects each text syntax error area and each error.

[0040] The first file block processing module further performs vocabulary frequency checking on each text file block to obtain the frequency of each corresponding vocabulary in each text file block, and takes the vocabulary meeting a specific frequency condition as a key marked vocabulary of the corresponding text file block, thereby performing content marking on the corresponding text file block.

[0041] Further, the second file block processing module performs picture checking on each picture file block, and performs picture repair and picture content marking on the corresponding picture file block according to the results of the picture checking, specifically including:

[0042] The second file block processing module performs picture pixel checking on each picture file block to determine all bad point pixels in each picture file block and object contour information in the picture frame.

[0043] Each picture file block is repaired one by one.

[0044] According to the object contour information, the object type in the picture frame is determined, and the picture content of the corresponding picture file block is marked according to the object type.

[0045] Further, the file block integration module recombines all text file blocks and picture file blocks that have completed repair according to the original data position, thereby restoring the corresponding file, specifically including:

[0046] The file block integration module recombines all text file blocks and picture file blocks according to the position of the start text code and the position of the end text code of each text file block, and the position of the first pixel and the position of the last pixel of each picture file block, thereby restoring the corresponding file.

[0047] Further, the file classification and storage module classifies and saves the restored file to the corresponding archive storage space according to the results of the text content marking and the picture content marking, specifically including:

[0048] The file classification and storage module gives the restored file a classification index word corresponding to the key mark word and the object type according to the results of the text content mark processing and the picture content mark processing, so as to realize multi-dimensional classification of the restored file, and saves the restored file and the corresponding classification index to the corresponding archive storage space.

[0049] Compared with the prior art, the artificial intelligence-based archive management system performs segmentation processing on the files uploaded by the user terminal to obtain a plurality of text file blocks and a plurality of picture file blocks, performs inspection processing on each text file block and each picture file block to obtain corresponding text content and picture content, and performs repair processing on the text file blocks and the picture file blocks; all the text file blocks and the picture file blocks are recombined to restore the corresponding files, and the restored files are classified and stored according to the text content and the picture content; the above method can identify and analyze each uploaded file in terms of text and picture, so as to realize reclassification of the files, ensure that the classified files are compatible with the classification system of the archive storage space, and improve the efficiency and accuracy of subsequent file searching from the archive storage space.

[0050] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be achieved and obtained by means of the structure particularly pointed out in the written description, claims, and drawings.

[0051] The technical solutions of the present application will be further described in detail below with the help of the drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0053] Figure 1 A structural schematic diagram of an artificial intelligence-based archive management system is provided. DETAILED DESCRIPTION

[0054] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0055] With reference to Figure 1 A structural schematic diagram of an artificial intelligence-based archive management system is provided in the embodiments of the present application. The artificial intelligence-based archive management system comprises:

[0056] A file sending source terminal identification module is configured to analyze and process a file upload request from a user terminal, and determine whether the user terminal has file upload permission.

[0057] A file receiving module is configured to connect with the user terminal in a predetermined file upload mode according to the determination result of the file upload permission, and receive a file uploaded by the user terminal.

[0058] A file splitting module is configured to split the received file into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the file, and determine the original data position of each text file block and each picture file block in the file.

[0059] A first file block processing module is configured to perform text content inspection processing on each text file block, and perform text repair processing and text content marking processing on the corresponding text file block according to the result of the text content inspection processing.

[0060] A second file block processing module is configured to perform picture inspection processing on each picture file block, and perform picture repair processing and picture content marking processing on the corresponding picture file block according to the result of the picture inspection processing.

[0061] A file block integration module is configured to recombine all the text file blocks and picture file blocks that have completed repair processing according to the original data position, so as to restore the corresponding file.

[0062] A file classification and storage module is configured to perform multi-dimensional classification on the restored file according to the results of the text content marking processing and the picture content marking processing, and save the file to a corresponding archive storage space.

[0063] The beneficial effects of the above technical solutions are that the artificial intelligence-based archive management system performs segmentation processing on the files uploaded by the user terminal, obtains a plurality of text file blocks and a plurality of picture file blocks, performs inspection processing on each text file block and each picture file block, obtains corresponding text content and picture content, and performs repair processing on the text file blocks and the picture file blocks; all the text file blocks and all the picture file blocks are recombined to restore the corresponding files, and the files are classified and stored in multiple dimensions according to the text content and the picture content; the above method can identify and analyze the text and the picture of each uploaded file, thereby realizing the reclassification of the files, ensuring that the classified files are compatible with the classification system of the archive storage space, and improving the efficiency and accuracy of subsequent file searching in the archive storage space.

[0064] Preferably, the file sending source terminal identification module is configured to analyze and process the file upload request from the user terminal to determine whether the user terminal has the file upload permission, and the determination specifically includes:

[0065] The file sending source terminal identification module extracts terminal identity information of the user terminal from the file upload request from the user terminal, compares the terminal identity information with a preset terminal identity information library, and determines that the user terminal has the file upload permission if the terminal identity information exists in the preset terminal identity information library; otherwise, it is determined that the user terminal does not have the file upload permission.

[0066] The beneficial effects of the above technical solutions are that the file sending source terminal identification module is used to identify and authenticate the identity of the user terminal to determine whether the user terminal has the file upload permission, thereby ensuring that only the allowed user terminal can have the file upload permission and ensuring the data security of the archive management system.

[0067] Preferably, the file receiving module is connected to the user terminal in a predetermined file upload mode according to the determination result of the file upload permission to receive the file uploaded by the user terminal, and the connection specifically includes:

[0068] When the user terminal does not have the file upload permission, the file receiving module is not connected to the user terminal.

[0069] When the user terminal has the file uploading permission, the file receiving module determines the average file uploading duration of the user terminal in the historical file uploading process according to the historical file uploading log information of the user terminal, takes the average file uploading duration as the connection life cycle between the user terminal and the file receiving module, so that the user terminal accesses to the file receiving module, and determines the average file uploading rate of the user terminal as half of the minimum of the maximum file uploading rate of the user terminal and the maximum file receiving rate of the file receiving module.

[0070] The above technical scheme has the beneficial effects that: by the above manner, the average file uploading duration of the user terminal in the historical file uploading process is taken as the connection life cycle between the user terminal and the file receiving module, so that the user terminal can be connected with the file receiving module within the connection life cycle, and the user terminal is disconnected with the file receiving module outside the connection life cycle, thereby avoiding that the user terminal always connects with the file receiving module to occupy the file uploading bandwidth of the file management system. In addition, half of the minimum of the maximum file uploading rate of the user terminal and the maximum file receiving rate of the file receiving module is determined as the average file uploading rate of the user terminal, so that the user terminal can efficiently upload the file to the file receiving module.

[0071] Preferably, the file splitting module splits the received file into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the file, and determines the original data position of each text file block and each picture file block in the file, specifically comprising:

[0072] When the file splitting module confirms that the user terminal completes uploading a complete file to the file receiving module, the file splitting module performs data content recognition on the file to determine the position of the start text code and the position of the end text code of each text file block in the file, and the position of the first pixel and the position of the last pixel of each picture file block;

[0073] According to the position of the start text code and the position of the end text code, all text file blocks contained in the file are extracted from the file, and the original data position of each text file block in the file is determined;

[0074] According to the position of the first pixel and the position of the last pixel, all picture file blocks contained in the file are extracted from the file, and the original data position of each picture file block in the file is determined.

[0075] The technical scheme has the beneficial effects that: the file splitting module splits the file into text file blocks and picture file blocks, so that each file block of the file can be accurately positioned and analyzed, thereby improving the identification accuracy of each file block.

[0076] Preferably, the file splitting module splits the file into text file blocks and picture file blocks according to the data content of the received file, and determines the original data position of each text file block and each picture file block in the file, and the method further comprises the following steps:

[0077] After the file splitting module splits the file into text file blocks and picture file blocks, it first determines whether each file block has the condition of splitting into text and splitting into a complete picture according to the position of the start text code and the position of the end text code of each text file block, and the position of the first pixel and the position of the last pixel of each picture file block, records this condition as a splitting abnormal condition, and if the splitting abnormal condition exists, locates the position points of the splitting edge head and tail ends of the splitting abnormal condition, then finds the remaining file blocks that are spliced with the file block abnormal edge of the current splitting abnormal condition according to the position points of the splitting edge head and tail ends of the splitting abnormal condition, re-splices the file blocks, and then performs the above step detection after splitting the re-spliced file blocks, until the text file blocks and the picture file blocks are ensured not to be split into text and not to be split into a complete picture, and the process is as follows:

[0078] In step S1, the position of the start text code and the position of the end text code of each text file block (optionally, the text file blocks can be converted into picture file blocks, and the text in the text file blocks can be converted into pixel points in the picture file blocks, and the pixel values of the pixel points in the places with text are marked as 1, and the pixel values of the pixel points in the places without text are marked as 0, and the position of the first pixel and the position of the last pixel of each text file block are the position of the start text code and the position of the end text code) and the position of the first pixel and the position of the last pixel of each picture file block are used to determine whether each file block has the condition of splitting into text and splitting into a complete picture,

[0079]

[0080] In the above formula (1), W(a) represents a decision value of whether the a-th file block has a case of being segmented into a character and a complete picture; V{} represents that if one or more of the expressions in the brackets are true, the overall result value is 1, otherwise the overall result value is 0; [X0(a), Y0(a)] represents a position point of a start text code or a position point of a first pixel of the a-th file block; [X(a), Y(a)] represents a position point of an end text code or a position point of a last pixel of the a-th file block; G{→} represents that if there is a pixel point between the left position point of the arrow and the right position point of the arrow, the overall result value is 1, otherwise the overall result value is 0;

[0081] If W(a) = 0, it represents that the a-th file block does not have a case of being segmented into a character and a complete picture.

[0082] If W(a) = 1, it represents that the a-th file block has a case of being segmented into a character and a complete picture.

[0083] In step S2, whether the segmented edge of the segmentation abnormal case coincides with the four edges of the original file is determined according to the position points of the start and end of the segmented edge of the segmentation abnormal case and the position points of the four vertices of the original file before segmentation by using the following formula (2), so as to avoid the case that the original file has a character being segmented.

[0084]

[0085] In the above formula (2), F b (i) represents a decision value of whether the i-th edge of the b-th file block having a segmentation abnormal case is segmented into a character and a complete picture and coincides with the four edges of the original file; [x b (i_1), y b (i_1)] represents a position point of a start of the i-th edge of the b-th file block having a segmentation abnormal case being segmented into a character and a complete picture; [x b (i_2), y b (i_2)] represents a position point of an end of the i-th edge of the b-th file block having a segmentation abnormal case being segmented into a character and a complete picture; (X k ,Y k ) represents a k-th vertex position point of the original file. represents that the value of k is taken from 1 to 4, and if one or more of the expressions in the brackets are true, the overall value is 1, otherwise the overall value is 0.

[0086] If F b(i) = 0, indicating that the i-th split-to-text and split-to-complete picture edge in the b-th file block with a split abnormality does not coincide with the four edges of the original file;

[0087] If F b (i) = 1, indicating that the i-th split-to-text and split-to-complete picture edge in the b-th file block with a split abnormality coincides with the four edges of the original file, the b-th file block with a split abnormality is included in the file block without a split abnormality;

[0088] Step S3, if the split edge of the split abnormality does not coincide with the four edges of the original file, then according to the position points of the split edge of the split abnormality, the remaining file blocks that are spliced with the abnormal edge of the file block of the current split abnormality are obtained by using the following formula (3),

[0089]

[0090] In the above formula (3), P(a) represents the control value of the a-th file block spliced with the i-th split-to-text and split-to-complete picture edge in the b-th file block with a split abnormality;

[0091] If P(a) = 0, the a-th file block is controlled not to be spliced with the i-th split-to-text and split-to-complete picture edge in the b-th file block with a split abnormality in any form;

[0092] If P(a) = 1, the a-th file block is controlled to be spliced with the i-th split-to-text and split-to-complete picture edge in the b-th file block with a split abnormality according to the corresponding overlapping coordinate points.

[0093] The beneficial effects of the above technical solutions are: by using the above formula (1), whether each file block is divided into text and complete picture is judged according to the positions of the start text code and the end text code of each text file block and the positions of the first pixel and the last pixel of each picture file block, so as to know whether there is a problem in the division, and the problem can be found and corrected in time; then by using the above formula (2), whether the division edge of the division abnormal condition coincides with the four edges of the original file is judged according to the position points of the head and tail of the division edge of the division abnormal condition and the four vertex position points of the original file before division, so as to avoid the case that the original file is divided into text, thereby further determining the division abnormal condition and ensuring the reliability of the system; finally, by using the above formula (3), the remaining file blocks that are spliced with the file block abnormal edge of the current division abnormal condition are obtained according to the position points of the head and tail of the division edge of the division abnormal condition, so as to ensure that the division position is a blank position and does not affect the subsequent steps, and the accuracy of the system is ensured.

[0094] Preferably, the first file block processing module performs text content inspection processing on each text file block, and according to the result of the text content inspection processing, performs text repair processing and text content marking processing on the corresponding text file block, specifically including:

[0095] The first file block processing module performs text syntax inspection processing and text error character inspection processing on each text file block to determine the text syntax error area and the error character existing in each text file block, and performs correction processing on each text syntax error area and each error character.

[0096] The first file block processing module also performs vocabulary frequency inspection processing on each text file block to obtain the appearance frequency of the corresponding vocabulary in each text file block, and takes the vocabulary satisfying a specific appearance frequency condition as a key marked vocabulary of the corresponding text file block, thereby performing content marking processing on the corresponding text file block.

[0097] The beneficial effects of the above technical solutions are: by the above-mentioned manner, the text content of each text file block can be corrected and marked, and the content correctness and traceability of the text file block are ensured.

[0098] Preferably, the second file block processing module performs picture inspection processing on each picture file block, and according to the result of the picture inspection processing, performs picture repair processing and picture content marking processing on the corresponding picture file block, specifically including:

[0099] The second file block processing module performs picture pixel inspection processing on each picture file block to determine all bad point pixels existing in each picture file block and object contour information existing in the picture screen.

[0100] All bad pixel points existing in each picture file are repaired one by one, wherein the bad pixel points can be but are not limited to pixel points with a luminance value lower than a preset luminance threshold or a resolution value lower than a preset resolution threshold; correspondingly, the repairing processing can be but is not limited to luminance or resolution repairing processing of the bad pixel points.

[0101] According to the object contour information, the type of the object existing in the picture frame is determined, and the picture content of the corresponding picture file block is marked according to the type of the object.

[0102] The beneficial effects of the above technical solutions are that the picture content of each picture file block can be corrected and marked, and the content correctness and traceability of the picture file block can be ensured.

[0103] Preferably, the file block integration module recombines all text file blocks and picture file blocks that have completed the repairing processing according to the original data positions, so as to restore the corresponding file, which specifically includes:

[0104] The file block integration module recombines all text file blocks and picture file blocks according to the positions of the start text code and the end text code of each text file block and the positions of the first pixel and the last pixel of each picture file block, so as to restore the corresponding file.

[0105] The beneficial effects of the above technical solutions are that the text file blocks and picture file blocks that have been corrected or repaired can be recombined according to their original positions in the file, so as to ensure the content correctness of the restored file.

[0106] Preferably, the file classification and storage module classifies and saves the restored file to the corresponding archival storage space according to the results of the text content marking processing and the picture content marking processing, specifically including:

[0107] The file classification and storage module assigns the restored file with a classification index word corresponding to the key marking word and the object type according to the results of the text content marking processing and the picture content marking processing, so as to realize multi-dimensional classification of the restored file, and saves the restored file and the corresponding classification index to the corresponding archival storage space.

[0108] The beneficial effects of the above technical solutions are that the restored file can be classified in multiple dimensions, so that the corresponding file can be accurately searched in the archival storage space from multiple aspects in the future.

[0109] From the content of the above embodiment, the file uploaded by the user terminal is segmented to obtain a plurality of text file blocks and a plurality of picture file blocks, each text file block and each picture file block is checked to obtain corresponding text content and picture content, and the text file blocks and the picture file blocks are repaired; all the text file blocks and all the picture file blocks are recombined to restore the corresponding file, and the file is classified and stored in multiple dimensions according to the text content and the picture content; the above method can identify and analyze the uploaded file in terms of text and picture to realize the reclassification of the file, ensure the compatibility of the classified file and the classification system of the archive storage space, and improve the efficiency and accuracy of subsequent file searching in the archive storage space.

[0110] Obviously, various modifications and changes can be made to the present application by those skilled in the art without departing from the spirit and scope of the present application. Thus, if these modifications and changes of the present application fall within the scope of the present application's claims and their equivalents, it is intended to include these modifications and changes in the present application.

Claims

1. An artificial intelligence-based archive management system, characterized by, It includes: File sending source terminal identification module for analyzing and processing file upload request from user terminal, judging whether the user terminal has file upload permission; File receiving module for connecting with the user terminal in a predetermined file upload mode according to the judgment result of the file upload permission, so as to receive the file uploaded by the user terminal; File segmentation module for segmenting the received file into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the file, and determining the original data position of each text file block and each picture file block in the file; First file block processing module for text content inspection processing of each text file block, and text repair processing and text content marking processing of the corresponding text file block according to the result of text content inspection processing; Second file block processing module for picture inspection processing of each picture file block, and picture repair processing and picture content marking processing of the corresponding picture file block according to the result of picture inspection processing; File block integration module for recombining all text file blocks and picture file blocks that have completed repair processing according to the original data position, so as to restore the corresponding file; File classification and storage module for multi-dimensional classification and saving the restored file to the corresponding archive storage space according to the results of the text content marking processing and the picture content marking processing; Wherein, the file segmentation module segments the received file into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the file, and determines the original data position of each text file block and each picture file block in the file, which specifically includes: When the file segmentation module confirms that the user terminal completes uploading a complete file to the file receiving module, the file segmentation module performs data content identification on the file, determines the position of the start text code and the end text code of each text file block in the file, and the position of the first pixel and the end pixel of each picture file block; According to the position of the start text code and the end text code, all text file blocks contained in the file are extracted from the file, and the original data position of each text file block in the file is determined; According to the position of the first pixel and the end pixel, all picture file blocks contained in the file are extracted from the file, and the original data position of each picture file block in the file is determined; Wherein, the file block integration module recombines all text file blocks and picture file blocks that have completed repair processing according to the original data position, so as to restore the corresponding file, which specifically includes: The file block integration module reassembles all the text file blocks and all the picture file blocks according to the position of the start text code and the position of the end text code of each text file block and the position of the first pixel and the position of the last pixel of each picture file block, so as to restore the corresponding file. 2.The artificial intelligence-based archive management system of claim 1, wherein: The file sending source terminal identification module is configured to analyze and process the file upload request from the user terminal to determine whether the user terminal has the file upload permission, and specifically comprises: The file sending source terminal identification module extracts the terminal identity information of the user terminal from the file upload request from the user terminal, compares the terminal identity information with the preset terminal identity information library, and if the terminal identity information exists in the preset terminal identity information library, determines that the user terminal has the file upload permission; otherwise, determines that the user terminal does not have the file upload permission. 3.The artificial intelligence-based archive management system of claim 2, wherein: The file receiving module is connected with the user terminal in a predetermined file upload mode according to the judgment result of the file upload permission to receive the file uploaded by the user terminal, and specifically comprises: When the user terminal does not have the file upload permission, the file receiving module is not connected with the user terminal; When the user terminal has the file upload permission, the file receiving module determines the average file upload duration of the user terminal in the historical file upload process according to the historical file upload log information of the user terminal, takes the average file upload duration as the connection life cycle between the user terminal and the file receiving module, so that the user terminal accesses to the file receiving module, and determines half of the minimum of the maximum file upload rate of the user terminal and the maximum file receiving rate of the file receiving module as the average file upload rate of the user terminal. 4.The artificial intelligence-based archive management system of claim 1, wherein: The file segmentation module segments the received file into a plurality of text file blocks and a plurality of picture file blocks according to the data content of the file, and determines the original data position of each text file block and each picture file block in the file, and further comprises: The file segmentation module first determines whether each file block is segmented into text and complete picture according to the position of the start text code and the position of the end text code of each text file block and the position of the first pixel and the position of the last pixel of each picture file block, records the situation as a segmentation abnormal situation, and if the segmentation abnormal situation exists, locates the position points of the first and last ends of the segmentation edge of the segmentation abnormal situation, then finds the remaining file blocks that are spliced with the file block abnormal edge of the current segmentation abnormal situation according to the position points of the first and last ends of the segmentation edge of the segmentation abnormal situation, re-splices the file blocks, and performs the above step detection again after segmenting the re-spliced file blocks, until the segmented several text file blocks and several picture file blocks ensure that no text is segmented and no complete picture is segmented, and the process is as follows: Step S1, according to the position of the start text code and the position of the end text code of each text file block and the position of the first pixel and the position of the last pixel of each picture file block, determine whether each file block is segmented into text and complete picture by using the following formula (1), (1) In the above formula (1), represents the determination value indicating whether or not the file block No. exists the case of being split into a character and a complete picture; represents that the overall result value is 1 if one or more of the expressions in the parentheses are true, and otherwise the overall result value is 0; represents the start text code position point or the first pixel position point of the file block No. ; represents the end text code position point or the last pixel position point of the file block No. ; represents that the overall result value is 1 if there is a pixel point not being 0 between the position point on the left of the arrow and the position point on the right of the arrow in the parentheses, and otherwise the overall result value is 0. If , indicates that the rd file block does not exist split to text and split to complete picture cases; If , it indicates that the first file block is split into text and complete picture; Step S2, according to the position points of the first and last ends of the segmentation edge of the segmentation abnormal situation and the four vertex position points of the original file before segmentation, determine whether the segmentation edge of the segmentation abnormal situation coincides with the four edges of the original file by using the following formula (2), to avoid the case that the original file is segmented into text, (2) In the above formula (2), represents the determination value of whether the edge of the i-th divided to text and the edge of the i-th divided to complete picture in the i-th file block in which the division abnormality exists coincides with the four edges of the original file; represents the value of the i-th vertex position point of the original file; represents the value of the i-th vertex position point of the original file; if the formula in the square bracket exists and one or more of the formulas in the square bracket are true, the overall value is 1, otherwise the overall value is 0.​​​​​​​​​​ If , indicates the th file block in which a split abnormality case exists, the th split to text and the split to the edge of a complete picture do not coincide with the four edges of the original file; If , the first segmented abnormality case file block is listed in the non-segmented abnormality case file block if the first segmented to text and the segmented to complete picture edge coincides with the four edges of the original file. ​ Step S3, if the segmentation edge of the segmentation abnormal situation does not coincide with the four edges of the original file, find the remaining file blocks that are spliced with the file block abnormal edge of the current segmentation abnormal situation according to the position points of the first and last ends of the segmentation edge of the segmentation abnormal situation by using the following formula (3), (3) In the above equation (3), denotes the th file block and the th file block with a split abnormality case, and denotes the control value of the split to the text and the split to the edge of the complete picture that are spliced. If , the control of the first file block and the first file block in the presence of the split abnormal situation of the first split into the text and the edge of the split into the complete picture does not carry out any form of splicing; If , the control unit controls the first file block to be combined with the first file block in which the division is abnormal, and the first division into characters and the division into the edge of a complete picture are combined according to the corresponding overlapping coordinate points.

5. The artificial intelligence-based archive management system of claim 1, wherein: The first file block processing module performs text content inspection processing on each text file block, and performs text repair processing and text content marking processing on the corresponding text file block according to the results of the text content inspection processing, specifically including: The first file block processing module performs text syntax inspection processing and text error character inspection processing on each text file block to determine the text syntax error area and the error character existing in each text file block; and corrects each text syntax error area and each error character; The first file block processing module also performs vocabulary frequency inspection processing on each text file block to obtain the appearance frequency of the corresponding vocabulary in each text file block, and takes the vocabulary that meets the specific appearance frequency condition as the key marked vocabulary of the corresponding text file block, thereby performing content marking processing on the corresponding text file block.

6. The artificial intelligence-based archive management system of claim 5, wherein: The second file block processing module performs picture inspection processing on each picture file block, and according to the result of the picture inspection processing, performs picture repair processing and picture content marking processing on the corresponding picture file block, specifically including: The second file block processing module performs picture pixel inspection processing on each picture file block, determines all bad point pixels existing in each picture file block and object contour information existing in the picture frame; Repair processing is performed on all bad point pixels existing in each picture file one by one; According to the object contour information, determine the object type existing in the picture frame; and according to the object type, mark the picture content of the corresponding picture file block.

7. The artificial intelligence-based archive management system of claim 6, wherein: The file classification and storage module classifies and saves the restored files to the corresponding archive storage space according to the results of the text content marking processing and the picture content marking processing, specifically including: The file classification and storage module assigns the restored files with classification index words corresponding to the key marking words and the object types according to the results of the text content marking processing and the picture content marking processing, thereby realizing multi-dimensional classification of the restored files, and saving the restored files and their corresponding classification index words to the corresponding archive storage space.

Citation Information

Patent Citations

  • File conversion method and device and file transmission system

    CN106776677A

  • Text file content pixelated conversion and restoration method

    CN110781185A