A method for optimizing the design and printing of file boxes based on electronic archive systems

By using multi-level clustering and matching methods in the electronic record system, the inconvenience of manual operation in the traditional record box design and printing process has been solved, realizing automated customized printing of record boxes and improving the efficiency and accuracy of record management.

CN121168043BActive Publication Date: 2026-05-26广东粤海粤西供水有限公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
广东粤海粤西供水有限公司
Filing Date
2025-09-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Traditional file box design and printing processes are inconvenient due to manual operation, resulting in inconsistent cover patterns and unsuitable sizes, making it difficult to meet personalized needs and affecting the efficiency and security of file management.

Method used

By employing a multi-level clustering method based on electronic archive systems, and combining clustering based on shape, size, cover pattern, and pattern layout position with archive format and semantic layout matching, automated customized printing of archive boxes can be achieved.

Benefits of technology

It improves the precision and personalization of file box design, enhances the efficiency and accuracy of file management, and provides an intelligent matching and feedback mechanism to meet users' personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168043B_ABST
    Figure CN121168043B_ABST
Patent Text Reader

Abstract

This invention relates to a method for optimizing the design and printing of file boxes based on an electronic archive system, belonging to the field of data processing. The method involves downloading a set of electronic archives with file box identifiers, performing file box shape and size clustering to obtain a first-level file clustering result; clustering the file box cover patterns within the first-level clustering result to obtain a second-level file clustering result; clustering the cover pattern layout positions within the second-level file clustering result to obtain a third-level file clustering result; initializing the electronic archive system based on the third-level file clustering result; receiving images of files to be packaged uploaded by the user; traversing the electronic archive system to match files and obtain matching file categories; when the matching file categories are the same, designing and printing file boxes based on the category file box identifier. This method solves the technical problem of low actual production efficiency and insufficient systematic management caused by the inability to achieve automated customized printing of file boxes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to a method for optimizing the design and printing of file boxes based on electronic file systems. Background Technology

[0002] In traditional record management processes, file boxes, as an important carrier of records, bear the crucial responsibility of protecting and storing them. However, traditional file boxes present numerous inconveniences in their design and printing, limiting the efficiency and quality of record management. The selection of cover designs for traditional file boxes typically requires manual processing, which not only increases workload but can also lead to inconsistent cover designs, affecting the overall aesthetics and the professionalism of record management. Furthermore, manual cover design selection struggles to meet the personalized needs of different users, limiting customized file box services. Secondly, the dimensions of traditional file boxes require manual cutting, making automated customization based on the actual needs of the records impossible. This not only results in inconsistent file box sizes, making unified management and storage difficult but also increases production costs and time. In practical applications, unsuitable file box sizes can also lead to damage or loss of records, posing security risks to record management.

[0003] Therefore, existing technologies often suffer from technical problems such as the inability to automate customized printing of file boxes, resulting in low actual production efficiency and insufficient systematic management. Summary of the Invention

[0004] This invention addresses the technical problem in existing technologies where file boxes cannot be automatically customized and printed, resulting in low actual production efficiency and insufficient systematic management. It provides a file box optimization design and printing method based on an electronic file system to solve this problem.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0006] This invention provides a method for optimizing the design and printing of file boxes based on an electronic archive system, comprising: downloading a set of electronic archives with file box identifiers; performing file box shape and size clustering to obtain a first-level file clustering result; performing file box cover pattern clustering on the first-level file clustering result to obtain a second-level file clustering result; performing cover pattern layout and position clustering on the second-level file clustering result to obtain a third-level file clustering result; initializing the electronic archive system according to the third-level file clustering result; receiving the file image to be packaged uploaded by the user terminal; traversing the electronic archive system to perform file matching to obtain the matching file categories; when the matching file categories are the same, designing and printing the file box according to the category file box identifier.

[0007] The beneficial effects of this invention are: by using an electronic archive system to achieve optimized design and printing of archive boxes, and by using three-level clustering, namely shape and size, cover pattern, and pattern layout and position, archives are accurately classified. This allows for the rapid matching of archive categories based on the user-uploaded images of the archives to be packaged, and the design and printing of archive boxes with corresponding markings. This achieves automated customized printing of archive boxes, improving the efficiency and accuracy of archive management and printing. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating a method for optimizing the design and printing of file boxes based on an electronic archive system, as provided by the present invention.

[0009] Figure 2 This invention provides a flowchart illustrating the process of matching file categories in a file box optimization design and printing method based on an electronic file system. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0012] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein. Example

[0013] like Figure 1As shown, this embodiment of the invention provides a method for optimizing the design and printing of file boxes based on an electronic file system, including:

[0014] S10: Download the collection of electronic archives with archive box identifiers, perform archive box shape and size clustering, and obtain the first-level archive clustering results.

[0015] S20: Perform file box cover pattern clustering on the first-level file clustering results to obtain the second-level file clustering results.

[0016] S30: Perform cover pattern layout position clustering on the secondary archive clustering results to obtain the tertiary archive clustering results.

[0017] S40: After initializing the electronic archive system based on the three-level archive clustering results, receive the archive image to be packaged uploaded by the user terminal, traverse the electronic archive system to perform archive matching, and obtain the matching archive category.

[0018] S50: When the matched files belong to the same category, design and print the file box according to the category file box identifier.

[0019] For example, in records management, electronic records often coexist with physical records. File boxes, as important containers for physical records, bear the responsibility of organizing, binding, and storing documents. Storing file box information in the electronic records system enables the associated management of physical and electronic records. This associated management helps ensure the integrity and consistency of the records and facilitates quick location and retrieval of the corresponding physical records when needed. In this solution, a series of electronic records labeled with file box identifiers are first downloaded from the electronic records system. These electronic records contain information on file boxes with different shapes, sizes, cover designs, and layouts. Next, these electronic records undergo a first-level clustering: shape and size clustering. Based on the shape and size characteristics of the file boxes, they are divided into different categories, resulting in a first-level clustering result. This step is to group file boxes with similar shapes together for easier subsequent processing. Then, based on the first-level clustering result, a second-level clustering is performed: cover design clustering. For each shape and size category of file boxes, further subdivisions can be made based on their cover designs, resulting in a second-level clustering result. This step aims to further categorize file boxes with similar cover designs. Subsequently, a third-level clustering, namely cover pattern layout and position clustering, is performed on the secondary archive clustering results. In this step, each category of archive boxes with similar cover patterns is further subdivided based on the layout and position of the pattern on the cover, resulting in a tertiary archive clustering result. This yields an archive box classification system with three features: shape and size, cover pattern, and pattern layout and position. The electronic archive system is then initialized based on this tertiary archive clustering result. The initialized system can quickly identify and match archive boxes based on different features. In practical application, when a user needs to encapsulate a document into an archive box, they can upload an image of the document to be encapsulated through the user's client. The uploaded image of the document to be encapsulated refers to the image file corresponding to the document the user wishes to encapsulate and print. This image file typically contains the specific content or identifying information of the document, such as text, charts, and photographs; it represents the archive entity that the user needs to manage, save, or archive. After receiving this image, the system automatically extracts its features, such as shape and pattern, and performs traversal and matching within the electronic archive system. Based on the extracted features, a multi-level matching process is performed within the electronic records system. First, shape and size are matched; then, cover patterns are matched; finally, the layout and position of the patterns are matched, until the record category that best matches the record to be packaged is found. When a match is found within the same record category, the system designs the record box based on the corresponding record box identifier. After the design is complete, the designed record box is printed so that the user can package the record inside. In summary, the entire process, through multi-level clustering, feature extraction, and matching of electronic records, achieves optimized design and printing of record boxes, improving the efficiency and accuracy of record management.

[0020] In a preferred embodiment, such as Figure 2 As shown, after initializing the electronic archive system based on the three-level archive clustering results, the system receives archive images to be packaged uploaded by the user, traverses the electronic archive system to perform archive matching, and obtains the matching archive categories, including:

[0021] After initializing the electronic archive system based on the three-level archive clustering results, the system receives archive images to be packaged uploaded by the user. It then traverses the electronic archive system to perform archive format layout matching to obtain format-matching archive categories. When the format-matching archive categories are the same, the format-matching archive category is output as the matching archive category. When the format-matching archive categories are multiple, the system traverses the format-matching archives to perform archive semantic layout matching to obtain semantically matching archive categories. When the semantically matching archive categories are the same, the semantically matching archive category is output as the matching archive category. When the semantically matching archive categories are multiple, the archive box identifier of the semantically matching archive category is sent to the user for selection, and feedback information is obtained, wherein the feedback information includes the matching archive category.

[0022] Optionally, in the electronic archive system, the system is first initialized based on the three-level archive clustering results. This step ensures that the archives in the system have been effectively classified and organized. Subsequently, the system receives the archive image to be packaged uploaded by the user. This image is the image file corresponding to the archive that the user wishes to package and print; it may contain the specific content or identification information of the archive. Afterward, the system begins to traverse the electronic archive system to perform archive format layout matching. This step mainly relies on the format layout features of the archives, such as shape, pattern, and arrangement, to find the electronic archive category that best matches the archive image to be packaged. Specifically, archive format layout matching refers to constructing a Structured Feature Alignment Network (SFAN) based on the three-level clustering results, calculating the multi-level format similarity between the user-uploaded image and the cluster centers, and dynamically fusing shape, pattern, and layout features. The specific calculation formula is as follows: ,;in, The standardized size difference is represented by j, which represents the number of semantic nodes extracted from the pattern. Characterizing the differences in color and texture after dimensionality reduction, while This is the layout Mahalanobis distance. During calculation, These refer to the dynamic weights of standardized size differences, color and texture differences after dimensionality reduction, and layout Mahalanobis distance, respectively. It can adaptively adjust according to the importance of features, such as =0.4, =0.3, =0.3. This refers to the maximum value of each dimension in the dataset, used to eliminate dimensional differences. The system will use the coordinates of the user-uploaded image. and the mean of the archival images of the cluster centers. Covariance Matrix To calculate the multi-level format similarity between the uploaded image and the cluster centers. Here, Characterize the color and texture features of the uploaded image, while Characterize the color and texture features of images within the same class. Also, consider the size of images uploaded from the same location. Image size within the same location class The differences in these dimensions include length, width, and height. A non-linear fusion function (reciprocal weighted sum) maps multidimensional differences to similarity scores in the 0-1 interval, avoiding the dominance of a single feature. If the layout variance of a certain type of file box is large, i.e. A higher diagonal value indicates a decrease. Weighting reduces the impact of layout. However, if a unique [structure / feature] exists... satisfy If they are of the same type, they are considered "the same type"; if they are multiple satisfy If the format matching file category is the same, meaning only one category highly matches the uploaded image in terms of format layout, then that category is output as the matching file category. However, if the format matching file category is multiple, meaning multiple categories have a certain degree of matching with the uploaded image in terms of format layout, then further semantic layout matching of the files will be performed. By comparing these features, the file category that is closest to the uploaded image in terms of format layout can be found, i.e., the format matching file category.

[0023] In the semantic layout matching stage of the archives, the semantic features of the archive content, such as the meaning and layout of elements like text, charts, and photos, are considered to further narrow down the matching range. A Semantic-Spatial Graph Attention Network (SSGAT) is constructed to extract local semantic features of the image, such as text regions and icon themes, and align them topologically with the semantic templates of candidate archive boxes. The SSGAT network can capture key semantic information in user-uploaded images and represent this information as nodes in a graph structure. These nodes not only contain semantic features but also reflect the spatial relationships between elements. The specific semantic graph construction process is as follows: [Image...] and candidate file boxes Extract semantic nodes separately, i.e. V represents the semantic nodes extracted from the image and candidate file boxes. It belongs to a key area; among them, For BERT text encoding or ResNet image feature extractors, used to extract key regions. Features The coordinates of the region center are given. Edge weights are determined by both the spatial distance and semantic similarity between nodes. , The edge weights represent the edge weights extracted from key regions, where i and j are the semantic node indices of the image and candidate file boxes, respectively. , representing the balance factor between semantic and spatial relationships; D refers to the diagonal length of the image, used to normalize spatial distance. Through the above steps, the extracted semantic feature graph structure is aligned with candidate file box templates, i.e., semantic templates representing different file categories. The purpose of alignment is to find the file category that is most semantically similar to the user-uploaded image. During alignment, the Gromov-Wasserstein distance is used to measure the similarity between the two graph structures. Specifically, this is quantified using the graph matching score formula, which is as follows: This formula non-linearly fuses differences in shape, pattern, and layout, mapping them to similarity scores in the 0-1 range. It can dynamically match user-uploaded images with pre-stored clustering templates, triggering printing or further semantic matching. Representation graph matching score, The semantic node number difference between the image and the candidate file boxes is represented, where GW is the Gromov-Wasserstein distance, which measures the alignment of the two graph structures. Semantic-spatial fusion modeling is performed through the above steps, capturing the correlation of semantic elements through graph structures, such as the relative positions of the title and logo. This distance metric considers the relative relationships between nodes in the graph structure, not just the features of the nodes themselves. By calculating the Gromov-Wasserstein distance between the graph structure of the user-uploaded image and the graph structure of the candidate file boxes, the semantic layout matching degree between them can be evaluated. Simultaneously, by analyzing the difference in the number of nodes between data items, the interference of local noise regions with the matching results is avoided. Subsequently, similarly, if a unique... satisfy If only one category matches the uploaded image semantically in layout, it is considered to be of the same type. If multiple categories still meet the criteria, the user is prompted to select one from the client's settings. In other words, if the semantically matched file category is the same (meaning only one category highly matches the uploaded image semantically in layout), then that category will be output as the matching file category. However, if the semantically matched file category is still multi-category (meaning multiple categories match the uploaded image semantically in layout), then the file box identifiers for these categories will be sent to the user for selection. The user can choose the most suitable category from the multiple matching categories as the matching file category and send this selection as feedback to the system. Finally, the system will determine the matching file category based on the user's feedback and design and print the file box accordingly.

[0024] In a preferred embodiment, when the format matching file category or the semantic matching file category is zero, an unfit identifier is fed back to the user terminal for the file image to be packaged.

[0025] Furthermore, if either the format matching or semantic matching of the file category yields a zero result—meaning no matching file category was found—the system will not directly proceed with the file encapsulation operation. In this case, a specific feedback mechanism is employed: a "no-fit" flag is added to the file image to be encapsulated. This flag explicitly informs the user that the current file image cannot match any file category in the system and therefore cannot be encapsulated. Finally, this file image with the "no-fit" flag is sent back to the user. Upon seeing this feedback, the user understands why their file image cannot be encapsulated and can make subsequent processing or adjustments accordingly. This process design ensures both the standardization of system operations and provides a user-friendly feedback mechanism, making the entire file encapsulation process more efficient and reliable.

[0026] In a preferred embodiment, file box shape and size clustering is performed to obtain first-level file clustering results, including:

[0027] Using length, width, and height as coordinate axes, a three-dimensional distribution space is constructed to distribute the shapes and sizes of the file boxes, obtaining a set of electronic file distribution coordinates. The number of clusters in this set of electronic file distribution coordinates is determined using the elbow rule. A clustering evaluation function is then constructed.

[0028] ,

[0029] in, Characterizes the cluster evaluation value, , and The weights representing length, width, and height , , , , and The sum of equals 1. , and The three-dimensional coordinates representing the center of the j-th cluster, and K representing the number of clusters. , and The coordinates representing the distribution of any electronic file in the j-th class. The distribution coordinate set of electronic archives of the j-th class is characterized; clustering optimization is performed on the distribution coordinate set of electronic archives based on the clustering evaluation function to obtain the first-level archive clustering result.

[0030] Specifically, to perform cluster analysis on the shapes and dimensions of a series of file boxes, a three-dimensional distribution space is constructed. This space uses the length, width, and height of the file boxes as three coordinate axes, forming a three-dimensional coordinate system. In this space, the shape and dimensions of each file box are mapped to a unique coordinate point, thus forming a set of electronic file distribution coordinates. This set visually displays the distribution of all file boxes in terms of shape and size. Further, the elbow rule, a statistical method, is used to determine how many clusters this coordinate set will be formed into. The elbow rule calculates the clustering effect under different numbers of clusters and plots a graph showing the relationship between clustering effect and the number of clusters. Then, an "elbow" point is found on the graph; the number of clusters corresponding to this point is generally considered optimal. Subsequently, a clustering evaluation function is constructed to quantify the quality of the clustering effect. Its value, i.e., the clustering evaluation value, depends on multiple factors, including the weights of length, width, and height. These weights reflect the importance given to different size dimensions during the clustering process and the distance between the cluster centers and the electronic file distribution coordinates. Specifically, the formula for calculating the clustering evaluation function is: The underlying logic of the formula design is a variant of K-means based on weighted Euclidean distance, which optimizes shape similarity classification by introducing weight coefficients for length, width, and height; among which, Characterizes the cluster evaluation value, , and The weights representing length, width, and height can be customized according to the needs of the scenario. For example, if the file box needs to fit a fixed shelf, the weights of length and width may be higher. By default, they can be based on the following... , , Select, , and The sum equals 1. Alternatively, in practical applications, cross-validation can be used to adjust the weights, selecting the combination with the highest classification accuracy. , and The three-dimensional coordinates representing the center of the j-th cluster, and K representing the number of clusters. , and The coordinates representing the distribution of any electronic file in the j-th class. This represents the set of coordinates representing the distribution of electronic archives in the j-th category. In other words, for each cluster (i.e., the j-th category), the three-dimensional coordinates of its cluster center are calculated (from...). , and (represented), and then calculate the distribution coordinates of each electronic file in that category (by...). , and The distance between the data points and the cluster centers is calculated, and these distances are weighted and summed according to their weights. This summation is then divided by the number of clusters, K, and the number of distribution coordinates for that type of electronic archive to obtain the cluster evaluation value for that cluster. Performing the same calculation on all clusters yields an overall cluster evaluation value. Subsequently, cluster optimization is performed on the set of electronic archive distribution coordinates based on this evaluation function. This process may involve multiple iterations, each trying different clustering schemes and calculating the corresponding cluster evaluation value. By comparing the cluster evaluation values ​​of different schemes, the optimal clustering scheme can be found; this scheme is the first-level archive clustering result. In summary, by constructing a three-dimensional distribution space, determining the number of clusters, constructing a cluster evaluation function, and implementing cluster optimization steps, cluster analysis of the shape and size of archive boxes has been achieved, providing strong support for subsequent archive management.

[0031] In a preferred embodiment, cluster optimization is performed on the electronic archive distribution coordinate set based on the clustering evaluation function to obtain the first-level archive clustering result, including:

[0032] Based on the number of clusters, several groups of cluster centers are initialized in the electronic archive distribution coordinate set; a distance evaluation function is constructed, wherein the distance evaluation function is used to calculate the Euclidean distribution distance between any two groups of cluster centers; the several groups of cluster centers are extracted and processed based on the cluster evaluation function to obtain several cluster evaluation values; the first cluster center with a smaller number of cluster evaluation values ​​and the second cluster center with a larger number of cluster evaluation values ​​are extracted; taking the first cluster center with a larger number of cluster evaluation values ​​as the target, the distance of the second cluster center with a larger number of cluster evaluation values ​​is adjusted based on the distance evaluation function to obtain updated cluster centers; when the optimal cluster center is no longer updated after a preset number of consecutive iterations, the first-level archive clustering result is obtained.

[0033] For example, after determining the number of clusters, the optimal cluster centers can be found through repeated iterations and adjustments of the electronic archive distribution coordinate set, thus obtaining the first-level archive clustering result. First, several groups of cluster centers are initialized in the electronic archive distribution coordinate set according to the previously determined number of clusters. These cluster centers form the basis for subsequent clustering operations, and their positions are continuously adjusted during the iteration process. Next, to quantify the distribution relationship between cluster centers, a distance evaluation function is constructed. This function, based on Euclidean distance, is used to calculate the distance between any two groups of cluster centers. This function provides an intuitive understanding of the relative positional relationship between cluster centers, providing a basis for subsequent optimization operations. Then, these initialized cluster centers are extracted and processed based on the previously constructed clustering evaluation function. Several clustering evaluation values ​​are obtained through calculation, reflecting the quality of the current clustering scheme. After obtaining the clustering evaluation values, the smaller first-number cluster centers and the larger second-number cluster centers are extracted from the several clustering evaluation values. The purpose of this step is to identify cluster centers with relatively poor clustering performance (higher evaluation values) but still room for optimization, and cluster centers with relatively good clustering performance (lower evaluation values) as references. Next, using the cluster centers with higher evaluation values ​​from the first set as targets, distance reduction adjustments are made to the cluster centers with higher evaluation values ​​from the second set based on the distance evaluation function. The purpose of this step is to improve the clustering effect by adjusting the positions of the cluster centers to make them more closely clustered around the target cluster center. After the above adjustments, the clustering evaluation values ​​are recalculated, and it is determined whether the optimal cluster center has stopped updating for a preset number of consecutive times. If this condition is met, it indicates that the current clustering scheme is relatively stable and the clustering effect has reached its optimum. At this point, the current clustering result can be output as the first-level archive clustering result. In summary, this process, through steps such as initializing cluster centers, constructing the distance evaluation function, calculating clustering evaluation values, extracting key cluster centers, performing distance reduction adjustments, and determining whether the optimal cluster center is stable, achieves clustering optimization of the electronic archive distribution coordinate set, ultimately obtaining the first-level archive clustering result.

[0034] In a preferred embodiment, the primary archive clustering results are further clustered using archive box cover patterns to obtain secondary archive clustering results, including:

[0035] The first-level archive clustering results are subjected to pairwise calculations to obtain several color histogram distances and several texture feature distances for the archive box cover patterns; principal component analysis is performed on the several color histogram distances and several texture feature distances to reduce dimensionality and obtain several dimensionality-reduced distance features; based on the dimensionality-reduced distance threshold and combined with the several dimensionality-reduced distance features, the first-level archive clustering results are subjected to archive box cover pattern clustering to obtain the second-level archive clustering results.

[0036] Optionally, cluster analysis of the cover patterns of the archive boxes is performed based on the results of the first-level archive clustering. To capture the similarities and differences between the cover patterns, two key features are used: color histogram distance and texture feature distance. The color histogram visually reflects the color distribution of the pattern, while texture features reveal the detailed structure and arrangement patterns. First, pairwise calculations are performed on each pair of cover patterns in the first-level archive cluster to calculate their color histogram distance and texture feature distance. This step aims to quantify the visual differences between the cover patterns, providing data support for subsequent clustering. However, directly using distance features for clustering may result in excessive dimensionality, affecting the efficiency and accuracy of clustering. Therefore, Principal Component Analysis (PCA) is used for dimensionality reduction. By performing PCA on several color histogram distances and several texture feature distances, the most representative dimensionality-reduced distance features can be extracted. These features retain most of the information in the original data while significantly reducing the dimensionality. After obtaining the dimensionality-reduced distance features, the next step is to perform second-level archive clustering. In this step, a dimensionality reduction distance threshold is set to determine which cover images are visually similar enough to be grouped together. Combining several dimensionality reduction distance features and the preset threshold, the primary archive clustering results are further subdivided, grouping archives with similar cover images together to obtain the secondary archive clustering results. In summary, by calculating the color histogram distance and texture feature distance of the cover images in the primary archive clustering results, performing principal component analysis for dimensionality reduction, and implementing clustering based on the dimensionality reduction distance threshold, secondary archive clustering was successfully achieved, further refining the archive classification results.

[0037] In a preferred embodiment, the cover pattern layout position clustering is performed on the secondary archive clustering results to obtain the tertiary archive clustering results, including:

[0038] Using a Gaussian mixture model, the secondary archive clustering results are traversed to model the center coordinates of the cover image, obtaining several center coordinates of the cover image, including: constructing a cover image layout position evaluation function:

[0039] ,

[0040] ,

[0041] in, Representation in two-dimensional coordinates The joint probability density of the observed center point of the cover pattern, M represents the number of preset cover pattern layout positions. The mixture weights, representing the m-th Gaussian distribution, indicate the proportion of that component's contribution to the overall data distribution. The covariance matrix representing the m-th Gaussian distribution describes the diffusion direction and extent of the layout. Diagonal elements... and These represent the variances in the z and v directions, respectively. Larger values ​​indicate a more dispersed distribution. Off-diagonal elements... Characterizing the correlation between z and v, The layout tends to spread along the diagonal. Tendency to concentrate in the center The mean vector representing the m-th Gaussian distribution represents the center coordinates of the layout pattern. Based on the cover pattern layout position evaluation function, the center coordinates of the cover patterns are optimized by traversing the secondary archive clustering results, resulting in several possible center coordinates for each cover pattern. The probability density function characterizes the two-dimensional Gaussian distribution, and the layout distribution characteristics of the m-th component are also represented.

[0042] Furthermore, to further cluster the cover pattern layout positions based on the secondary archive clustering results to obtain tertiary archive clustering results, a Gaussian Mixture Model (GMM) is used to model and analyze the center coordinates of the cover patterns. Specifically, an evaluation function is constructed to describe the joint probability density of observing the center point of the cover pattern in two-dimensional coordinates. This function is based on the Gaussian Mixture Model, where M represents the preset number of cover pattern layout positions. Its specific form is as follows:

[0043] ,

[0044] This formula uses a Gaussian mixture model (GMM) to model the center coordinates of the cover pattern, which can capture the directionality and dispersion of the layout distribution, and realize the quantification of the probability of position matching. Representation in two-dimensional coordinates The joint probability density of the observed center point of the cover pattern, M represents the number of preset cover pattern layout positions. The mixture weights characterize the m-th Gaussian distribution, representing the proportion of the m-th Gaussian distribution's contribution to the entire data distribution. The covariance matrix representing the m-th Gaussian distribution describes the diffusion direction and extent of the layout. Diagonal elements... and These represent the variances in the z and v directions, respectively. Larger values ​​indicate a more dispersed distribution. Off-diagonal elements... Characterizing the correlation between z and v, The layout tends to spread along the diagonal. Tendency to concentrate in the center The mean vector representing the m-th Gaussian distribution indicates the center coordinates of this layout pattern. Simultaneously, after obtaining the secondary archival clustering results, the cover patterns in each cluster are traversed, and their center coordinates are extracted. The constructed evaluation function is used to optimize the center coordinates of the cover patterns in each secondary cluster. This typically involves iteratively adjusting the parameters of the Gaussian mixture model to maximize the likelihood function of the observed data. Optimization methods such as the expectation-maximization algorithm can be used during the optimization process. After optimization, several center coordinates of the cover patterns are obtained, representing the center positions of different layout patterns.

[0045] For example, suppose in a layout clustering of file box cover patterns, the dataset contains 100 pattern center coordinates. Furthermore, we aim to divide these patterns into 3 clusters, i.e., M=3. The following are the specific steps and calculation process of the EM algorithm based on the Gaussian Mixture Model (GMM).

[0046] First, we perform initialization, assuming the initial blending weights are: , , These weights represent the initial assumption that 40% of the data points belong to the first Gaussian distribution, 30% to the second, and 30% to the third. Subsequently, for each data point... Calculate the posterior probability that it belongs to each Gaussian distribution, i.e. These posterior probabilities represent the probability that each data point belongs to each Gaussian distribution.

[0047] For example, for the first data point ,but Similarly, calculate and Next, update the weights of each Gaussian distribution, assuming the calculation result is: , , Then repeat the calculation and update steps until the weights converge. For example, after multiple iterations, the final weights are: , , In other words, the final weights indicate that 60% of the patterns belong to the first Gaussian distribution (e.g., a centered layout), 25% to the second, and 15% to the third. Through the above steps, a Gaussian Mixture Model (GMM) is used to cluster the layout positions of the file box cover patterns. By calculating and updating the iterations, the weights of each Gaussian distribution are finally obtained, representing the distribution of different layout patterns. Ultimately, the layout positions of the cover patterns are modeled and optimized, thereby achieving intelligent design and printing of file box cover patterns.

[0048] In a preferred embodiment, based on the cover pattern layout position evaluation function, the secondary archive clustering results are traversed to optimize the center coordinates of the cover pattern, resulting in several cover pattern center coordinates, including:

[0049] Extract the cover pattern of the first file box from the secondary file clustering results; construct the posterior probability function: In the E-step of the EM algorithm, the probability that a data point belongs to the m-th Gaussian distribution is calculated. This allows for dynamic allocation of data points to various Gaussian components, guiding parameter updates. The posterior probability of data point i with respect to the m-th Gaussian distribution after the t-th iteration is represented by . The weights characterizing the m-th Gaussian distribution at the t-th iteration. The coordinates of the center of the cover pattern representing the m-th Gaussian distribution at the t-th iteration. The covariance matrix of the m-th Gaussian distribution at the t-th iteration The sum of the posterior probabilities of data point i with respect to M Gaussian distributions after the t-th iteration is represented, with t initially equal to 0; the weight update formula is constructed as follows: ,in, The (t+1)th iteration represents the m-th Gaussian distribution, and n represents the number of data points for the cover pattern. The Gaussian distribution mixture weights are initialized to obtain the initial weights of the Gaussian mixture distribution, wherein the initial weights of the Gaussian mixture distribution are equal to 1 / M. The cover pattern layout position evaluation function, the posterior probability function, and the weight update formula are iterated until the weight change of any Gaussian distribution for a consecutive preset number of iterations is not greater than or equal to the weight fluctuation threshold, or the number of iterations meets the preset number of iterations. The coordinate of the center of the cover pattern of the first file box is then set as the coordinate of the center of the cover pattern and added to the coordinates of the center of the several cover patterns.

[0050] Specifically, the cover pattern data of the first file box is extracted from the secondary file clustering results. These data points represent the cover pattern in a certain space or feature space, and are used for subsequent analysis and processing. To model and analyze this cover pattern data, a posterior probability function is constructed. This function is based on a Gaussian Mixture Model (GMM) to describe the probability that a data point belongs to different Gaussian distributions (i.e., clusters). Several key parameters are included in this function: the posterior probability of data point i with respect to the m-th Gaussian distribution, the weight of the m-th Gaussian distribution, the center coordinates of the cover pattern, and the covariance matrix. These parameters collectively determine the likelihood of a data point being assigned to different clusters. The specific formula for the posterior probability function is as follows: ,in, The posterior probability of data point i with respect to the m-th Gaussian distribution after the t-th iteration is represented by . The weights characterizing the m-th Gaussian distribution at the t-th iteration. The coordinates of the center of the cover pattern representing the m-th Gaussian distribution at the t-th iteration. The covariance matrix of the m-th Gaussian distribution at the t-th iteration Let t represent the sum of the posterior probabilities of data point i with respect to M Gaussian distributions after the t-th iteration, with t initially equal to 0. Then, to update the weights of these Gaussian distributions, a weight update formula is constructed. This formula calculates the weight of each Gaussian distribution in the next iteration based on the data point allocation (i.e., posterior probability) of the current iteration. Weight updating is a crucial step in the iteration process, enabling the model to gradually converge to the optimal solution. The specific weight update formula is: ,in, Let n represent the (t+1)th iteration of the m-th Gaussian distribution, where n represents the number of data points for the cover pattern. Before starting the iteration, the initial weights of the Gaussian mixture distribution need to be initialized. Specifically, the initial weights can be set to 1 / M, where M is the number of Gaussian distributions (or clusters). This is to ensure that each cluster has equal importance at the start of the iteration. Next, the iteration is performed iteratively based on the cover pattern layout position evaluation function, the posterior probability function, and the weight update formula. In each iteration, the weights and center coordinates of the Gaussian distributions are updated according to the current data point allocation. The iteration process continues until one of two stopping conditions is met: first, in a preset number of iterations, the weight change of any Gaussian distribution is not greater than or equal to the weight fluctuation threshold, indicating that the model has converged to a stable state; second, the number of iterations reaches a preset maximum number, which is to prevent the iteration process from continuing indefinitely. When the iteration stops, the center coordinates of the cover pattern of the Gaussian distribution with the largest weight are output as the optimal center position of the first file box cover pattern, and this is added to the set of center coordinates of several cover patterns. This process is repeated for each file box in the secondary file clustering results, resulting in a series of optimal center coordinates for the cover patterns. In summary, the entire process is an iterative optimization based on a Gaussian mixture model, aiming to determine the optimal center position of the cover pattern using mathematical and statistical methods.

[0051] The present invention provides a method for optimizing the design and printing of file boxes based on an electronic archive system, which has at least the following technical effects:

[0052] 1. By performing clustering based on the shape and size of the file boxes, the cover patterns, and the layout of the cover patterns, a multi-dimensional and detailed classification of electronic file collections is achieved. This multi-dimensional clustering method not only improves the accuracy and personalization of file box design but also makes file storage and management more efficient and orderly. By comprehensively considering the physical dimensions of the file boxes, the cover patterns, and their layout, file box design solutions that better meet users' actual needs can be provided.

[0053] 2. After the electronic archive system is initialized, it can receive images of archives to be packaged uploaded by users and perform archive matching throughout the electronic archive system. Through a dual mechanism of archive format layout matching and archive semantic layout matching, the category of the archives to be packaged can be determined more accurately. When the matching result is multiple categories, the archive box identifier can also be sent to the user for selection, realizing an intelligent matching and feedback mechanism. This mechanism not only improves the accuracy and efficiency of archive packaging but also enhances user participation and satisfaction.

[0054] 3. A Gaussian mixture model is innovatively introduced into the cover pattern layout position clustering. By constructing an evaluation function for the cover pattern layout position and modeling and optimizing the center coordinates of the cover patterns, the center coordinates of several cover patterns can be accurately determined. This method not only improves the accuracy and rationality of the cover pattern layout position but also provides strong data support for subsequent file box design and printing. Furthermore, the parameters of the Gaussian mixture model can be continuously optimized through iterative iteration and weight update formulas, further improving the accuracy and stability of the cover pattern layout position clustering.

[0055] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for optimizing the design and printing of file boxes based on an electronic archive system, characterized in that, include: Download the collection of electronic archives with archive box identifiers, perform archive box shape and size clustering, and obtain the first-level archive clustering results; The primary archive clustering results are then used to cluster archive box cover patterns to obtain secondary archive clustering results; The cover pattern layout position clustering is performed on the secondary archive clustering results to obtain the tertiary archive clustering results; After initializing the electronic archive system based on the three-level archive clustering results, the system receives archive images to be packaged uploaded by the user, traverses the electronic archive system to perform archive matching, and obtains the matching archive category. When the matched files belong to the same category, the file box is designed and printed according to the category file box identifier; Specifically, after initializing the electronic archive system based on the three-level archive clustering results, the system receives archive images to be packaged uploaded by the user, traverses the electronic archive system to perform archive matching, and obtains the matching archive categories, including: After initializing the electronic archive system based on the three-level archive clustering results, the system receives archive images to be packaged uploaded by the user, traverses the electronic archive system to perform archive format layout matching, and obtains the format-matched archive category. When the format matching file belongs to the same category, the format matching file category is output as the matching file category; When the format matching file belongs to multiple categories, the format matching files are traversed to perform semantic layout matching to obtain the semantic matching file category. When the semantically matched file categories are the same, the semantically matched file category is output as the matched file category; When the semantically matched file category is multiple, the file box identifier of the semantically matched file category is sent to the user terminal for selection, and feedback information is obtained, wherein the feedback information includes the matched file category.

2. The method as described in claim 1, characterized in that, Also includes: When the format matching file category or the semantic matching file category is zero, an unfit identifier is generated for the file image to be packaged and fed back to the user terminal.

3. The method as described in claim 1, characterized in that, Perform file box shape and size clustering to obtain first-level file clustering results, including: Using length, width, and height as coordinate axes, a three-dimensional distribution space is constructed to distribute the shape and size of the file boxes, thereby obtaining a set of electronic file distribution coordinates; The number of clusters in the electronic archive distribution coordinate set is determined using the elbow rule; Constructing a clustering evaluation function: , in, Characterizes the cluster evaluation value, , and The weights representing length, width, and height , and The three-dimensional coordinates representing the center of the j-th cluster, and K representing the number of clusters. , and The coordinates representing the distribution of any electronic file in the j-th class. Characterizes the set of coordinates representing the distribution of the j-th type of electronic archives. , , and The sum of them equals 1; Clustering optimization is performed on the electronic archive distribution coordinate set based on the clustering evaluation function to obtain the first-level archive clustering result.

4. The method as described in claim 3, characterized in that, Clustering optimization is performed on the electronic archive distribution coordinate set based on the clustering evaluation function to obtain the first-level archive clustering results, including: Based on the number of clusters, several groups of cluster centers are initialized in the electronic archive distribution coordinate set; Construct a distance evaluation function, wherein the distance evaluation function is used to calculate the Euclidean distribution distance between any two sets of cluster centers; Extract the cluster centers from the aforementioned clusters, process them based on the clustering evaluation function, and obtain several clustering evaluation values. Extract the cluster centers with the smaller first number and the cluster centers with the larger second number from the plurality of cluster evaluation values; Using the first cluster center with the larger number of clusters as the target, the second cluster center with the larger number of clusters is adjusted by reducing the distance based on the distance evaluation function to obtain updated cluster centers; The first-level archive clustering result is obtained when the optimal cluster center is no longer updated after a preset number of consecutive iterations.

5. The method as described in claim 1, characterized in that, The primary archive clustering results are then used to cluster archive box cover patterns to obtain secondary archive clustering results, including: The first-level archive clustering results are used to perform pairwise calculations to obtain several color histogram distances and several texture feature distances for the cover patterns of the archive boxes; Principal component analysis is performed to reduce the dimensionality of the color histogram distances and the texture feature distances to obtain several dimensionality-reduced distance features; Based on the dimensionality reduction distance threshold and combined with the aforementioned dimensionality reduction distance features, the primary archive clustering results are used to cluster the archive box cover patterns to obtain the secondary archive clustering results.

6. The method as described in claim 1, characterized in that, The secondary archive clustering results are then subjected to cover pattern layout position clustering to obtain tertiary archive clustering results, including: Using a Gaussian mixture model, the secondary archive clustering results are traversed to model the center coordinates of the cover image, resulting in several center coordinates of the cover image, including: Construct a function to evaluate the layout and position of the cover image: , , in, Representation in two-dimensional coordinates The joint probability density of the observed center point of the cover pattern, M, represents the number of preset cover pattern layout positions. The mixture weights characterize the m-th Gaussian distribution, representing the proportion of the m-th Gaussian distribution's contribution to the entire data distribution. The covariance matrix representing the m-th Gaussian distribution describes the diffusion direction and extent of the layout. Diagonal elements... and These represent the variances in the z and v directions, respectively. Larger values ​​indicate a more dispersed distribution. Off-diagonal elements... Characterizing the correlation between z and v, The layout tends to spread along the diagonal. Tendency to concentrate in the center The mean vector representing the m-th Gaussian distribution represents the center coordinates of this layout pattern. The probability density function characterizes the two-dimensional Gaussian distribution, and the layout distribution characteristics of the m-th component are also represented. Based on the cover pattern layout position evaluation function, the secondary archive clustering results are traversed to optimize the center coordinates of the cover pattern and obtain several cover pattern center coordinates.

7. The method as described in claim 6, characterized in that, Based on the cover pattern layout position evaluation function, the secondary archive clustering results are traversed to optimize the center coordinates of the cover pattern, resulting in several cover pattern center coordinates, including: Extract the cover pattern of the first archive box from the secondary archive clustering results; Construct the posterior probability function: , in, The posterior probability of data point i with respect to the m-th Gaussian distribution after the t-th iteration is represented by . The weights characterizing the m-th Gaussian distribution at the t-th iteration. The coordinates of the center of the cover pattern representing the m-th Gaussian distribution at the t-th iteration. The covariance matrix of the m-th Gaussian distribution at the t-th iteration The sum of the posterior probabilities of data point i with respect to M Gaussian distributions after the t-th iteration is given, where t is initially equal to 0. The probability density function characterizes the two-dimensional Gaussian distribution, and the layout distribution characteristics of the m-th component are also represented. Construct the weight update formula: , in, The m-th Gaussian distribution is represented by the (t+1)-th iteration, and n represents the number of data points for the cover pattern. The Gaussian mixture weights are initialized to obtain the initial weights of the Gaussian mixture distribution, wherein the initial weights of the Gaussian mixture distribution are equal to 1 / M; The system iterates according to the cover pattern layout position evaluation function, the posterior probability function and the weight update formula until the weight change of any Gaussian distribution for a consecutive preset number of times is not greater than or equal to the weight fluctuation threshold, or the number of iterations meets the preset number of times. Then, the coordinate of the center of the cover pattern of the first file box is set as the coordinate of the center of the cover pattern and added to the coordinates of the center of the cover pattern.