Real-time paper fiber intelligent analysis system based on large model visual extraction
Through multimodal fusion module and large-model visual extraction technology, high-precision, real-time detection and analysis of paper fibers are achieved, the problem of inefficiency in traditional methods is solved, the detection accuracy and adaptability are improved, and the efficient operation of industrial production is supported.
Patent Information
- Application Number
- CN202510418285.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is inefficient in paper fiber detection, making it difficult to refinely identify local detailed features of complex objects. In addition, multimodal large models lack local interpretability during paper fiber detection, and cannot meet the real-time and high-precision detection requirements.
The multimodal fusion module is adopted, combining image encoder and text encoder for multi-scale slicing and feature vectorization, and feature fusion and reasoning are used to use the extensible knowledge base and RAG mechanism to make final decisions through large models.
It improves the accuracy of paper fiber detection and the ability to adapt to complex production environments, can quickly adapt to new scenarios, provide efficient and accurate fiber analysis and processing suggestions, and improves industrial production efficiency and automation level.
Smart Images

Figure CN120296667A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mechanical technology, and particularly to a real-time paper fiber intelligent analysis system based on large model vision extraction. Background Art
[0002] With the booming development of industries such as papermaking and packaging, the demand for the detection of paper fiber quality and composition is increasing day by day. In actual production, the quality of paper fibers is directly related to the key performance of the final product. In the papermaking industry, for example, the quality of paper fibers determines the physical strength, burst resistance, and smoothness of paper; in the fields of packaging and printing, its structure and distribution affect printing adaptability and packaging strength. Therefore, accurate detection and analysis of paper fibers are crucial.
[0003] However, there are many problems with current detection methods. Traditional methods mostly rely on manual sampling or single-modal sensors, such as microscopic images. The detection efficiency is low. In complex production environments, it is difficult to finely identify the morphology of paper fibers, and it is easily interfered, resulting in frequent misjudgments. Some object detection algorithms based on computer vision and deep learning, although having certain achievements in image object detection, are unable to handle complex objects like paper fibers when it comes to finely identifying their diverse internal features, such as fiber thickness, distribution, and origin. Existing multi-modal large models, such as CLIP, although having the ability to process multiple-modal data, are difficult to effectively capture the detailed features of local regions when dealing with tasks like paper fibers that require attention to local details, lacking local interpretability and pertinence, and unable to meet the industry's requirements for real-time, high-precision detection and analysis of paper fibers. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and to propose a real-time paper fiber intelligent analysis system based on large model vision extraction.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A real-time paper fiber intelligent analysis system based on large model vision extraction, comprising:
[0007] A multi-modal fusion module, the multi-modal fusion module includes an image encoder and a text encoder. The image encoder is used to divide the input paper fiber image into multiple scales to generate multiple groups of sub-images, and to vectorize the features of the sub-images at different scales; the text encoder is used to convert the semantic descriptions related to paper fibers into text features; through this module, the complementary relationship between vision and language is realized, improving the detection accuracy of paper fibers and the adaptability to complex production environments;
[0008] An extensible knowledge base module for storing information covering different pulp types, different fiber processing technologies, detection standards, and industry experience, supporting the retrieval and update of multi-dimensional attributes of products, and featuring sustainable incremental expansion;
[0009] A visual content relevance processing module that calculates the cosine similarity of image and text features based on the features output by the image encoder and the text encoder, selects the top-k groups of the most discriminative image features according to the variance of the similarity distribution vector, and then performs weighted fusion of the selected local image features and the corresponding text features to form a multi-modal feature representation with high relevance;
[0010] A multi-modal large model inference and RAG module that inputs the high-confidence features screened and merged by visual content relevance into the large model for preliminary inference to obtain information on the type or defects of paper fibers; with the help of the RAG mechanism, it performs a secondary retrieval match between the model output and the knowledge base to find more refined processing or improvement suggestions, and then merges the retrieval results with the preliminary inference results to output the final decision or diagnostic report.
[0011] Preferably, in the multi-modal fusion processing module, the scales at which the image encoder segments the paper fiber image include, but are not limited to, 4×4, 2×2, and 1×1. Global background information is retained and local details are extracted through multi-scale segmentation, and the segmentation ratio is designed according to the data acquisition quality and microscopy level in the industrial field.
[0012] Preferably, the information stored in the extensible knowledge base module includes the compositional characteristics, mechanical property indicators, applicable scenarios of different fiber types, as well as the standards and operating procedures for production, detection, and post-processing, and can incrementally add newly detected scenarios or new fiber types in the production line or laboratory environment.
[0013] Preferably, in the visual content relevance processing module, the formula for calculating the cosine similarity of image and text features is: where A is the image feature vector and B is the text feature vector.
[0014] Preferably, in the multi-modal large model inference and RAG module, the large model deeply fuses images and texts based on technologies such as contrastive learning and self-attention mechanisms to infer relevant information about paper fibers; after the model inference is completed, the RAG mechanism performs a retrieval on the key inference results, searches for knowledge or expert experience cases in the field of paper fibers within the system, and inserts the retrieved refined knowledge entries into the inference context for the large model to perform the final comprehensive inference or summary output.
[0015] Preferably, it further includes a modular deployment module. Each functional module of the system is designed in a loosely coupled manner, flexibly split and combined according to the actual production line requirements, and integrated with the existing industrial control or laboratory management system through containerization or microservices deployment methods.
[0016] Preferably, the analysis system uses a real-time paper fiber intelligent analysis method based on large model vision extraction, including the following steps:
[0017] S1. Image decomposition: Perform multi-scale decomposition on the input paper fiber image to generate sub-images with different resolutions and local ranges.
[0018] S2. Feature extraction and screening: Use a pre-trained image encoder to vectorize the features of sub-images at different scales, use a pre-trained text encoder to convert the semantic descriptions related to paper fibers into text features, calculate the cosine similarity of the image-text features, and select the Top-k groups of the most discriminative image features according to the variance of the similarity distribution vector.
[0019] S3. Feature fusion: Perform weighted fusion on the selected local image features and the corresponding text features to form a multi-modal feature representation with high correlation.
[0020] S4. Model inference and retrieval: Input the fused features into the large model for preliminary inference to obtain the category or defect information of the paper fibers, use the RAG mechanism to perform secondary retrieval and matching of the model output with the knowledge base, merge the retrieval results and the preliminary inference results, and output the final decision or diagnostic report.
[0021] Preferably, before the image decomposition, it further includes the step of initializing the system.
[0022] Preferably, after the model inference and retrieval step, it further includes the step of adjusting the production process or further researching the paper fibers according to the final decision or diagnostic report, and feeding the newly obtained knowledge or cases back to the extensible knowledge base module for updating.
[0023] The present invention has the following beneficial effects:
[0024] 1. Through multi-scale visual decomposition, the present invention cuts the paper fiber image into multiple scales, such as 4×4, 2×2, 1×1, which not only retains the global background information but also can accurately extract local detail features such as fiber breakage and impurities. In the process of extracting visual content relevance, key regions are screened out by calculating the similarity of image-text features, realizing the accurate alignment of microscopic features and text descriptions. Compared with the traditional method that relies on single visual information, it greatly improves the ability to capture tiny fiber features, effectively improves the accuracy of fiber detection and analysis, and reduces detection errors and missed detections.
[0025] 2. The system core adopts a multi-modal fusion structure, which combines an image encoder, a text encoder, a knowledge base, and large model reasoning to fuse various modal information. Its extensible knowledge base covers different pulp types, fiber processing technologies, detection standards, and industry experience, and supports incremental updates, enabling it to quickly adapt to open scenarios such as new types of pulp, recycled fibers, or other composite materials. In the face of changes in factors such as raw materials, temperature, and humidity in the production environment, the system quickly retrieves and matches solutions from the knowledge base through the RAG mechanism, demonstrating strong environmental adaptability and the ability to adaptively learn new scenarios and knowledge.
[0026] 3. To meet the high requirements of industrial production lines for detection efficiency and stability, the system adopts a lightweight and parallelizable design concept in three key modules: image decomposition, correlation selection, and combined reasoning, enabling quasi-real-time analysis of large volumes of image data. With just one image, multi-task analysis can be completed in an open environment, reducing the dependence on a large amount of labeled data and the deployment difficulty and maintenance cost. At the same time, through containerized or microservices-based modular deployment methods, it is easy to integrate with existing industrial control or laboratory management systems, providing enterprises with efficient and convenient quality control tools and improving industrial production efficiency.
[0027] 4. This invention breaks through the limitation of traditional detection that only focuses on the detection link. By deeply integrating a multi-modal large model with an extensible knowledge base, an integrated closed-loop of "detection - diagnosis - decision" is constructed. The system can not only detect fiber abnormalities, defects, or impurities, but also automatically recommend corresponding treatment solutions based on historical experience, industry standards, and expert strategies in the knowledge base, such as optimizing the production process and adjusting process parameters. Through continuous iteration and incremental learning, the knowledge reserve and diagnostic ability of the system are continuously improved, providing strong support for enterprises in digital transformation and intelligent upgrading, and comprehensively enhancing the automation level and competitiveness of the industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is the system block diagram of a real-time paper fiber intelligent analysis system based on large model vision extraction proposed by the present invention;
[0029] Figure 2 is the method flow chart of a real-time paper fiber intelligent analysis system based on large model vision extraction proposed by the present invention;
[0030] Figure 3 is the schematic diagram of image encoder image decomposition proposed by the present invention;
[0031] Figure 4 is the schematic diagram of the text encoder principle proposed by the present invention;
[0032] Figure 5 is the schematic diagram of the vectorization of sub-graph features of the image encoder proposed by the present invention;
[0033] Figure 6 This is the schematic diagram of feature fusion proposed by the present invention;
[0034] Figure 7 This is the schematic diagram of the merger of large model inference and retrieval proposed by the present invention. Specific implementation manners
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0036] Refer to Figure 1-7 , a real-time paper fiber intelligent analysis system based on large model visual extraction, including:
[0037] A multi-modal fusion module, the multi-modal fusion module includes an image encoder and a text encoder. The image encoder is used to divide the input paper fiber image at multiple scales to generate multiple groups of sub-images, and perform feature vectorization on the sub-images at different scales; the text encoder is used to convert the semantic description related to the paper fiber into text features; through this module, the complementary relationship between vision and language is realized, and the detection accuracy of paper fibers and the adaptability to complex production environments are improved; in the multi-modal fusion processing module, the scales at which the image encoder divides the paper fiber image include but are not limited to 4×4, 2×2, 1×1. By multi-scale segmentation, global background information is retained and local details are extracted, and the segmentation ratio is designed according to the data acquisition quality and microscopic level of the industrial site.
[0038] An extensible knowledge base module, which is used to store information covering different pulp types, different fiber processing technologies, detection standards and industry experience, support the retrieval and update of multi-dimensional attributes of products, and has the characteristics of sustainable incremental expansion; the information stored in the extensible knowledge base module includes the component characteristics, mechanical property indexes, applicable scenarios of different fiber types, as well as the standards and operating procedures for production, detection and post-processing, and can incrementally add new scenarios or new fiber types detected in the production line or laboratory environment.
[0039] A visual content relevance processing module, based on the features output by the image encoder and the text encoder, calculates the cosine similarity of the text and image features, and selects the Top-k groups of the most discriminative image features according to the variance of the similarity distribution vector, and then weights and fuses the selected local image features with the corresponding text features to form a multi-modal feature representation with high relevance; in the visual content relevance processing module, the formula for calculating the cosine similarity of the text and image features is: where A is the image feature vector and B is the text feature vector.
[0040] The multimodal large model reasoning and RAG module inputs the highly confident features that have been screened and merged for visual content relevance into the large model for preliminary reasoning to obtain information about the category or defects of paper fibers. With the help of the RAG mechanism, the model output is retrieved and matched with the knowledge base for a second time to find more refined processing or improvement suggestions, and then the retrieval result is merged with the preliminary reasoning result to output the final decision or diagnostic report. In the multimodal large model reasoning and RAG module, the large model deeply fuses the text and images based on techniques such as contrastive learning and self-attention mechanism to infer relevant information about paper fibers. After the model reasoning is completed, the RAG mechanism retrieves the key reasoning results, searches for knowledge in the field of paper fibers or expert experience cases within the system, and inserts the retrieved refined knowledge entries into the reasoning context for the large model to perform final comprehensive reasoning or summary output.
[0041] A computer-readable storage medium is coupled to the above-mentioned modules and stores a computer program thereon.
[0042] It further includes a modular deployment module. Each functional module of the system is designed in a loosely coupled manner, flexibly split and combined according to the actual production line requirements, and integrated with the existing industrial control or laboratory management system through containerization or microservices deployment methods.
[0043] The analysis system uses a real-time intelligent analysis method for paper fibers based on large model vision extraction, including the following steps:
[0044] S1. Image decomposition: The input paper fiber image is decomposed at multiple scales to generate sub-images with different resolutions and local ranges.
[0045] S2. Feature extraction and screening: The pre-trained image encoder is used to vectorize the features of sub-images at different scales, and the pre-trained text encoder is used to convert the semantic descriptions related to paper fibers into text features. The cosine similarity of the text and image features is calculated, and the Top-k groups of the most discriminative image features are selected according to the variance of the similarity distribution vector.
[0046] S3. Feature fusion: The selected local image features are weighted and fused with the corresponding text features to form a multi-modal feature representation with high correlation.
[0047] S4. Model reasoning and retrieval: The fused features are input into the large model for preliminary reasoning to obtain information about the category or defects of paper fibers. With the help of the RAG mechanism, the model output is retrieved and matched with the knowledge base for a second time, and the retrieval result is merged with the preliminary reasoning result to output the final decision or diagnostic report.
[0048] Before the image decomposition, it also includes the step of initializing the system. After the model inference and retrieval step, it also includes the step of adjusting the production process or further studying the paper fibers according to the final decision or diagnostic report, and feeding back the newly obtained knowledge or cases to the extensible knowledge base module for updating.
[0049] Example 1
[0050] Accuracy test for different fiber type identifications
[0051] I. Experimental preparation
[0052] Prepare a variety of known types of paper fiber samples in the laboratory, including 20 samples each of softwood pulp fibers, hardwood pulp fibers, and waste paper pulp fibers. Use a high-resolution microscope to collect fiber images with an image resolution of 2000×2000 pixels to ensure that the fiber details are clearly presented in the images.
[0053] II. Experimental process
[0054] Input the collected images into a real-time paper fiber intelligent analysis system based on large model visual extraction proposed by the present invention. The image decomposition module of the system performs 4×4, 2×2, and 1×1 scale cuts on the images. The image encoder extracts features from sub-images of different scales, and the text encoder converts semantic descriptions such as "softwood pulp fibers", "hardwood pulp fibers", and "waste paper pulp fibers" into text features. The visual content correlation processing module calculates the similarity of text and image features, screens key features, and then the multi-modal large model infers the fiber type.
[0055] III. Experimental results
[0056] After testing 60 samples, the system correctly identified 19 softwood pulp fibers, 18 hardwood pulp fibers, and 19 waste paper pulp fibers, with an overall identification accuracy rate reaching 93.3%. While using the traditional detection method based on single visual information, it could only correctly identify 15 softwood pulp fibers, 13 hardwood pulp fibers, and 14 waste paper pulp fibers, with an accuracy rate of 70%. The comparison data is shown in the following table:
[0057] Fiber type / Identification type This system Traditional method Softwood pulp fiber 19 15 Hardwood pulp fiber 18 13 Wastepaper pulp fiber 19 14
[0058] This example shows that through multi-scale visual decomposition and visual content correlation extraction, the system can accurately capture the microscopic features of fibers and align them with text descriptions, with much higher accuracy in fiber type identification than traditional methods, effectively improving the detection accuracy and demonstrating the advantages of the system in fine-grained detection.
[0059] Example 2
[0060] Adaptability test for simulating production environment changes
[0061] I. Experiment Preparation
[0062] Simulate the changes in the papermaking production environment in the laboratory, and set three variables: temperature, humidity, and the ratio of fiber raw materials. Prepare 10 groups of the same softwood pulp fiber samples, and process the fiber samples under different temperatures (20°C, 25°C, 30°C), different humidities (40%, 50%, 60%), and different mixing ratios of waste paper pulp (0%, 5%, 10%). Use professional environment simulation equipment to ensure the stability of environmental parameters.
[0063] II. Experiment Process
[0064] Collect images of each group of processed fiber samples and input them into the analysis system. The system processes them according to the multi-modal fusion architecture and key algorithm processes, uses an extensible knowledge base and RAG retrieval mechanism to analyze the characteristic changes of fibers under different environments, such as changes in fiber strength and impurity content, and gives corresponding processing suggestions.
[0065] III. Experiment Results
[0066] When the temperature rises, the humidity increases, and the mixing ratio of waste paper pulp increases, the traditional detection method cannot accurately judge the changes in fiber characteristics and cannot give effective processing suggestions. However, this system can accurately identify problems such as a decrease in fiber strength and an increase in impurity content, and retrieve processing suggestions such as adjusting the drying time and increasing the intensity of the impurity removal process from the knowledge base. The specific data is as follows in the table:
[0067]
[0068] This example shows that when faced with changes in the simulated production environment, this system, with the help of the fusion of the multi-modal large model and the knowledge base and the RAG retrieval mechanism, demonstrates good adaptability, can accurately analyze problems and provide effective processing suggestions, demonstrating the beneficial effects of the strong adaptability of the system and providing decision-making support.
[0069] Example 3
[0070] System Processing Efficiency and Stability Test
[0071] I. Experiment Preparation
[0072] Build a simulated image data batch processing environment in the laboratory, prepare 1000 different paper fiber images, and there are three types of image resolutions: 1000×1000, 1500×1500, and 2000×2000, simulating the images obtained by different acquisition devices in actual production.
[0073] II. Experiment Process
[0074] Input these 1000 images into the system one by one, and the system processes the images according to the design concept of being lightweight and capable of parallel processing. Record the time taken by the system to process each image and the number of errors that occur during the continuous processing of 1000 images.
[0075] III. The average time taken by the system to process 1000×1000 resolution images is 0.2 seconds per image, 0.3 seconds per image for 1500×1500 resolution images, and 0.4 seconds per image for 2000×2000 resolution images. Only 2 errors occurred during the continuous processing of 1000 images, and the error rate was 0.2%. Compared with other similar detection systems, when processing images of the same quantity and resolution, the average time taken is between 0.5 - 1 second per image, and the error rate is between 5% - 10%. The data comparison is as follows in the table:
[0076]
[0077] This embodiment shows that when the system processes a large number of paper fiber images with different resolutions, it has a fast processing speed and high stability, demonstrating the beneficial effect of being efficient and convenient in industrial applications, and can meet the requirements of real-time detection on the production line.
[0078] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A real-time paper fiber intelligent analysis system based on large model visual extraction, characterized in that Comprising: A multimodal fusion module, which includes an image encoder and a text encoder. The image encoder is used to segment the input paper fiber image at multiple scales, generate multiple groups of sub-images, and vectorize the features of sub-images at different scales; the text encoder is used to convert the semantic description related to paper fibers into text features; through this module, the complementary relationship between vision and language is realized, and the detection accuracy of paper fibers and the adaptability to complex production environments are improved; An extensible knowledge base module, which is used to store information covering different pulp types, different fiber processing technologies, detection standards and industry experience, support the retrieval and update of multi-dimensional attributes of products, and has the characteristics of sustainable incremental expansion; A visual content relevance processing module, based on the features output by the image encoder and the text encoder, calculates the cosine similarity of the text and image features, and selects the Top-k groups of the most discriminative image features according to the variance of the similarity distribution vector. Then, the selected local image features are weighted and fused with the corresponding text features to form a multi-modal feature representation with high correlation; A multimodal large model inference and RAG module, which inputs the highly confident features screened and merged by visual content relevance into the large model for preliminary inference to obtain the category or defect information of paper fibers; With the help of the RAG mechanism, the model output is retrieved and matched with the knowledge base for the second time to find more refined processing or improvement suggestions, and then the retrieval result is merged with the preliminary inference result to output the final decision or diagnostic report; A computer-readable storage medium, which is coupled to the above-mentioned modules and stores a computer program thereon.
2. The real-time paper fiber intelligent analysis system based on large model vision extraction according to claim 1, wherein In the multimodal fusion processing module, the scales at which the image encoder segments the paper fiber image include but are not limited to 4×4, 2×2, 1×1. Global background information is retained and local details are extracted through multi-scale segmentation, and the segmentation ratio is designed according to the data acquisition quality and microscopic level of the industrial site.
3. The real-time paper fiber intelligent analysis system based on large model vision extraction according to claim 1, wherein, The information stored in the extensible knowledge base module includes the component characteristics, mechanical property indexes, applicable scenarios of different fiber types, as well as the standards and operating procedures for production, detection and post-processing, and can incrementally add new scenarios or new fiber types detected in the production line or laboratory environment.
4. The real-time paper fiber intelligent analysis system based on large model visual extraction according to claim 1, characterized in that, In the visual content relevance processing module, the formula for calculating the cosine similarity of image and text features is as follows: where A is the image feature vector and B is the text feature vector.
5. The real-time paper fiber intelligent analysis system based on large model vision extraction according to claim 1, characterized in that In the multimodal large model inference and RAG module, the large model deeply fuses text and images based on technologies such as contrastive learning and self-attention mechanism to infer relevant information about paper fibers; after the model inference is completed, the RAG mechanism performs a retrieval on the key inference results, searches for knowledge or expert experience cases in the field of paper fibers in the system, and inserts the retrieved refined knowledge entries into the inference context for the large model to perform final comprehensive inference or summary output.
6. The real-time paper fiber intelligent analysis system based on large model visual extraction according to claim 1, characterized in that, It also includes a modular deployment module. Each functional module of the system is designed in a loosely coupled manner, flexibly split and combined according to the actual production line requirements, and integrated with the existing industrial control or laboratory management system through containerization or microservice deployment methods.
7. The real-time paper fiber intelligent analysis system based on large model vision extraction according to claim 1, wherein The analysis system uses a real-time paper fiber intelligent analysis method based on large model vision extraction, including the following steps: S1. Image decomposition: perform multi-scale decomposition on the input paper fiber image to generate sub-images with different resolutions and local ranges; S2. Feature extraction and screening: use a pre-trained image encoder to vectorize the features of sub-images at different scales, use a pre-trained text encoder to convert semantic descriptions related to paper fibers into text features, calculate the cosine similarity of the image-text features, and select the top-k groups of most discriminative image features according to the variance of the similarity distribution vector; S3. Feature fusion: perform weighted fusion on the selected local image features and the corresponding text features to form a multi-modal feature representation with high correlation; S4. Model inference and retrieval: input the fused features into the large model for preliminary inference to obtain the category or defect information of the paper fiber, use the RAG mechanism to perform secondary retrieval and matching on the model output and the knowledge base, merge the retrieval results and the preliminary inference results, and output the final decision or diagnostic report.
8. The real-time paper fiber intelligent analysis system based on large model vision extraction according to claim 7, characterized in that, Before the image decomposition, it also includes the step of initializing the system.
9. An intelligent real-time paper fiber analysis system based on large model visual extraction according to claim 7, characterized in that After the model inference and retrieval step, it also includes the step of adjusting the production process or further studying the paper fiber according to the final decision or diagnostic report, and feeding back the newly obtained knowledge or cases to the extensible knowledge base module for updating.
Citation Information
Cited By
Image classification method and device, equipment and medium
CN121074917A
Image classification method, apparatus, device, and medium
CN121074917B