Enterprise multi-mode image intelligent identification processing system

By adopting a three-tier architecture of generating multimodal image association maps and feature extraction modules, the dynamic association problem of multimodal image recognition technology in complex scenarios is solved, enabling accurate source tracing and automatic business suggestions, and supporting cross-scenario model reuse and hardware optimization.

CN121543789AActive Publication Date: 2026-02-17QINGDAO QIANYUAN JIUZE ELECTRONIC INFORMATION TECHNOLOGY CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511600712.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-17
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

Existing multimodal image recognition technologies cannot establish dynamic relationships in complex scenarios, resulting in high cross-stage tracing errors. The recognition results need to be manually broken down into business actions, and the models cannot be optimized with business iterations, leading to a break in the chain of recognition, decision-making, and execution.

Method used

The system employs an association mapping module to generate multimodal image association maps, corrects sample weight biases through a spatiotemporal weight scene correction algorithm, and combines a three-level architecture of the feature extraction module with business suggestion generation from the recognition and decision suggestion module, supporting cross-scenario model reuse and data security protection.

Benefits of technology

It achieves accurate association of multimodal images, reduces source tracing errors, automatically converts them into actionable business suggestions, supports model iterative optimization, and reduces manual intervention and hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543789A_ABST
    Figure CN121543789A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of multi-modal image recognition, and discloses an enterprise multi-modal image intelligent recognition processing system, which is characterized in that multiple types of data sources are accessed through an association mapping module, and image format standardization conversion is automatically completed; based on image metadata and a business process template, a double-domain system of a general label and a scene label is constructed, then the problem of non-uniform long-tail scene samples is solved by using a space-time weight scene correction algorithm, and finally a multi-modal image association map is generated through a label evolution engine. Meanwhile, the feature extraction module excavates associated features among different images based on the map to form a single-mode and associated feature double-layer system, so that design drawings, process images, quality control drawings and the like in the manufacturing industry can be accurately associated, and traceability errors are greatly reduced; the optimization module adapts to business requirements in advance in combination with LSTM time sequence prediction based on hour-level fast circulation, week-level deep circulation, iterative feature extraction parameters, a recognition model and a decision rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multimodal image recognition technology, specifically a multimodal image intelligent recognition and processing system for enterprises. Background Technology

[0002] Multimodal image recognition technology is a crucial supporting technology for enterprise intelligent transformation, widely used in manufacturing quality inspection, retail merchandise management, and financial document verification. As business complexity increases, enterprises have growing demands for its cross-process collaboration and adaptability for implementation. However, the following technical challenges remain in its application in complex scenarios: Image data in enterprise operations generally exhibits both temporal and logical correlation characteristics. For example, in manufacturing, design drawings, process images, and quality inspection images need to be correlated according to the production process timeline, while in retail, product warehousing images, shelf display images, and near-expiry scan images need to be correlated according to the product flow logic. Existing technologies can only achieve feature overlay of single-modal or static multi-source images and cannot establish dynamic correlation relationships. This leads to a high error rate when tracing cross-stage problems due to isolated analysis of image data, making it difficult to meet the business needs of enterprises for full-process traceability.

[0003] The existing system can only output identification conclusions, such as whether a part has a crack or a product is nearing its expiration date, but it cannot directly translate these conclusions into actionable steps. Enterprises need to invest a lot of manpower to break down the identification results into actionable tasks, such as manually generating repair work orders and adjusting inventory parameters. The intermediate conversion process is prone to errors in business execution due to information omissions. At the same time, the system cannot receive feedback from business execution to optimize the identification model, creating a broken chain in the identification, decision-making, execution, and optimization process. The model accuracy cannot be continuously improved with business iterations, and identification failure is likely to occur after long-term use. Summary of the Invention

[0004] The purpose of this invention is to provide an enterprise multimodal image intelligent recognition and processing system to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an enterprise multimodal image intelligent recognition and processing system, the system comprising: Association Mapping Module: Integrates with image data sources, sets general and scene labels; uses a spatiotemporal weighted scene correction algorithm to correct the weight allocation bias of long-tail scene samples; generates image association maps through a label evolution engine; Feature extraction module: mines the correlation features between different images; designs a three-level feature extraction architecture of edge, cloud, and collaborative layer, allocates edge and cloud computing power according to scene complexity; dynamically adjusts the weights of correlation features; Recognition and Decision Suggestion Module: Set up a scene recognition model library to recognize images, convert the recognition results into business suggestions, predict the effect of decision-making, and conduct manual review of high-risk decisions; Data security protection module: encrypts the transmission and storage of images and sets access permissions; retains associated tags when desensitizing sensitive information in images. Visualization module: Set up a visual dashboard to present key information throughout the entire process; Optimization module: Sets up a dual optimization mechanism of fast loop and deep loop for iterating feature extraction parameters, recognition model and decision rules, and adjusts resource priority according to the enterprise's business time period; Adaptation module: Through a standardized interface library and configuration engine, it is used for bidirectional data interaction with the enterprise's core business systems; it supports cross-scenario model reuse.

[0006] Furthermore, the association mapping module is specifically as follows: It supports access to image data sources including industrial cameras, document scanners, enterprise cloud storage, and mobile devices, automatically identifies image formats, performs standardized conversion, and outputs standardized image data. A dual-domain tagging system of general tags and scene tags is established based on image metadata and enterprise business process templates. General tags are generated into basic semantic tags through CLIP cross-modal pre-trained models, while scene tags are embedded into vertical industry knowledge graphs to achieve targeted enhancement. At the same time, a spatiotemporal weighted scene correction algorithm is adopted to correct the weight allocation bias caused by the uneven distribution of long-tail scene samples. The label evolution engine performs self-supervised learning cluster analysis on historical labels, automatically generates association rules between labels, and updates the label system monthly, ultimately generating a multimodal image association map.

[0007] Furthermore, the feature extraction module is specifically as follows: Based on the multimodal image association map and standardized image data of the association mapping module, character semantic features, geometric topological features, and temperature gradient features are extracted from multimodal images including text documents, three-dimensional product images, and infrared thermal images, respectively. At the same time, the association features between different images are mined based on the multimodal image association map, forming a two-layer feature system of single-modal features and association features. A three-level feature extraction architecture is adopted. The edge uses the MobileNet-v3 network for real-time feature extraction; the cloud uses the DeepSeekR1 generative AI model for cross-modal deep inference; the collaboration layer performs feature fusion based on the association logic of the association mapping module and outputs a unified feature vector; the cloud computing power is allocated according to the complexity of the scene, and the feature extraction is completed independently by the edge in low-complexity scenes, while cloud collaboration is initiated in high-complexity scenes; the weight of the associated features is dynamically adjusted according to the needs of industry scenarios, and a two-layer feature system is output.

[0008] Furthermore, the identification and decision suggestion module is specifically as follows: The two-layer feature system of the feature extraction module is invoked to set up a subdivided scene recognition model library including manufacturing, finance and retail industries for image recognition; Based on the recognition results, the system calls upon an industry-specific decision rule library. Through the recognition and decision mapping engine, which combines recognition results, decision rules, and business action mapping logic, the recognition results are transformed into actionable business suggestions. At the same time, a decision effect prediction model is built based on historical execution data to evaluate the business benefits of different decisions. High-risk decisions are automatically subject to manual review. The generated business suggestions directly connect to the enterprise's ERP and MES systems to trigger business actions, and are also synchronized to the data security protection module for data security processing; a standardized execution interface library is developed.

[0009] Furthermore, the data security protection module is specifically as follows: Upon receiving the recognition results and decision suggestions from the recognition and decision suggestion module, as well as the associated image data from the association mapping module, the image transmission is encrypted using the national cryptographic SM4 algorithm, image storage is encrypted using blockchain distributed storage, and image access is controlled through RBAC and ABAC dual-dimensional permission management. Sensitive information in images is desensitized in real time, and the association tags generated by the association mapping module are retained during the desensitization process, so that cross-stage analysis can still be carried out through the association map after desensitization; linkage protection of associated data is implemented, and when a certain associated image triggers permission verification, its associated images are verified synchronously. Record the time, user, and purpose of each image access and generate an audit report. The processed security data is then synchronized to the visualization module.

[0010] Furthermore, the visualization module is specifically as follows: The identification results and decision suggestions of the identification and decision suggestion module, the security image data processed by the data security protection module, and the execution status fed back by the enterprise business system are integrated and presented centrally through a visual dashboard, so that users can intuitively obtain key information of the whole process. It allows users to click on any node to trace back the original image, association map, and execution log of the recognition and decision suggestion module, forming a full-link traceable view; it analyzes business feedback text through NLP and filters invalid feedback, and supports one-click feedback data return function; It integrates the results of LSTM time series forecasting models and displays changes in business needs over the next three months on a dashboard, thereby assisting enterprises in planning ahead; it also supports filtering data by conditions.

[0011] Furthermore, the optimization module is specifically as follows: The system monitors the operating status, including image processing time, model recognition accuracy, and hardware resource utilization, and simultaneously receives business feedback data from the visualization module to automatically identify optimization directions. A dual-loop optimization mechanism is constructed. The fast loop operates on an hourly basis, adjusting the recognition threshold through reinforcement learning. The deep loop operates on a weekly basis, using simulated annealing to optimize the network structure. A multi-objective optimization function is set in conjunction with business KPIs, and the feature extraction parameters of the feature extraction module and the recognition model and decision rules of the recognition and decision suggestion module are iteratively updated periodically. The model predicts changes in business demand based on the LSTM time series forecasting model and completes model adaptation in advance; a business peak resource scheduling mechanism is introduced to adjust resource priority according to the enterprise's business time period, and the optimized parameters and rules are synchronized to the adaptation module.

[0012] Furthermore, the adaptation module is specifically as follows: Based on the decision suggestions from the identification and decision suggestion module and the optimization parameters from the optimization module, industry-specific business closed-loop templates are provided, including those for manufacturing, finance, and retail, to meet the core business process needs of different industries. It features a predefined standardized interface library and a visual configuration engine, built-in data mapping rule templates, and supports drag-and-drop configuration of interactive logic. This enables bidirectional data interaction with the enterprise's core business systems. On one hand, it pushes the identification results and decision suggestions from the identification and decision suggestion module to the business system to trigger actions. On the other hand, it retrieves the execution results from the business system and synchronizes them to the visualization module for feedback presentation. It supports cross-scenario model reuse and allows enterprises to adjust the closed-loop logic according to their own needs. At the same time, it combines the resource scheduling strategy of the optimization module to support hardware cost optimization, ultimately forming an end-to-end technology link.

[0013] The beneficial effects of this invention are as follows: 1. This invention connects to various data sources such as industrial cameras, document scanners, and mobile devices through an association mapping module, automatically completing image format standardization conversion. Based on image metadata and business process templates, it constructs a dual-domain system of general tags and scene tags, then uses a spatiotemporal weighted scene correction algorithm to solve the problem of uneven sample distribution in long-tail scenes, and finally generates a multimodal image association map through a tag evolution engine. At the same time, the feature extraction module relies on this map to mine the association features between different images, forming a dual-layer system of single-modality and association features, which can accurately associate manufacturing design drawings, process images, quality inspection images, etc., significantly reducing traceability errors and meeting the needs of enterprises for full-process traceability. For example, the accuracy of tracing the source of production line failures in manufacturing can be significantly improved.

[0014] 2. The recognition and decision suggestion module of this invention calls a two-layer feature system. After completing image recognition through a subdivided scene recognition model library, it calls an industry-specific decision rule library to generate executable business suggestions, which directly connect to ERP and MES systems to trigger actions. It also supports hardware linkage with industrial cameras, robotic arms and other devices. The visualization module can analyze business feedback text through NLP, filter invalid information and send it back with one click. The optimization module relies on hourly rapid loops and weekly deep loops to iterate feature extraction parameters, recognition models and decision rules, and combine LSTM time series prediction to adapt to business needs in advance. It eliminates the step of manually disassembling the recognition results, avoids information omissions, and allows the model to be continuously optimized with business iterations. For example, the accuracy of the near-expiry product recognition model in the retail industry can remain stable at a high level for a long time.

[0015] 3. The data security protection module of this invention uses the national cryptographic SM4 algorithm to encrypt image transmission, and uses blockchain distributed storage to encrypt images. It then uses RBAC and ABAC dual-dimensional permission management to control access. When desensitizing sensitive information, it retains the associated tags to ensure that the associated graph analysis can still be performed after desensitization. It also records image call information to generate audit reports to ensure compliance. The adaptation module provides industry-specific business closed-loop templates such as "equipment inspection-identification-repair-re-inspection" in the manufacturing industry and "goods warehousing-inventory-near-expiry identification-replenishment" in the retail industry. It has a predefined standardized interface library and a visual configuration engine, supports drag-and-drop configuration interaction logic, can bidirectionally connect to the enterprise's core business system, and can reuse models across scenarios. For example, the security model can be used for logistics after a small number of samples are fine-tuned. Combined with resource scheduling strategies, it optimizes hardware costs, reduces the access threshold for enterprises of different sizes, and shortens the system deployment cycle. Attached Figure Description

[0016] Figure 1 This is a flowchart of the enterprise's multimodal image intelligent recognition and processing system. Figure 2 This is a flowchart of the multimodal feature extraction and recognition decision-making process of the present invention; Figure 3 This is a flowchart of the dual-loop optimization process of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] like Figures 1 to 3 As shown, this embodiment of the invention provides an enterprise multimodal image intelligent recognition and processing system, which includes: Association Mapping Module: Integrates with image data sources, sets general and scene labels; uses a spatiotemporal weighted scene correction algorithm to correct the weight allocation bias of long-tail scene samples; generates image association maps through a label evolution engine; Feature extraction module: mines the correlation features between different images; designs a three-level feature extraction architecture of edge, cloud, and collaborative layer, allocates edge and cloud computing power according to scene complexity; dynamically adjusts the weights of correlation features; Recognition and Decision Suggestion Module: Set up a scene recognition model library to recognize images, convert the recognition results into business suggestions, predict the effect of decision-making, and conduct manual review of high-risk decisions; Data security protection module: encrypts the transmission and storage of images and sets access permissions; retains associated tags when desensitizing sensitive information in images. Visualization module: Set up a visual dashboard to present key information throughout the entire process; Optimization module: Sets up a dual optimization mechanism of fast loop and deep loop for iterating feature extraction parameters, recognition model and decision rules, and adjusts resource priority according to the enterprise's business time period; Adaptation module: Through a standardized interface library and configuration engine, it is used for bidirectional data interaction with the enterprise's core business systems; it supports cross-scenario model reuse.

[0019] The association mapping module is specifically as follows: The image data source access supports GigEVision industrial cameras, USB 3.0 document scanners (supporting CAD / TIFF formats), and enterprise private cloud S3 interfaces. It automatically identifies the color gamut, resolution, and modal type (2D / 3D / infrared) of the original image, standardizes and converts it into PNG format with timestamps and device IDs. The 3D model retains STL geometric features, the infrared image retains temperature metadata, and outputs standardized image data, providing a unified format foundation for subsequent data processing. A dual-domain tagging system of general tags and scene tags is established based on image metadata (filename, acquisition device, timestamp) and enterprise business process templates. General tags generate basic semantic tags through CLIP cross-modal pre-trained model, and scene tags are embedded in vertical industry knowledge graphs (smart manufacturing equipment fault graph, security abnormal behavior graph, etc.) to achieve targeted enhancement. At the same time, a spatiotemporal weighted scene correction algorithm is adopted to solve the weight allocation bias caused by uneven sample distribution in long-tail scenes. The tag evolution engine performs self-supervised learning cluster analysis on historical tags and automatically generates association rules between tags, such as associating equipment oil leakage with abnormal temperature. The tag system is updated monthly, and finally a multimodal image association map is generated, supporting multi-dimensional association logic such as production line in the manufacturing industry and business order number in the financial industry.

[0020] Spatiotemporal weighting scenario correction formula: ; In the formula: It represents the final weight of the image association mapping, which is used to construct multimodal image association maps and determine the association priority of different images; The images were determined to be strongly correlated and were given priority for inclusion in cross-process source tracing analysis; The images are marked as weakly associated and need to be validated again in conjunction with business rules to resolve the association bias caused by uneven sample distribution in long-tail scenarios. It represents the basic weight of spatiotemporal correlation, calculated based on metadata such as image acquisition timestamps and device locations, and reflects the strength of the natural temporal and spatial correlation between images; This represents the scene deviation correction coefficient, with a value ranging from 0.1 to 0.9, determined by business priority. The industry sample proportion is dynamically calculated, such as in manufacturing production line traceability scenarios. Retail Industry Review Scenarios ; It represents the scene label weight, which is embedded in the vertical industry knowledge graph generation, such as the equipment failure label weight in the smart manufacturing scene and the abnormal behavior label weight in the security scene; Represents the basic weights of general labels, generated by the CLIP cross-modal pre-trained model, reflecting the strength of general semantic associations in an image, such as the association weights of basic labels like parts and cracks.

[0021] The feature extraction module is specifically as follows: Based on the multimodal image association map and standardized image data of the association mapping module, character semantic features, geometric topological features, temperature gradients and other exclusive features are extracted from multimodal images such as text documents, 3D product images and infrared thermal images. At the same time, the association features between different images are mined by relying on the multimodal image association map, such as the spatial correspondence between abnormal welding temperature in the process and the location of quality inspection cracks, forming a two-layer feature system of single-modal features and association features. A three-level feature extraction architecture is adopted. The edge uses the MobileNet-v3 lightweight network for real-time feature extraction, with an edge processing latency of no more than 50 milliseconds. The cloud uses the DeepSeekR1 generative AI model for cross-modal deep inference. The collaboration layer performs feature fusion based on the association logic of the association mapping module and outputs a unified feature vector. Cloud computing power is allocated according to the complexity of the scenario. In low-complexity scenarios such as routine inspections, feature extraction is completed independently by the edge. In high-complexity scenarios such as root cause analysis of faults, cloud collaboration is initiated. The scenario complexity is dynamically determined based on the number of image modalities (≤2 single modalities are considered low, >2 are considered high) and the complexity of defect types (e.g., crack detection in manufacturing is considered high when it contains more than 3 types). The edge independently processes scenarios with single modalities and ≤2 defect types (e.g., routine inspection). When cross-modal correlation features are detected (e.g., spatial correspondence between process diagrams and quality inspection diagrams) or the recognition confidence is <0.7, cloud collaboration is triggered via the MQTT protocol, with a collaboration latency ≤150ms (based on a 5G network). By combining industry scenario requirements, such as increasing the weight of associated features in manufacturing traceability scenarios and decreasing the weight of associated features in financial document recognition scenarios, the weight of associated features is dynamically adjusted to output a two-layer feature system.

[0022] Formula for assigning weights to related features: ; In the formula: This represents the final weight of the cross-modal association features, used in a two-layer feature system of single-modal features and association features to determine the contribution of association features in the recognition model; This represents the weighting percentage of single-modal features, ranging from 0.3 to 0.7, and is dynamically adjusted based on industry scenarios, such as document recognition in the financial industry. Manufacturing equipment testing scenarios ; This represents the single-modal feature weight, corresponding to the exclusive feature weights of a single modality such as text, 3D image, and infrared image, such as the character semantic feature weights of a text document and the temperature gradient feature weights of an infrared image. It represents the weight of cross-image association features, calculated based on the association map of the association mapping module, reflecting the association strength between different images, such as the spatial correspondence weight between the welding process image and the quality inspection crack image.

[0023] The identification and decision suggestion module is specifically as follows: The two-layer feature system of the feature extraction module is invoked to set up a recognition model library for subdivided scenarios such as manufacturing, finance and retail. Enterprises can customize training models through a visual interface to complete image recognition tasks, and only fifty to one hundred industry samples are required. Based on the recognition results, the industry-specific decision rule library is invoked. Through the recognition and decision mapping engine that combines recognition results, decision rules, and business action mapping logic, the decision suggestion generation time is less than 100 milliseconds. For example, when the crack length of a part is identified as 2 mm, the shutdown and maintenance of welding equipment on production line 3 is automatically output, and maintenance work order number W2024052001 is generated, which transforms the recognition results into executable business suggestions. The identification and decision mapping engine has built-in industry decision rules: In manufacturing scenarios, when a crack length of >1mm is detected in a part and it is located in a load-bearing structure (confidence level ≥0.85), an MES work order containing the equipment number and maintenance level (Level I: immediate shutdown) is automatically generated; if the decision involves production line shutdown for more than 2 hours or estimated loss of >100,000 yuan, a manual review process is triggered, and after the review is passed, it is pushed to the ERP system via RESTful API.

[0024] At the same time, a decision-making effect prediction model is established based on historical execution data to evaluate the business benefits of different decisions, such as the impact of maintenance plans on production efficiency. High-risk decisions (involving production stoppages) will automatically trigger manual review. The generated business suggestions directly connect to the enterprise's ERP and MES systems to trigger business actions, and are also synchronized to the data security protection module for data security processing; a standardized execution interface library is developed based on the standardized interface library of the adaptation module.

[0025] Formula for verifying the confidence level of decision recommendations: ; In the formula: This represents the overall confidence level of the decision recommendation, used to determine whether manual review is triggered. Its value ranges from 0 to 1. Execute automatically at any time. Manual review is triggered at any time; This represents the confidence level of the image recognition result. It is the original confidence level output by the recognition model of the recognition and decision suggestion module. For example, the confidence level of recognizing a crack length of 2mm on a part is 0.92. This represents the weight of the recognition result, with a value ranging from 0.4 to 0.6, determined by the complexity of the recognition task, such as complex defect recognition. Simple target recognition ; This represents the confidence level of the decision rule matching, which is the degree of matching between the identification result and the industry decision rule library. For example, the confidence level of matching between the crack length of 2mm and the shutdown and maintenance rule is 0.88. This represents the weight of the decision rule, with a value ranging from 0.4 to 0.6. When the sum is 1, the rule complexity is high. When the rules are simple .

[0026] The data security protection module is specifically as follows: Upon receiving the identification results and decision suggestions from the identification and decision suggestion module, as well as the associated image data from the association mapping module, the image transmission is encrypted using the national cryptographic SM4 algorithm, and the image storage is encrypted using blockchain distributed storage. Image access is controlled through dual-dimensional permission management using RBAC and ABAC, allowing only quality inspectors to view defective images and preventing them from downloading them. Sensitive information in images is anonymized in real time, while retaining associated tags during the anonymization process. Based on tests in two core scenarios (1000 sets of associated image samples each) of manufacturing production line traceability and financial document verification, the accuracy rate of association analysis before anonymization was 98.5%, and the accuracy rate after anonymization was ≥97.5%, with a decrease of ≤1%, ensuring that cross-process analysis can still be performed through the association map after anonymization. Linked data protection is implemented. When an associated image triggers permission verification, its associated images are verified simultaneously to prevent attackers from deducing highly sensitive data through low-privilege images. Record the time, user, and purpose of each image call and generate an audit report. The processed security data is synchronized to the visualization module for visualization presentation, while ensuring that the entire data flow meets compliance audit requirements.

[0027] The visualization module is specifically as follows: The system integrates the identification results and decision suggestions from the identification and decision suggestion module, the security image data processed by the data security protection module, and the execution status feedback from the enterprise business system, such as the dispatch of maintenance work orders and the completion of equipment maintenance. This information is then centrally presented through a visual dashboard, allowing users to intuitively obtain key information throughout the entire process. It allows users to click on any node (such as decision suggestions) to trace back the original image, association map, and execution log of the recognition and decision suggestion module, forming a full-link traceable view; it analyzes business feedback text through NLP and filters invalid feedback, supports one-click feedback data return function, and automatically synchronizes feedback data to the optimization module for adaptive optimization after the user confirms the execution result and the defect has been fixed. This module integrates the results of the LSTM time-series forecasting model (trained based on the company's business needs data over the past two years, with a prediction accuracy of ≥92%) developed in this module, displaying changes in business needs over the next three months on the dashboard to help companies plan ahead; the model parameters are synchronized to the optimization module for reuse; it supports filtering data based on conditions such as confidence level not less than 90% and specific business processes, and the export format is compatible with common types such as Excel and PDF, and supports the standard import interface of the enterprise reporting system to achieve seamless data integration with the enterprise reporting system.

[0028] The optimization module is specifically as follows: The system monitors the image processing time, model recognition accuracy, and hardware resource utilization in real time, and receives business feedback data from the visualization module. It automatically identifies optimization directions, such as updating the decision rule base of the recognition and decision suggestion module when the decision suggestion is invalid, and adjusting the feature weights of the feature extraction module when the defect is still identified after repair. A dual-loop optimization mechanism is constructed. The fast loop is based on the real-time monitoring of the false negative rate (>5%) or false positive rate (>3%), and the recognition threshold is adjusted in 0.05 steps through reinforcement learning (e.g., the crack detection threshold is reduced from 0.85 to 0.80). The deep loop uses simulated annealing algorithm to optimize the number of bottleneck layer channels of MobileNet-v3 at the edge (adjustment range ±20%). A multi-objective function is constructed by combining the downtime of the manufacturing production line and the recognition accuracy (target ≥95%). The feature extraction parameters of the feature extraction module and the recognition model and decision rules of the recognition and decision suggestion module are iteratively updated regularly.

[0029] The results of the LSTM time-series prediction model developed by the visualization module are reused to predict changes in business demand (such as the increase in demand for package recognition before the peak season in the retail industry), and the model adaptation is completed in advance to avoid system performance fluctuations during business peaks. A business peak resource scheduling mechanism is introduced to dynamically allocate edge / cloud computing power according to the enterprise's business time period (such as the morning quality inspection peak in the manufacturing industry and the evening inventory low point in the retail industry). During the low point period, the cloud computing power occupation is reduced and redundant edge devices are shut down to reduce hardware energy consumption and rental costs. This strategy is synchronized to the adaptation module for hardware cost optimization.

[0030] Fast loop recognition threshold adjustment formula: ; In the formula: This represents the adjusted image recognition threshold, used to control the accuracy of the recognition results, such as the confidence threshold for defect recognition and the IOU threshold for target detection. This indicates the original recognition threshold before adjustment, which is either the initial system configuration or the threshold after the last optimization. The default value is preset according to the industry scenario, such as the default value for defect recognition in manufacturing. ; This indicates the threshold adjustment step size, ranging from 0.01 to 0.05, determined by business sensitivity. High-risk scenarios such as production shutdowns and maintenance are also considered. Low-risk scenarios such as inventory checks ; This represents the target identification indicator, which is the business target value preset by the enterprise, such as a false negative rate of ≤5% and an identification accuracy rate of ≥95%. This indicates the current identification metric, which is the actual value monitored by the system in real time, such as the current false negative rate of 6.2% and the current accuracy rate of 92.3%.

[0031] The adaptation module is specifically as follows: Based on the decision suggestions from the identification and decision suggestion module and the optimization parameters from the optimization module, we provide industry-specific business closed-loop templates for manufacturing ("equipment inspection-identification-repair-re-inspection") and retail ("goods warehousing-inventory-near-expiry identification-replenishment"), adapting to the core business process needs of different industries. Based on a predefined standardized interface library and a visual configuration engine, a dedicated hardware linkage protocol is developed, supporting plug-and-play functionality for twelve types of enterprise hardware, including industrial cameras, robotic arms, and alarm devices. This enables direct linkage between recognition results and hardware actions, avoiding functional duplication across modules. Built-in data mapping rule templates support drag-and-drop configuration of interactive logic, enabling bidirectional data interaction with core enterprise business systems (MES, ERP, OA). On one hand, it pushes the recognition results and decision suggestions from the recognition and decision suggestion module to the business system to trigger actions; on the other hand, it retrieves the execution results from the business system and synchronizes them to the visualization module for feedback presentation. It supports cross-scenario model reuse. When a model trained in a security scenario is migrated to a logistics scenario, only fifty to one hundred logistics scenario samples are needed for fine-tuning to meet the business accuracy requirements (≥95%), improving model reuse efficiency. It also supports enterprises to adjust the closed-loop logic according to their own needs, such as adding a rework step for unqualified products, to adapt to the business characteristics of different industries and enterprises of different sizes. At the same time, combined with the time-based computing power scheduling and redundant equipment management strategy of the optimization module, it can optimize hardware energy consumption and leasing costs, ultimately forming an end-to-end technology link.

[0032] The standardized interface library includes RESTful API (supports JSON format, response latency ≤200ms), ModbusRTU (for industrial camera triggering) and OPCUA (for MES system integration), with built-in data mapping templates (such as defect coordinates → robotic arm positioning instructions). It supports drag-and-drop configuration to achieve second-level linkage between recognition results (such as 'abnormality at workstation A-3 on the production line') and hardware actions (such as the robotic arm marking the defect location). It is compatible with more than 90% of mainstream industrial protocols (based on the IEEE 2030.5 standard).

[0033] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0034] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multimodal image intelligent recognition and processing system for enterprises, characterized in that: The system includes: Association Mapping Module: Integrates with image data sources, sets general and scene labels; uses a spatiotemporal weighted scene correction algorithm to correct the weight allocation bias of long-tail scene samples; generates image association maps through a label evolution engine; Feature extraction module: mines the correlation features between different images; designs a three-level feature extraction architecture of edge, cloud, and collaborative layer, allocates edge and cloud computing power according to scene complexity; dynamically adjusts the weights of correlation features; Recognition and Decision Suggestion Module: Set up a scene recognition model library to recognize images, convert the recognition results into business suggestions, predict the effect of decision-making, and conduct manual review of high-risk decisions; Data security protection module: encrypts the transmission and storage of images and sets access permissions; retains associated tags when desensitizing sensitive information in images. Visualization module: Set up a visual dashboard to present key information throughout the entire process; Optimization module: Sets up a dual optimization mechanism of fast loop and deep loop for iterating feature extraction parameters, recognition model and decision rules, and adjusts resource priority according to the enterprise's business time period; Adaptation module: Through a standardized interface library and configuration engine, it is used for bidirectional data interaction with the enterprise's core business systems; it supports cross-scenario model reuse.

2. The enterprise multimodal image intelligent recognition and processing system according to claim 1, characterized in that: The specific details of the association mapping module are as follows: It supports access to image data sources including industrial cameras, document scanners, enterprise cloud storage, and mobile devices, automatically identifies image formats, performs standardized conversion, and outputs standardized image data. A dual-domain tagging system of general tags and scene tags is established based on image metadata and enterprise business process templates. General tags are generated into basic semantic tags through CLIP cross-modal pre-trained models, while scene tags are embedded into vertical industry knowledge graphs to achieve targeted enhancement. At the same time, a spatiotemporal weighted scene correction algorithm is adopted to correct the weight allocation bias caused by the uneven distribution of long-tail scene samples. The label evolution engine performs self-supervised learning cluster analysis on historical labels, automatically generates association rules between labels, and updates the label system monthly, ultimately generating a multimodal image association map.

3. The enterprise multimodal image intelligent recognition and processing system according to claim 2, characterized in that: The feature extraction module is as follows: Based on the multimodal image association map and standardized image data of the association mapping module, character semantic features, geometric topological features, and temperature gradient features are extracted from multimodal images including text documents, three-dimensional product images, and infrared thermal images, respectively. At the same time, the association features between different images are mined based on the multimodal image association map, forming a two-layer feature system of single-modal features and association features. A three-level feature extraction architecture is adopted, with MobileNet-v3 network used at the edge for real-time feature extraction; The cloud-based DeepSeekR1 generative AI model is used for cross-modal deep inference; the collaboration layer performs feature fusion based on the association logic of the association mapping module and outputs a unified feature vector. Cloud computing power is allocated based on the complexity of the scenario. In low-complexity scenarios, feature extraction is completed independently by the edge, while in high-complexity scenarios, cloud collaboration is initiated. The weights of associated features are dynamically adjusted in combination with industry scenario requirements to output a two-layer feature system.

4. The enterprise multimodal image intelligent recognition and processing system according to claim 3, characterized in that: The identification and decision suggestion module is as follows: The two-layer feature system of the feature extraction module is invoked to set up a subdivided scene recognition model library including manufacturing, finance and retail industries for image recognition; Based on the recognition results, the system calls upon an industry-specific decision rule library. Through the recognition and decision mapping engine, which combines recognition results, decision rules, and business action mapping logic, the recognition results are transformed into actionable business suggestions. At the same time, a decision effect prediction model is built based on historical execution data to evaluate the business benefits of different decisions. High-risk decisions are automatically subject to manual review. The generated business suggestions directly connect to the enterprise's ERP and MES systems to trigger business actions, and are also synchronized to the data security protection module for data security processing. Develop a standardized execution interface library.

5. The enterprise multimodal image intelligent recognition and processing system according to claim 4, characterized in that: The data security protection module is specifically as follows: Upon receiving the recognition results and decision suggestions from the recognition and decision suggestion module, as well as the associated image data from the association mapping module, the image transmission is encrypted using the national cryptographic SM4 algorithm, image storage is encrypted using blockchain distributed storage, and image access is controlled through RBAC and ABAC dual-dimensional permission management. Sensitive information in images is desensitized in real time, and the association tags generated by the association mapping module are retained during the desensitization process, so that cross-stage analysis can still be carried out through the association map after desensitization; linkage protection of associated data is implemented, and when a certain associated image triggers permission verification, its associated images are verified synchronously. Record the time, user, and purpose of each image access and generate an audit report. The processed security data is then synchronized to the visualization module.

6. The enterprise multimodal image intelligent recognition and processing system according to claim 5, characterized in that: The visualization module is as follows: The identification results and decision suggestions of the identification and decision suggestion module, the security image data processed by the data security protection module, and the execution status fed back by the enterprise business system are integrated and presented centrally through a visual dashboard, so that users can intuitively obtain key information of the whole process. It allows users to click on any node to trace back the original image, association map, and execution log of the recognition and decision suggestion module, forming a full-link traceable view; it analyzes business feedback text through NLP and filters invalid feedback, and supports one-click feedback data return function; Integrate the results of the LSTM time series prediction model and display the changes in business demand over the next three months on the dashboard to help enterprises plan ahead; Supports filtering data based on conditions.

7. The enterprise multimodal image intelligent recognition and processing system according to claim 6, characterized in that: The optimization module is as follows: The system monitors the operating status, including image processing time, model recognition accuracy, and hardware resource utilization, and simultaneously receives business feedback data from the visualization module to automatically identify optimization directions. A dual-loop optimization mechanism is constructed. The fast loop operates on an hourly basis, adjusting the recognition threshold through reinforcement learning. The deep loop operates on a weekly basis, using simulated annealing to optimize the network structure. A multi-objective optimization function is set in conjunction with business KPIs, and the feature extraction parameters of the feature extraction module and the recognition model and decision rules of the recognition and decision suggestion module are iteratively updated periodically. The model predicts changes in business demand based on the LSTM time series forecasting model and completes model adaptation in advance; a business peak resource scheduling mechanism is introduced to adjust resource priority according to the enterprise's business time period, and the optimized parameters and rules are synchronized to the adaptation module.

8. The enterprise multimodal image intelligent recognition and processing system according to claim 7, characterized in that: The specific adaptation module is as follows: Based on the decision suggestions from the identification and decision suggestion module and the optimization parameters from the optimization module, industry-specific business closed-loop templates are provided, including those for manufacturing, finance, and retail, to meet the core business process needs of different industries. It features a predefined standardized interface library and a visual configuration engine, built-in data mapping rule templates, and supports drag-and-drop configuration of interactive logic. This enables bidirectional data interaction with the enterprise's core business systems. On one hand, it pushes the identification results and decision suggestions from the identification and decision suggestion module to the business system to trigger actions. On the other hand, it retrieves the execution results from the business system and synchronizes them to the visualization module for feedback presentation. It supports cross-scenario model reuse and allows enterprises to adjust the closed-loop logic according to their own needs. At the same time, it combines the resource scheduling strategy of the optimization module to support hardware cost optimization, ultimately forming an end-to-end technology link.

Citation Information

Patent Citations

  • Enterprise portrait label intelligent generation method

    CN119168504A

  • Non-performing asset cross-scene question and answer framework based on knowledge graph

    CN120216706A

  • Intelligent office system and method based on multi-modal large model

    CN120450631A

  • Intelligent identification and management system and method for key commodities of vegetable basket

    CN120672188A

  • Industrial graph construction method and system based on scene type marketing

    CN120687621A