AI-based intelligent identification method and system for multimodal anomalies in document printers

By collecting multimodal data through a sensor array and combining it with an AI method of dynamic weight allocation, the problem of multi-dimensional coverage and weight adaptation in printer anomaly detection is solved, achieving high-precision anomaly identification and early warning, and ensuring stable printer operation.

CN121580176BActive Publication Date: 2026-05-05BEIJING CGPRINTECH TECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CGPRINTECH TECHNOLOGY CO LTD
Filing Date
2026-01-27
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing printer anomaly detection technologies suffer from problems such as limited detection dimensions, lack of dynamic adaptability in weight allocation, and insufficient accuracy in time-series prediction, making it difficult to meet the requirements for accuracy, real-time performance, and proactive early warning in complex scenarios.

Method used

An AI-based intelligent identification method for multimodal anomalies in document printers is adopted. Multimodal data is collected by a sensor group combined with the internal temperature of the printer. The data is preprocessed and features are extracted. Multimodal feature vectors are generated using a dynamic weight allocation method and then input into an improved LSTM model for anomaly identification and early warning.

Benefits of technology

It enables multi-dimensional and targeted multimodal data acquisition, improving data quality and the accuracy of anomaly identification. It can accurately identify early hidden anomalies, avoid missed or false judgments, provide sufficient time for fault handling, and prevent printing interruptions and equipment damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580176B_ABST
    Figure CN121580176B_ABST
Patent Text Reader

Abstract

This invention proposes an AI-based intelligent multimodal anomaly identification method and system for document printers, belonging to the field of printer anomaly identification technology. The intelligent multimodal anomaly identification method for document printers includes: using a sensor array combined with a printer internal temperature comprehensive acquisition strategy to collect multimodal data information of the printer in real time; preprocessing the multimodal data information to obtain preprocessed multimodal data information; extracting features from the multimodal data information; dynamically adjusting the weight values ​​corresponding to the multimodal data information using a dynamic weight allocation method combined with the current anomaly tendency to obtain the weight value corresponding to each multimodal data information, and generating a multimodal feature vector; inputting the multimodal feature vector into an improved LSTM model to identify printer anomalies and obtain anomaly identification results; and issuing anomaly warnings based on the anomaly identification results. The system includes modules corresponding to the steps of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes an AI-based intelligent identification method and system for multimodal anomalies in document printers, belonging to the field of printer anomaly identification technology. Background Technology

[0002] In scenarios such as office automation, home office, and industrial production support, document printers serve as core information output devices, and their operational stability directly impacts work efficiency and business progress. With the increasing complexity of printer functions (such as integrated copying, scanning, and wireless transmission) and the significant increase in usage frequency, printers are prone to various abnormal malfunctions, such as blurry prints due to printhead clogging, shutdowns caused by internal circuit overheating, paper waste due to paper feed mechanism jamming, and false status alarms caused by sensor malfunctions. If these anomalies are not identified and warned of in a timely manner, they can lead to minor issues like print job interruptions and increased equipment maintenance costs, or more serious problems such as overheating causing safety hazards or indirect economic losses due to delays in printing critical business documents.

[0003] Existing printer anomaly detection technologies generally suffer from problems such as limited detection dimensions, lack of dynamic adaptability in weight allocation, and insufficient accuracy in time-series prediction. These shortcomings make it difficult to meet the printer's requirements for accuracy, real-time performance, and proactive early warning in complex scenarios. Therefore, there is an urgent need for an intelligent identification method that can integrate multimodal data, dynamically adjust data weights, and achieve high-precision anomaly prediction based on an improved model. This method aims to overcome the deficiencies of existing technologies and ensure the stable and efficient operation of printers. Summary of the Invention

[0004] This invention provides an AI-based intelligent identification method and system for multimodal anomalies in document printers, to solve the technical problems existing in the prior art. The technical solution adopted is as follows:

[0005] An AI-based intelligent identification method for multimodal anomalies in document printers, comprising:

[0006] The printer uses a sensor array combined with a comprehensive temperature acquisition strategy to collect multimodal data information of the printer in real time, and preprocesses the multimodal data information to obtain preprocessed multimodal data information.

[0007] Feature extraction is performed on the multimodal data information, and the weight values ​​corresponding to the multimodal data information are dynamically adjusted by using a dynamic weight allocation method combined with the current abnormal tendency. The weight value corresponding to each multimodal data information is obtained, and a multimodal feature vector is generated.

[0008] The multimodal feature vectors are input into the improved LSTM model to perform anomaly identification on the printer, obtain the anomaly identification results, and issue anomaly warnings based on the anomaly identification results.

[0009] Furthermore, a sensor array combined with a printer internal temperature comprehensive acquisition strategy is used to collect multimodal data information of the printer in real time, and the multimodal data information is preprocessed to obtain preprocessed multimodal data information, including:

[0010] A high-definition industrial camera is installed at the printer's paper output port to capture real-time image data of the printed parts corresponding to the printed documents.

[0011] Vibration sensors are used to collect the vibration amplitude of the printer mechanism in real time during operation.

[0012] The printer uses an internal temperature sensor array combined with a comprehensive internal temperature acquisition strategy to collect the printer's internal temperature in real time during operation.

[0013] The printer uses an internal pressure sensor to monitor the toner cartridge level in real time.

[0014] The temperature and humidity sensors installed on the outside of the printer are used to collect real-time temperature and humidity data of the printer's environment.

[0015] The image data of the printed parts, vibration amplitude, overall temperature inside the printer, toner cartridge balance, and temperature and humidity data of the printer's environment are subjected to noise reduction processing to obtain pre-processed multimodal data information.

[0016] Furthermore, the comprehensive temperature acquisition strategy for the printer's internal environment includes:

[0017] The temperature sensors included in the internal temperature sensor group are used to collect real-time temperature data of the inner wall of the middle section of the paper feed channel and the outlet temperature data of the cooling fan.

[0018] Based on the temperature values ​​of the inner wall temperature and the cooling fan outlet temperature collected at each temperature sampling point, the temperature standard deviation σ of the inner wall temperature and the cooling fan outlet temperature for each sampling point is obtained. x ;

[0019] The weight values ​​w corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data are set using the temperature standard deviation of the middle section inner wall temperature data and the cooling fan outlet temperature data corresponding to each temperature acquisition.

[0020] The printer's internal comprehensive temperature during the current printer operation is obtained by using the temperature values ​​corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data, combined with their corresponding weight values, to perform weighted average processing.

[0021] Furthermore, feature extraction is performed on the multimodal data information, and the weight values ​​corresponding to the multimodal data information are dynamically adjusted using a dynamic weight allocation method combined with the current anomaly tendency to obtain the weight value corresponding to each multimodal data information, including:

[0022] The multimodal data information is classified and features are extracted to obtain the feature vector corresponding to each data type contained in the multimodal data information;

[0023] The feature vectors corresponding to each data type contained in the multimodal data information are normalized to obtain the normalized feature vectors corresponding to each data type.

[0024] The current anomaly tendency for each data type during printer operation is determined by using the normalized feature vectors corresponding to each data type.

[0025] Based on the normalized feature vector corresponding to each data type and the current anomaly tendency, the weight values ​​corresponding to the multimodal data information are dynamically adjusted to obtain the weight value corresponding to each multimodal data information.

[0026] Further, the multimodal data information is classified and features are extracted to obtain feature vectors corresponding to each data type contained in the multimodal data information, including:

[0027] The preprocessed multimodal data information is retrieved and classified to obtain various data types corresponding to the multimodal data information. These data types include printed image data, printer operating parameters, and printer environmental parameters. The printer operating parameters include vibration amplitude, overall internal temperature of the printer, and toner cartridge balance. The printer environmental parameters include temperature and humidity data of the printer's environment.

[0028] An improved ResNet-18 model is used to perform image recognition on printed image data. The average proportions of blurred, missing, and misaligned areas are obtained for all printed image data. Then, a feature vector S=[S] is generated based on these average proportions. 01 S 02 S 03 ]; where S 01 S 02 and S 03 These represent the average area percentage of blurred areas, the average area percentage of missing areas, and the average area percentage of misaligned areas, respectively.

[0029] The current vibration amplitude, printer internal temperature and toner cartridge balance are retrieved, and the vibration amplitude, printer internal temperature and toner cartridge balance are normalized to obtain the normalized current vibration amplitude, printer internal temperature and toner cartridge balance.

[0030] The feature vector Y=[Y] is generated by integrating the current vibration amplitude, the overall internal temperature of the printer, and the remaining toner in the toner cartridge to produce the printer's operating parameters. 01 Y 02 Y 03 ], where Y 01 Y 02 and Y 03 These represent the normalized vibration amplitude, the printer's internal temperature, and the remaining toner in the toner cartridge, respectively.

[0031] Retrieve the temperature and humidity data of the current printer environment, and normalize the temperature and humidity data of the printer environment to obtain the normalized temperature and humidity values ​​of the current printer environment.

[0032] By integrating the normalized temperature and humidity values ​​of the current printer environment, a feature vector R=[R] is generated corresponding to the environmental parameters. 01 R 02 ], where R 01 and R 02 These represent the normalized temperature and humidity values, respectively.

[0033] Furthermore, the structure of the improved ResNet-18 model is as follows:

[0034] The input layer and preprocessing layer include an image input sub-layer, a normalization sub-layer, and a printable effective area mask generation sub-layer;

[0035] The initial convolutional layer is used to extract low-level texture features of the image using a 3×3 convolutional kernel; the BN layer stabilizes the feature distribution, and ReLU activation enhances the effective feature response;

[0036] The dual-branch parallel processing layer includes an original feature extraction branch and a region area statistics branch; wherein, the original feature extraction branch is used to divide the printed image data into feature maps and obtain multiple feature maps corresponding to the printed image data;

[0037] The regional area statistics branch includes a sub-branch for fuzzy region positioning and area ratio calculation, a sub-branch for missing printing region positioning and area ratio calculation, a sub-branch for misaligned region positioning and area ratio calculation, and a global statistics module;

[0038] The sub-branch for locating and calculating the area ratio of blurred regions is used to identify low-contrast blurred regions and calculate the proportion of blurred regions in a single image to the effective printing area. The sub-branch for locating and calculating the area ratio of missing areas is used to identify missing areas in dark areas without toner and calculate the proportion of missing areas in a single image. The sub-branch for locating and calculating the area ratio of misaligned regions is used to identify misaligned regions with offset edges of printed content and their corresponding proportions. The global statistics module is used to calculate and obtain the average area ratio of blurred regions, the average area ratio of missing areas, and the average area ratio of misaligned areas.

[0039] Furthermore, the current anomaly tendency for each data type during printer operation is determined using the normalized feature vector corresponding to each data type, including:

[0040] Retrieve the feature vector corresponding to each data type of the multimodal data collected each time;

[0041] The L1 norm of the feature vector corresponding to each data type is obtained by using the feature vector corresponding to each data type of multimodal data collected each time.

[0042] Based on the L1 norm corresponding to the feature vector of each data type, obtain the standard deviation of the L1 norm for each data type;

[0043] The L2 norm of each data type's feature vector is obtained by using the feature vector corresponding to each data type collected in each multimodal data collection.

[0044] The abnormal tendency determination coefficient K for each data type is obtained based on the L1 norm standard deviation and L2 norm corresponding to the feature vector of each data type.

[0045] Furthermore, based on the normalized feature vector corresponding to each data type and the current anomaly tendency, the weight values ​​corresponding to the multimodal data information are dynamically adjusted to obtain the weight values ​​corresponding to each multimodal data information, including:

[0046] Retrieve the abnormal tendency judgment coefficient K for each data type and the corresponding initial weight value for each data type;

[0047] The initial weight values ​​for each data type are adjusted using the abnormal tendency judgment coefficient K corresponding to each data type, and the adjusted weight values ​​for each data type are obtained.

[0048] The dimension is unified by using the feature vectors corresponding to each data type contained in the multimodal data information to obtain the dimension-unified feature vectors corresponding to each data type.

[0049] The feature vectors corresponding to each data type with unified dimensions are fused with their corresponding dynamic weight values ​​to generate a fused multimodal feature vector.

[0050] Furthermore, the multimodal feature vector is input into the improved LSTM model to perform anomaly identification on the printer, and anomaly identification results are obtained. Based on the anomaly identification results, anomaly warnings are issued, including:

[0051] The modal feature vectors are input into the improved LSTM model;

[0052] The probability of an anomaly occurrence for each anomaly is obtained by using the output of the improved LSTM model.

[0053] Compare the probability of occurrence of each abnormal situation with its corresponding probability threshold;

[0054] An anomaly warning is issued when the probability of the anomaly occurring reaches or exceeds its corresponding probability threshold.

[0055] An AI-based intelligent multimodal anomaly recognition system for document printers, comprising:

[0056] The data acquisition module is used to collect multimodal data information of the printer in real time by using a sensor group combined with the printer's internal temperature comprehensive acquisition strategy, and to preprocess the multimodal data information to obtain preprocessed multimodal data information.

[0057] The dynamic weight adjustment module is used to extract features from the multimodal data information, dynamically adjust the weight values ​​corresponding to the multimodal data information using a dynamic weight allocation method combined with the current abnormal tendency, obtain the weight value corresponding to each multimodal data information, and generate a multimodal feature vector.

[0058] The anomaly warning module is used to input the multimodal feature vector into the improved LSTM model to perform anomaly recognition on the printer, obtain the anomaly recognition result, and issue anomaly warning based on the anomaly recognition result.

[0059] Beneficial effects of this invention:

[0060] This invention proposes an AI-based multimodal anomaly intelligent identification method and system for document printers. By combining a sensor array with a comprehensive temperature acquisition strategy, it achieves multi-dimensional and targeted multimodal data acquisition, covering key printer operating status information and avoiding the dimensional limitations of single-modal data. Simultaneously, the preprocessing stage effectively eliminates data noise and format deviations, significantly improving data quality and consequently enhancing the reliability and accuracy of subsequent feature extraction and model prediction. This effectively solves the detection bias problems caused by incomplete data coverage and poor data quality in existing technologies. Furthermore, by dynamically allocating weights based on current anomaly trends, the weights of the multimodal data are adapted to the printer's real-time operating status, allowing data more indicative of anomaly identification to contribute more significantly and avoiding the shortcomings of fixed weights that cannot adapt to different anomaly scenarios. Moreover, the multimodal feature vectors generated based on the dynamic weight allocation method can more accurately and comprehensively characterize the printer's operating status and anomaly correlation features, effectively improving the ability of multimodal feature vectors to represent anomaly states. Finally, the model's anomaly prediction results have higher accuracy and reliability, accurately identifying early latent anomalies and avoiding missed or false positives, effectively maximizing the accuracy and stability of anomaly prediction. Based on accurate anomaly prediction results, early warnings can be triggered at the initial stage of anomalies, allowing maintenance personnel sufficient time to handle faults and avoiding printing interruptions and equipment damage caused by the escalation of anomalies. Attached Figure Description

[0061] Figure 1 This is a flowchart of the method described in this invention;

[0062] Figure 2 This is a system block diagram of the system described in this invention. Detailed Implementation

[0063] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0064] This embodiment proposes an AI-based intelligent identification method for multimodal anomalies in document printers, such as... Figure 1 As shown, the intelligent identification method for multimodal anomalies in document printers includes:

[0065] The printer uses a sensor array combined with a comprehensive temperature acquisition strategy to collect multimodal data information of the printer in real time, and preprocesses the multimodal data information to obtain preprocessed multimodal data information.

[0066] Feature extraction is performed on the multimodal data information, and the weight values ​​corresponding to the multimodal data information are dynamically adjusted by using a dynamic weight allocation method combined with the current abnormal tendency. The weight value corresponding to each multimodal data information is obtained, and a multimodal feature vector is generated.

[0067] The multimodal feature vectors are input into the improved LSTM model to perform anomaly identification on the printer, obtain the anomaly identification results, and issue anomaly warnings based on the anomaly identification results.

[0068] The working principle of the above technical solution is as follows: A sensor array, combined with a comprehensive internal temperature acquisition strategy for the printer, collects multimodal data information in real time during printer operation. The collected multimodal data is preprocessed, including but not limited to noise reduction, data standardization, and missing value completion, to eliminate data interference and format differences, resulting in high-quality preprocessed multimodal data. Feature extraction is performed on the preprocessed multimodal data to uncover key features related to abnormal states in each modality. A dynamic weight allocation method is adopted, adjusting the weight values ​​of each modality in real time based on the abnormal tendencies exhibited by the printer's current operating state, highlighting the contribution of data with stronger indicative power of the current abnormal tendencies. Based on the extracted features and dynamically adjusted weight values, a multimodal feature vector that accurately reflects the printer's operating state is generated. This multimodal feature vector is input into an improved LSTM model. Through temporal correlation analysis of the multimodal features, the model predicts future abnormalities in the printer's operating state and outputs the prediction results. Based on the prediction results, a corresponding early warning mechanism is triggered to complete the warning output.

[0069] The above technical solution achieves the following effects: By combining sensor arrays with a comprehensive temperature acquisition strategy, multi-dimensional and targeted multi-modal data acquisition is realized, covering key printer operating status information and avoiding the dimensional limitations of single-modal data. Simultaneously, the preprocessing stage effectively eliminates data noise and format deviations, significantly improving data quality and consequently enhancing the reliability and accuracy of subsequent feature extraction and model prediction. This effectively solves the detection bias problems caused by incomplete data coverage and poor data quality in existing technologies. Furthermore, by dynamically allocating weights based on current anomaly trends, the weights of multi-modal data are adapted to the real-time operating status of the printer, allowing data more indicative of anomaly identification to contribute more significantly and avoiding the shortcomings of fixed weights that cannot adapt to different anomaly scenarios. Moreover, the multi-modal feature vectors generated based on the dynamic weight allocation method can more accurately and comprehensively characterize the printer's operating status and anomaly correlation features, effectively improving the ability of multi-modal feature vectors to represent anomaly states. Finally, the model's anomaly prediction results have higher accuracy and reliability, accurately identifying early latent anomalies, avoiding missed and false positives, and maximizing the accuracy and stability of anomaly prediction. Based on accurate anomaly prediction results, early warnings can be triggered at the initial stage of anomalies, allowing maintenance personnel sufficient time to handle faults and avoiding printing interruptions and equipment damage caused by the escalation of anomalies.

[0070] In one embodiment of the present invention, a sensor array combined with a printer internal temperature comprehensive acquisition strategy is used to collect multimodal data information of the printer in real time, and the multimodal data information is preprocessed to obtain preprocessed multimodal data information, including:

[0071] A high-definition industrial camera is installed at the printer's paper output port to capture the image data of the printed document in real time. The high-definition industrial camera uses 2 megapixels and a frame rate of 15fps to capture the image data of the printed document. The image of the printed document captured by the high-definition industrial camera covers the text and graphic areas, with a resolution of 1920×1080 and a JPEG format.

[0072] The vibration amplitude during the operation of the printer mechanism is collected in real time using a vibration sensor; wherein the sampling frequency of the vibration sensor is 10Hz.

[0073] The printer uses an internal temperature sensor group combined with a comprehensive internal temperature acquisition strategy to collect the printer's internal temperature in real time during operation; wherein the internal temperature sensor group has a sampling frequency of 10Hz.

[0074] The printer uses an internal pressure sensor to monitor the toner cartridge level in real time; the pressure sensor has a sampling frequency of 1Hz.

[0075] The temperature and humidity sensors installed on the outside of the printer are used to collect real-time temperature and humidity data of the printer's environment.

[0076] The printed image data, vibration amplitude, printer internal temperature, toner cartridge balance, and ambient temperature and humidity data are subjected to noise reduction processing to obtain preprocessed multimodal data information. The printed image data, vibration amplitude, printer internal temperature, toner cartridge balance, and ambient temperature and humidity data constitute the multimodal data information.

[0077] The working principle of the above technical solution is as follows: Data acquisition is achieved through a sensor group deployed at specific locations combined with parameter configuration. A 2-megapixel, 15fps high-definition industrial camera is installed at the printer's paper output outlet to acquire complete image data of the printed parts covering text and graphics areas at a resolution of 1920×1080 in JPEG format. A vibration sensor with a sampling frequency of 10Hz is deployed at the printer core to collect the vibration amplitude of the core in real time. An internal temperature sensor group with a sampling frequency of 10Hz is used in conjunction with a comprehensive internal temperature acquisition strategy to collect the comprehensive internal temperature of the printer in real time. A pressure sensor with a sampling frequency of 1Hz is set inside the printer to collect the toner level in the toner cartridge in real time. Temperature and humidity sensors are deployed outside the printer to collect the temperature and humidity data of the environment in which the equipment is located in real time. Finally, multimodal data information consisting of printed part image data, vibration amplitude, internal comprehensive temperature, toner level, and ambient temperature and humidity data is obtained. The above multimodal data information is uniformly processed for noise reduction to eliminate interference signals that may exist during the data acquisition process and remove invalid interference components from the data, finally obtaining preprocessed multimodal data information.

[0078] The above technical solution achieves the following effects: This embodiment utilizes targeted data acquisition methods with multiple types of sensors to achieve comprehensive multi-dimensional data coverage of printer output quality, equipment operating status, consumable status, and environmental influencing factors (ambient temperature and humidity). This breaks through the limitations of single-modality or partial data acquisition, effectively improving the completeness of capturing key status information related to printer operation and avoiding the omission of abnormal correlation features due to missing data dimensions. Simultaneously, the parameter configuration of the high-definition industrial camera ensures clear details in the printed image and meets real-time acquisition standards, accurately reflecting print quality-related characteristics. The specific acquisition frequencies of each sensor are used to adapt to the changing characteristics of different data, ensuring real-time capture of high-frequency changing data (e.g., vibration and temperature) while avoiding redundant acquisition of low-frequency changing data (e.g., toner balance). Furthermore, the printed image covers text and graphic areas, ensuring no key areas are missed in data acquisition, thus guaranteeing the accuracy and completeness of multi-modal data from a parameter perspective. Furthermore, in this embodiment, through unified noise reduction preprocessing, interference components in multimodal data are effectively filtered out, reducing the impact of image noise and signal clutter on data authenticity, minimizing invalid data interference, and effectively improving data purity and reliability. Simultaneously, the multimodal data preprocessed using the above method in this embodiment has a unified format and stable quality, avoiding subsequent feature extraction deviations caused by interference factors in the original data. This provides high-quality data support for subsequent dynamic weight allocation, multimodal feature vector generation, and improved LSTM model prediction, effectively reducing the anomaly identification error rate caused by data quality issues. The collected multimodal data is directly related to the core of printer anomaly, and the preprocessed data quality meets standards, accurately matching the data requirements for subsequent feature extraction. This ensures that multimodal data can effectively participate in dynamic weight calculation and model prediction, thereby effectively improving the conversion efficiency from data to anomaly identification results in the overall technical solution.

[0079] In one embodiment of the present invention, the printer internal temperature comprehensive acquisition strategy includes:

[0080] The temperature sensors included in the internal temperature sensor group are used to collect real-time temperature data of the inner wall of the middle section of the paper feed channel and the outlet temperature data of the cooling fan.

[0081] Based on the temperature values ​​of the inner wall temperature and the cooling fan outlet temperature collected at each temperature sampling point, the temperature standard deviation σ of the inner wall temperature and the cooling fan outlet temperature for each sampling point is obtained. x ;

[0082] The weight values ​​w corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data are set using the temperature standard deviation of the middle section inner wall temperature data and the cooling fan outlet temperature data corresponding to each temperature acquisition.

[0083] The weight values ​​corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data are obtained by the following formula: Where w represents the weight value corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data; w0 represents the preset initial weight value corresponding to the middle section inner wall temperature data and the cooling fan outlet temperature data; σ x σ represents the standard deviation of the temperature data corresponding to the temperature data of the inner wall of the middle section and the temperature data of the cooling fan outlet at the current temperature acquisition time. mx This represents the maximum standard deviation of the temperature data corresponding to the mid-section inner wall temperature and the cooling fan outlet temperature data that have occurred during the temperature acquisition process prior to the current temperature acquisition; x max and x min This indicates the maximum and minimum temperature values ​​corresponding to the mid-section inner wall temperature data and the cooling fan outlet temperature data that occur during printer operation;

[0084] The printer's internal comprehensive temperature during the current printer operation is obtained by using the temperature values ​​corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data, combined with their corresponding weight values, to perform weighted average processing.

[0085] The working principle of the above technical solution is as follows: Through an internal temperature sensor array, temperature data from three key locations during printer operation are collected, specifically the temperature data of the inner wall of the middle section of the paper feed path and the temperature data of the cooling fan outlet. This enables temperature monitoring of components inside the printer that are sensitive to temperature and strongly correlated with the equipment's operating status. For each set of temperature data collected—the inner wall of the middle section and the cooling fan outlet—the standard deviation σx of the temperature at that time is calculated to quantify the fluctuation difference between the two sets of temperature data and reflect the stability of the current temperature distribution. Based on a preset initial weight value w0 combined with the temperature standard deviation σx at the current collection time... x The maximum standard deviation of temperature during historical data collection σ mx And the maximum temperature value x of the two sets of temperature data during printer operation. max With minimum temperature value x min The system calculates the weight values ​​w corresponding to the two sets of current temperature data, thereby dynamically adjusting the weights based on real-time temperature fluctuations and historical temperature extremes. The specific values ​​of the two currently collected temperature data sets are then weighted with their respective dynamic weight values ​​w, and the weighted results are averaged to obtain the overall internal temperature of the printer, reflecting the overall temperature status inside the printer.

[0086] The above technical solution achieves the following results: This embodiment focuses on collecting data from two core temperature-sensitive components: the inner wall of the middle section of the paper feed channel and the outlet of the cooling fan. This avoids invalid data collection from non-critical areas. Simultaneously, it comprehensively covers key temperature dimensions within the printer that affect operational stability, ensuring that temperature data accurately correlates with the core operating status of the equipment. This provides highly targeted temperature baseline data for subsequent anomaly identification. The temperature standard deviation σ is used to... x (Used to reflect real-time temperature fluctuations), historical standard deviation maximum σ mx (Used to reflect the limits of temperature fluctuation) and extreme temperature values ​​x max / x min The weight adjustment (used to reflect the boundaries of the temperature range) breaks the limitations of fixed weights, enabling the weight allocation to adapt to temperature change characteristics in real time, thereby effectively improving the sensitivity to abnormal temperature trends. The weighted averaging process combines the initial importance (w0) and real-time status (dynamic weight w) of two sets of temperature data. It retains the basic weight differences of different component temperatures while incorporating real-time temperature fluctuations and extreme value information, avoiding the problem of misjudging the overall temperature caused by temperature deviations at a single location or local temperature anomalies. At the same time, it makes the calculated internal comprehensive temperature more accurate, thus comprehensively reflecting the overall temperature status inside the printer and effectively reducing anomaly identification bias caused by inaccurate temperature characterization. Furthermore, through the above-mentioned accurate and dynamic internal comprehensive temperature data method, it can better adapt to the subsequent dynamic weight allocation mechanism at the multimodal data level, while increasing its compatibility with the prediction needs of improved LSTM models. Its data quality directly and effectively improves the effectiveness of the temperature dimension in multimodal data, avoiding multimodal feature vector bias caused by temperature data distortion or partiality, thereby providing reliable temperature dimension support for the accuracy of overall anomaly prediction and enhancing the accuracy and reliability of subsequent anomaly identification.

[0087] In one embodiment of the present invention, feature extraction is performed on the multimodal data information, and the weight values ​​corresponding to the multimodal data information are dynamically adjusted using a dynamic weight allocation method combined with the current anomaly tendency to obtain the weight value corresponding to each multimodal data information, including:

[0088] The multimodal data information is classified and features are extracted to obtain the feature vector corresponding to each data type contained in the multimodal data information;

[0089] The feature vectors corresponding to each data type contained in the multimodal data information are normalized to obtain the normalized feature vectors corresponding to each data type.

[0090] The current anomaly tendency for each data type during printer operation is determined by using the normalized feature vectors corresponding to each data type.

[0091] Based on the normalized feature vector corresponding to each data type and the current anomaly tendency, the weight values ​​corresponding to the multimodal data information are dynamically adjusted to obtain the weight value corresponding to each multimodal data information.

[0092] The working principle of the above technical solution is as follows: For multimodal data information such as printed image data, vibration amplitude, internal comprehensive temperature, toner cartridge balance, and ambient temperature and humidity data, the data is classified and processed according to data type. Feature extraction is performed on each data type to obtain a unique feature vector for each type, achieving accurate separation and extraction of features from different data types. The feature vectors corresponding to each data type are normalized using standardization algorithms (such as Min-Max normalization and Z-Score normalization) to map feature vectors of different data types to a unified numerical range, eliminating the inconsistency in feature value spans caused by differences in physical dimensions, resulting in normalized feature vectors. Based on the normalized feature vectors of each data type, the deviation of their numerical change trends from the normal range is analyzed to determine the abnormal tendency of each data type under the current printer operating state. Using the normalized feature vector reflecting feature intensity and its corresponding current abnormal tendency reflecting abnormal correlation as dual criteria, a weight adjustment rule is established to adjust the weight values ​​corresponding to each type of multimodal data information in real time.

[0093] The above technical solution achieves the following effects: By extracting features according to data type classification, it avoids feature overlap between different data types, accurately captures the unique attributes of each data type, retains the unique representational value of each data type for anomaly identification, and provides accurate feature data for subsequent anomaly tendency judgment, thereby effectively reducing anomaly misjudgments caused by feature overlap. Normalization effectively eliminates the inconsistency in feature value spans caused by differences in physical dimensions between different data types, ensuring that the feature vectors of all types of data are in the same numerical dimension (e.g., all mapped to the [0,1] interval), avoiding bias in subsequent weight calculations due to differences in dimensions, and effectively improving the consistency of subsequent multimodal feature fusion. Determining anomaly tendency based on the normalized feature vectors accurately locates the real-time state deviation of each data type, avoiding fuzzy judgments of the overall anomaly tendency of multimodal data, thereby effectively improving the targeting of weight allocation. By combining a dual dynamic adjustment mechanism of feature vectors and anomaly tendency, the limitations of static fixed weights can be effectively avoided. This allows the weights of multimodal data to flexibly change according to the real-time operating status of the printer, ensuring a high degree of match between weight allocation and the current anomaly state. This prevents non-critical data from being overweighted, thus masking anomaly features and effectively improving the effectiveness of weight allocation. Accurate classification features, normalized consistent features, targeted anomaly tendency judgment, and dynamically adapted weights collectively provide high-quality input for subsequent multimodal feature vector generation. This enables the generated multimodal feature vectors to more accurately reflect the real-time anomaly state of the printer, avoiding feature vector distortion caused by low feature quality and poor weight adaptation. This provides more reliable multimodal feature support for the anomaly prediction of the improved LSTM model, indirectly and effectively improving the accuracy of subsequent anomaly prediction.

[0094] In one embodiment of the present invention, the multimodal data information is classified and its features are extracted to obtain a feature vector corresponding to each data type contained in the multimodal data information, including:

[0095] The preprocessed multimodal data information is retrieved and classified to obtain various data types corresponding to the multimodal data information. These data types include printed image data, printer operating parameters, and printer environmental parameters. The printer operating parameters include vibration amplitude, overall internal temperature of the printer, and toner cartridge balance. The printer environmental parameters include temperature and humidity data of the printer's environment.

[0096] An improved ResNet-18 model is used to perform image recognition on printed image data. The average proportions of blurred, missing, and misaligned areas are obtained for all printed image data. Then, a feature vector S=[S] is generated based on these average proportions. 01 S 02 S 03 ]; where S 01 S 02 and S 03 These represent the average percentage of blurred area, the average percentage of missing area, and the average percentage of misaligned area, respectively. Furthermore, the percentage of blurred area refers to the ratio of the blurred area to the total area of ​​each printed document; the percentage of missing area refers to the ratio of the missing area to the total area of ​​each printed document; and the percentage of misaligned area refers to the ratio of the misaligned area to the total area of ​​each printed document.

[0097] The current vibration amplitude, printer internal temperature and toner cartridge balance are retrieved, and the vibration amplitude, printer internal temperature and toner cartridge balance are normalized to obtain the normalized current vibration amplitude, printer internal temperature and toner cartridge balance.

[0098] The feature vector Y=[Y] is generated by integrating the current vibration amplitude, the overall internal temperature of the printer, and the remaining toner in the toner cartridge to produce the printer's operating parameters. 01 Y 02 Y 03 ], where Y 01 Y 02 and Y 03 These represent the normalized vibration amplitude, the printer's internal temperature, and the remaining toner in the toner cartridge, respectively.

[0099] Retrieve the temperature and humidity data of the current printer environment, and normalize the temperature and humidity data of the printer environment to obtain the normalized temperature and humidity values ​​of the current printer environment.

[0100] By integrating the normalized temperature and humidity values ​​of the current printer environment, a feature vector R=[R] is generated corresponding to the environmental parameters. 01 R 02 ], where R 01 and R 02 These represent the normalized temperature and humidity values, respectively.

[0101] The structure of the improved ResNet-18 model is as follows:

[0102] The input layer and preprocessing layer include an image input sublayer, a normalization sublayer, and a printable effective region mask generation sublayer. Specifically, the image input sublayer receives the printed image data; the normalization sublayer maps the image pixel values ​​from [0,255] to [-1,1] (based on the ImageNet pre-trained mean μ and standard deviation σ) to eliminate dimensional differences and accelerate model convergence; the printable effective region mask generation sublayer segments the printed content area from the blank background using the Otsu thresholding method to generate a 224×224 binary mask M (M=1 for the effective region, M=0 for the background), limiting the scope of subsequent abnormal region statistics (only abnormalities within the effective region are counted).

[0103] The initial convolutional layer is used to extract low-level texture features of the image (such as paper edges and text outlines) using 3×3 convolutional kernels (64 kernels), with a stride of 2 to achieve the first downsampling (224→112); the BN layer stabilizes the feature distribution, and ReLU activation enhances the effective feature response;

[0104] The dual-branch parallel processing layer includes an original feature extraction branch and a region area statistics branch; wherein, the original feature extraction branch is used to divide the printed image data into feature maps and obtain multiple feature maps corresponding to the printed image data;

[0105] The regional area statistics branch includes a sub-branch for fuzzy region positioning and area ratio calculation, a sub-branch for missing printing region positioning and area ratio calculation, a sub-branch for misaligned region positioning and area ratio calculation, and a global statistics module;

[0106] The sub-branch for locating and calculating the area ratio of blurred regions is used to identify low-contrast blurred regions and calculate the proportion of blurred regions in a single image to the effective printing area. The sub-branch for locating and calculating the area ratio of missing areas is used to identify missing areas in dark areas without toner and calculate the proportion of missing areas in a single image. The sub-branch for locating and calculating the area ratio of misaligned regions is used to identify misaligned regions with offset edges of printed content and their corresponding proportions. The global statistics module is used to calculate and obtain the average area ratio of blurred regions, the average area ratio of missing areas, and the average area ratio of misaligned areas.

[0107] The working principle of the above technical solution is as follows: take the pre-processed multimodal data information and divide it into three types of data according to the data function attributes: printout image data, which directly reflects the printout quality; printer operating parameters, which include vibration amplitude, printer internal comprehensive temperature, and toner cartridge balance, which reflect the core operating status of the equipment; and printer environmental parameters containing ambient temperature and humidity data, which reflect external influencing factors, thus clarifying the attribution and scope of each type of data.

[0108] Feature extraction of printed image data using an improved ResNet-18 model:

[0109] Input layer and preprocessing: The image input sublayer receives the image data of the printed part; based on the ImageNet pre-trained mean μ and standard deviation σ, the normalization sublayer maps the pixel values ​​from [0,255] to [-1,1] to eliminate the difference in dimensions; the printing effective area mask generation sublayer divides the printed content area and the blank background through the Otsu thresholding method to generate a 224×224 binary mask. Specifically, M=1 is the effective area and M=0 is the background, which limits the statistical range of abnormal areas.

[0110] Initial convolutional layer processing: 3×3 convolutional kernels (64 kernels) are used to extract low-level texture features of the image. The low-level texture features of the image include, but are not limited to, paper edges and text outlines. A stride of 2 is used to achieve downsampling. The BN layer is combined with ReLU activation to enhance the effective feature response.

[0111] The dual-branch parallel processing is as follows: the original feature extraction branch divides the image data into feature maps to complete deep feature capture; the region area statistics branch locates the blurred region corresponding to low contrast, the missing area corresponding to the toner-free dark area, and the misaligned region corresponding to the content edge offset through three sub-branches, calculates the proportion of the three types of regions in the effective area in a single image, and then calculates the average area proportion of the three types of regions through the global statistics module.

[0112] Feature vector generation: Using the average proportion of blurred area, the average proportion of missing area, and the average proportion of misaligned area as elements, generate a feature vector S=[S 01 S 02 S 03 ].

[0113] Printer operating parameter feature extraction: The current vibration amplitude, printer internal temperature, and toner cartridge balance are retrieved and normalized to eliminate differences in physical dimensions. Using the normalized data as elements, an operating parameter feature vector Y=[Y] is generated. 01 Y 02 Y 03 ].

[0114] Printer environmental parameter feature extraction: Retrieve current ambient temperature and humidity data, normalize the temperature and humidity values, and use the normalized data as elements to generate an environmental parameter feature vector R=[R 01 R 02 ].

[0115] The above technical solution achieves the following effects: It categorizes multimodal data according to the functional logic of output quality, equipment operation, and external environment, clarifying the core role and attribution of each data type, avoiding feature overlap between different functional types, laying the foundation for subsequent precise categorization processing, and ensuring a high degree of match between the feature extraction direction and data function for each type of data, effectively improving the targeting of feature extraction. By improving the input preprocessing layer of the ResNet-18 model through standardization, it accelerates model convergence. Simultaneously, it utilizes the Otsu thresholding method to generate effective region masks, only counting anomalies within the printed content area, effectively avoiding interference from blank backgrounds in calculating the proportion of anomaly regions and reducing errors caused by invalid data. The dual-branch parallel structure captures deep texture features of the image through the original feature extraction branch and accurately locates three core printing quality anomalies—blurring, missing prints, and misalignment—through the region area statistics branch. Furthermore, the average value output by the global statistics module comprehensively reflects the overall quality status of multiple frames, enabling the generated feature vector S to comprehensively and accurately characterize the quality anomaly features of the printed image, effectively avoiding feature bias caused by a single frame of data or a single anomaly type. Meanwhile, the printer's operating parameters and environmental parameters have been normalized, effectively eliminating the differences in feature values ​​caused by different physical dimensions. This ensures that the elements of the feature vectors are in a unified numerical dimension, guaranteeing fair comparison of various features during subsequent dynamic weight allocation and avoiding misjudgments of feature contribution due to differences in numerical range. Furthermore, all three types of feature vectors use key quantitative indicators as elements, which effectively improves the dimensional clarity and physical meaning of the feature vectors, thereby effectively and directly reflecting the core state of their corresponding data types and thus effectively improving the feature vectors' ability to represent anomalies.

[0116] On the other hand, the improved ResNet-18 model optimizes its structure for the characteristics of printed image data. The 3×3 convolutional kernels and downsampling design of the initial convolutional layer adapt to the image resolution requirements. The dual-branch parallel processing enables simultaneous feature extraction and anomaly statistics, effectively avoiding efficiency losses caused by staged processing. The input layer is standardized based on ImageNet pre-trained parameters, reducing the difficulty of model training and accelerating convergence speed, thereby effectively reducing model debugging costs and enabling the model to quickly adapt to the feature extraction requirements of printer images, effectively improving the overall data processing efficiency. The three types of feature vectors generated by accurate classification and high-quality feature extraction provide clear and reliable feature inputs for subsequent dynamic weight allocation. That is, each type of vector can accurately reflect the anomaly association attributes of the corresponding data type. For example, the feature vector S of printed image data reflects printing quality anomalies, the feature vector Y of operating parameters reflects equipment operating anomalies, and the feature vector R of environmental parameters reflects environmental influences, effectively avoiding the problem of weight allocation deviations caused by feature ambiguity or distortion. At the same time, the quantified feature elements facilitate subsequent combination with anomaly tendency analysis, providing a clear basis for dynamic weight adjustment and effectively improving the accuracy and reliability of anomaly prediction in the subsequent improved LSTM model.

[0117] One embodiment of the present invention utilizes the normalized feature vector corresponding to each data type to determine the current anomaly tendency corresponding to each data type during printer operation, including:

[0118] Retrieve the feature vector corresponding to each data type of the multimodal data collected each time;

[0119] The L1 norm of the feature vector corresponding to each data type is obtained by using the feature vector corresponding to each data type of multimodal data collected each time.

[0120] Based on the L1 norm corresponding to the feature vector of each data type, obtain the standard deviation of the L1 norm for each data type;

[0121] The L2 norm of each data type's feature vector is obtained by using the feature vector corresponding to each data type collected in each multimodal data collection.

[0122] The abnormal tendency determination coefficient K for each data type is obtained based on the L1 norm standard deviation and L2 norm corresponding to the feature vector of each data type.

[0123] The abnormal tendency determination coefficient is obtained by the following formula: Where K represents the anomaly tendency judgment coefficient for each data type; n represents the number of data acquisitions the printer has performed; L 01i and L 02iLet L represent the L1 norm and L2 norm of the feature vector for each data type corresponding to the i-th data acquisition, respectively; 01b Let L1 norm be the standard deviation of the feature vector for each data type corresponding to n data collections.

[0124] The working principle of the above technical solution is as follows: when collecting multimodal data information each time, the normalized feature vectors of the three types of data corresponding to the printed image data, printer operating parameters, and printer environmental parameters are retrieved, namely the feature vector S of the printed image data, the feature vector Y of the operating parameters, and the feature vector R of the environmental parameters. This ensures that the obtained feature vectors correspond accurately to each data collection cycle, providing real-time and matching basic data for subsequent calculations.

[0125] For each type of data's feature vector, its L1 norm is calculated, which is the sum of the absolute values ​​of all elements in the vector. The L1 norm quantifies the overall numerical scale of the feature vector and is robust to slight noise in the data, providing a preliminary indication of the degree of deviation of the feature vector from the normal baseline state. The L1 norm (L1) of each data type is calculated based on n data acquisitions. 01i (i=1 to n), calculate the standard deviation of the L1 norm of this type of data (L... 01b The standard deviation can characterize the fluctuation range of the L1 norm during multiple data collections. Larger fluctuations indicate poorer stability of the feature vector, indirectly reflecting a higher likelihood of anomalous tendencies. For feature vectors of each data type, the L2 norm (the square root of the sum of squares of all elements) is calculated. The L2 norm is more sensitive to larger numerical deviations in the feature vector, accurately capturing significant anomalous elements outside the normal range, thus supplementing the L1 norm's robustness to minor deviations but insufficient sensitivity to significant anomalies. The L1 norm (L...) from n data collections... 01i L2 norm (L) 02i ) and L1 norm standard deviation (L 01b Substituting the values ​​into the specified formula, the anomaly tendency determination coefficient K for each data type is obtained through comprehensive calculation. The formula quantifies the degree of anomaly tendency for each data type by integrating norm statistics and fluctuation characteristics collected multiple times.

[0126] The above technical solution achieves the following effects: by using mathematical indicators such as L1 norm, L2 norm, and L1 norm standard deviation, and by deriving the determination coefficient K through formulaic calculation, it completely replaces subjective experience judgment and avoids judgment bias caused by human factors. At the same time, the anomaly tendency determination coefficient K is output in a quantitative form, allowing direct comparison of the anomaly tendencies of different data types. For example, if the anomaly tendency determination coefficient K of the operating parameter is higher than that of the environmental parameter, it indicates that the anomaly tendency of the operating parameter is more prominent, providing a clear and quantifiable decision-making basis for subsequent dynamic weight allocation, thereby effectively improving the objectivity and scientific nature of the judgment process. Combining the strong noise resistance of the L1 norm, which reflects overall deviation, with the L2 norm, which is sensitive to significant anomalies and captures strong local deviations, this approach avoids false triggering of anomaly detection by minor noise while accurately identifying significant anomalous elements in the feature vector. Furthermore, by introducing the L1 norm standard deviation, it integrates single-time feature states with historical fluctuation trends. Even if a single norm deviation is small, but historical fluctuations are significant (e.g., large standard deviation), the potential anomaly tendency can still be reflected through the anomaly tendency judgment coefficient. This comprehensively covers both static deviation and dynamic fluctuation anomaly signals, effectively reducing missed and false detections of anomaly tendencies. After each data acquisition, the L1 norm, L2 norm, and standard deviation are recalculated, and the anomaly tendency judgment coefficient K is updated in real time. This ensures that the anomaly tendency judgment coefficient K closely follows the changes in data type status during printer operation, avoiding the judgment lag caused by using fixed thresholds or historical static data. This achieves real-time tracking and dynamic judgment of anomaly tendencies, while also adapting to changes in data status under different printer operating conditions.

[0127] On the other hand, the quantified anomaly tendency determination coefficient K provides direct data input for adjusting weights based on anomaly tendencies. No additional conversion is required; the direction and magnitude of weight adjustment can be directly determined based on the magnitude of the anomaly tendency determination coefficient, avoiding weight adjustment deviations caused by ambiguous descriptions of anomalies. Simultaneously, the quantified attribute of the anomaly tendency determination coefficient allows weight allocation to be logically correlated, effectively improving the automation and accuracy of subsequent dynamic weight allocation and strengthening the overall process coherence of the technical solution. The noise resistance characteristics of the L1 norm and the statistical analysis of fluctuations using standard deviation effectively resist accidental noise during data acquisition (such as slight fluctuations in ambient temperature and humidity), avoiding false anomalies caused by noise interference. Furthermore, through statistical calculations of n collected data, the impact of single anomalous data (such as erroneous feature vectors caused by momentary sensor malfunctions) on the determination results is reduced, effectively improving the stability and robustness of anomaly tendency determination and ensuring reliable output results.

[0128] In one embodiment of the present invention, the weight values ​​corresponding to multimodal data information are dynamically adjusted based on the normalized feature vectors corresponding to each data type and the current anomaly tendency, to obtain the weight value corresponding to each multimodal data information, including:

[0129] Retrieve the abnormal tendency judgment coefficient K for each data type and the corresponding initial weight value for each data type;

[0130] The initial weight values ​​for each data type are adjusted using the abnormal tendency judgment coefficient K corresponding to each data type, and the adjusted weight values ​​for each data type are obtained.

[0131] The adjusted weight value is obtained using the following formula: , where w t This represents the adjusted weight value for each data type; w d Indicates the weight value before adjustment; K represents the anomaly tendency judgment coefficient corresponding to each data type; w max and w min These represent the maximum and minimum weights after weight adjustment for each data type, respectively.

[0132] The dimension is unified by using the feature vectors corresponding to each data type contained in the multimodal data information to obtain the dimension-unified feature vectors corresponding to each data type.

[0133] The feature vectors corresponding to each data type with unified dimensions are fused with their corresponding dynamic weight values ​​to generate a fused multimodal feature vector.

[0134] The working principle of the above technical solution is as follows: It retrieves the anomaly tendency judgment coefficient K for each data type calculated in the previous stage, namely, printed image data, printer operating parameters, and printer environmental parameters. Simultaneously, it retrieves the preset initial weight values ​​for each type of data to ensure that the input parameters for weight adjustment accurately match the data type, providing a basis for dynamic adjustment. The initial weight values ​​and anomaly tendency judgment coefficients for each data type, along with the preset maximum value w for that type of data, are then used to determine the anomaly tendency judgment coefficients. max and minimum value w min Substituting into the formula above in this embodiment, the adjusted weight value w is calculated. t This enables dynamic linkage between weight values ​​and data type anomaly tendencies, while simultaneously allowing for the preset maximum value w. max and minimum value w min Constraining the weight range can prevent imbalances caused by excessively large or small weights. Simultaneously, for different types of feature vectors in multimodal data, dimensionality unification is performed on the feature vectors through zero-padding or feature mapping. By adjusting various feature vectors to the same dimension, fusion bias caused by differences in the original dimensions is effectively eliminated, resulting in dimension-unified feature vectors. Furthermore, the dimension-unified feature vectors for each data type are then paired with their corresponding dynamic weight values ​​w. tWeighted fusion operations are performed to integrate the feature information of various types of data, and finally a fused multimodal feature vector is generated, realizing the effective aggregation of multimodal data features.

[0135] The above technical solution achieves the following results: by combining the anomaly tendency judgment coefficient K with the weight adjustment formula, the weight values ​​can respond in real time to changes in the abnormal state of data types. Specifically, data types with a high anomaly tendency receive higher weights, enhancing their contribution in subsequent fusion; data types with a low anomaly tendency have appropriately lower weights to avoid non-critical features consuming excessive weight resources. Compared to fixed weights, this method completely overcomes the limitation of static weights being unable to adapt to real-time abnormal states, allowing weight allocation to highly match the dynamic abnormal characteristics of printer operation, effectively improving the targeting and effectiveness of weights.

[0136] The introduction of maximum and minimum weights after weight adjustment for each data type creates clear constraints on the adjusted weight values. This prevents a certain type of data from having an excessively high anomaly tendency judgment coefficient, causing its weight to far exceed the reasonable range (e.g., monopolizing the dominance of fusion features and masking other potential anomalies). It also avoids a weight approaching zero due to an excessively low anomaly tendency judgment coefficient (e.g., completely ignoring potential latent anomaly features in low anomaly tendency data). This ensures that the weights of all types of data are always within the effective range, maintaining the balance of multimodal data feature fusion and effectively improving the overall rationality and stability of weight allocation.

[0137] Meanwhile, the dimensionality unification process eliminates the dimensional barriers between different types of feature vectors, avoiding the loss or misalignment of feature information during fusion due to dimensional differences, and ensuring that various feature vectors can be weighted in the same dimensional space. Furthermore, the integration of dynamic weights allows the fusion process to focus on the key features of data with high anomaly tendency while retaining the basic features of data with low anomaly tendency. This highlights core anomaly information without overlooking potential correlation features, significantly improving the information completeness and representation accuracy of the fused multimodal feature vectors. Moreover, the fused multimodal feature vectors combine the anomaly-oriented nature of dynamic weights with the consistency of dimensional unification. Their feature information accurately reflects the real-time anomaly status of the printer and can be directly adapted to the input requirements of the improved LSTM model through format standardization and dimensional unification, without requiring additional data conversion processing. Simultaneously, the high-quality fused feature vectors reduce prediction bias caused by feature ambiguity and dimensional chaos, providing reliable input for the model to accurately capture the temporal correlation features of multimodal data, indirectly and effectively improving the accuracy and efficiency of anomaly prediction. The combination of dynamic weight adjustment and dimensional unification solves the core problems of large differences in multimodal data features and uneven contributions. Dynamic weight allocation highlights the unique value of various data types under different anomaly scenarios, while dimensional unification enables the effective aggregation of various data features. This transforms multimodal data from a simple superposition of single dimensions into a synergistic and complementary feature system, fully releasing the comprehensive advantages of multimodal data in anomaly identification and reducing anomaly misses caused by the one-sidedness of single-modal features. The anomaly tendency judgment coefficient K itself possesses noise resistance due to its L1 norm-based noise immunity and multiple data collection and statistical analysis, and its driving weight adjustment effectively resists random noise interference. Simultaneously, the constraints on the maximum and minimum weight values ​​after weight adjustment for each data type further reduce the impact of anomaly data on the weights, ensuring that the weight adjustment remains within a stable and controllable range, effectively improving the robustness of the overall weight allocation mechanism.

[0138] In one embodiment of the present invention, the multimodal feature vector is input into an improved LSTM model to perform anomaly identification on a printer, and an anomaly identification result is obtained. An anomaly warning is then issued based on the anomaly identification result, including:

[0139] The modal feature vector is input into the improved LSTM model; the probability of occurrence of each abnormal situation is obtained through the output of the improved LSTM model; the probability of occurrence of each abnormal situation is compared with its corresponding probability threshold; when the probability of occurrence of the abnormal situation reaches or exceeds its corresponding probability threshold, an abnormality warning is issued.

[0140] The improved LSTM model is obtained using the following formula:

[0141] The input layer is used to receive the fused multimodal feature vectors.

[0142] The core temporal feature extraction layer is used to capture short, medium and long temporal dependencies through three layers of LSTM with residual connections, while solving the gradient vanishing problem and ensuring the effective transmission of long temporal information.

[0143] The temporal feature enhancement layer is used to dynamically and effectively improve the weights of key time steps by using the features of the last time step of the LSTM as the query and suppressing irrelevant temporal interference through ScaledDot-Product attention.

[0144] The feature compression and integration layer is used to reduce the dimension from 512 to 256 after global pooling of weighted time series features and passing them through a fully connected layer, thereby integrating long and short-term time series information and reducing the amount of subsequent computation.

[0145] The multi-time-point prediction layer uses five independent prediction heads combined with the Softmax activation function to output the probability of occurrence of 12 anomaly types for each minute within the next 5 minutes. The 12 anomaly types include:

[0146] Printhead clogging causes an abnormally high percentage of blurred areas in the printed image data.

[0147] A clogged toner cartridge nozzle causes an abnormally high percentage of the area of ​​missing printouts in the corresponding printed image data.

[0148] A paper feed mechanism misalignment fault results in an abnormally high percentage of the misaligned area in the corresponding printed image data.

[0149] Wear and tear on the machine mechanism gears will cause an abnormally high vibration amplitude in the corresponding printer operating parameters;

[0150] The cooling fan failure caused an abnormally high overall internal temperature in the printer's operating parameters.

[0151] The toner cartridge is low on ink, meaning the remaining toner level in the cartridge is below the normal threshold according to the printer's operating parameters.

[0152] The paper feed path component is overheating, causing an abnormal increase in the overall internal temperature corresponding to the printer's operating parameters.

[0153] The aging of the machine core bearings causes an abnormally high vibration amplitude in the corresponding printer operating parameters.

[0154] The paper jamming problem is caused by dampness, which corresponds to an abnormally high humidity level in the printer's environmental parameters, accompanied by abnormal vibration of the printer mechanism.

[0155] The circuit module overheating fault is caused by an abnormally high overall internal temperature corresponding to the printer's operating parameters, combined with a high ambient temperature.

[0156] Printhead positioning deviation fault, corresponding to an abnormally high percentage of misaligned area in the printed image data;

[0157] The toner cartridge is leaking ink due to poor sealing, and the remaining toner level in the cartridge drops abnormally and rapidly according to the printer's operating parameters.

[0158] The output layer is used to output the anomaly prediction results, that is, the probability of occurrence of each anomaly type. For example, the probability of occurrence of different anomaly types is different. Some anomaly types have an occurrence probability of 3%, some have an occurrence probability of 70%, and some have an occurrence probability of 0. The output shows the probability distribution of 12 anomaly types. When the anomaly occurrence probability is the highest, or exceeds its corresponding probability threshold, it indicates that the probability of this type of failure is very high. The anomaly occurrence probability is then output for further analysis and early warning.

[0159] The working principle of the above technical solution is as follows: The fused multimodal feature vectors are input into the input layer of the improved LSTM model; the core temporal feature extraction layer uses a three-layer LSTM network with residual connections to capture the short, medium, and long temporal dependencies of the multimodal data respectively. At the same time, the residual connections solve the gradient vanishing problem that traditional LSTM is prone to in long-term training, ensuring that the running state information in the long period can be effectively transmitted to subsequent layers. The temporal feature enhancement layer uses the features of the last time step of the LSTM as the query, and dynamically calculates the importance weight of each time step feature through the ScaledDot-Product attention mechanism, effectively improving the contribution of key time step features closely related to anomalies and effectively suppressing the interference intensity of irrelevant time series; the feature compression and integration layer performs global pooling on the weighted temporal features, and then reduces the feature dimension from 512 dimensions to 256 dimensions through a fully connected layer, while retaining the core information of long and short-term time series and reducing the amount of subsequent computation. The multi-time-point prediction layer uses five independent fully connected prediction heads combined with the Softmax activation function to output the probability of occurrence of 12 types of anomalies for each minute within the next 5 minutes. After the output layer outputs the probability results of all anomalies, it compares the probability of occurrence of each type of anomaly with a preset probability threshold. If the probability of occurrence of a certain type of anomaly reaches or exceeds the threshold, the corresponding anomaly warning is triggered.

[0160] The above technical solutions offer the following advantages: The use of a 3-layer LSTM design with residual connections overcomes the limitations of traditional LSTMs in transmitting long-term time-series information. It accurately captures sudden anomalies within a short timeframe and effectively extracts medium- to long-term trend anomalies, achieving comprehensive coverage of short, medium, and long-term time-series dependencies. Simultaneously, it avoids missed anomaly detection due to incomplete time-series capture. The ScaledDot-Product attention mechanism, through dynamic weight allocation, focuses the model's attention on time-step features that play a crucial role in anomaly detection, thereby weakening the interference of irrelevant time-series data and effectively improving the model's sensitivity to latent anomaly features, thus reducing misjudgments caused by feature confounding. The multi-time-point prediction head design enables refined predictions every minute for the next 5 minutes, reserving sufficient response time for operations and maintenance, and solving the problems of single-prediction and response lag in traditional prediction methods. The probabilistic results output by the Softmax activation function, combined with threshold judgment, quantify the probability of anomaly occurrence, avoiding subjective judgment bias. Furthermore, the classification prediction of 12 anomalies clearly defines the anomaly type, effectively improving fault location accuracy. The feature compression integration layer reduces the dimensionality from 512 to 56 dimensions, which can significantly reduce the subsequent computational load of the model without losing core information, effectively improve the prediction speed, and adapt to the needs of real-time printer monitoring. The combination of residual connection and attention mechanism not only enhances the model's adaptability to complex time series data, but also reduces gradient problems in the training process, effectively improving the overall stability and generalization ability of the model.

[0161] This invention proposes an AI-based intelligent identification system for multimodal anomalies in document printers, such as... Figure 1 As shown, the document printer multimodal anomaly intelligent identification system includes:

[0162] The data acquisition module is used to collect multimodal data information of the printer in real time by using a sensor group combined with the printer's internal temperature comprehensive acquisition strategy, and to preprocess the multimodal data information to obtain preprocessed multimodal data information.

[0163] The dynamic weight adjustment module is used to extract features from the multimodal data information, dynamically adjust the weight values ​​corresponding to the multimodal data information using a dynamic weight allocation method combined with the current abnormal tendency, obtain the weight value corresponding to each multimodal data information, and generate a multimodal feature vector.

[0164] The anomaly warning module is used to input the multimodal feature vector into the improved LSTM model to perform anomaly recognition on the printer, obtain the anomaly recognition result, and issue anomaly warning based on the anomaly recognition result.

[0165] The working principle of the above technical solution is as follows: A sensor array, combined with a comprehensive internal temperature acquisition strategy for the printer, collects multimodal data information in real time during printer operation. The collected multimodal data is preprocessed, including but not limited to noise reduction, data standardization, and missing value completion, to eliminate data interference and format differences, resulting in high-quality preprocessed multimodal data. Feature extraction is performed on the preprocessed multimodal data to uncover key features related to abnormal states in each modality. A dynamic weight allocation method is adopted, adjusting the weight values ​​of each modality in real time based on the abnormal tendencies exhibited by the printer's current operating state, highlighting the contribution of data with stronger indicative power of the current abnormal tendencies. Based on the extracted features and dynamically adjusted weight values, a multimodal feature vector that accurately reflects the printer's operating state is generated. This multimodal feature vector is input into an improved LSTM model. Through temporal correlation analysis of the multimodal features, the model predicts future abnormalities in the printer's operating state and outputs the prediction results. Based on the prediction results, a corresponding early warning mechanism is triggered to complete the warning output.

[0166] The above technical solution achieves the following effects: By combining sensor arrays with a comprehensive temperature acquisition strategy, multi-dimensional and targeted multi-modal data acquisition is realized, covering key printer operating status information and avoiding the dimensional limitations of single-modal data. Simultaneously, the preprocessing stage effectively eliminates data noise and format deviations, significantly improving data quality and consequently enhancing the reliability and accuracy of subsequent feature extraction and model prediction. This effectively solves the detection bias problems caused by incomplete data coverage and poor data quality in existing technologies. Furthermore, by dynamically allocating weights based on current anomaly trends, the weights of multi-modal data are adapted to the real-time operating status of the printer, allowing data more indicative of anomaly identification to contribute more significantly and avoiding the shortcomings of fixed weights that cannot adapt to different anomaly scenarios. Moreover, the multi-modal feature vectors generated based on the dynamic weight allocation method can more accurately and comprehensively characterize the printer's operating status and anomaly correlation features, effectively improving the ability of multi-modal feature vectors to represent anomaly states. Finally, the model's anomaly prediction results have higher accuracy and reliability, accurately identifying early latent anomalies, avoiding missed and false positives, and maximizing the accuracy and stability of anomaly prediction. Based on accurate anomaly prediction results, early warnings can be triggered at the initial stage of anomalies, allowing maintenance personnel sufficient time to handle faults and avoiding printing interruptions and equipment damage caused by the escalation of anomalies.

[0167] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An AI-based intelligent identification method for multimodal anomalies in document printers, characterized in that, The document printer multimodal anomaly intelligent identification method includes: The printer uses a sensor array combined with a comprehensive temperature acquisition strategy to collect multimodal data information of the printer in real time, and preprocesses the multimodal data information to obtain preprocessed multimodal data information. Feature extraction is performed on the multimodal data information, and the weight values ​​corresponding to the multimodal data information are dynamically adjusted by using a dynamic weight allocation method combined with the current abnormal tendency. The weight value corresponding to each multimodal data information is obtained, and a multimodal feature vector is generated. The multimodal feature vectors are input into the improved LSTM model to perform anomaly identification on the printer, obtain the anomaly identification results, and issue anomaly warnings based on the anomaly identification results. The process involves feature extraction from the multimodal data information, dynamically adjusting the weight values ​​corresponding to each multimodal data information based on the current anomaly tendency using a dynamic weight allocation method, obtaining the weight value corresponding to each multimodal data information, and generating a multimodal feature vector, including: The multimodal data information is classified and features are extracted to obtain the feature vector corresponding to each data type contained in the multimodal data information; the feature vector corresponding to each data type contained in the multimodal data information is normalized to obtain the normalized feature vector corresponding to each data type. Retrieve the feature vector corresponding to each data type of the multimodal data information collected each time, and use the feature vector corresponding to each data type of the multimodal data information collected each time to obtain the L1 norm of the feature vector corresponding to each data type; based on the L1 norm corresponding to the feature vector of each data type, obtain the standard deviation of the L1 norm of each data type. The L2 norm of each data type is obtained by using the feature vector corresponding to each data type collected in each multimodal data collection; the abnormal tendency determination coefficient K of each data type is obtained based on the L1 norm standard deviation and L2 norm of the feature vector of each data type. Retrieve the abnormal tendency judgment coefficient K and the initial weight value of each data type; adjust the initial weight value of each data type using the abnormal tendency judgment coefficient K, and obtain the adjusted weight value of each data type. The dimension of the feature vectors corresponding to each data type in the multimodal data information is unified to obtain the dimension-unified feature vectors for each data type. The dimension-unified feature vectors for each data type are then fused with their corresponding dynamic weight values ​​to generate the fused multimodal feature vectors.

2. The intelligent identification method for multimodal anomalies in document printers according to claim 1, characterized in that, A sensor array combined with a printer internal temperature comprehensive acquisition strategy is used to collect multimodal data information of the printer in real time, and the multimodal data information is preprocessed to obtain preprocessed multimodal data information, including: A high-definition industrial camera is installed at the printer's paper output port to capture real-time image data of the printed parts corresponding to the printed documents. Vibration sensors are used to collect the vibration amplitude of the printer mechanism in real time during operation. The printer uses an internal temperature sensor array combined with a comprehensive internal temperature acquisition strategy to collect the printer's internal temperature in real time during operation. The printer uses an internal pressure sensor to monitor the toner cartridge level in real time. The temperature and humidity sensors installed on the outside of the printer are used to collect real-time temperature and humidity data of the printer's environment. The image data of the printed parts, vibration amplitude, overall temperature inside the printer, toner cartridge balance, and temperature and humidity data of the printer's environment are subjected to noise reduction processing to obtain pre-processed multimodal data information.

3. The intelligent identification method for multimodal anomalies in document printers according to claim 2, characterized in that, The printer's internal temperature comprehensive acquisition strategy includes: The temperature sensors included in the internal temperature sensor group are used to collect real-time temperature data of the inner wall of the middle section of the paper feed channel and the outlet temperature data of the cooling fan. Based on the temperature values ​​of the inner wall temperature and the cooling fan outlet temperature collected at each temperature sampling point, the temperature standard deviation σ of the inner wall temperature and the cooling fan outlet temperature for each sampling point is obtained. x ; The weight values ​​w corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data are set using the temperature standard deviation of the middle section inner wall temperature data and the cooling fan outlet temperature data corresponding to each temperature acquisition. The printer's internal comprehensive temperature during the current printer operation is obtained by using the temperature values ​​corresponding to the current middle section inner wall temperature data and the cooling fan outlet temperature data, combined with their corresponding weight values, to perform weighted average processing.

4. The intelligent identification method for multimodal anomalies in document printers according to claim 1, characterized in that, The multimodal data information is classified and its features are extracted to obtain the feature vector corresponding to each data type contained in the multimodal data information, including: The preprocessed multimodal data information is retrieved and classified to obtain various data types corresponding to the multimodal data information. These data types include printed image data, printer operating parameters, and printer environmental parameters. The printer operating parameters include vibration amplitude, overall internal temperature of the printer, and toner cartridge balance. The printer environmental parameters include temperature and humidity data of the printer's environment. An improved ResNet-18 model is used to perform image recognition on printed image data. The average proportions of blurred, missing, and misaligned areas are obtained for all printed image data. Then, a feature vector S=[S] is generated based on these average proportions. 01 S 02 S 03 ]; where S 01 S 02 and S 03 These represent the average area percentage of blurred areas, the average area percentage of missing areas, and the average area percentage of misaligned areas, respectively. The current vibration amplitude, printer internal temperature and toner cartridge balance are retrieved, and the vibration amplitude, printer internal temperature and toner cartridge balance are normalized to obtain the normalized current vibration amplitude, printer internal temperature and toner cartridge balance. The feature vector Y=[Y] is generated by integrating the current vibration amplitude, the overall internal temperature of the printer, and the remaining toner in the toner cartridge to produce the printer's operating parameters. 01 Y 02 Y 03 ], where Y 01 Y 02 and Y 03 These represent the normalized vibration amplitude, the printer's internal temperature, and the remaining toner in the toner cartridge, respectively. Retrieve the temperature and humidity data of the current printer environment, and normalize the temperature and humidity data of the printer environment to obtain the normalized temperature and humidity values ​​of the current printer environment. By integrating the normalized temperature and humidity values ​​of the current printer environment, a feature vector R=[R] is generated corresponding to the environmental parameters. 01 R 02 ], where R 01 and R 02 These represent the normalized temperature and humidity values, respectively.

5. The intelligent identification method for multimodal anomalies in document printers according to claim 1, characterized in that, The multimodal feature vectors are input into the improved LSTM model to perform anomaly detection on the printer, and anomaly detection results are obtained. Based on the anomaly detection results, anomaly warnings are issued, including: The modal feature vectors are input into the improved LSTM model; The probability of an anomaly occurrence for each anomaly is obtained by using the output of the improved LSTM model. Compare the probability of occurrence of each abnormal situation with its corresponding probability threshold; An anomaly warning is issued when the probability of the anomaly occurring reaches or exceeds its corresponding probability threshold.

6. An AI-based multimodal anomaly intelligent recognition system for document printers, characterized in that, The document printer multimodal anomaly intelligent recognition system includes: The data acquisition module is used to collect multimodal data information of the printer in real time by using a sensor group combined with the printer's internal temperature comprehensive acquisition strategy, and to preprocess the multimodal data information to obtain preprocessed multimodal data information. The dynamic weight adjustment module is used to extract features from the multimodal data information, dynamically adjust the weight values ​​corresponding to the multimodal data information using a dynamic weight allocation method combined with the current abnormal tendency, obtain the weight value corresponding to each multimodal data information, and generate a multimodal feature vector. An anomaly warning module is used to input the multimodal feature vector into the improved LSTM model to perform anomaly recognition on the printer, obtain anomaly recognition results, and issue anomaly warnings based on the anomaly recognition results. The process involves feature extraction from the multimodal data information, dynamically adjusting the weight values ​​corresponding to each multimodal data information based on the current anomaly tendency using a dynamic weight allocation method, obtaining the weight value corresponding to each multimodal data information, and generating a multimodal feature vector, including: The multimodal data information is classified and features are extracted to obtain the feature vector corresponding to each data type contained in the multimodal data information; the feature vector corresponding to each data type contained in the multimodal data information is normalized to obtain the normalized feature vector corresponding to each data type. Retrieve the feature vector corresponding to each data type of the multimodal data information collected each time, and use the feature vector corresponding to each data type of the multimodal data information collected each time to obtain the L1 norm of the feature vector corresponding to each data type; based on the L1 norm corresponding to the feature vector of each data type, obtain the standard deviation of the L1 norm of each data type. The L2 norm of each data type is obtained by using the feature vector corresponding to each data type collected in each multimodal data collection; the abnormal tendency determination coefficient K of each data type is obtained based on the L1 norm standard deviation and L2 norm of the feature vector of each data type. Retrieve the abnormal tendency judgment coefficient K and the initial weight value of each data type; adjust the initial weight value of each data type using the abnormal tendency judgment coefficient K, and obtain the adjusted weight value of each data type. The dimension of the feature vectors corresponding to each data type in the multimodal data information is unified to obtain the dimension-unified feature vectors for each data type. The dimension-unified feature vectors for each data type are then fused with their corresponding dynamic weight values ​​to generate the fused multimodal feature vectors.

Citation Information

Patent Citations

  • Printer abnormal state diagnosis method and system

    CN116755648A