Storage chip defect detection method and system based on deep learning
The memory chip defect detection system based on deep learning constructs a heterogeneous frequency disturbance sensing module and a feature coupling map to identify and model memory chip defects in edge devices. This solves the problem of difficulty in real-time identification of dynamic defects in existing technologies and enables proactive adjustment and efficient protection of chip status.
Patent Information
- Application Number
- CN202511096470.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to identify dynamic defects in memory chips in edge devices in real time, especially hysteresis, metastable, or clustered evolution defects that occur under complex conditions such as frequent write operations, high-temperature environments, and sudden power outages, leading to problems such as video frame loss and file read/write failures.
A deep learning-based memory chip defect detection system is adopted, including a heterogeneous frequency disturbance sensing module, a feature coupling map construction module, a micro-deviation anomaly extraction network module, a behavior evolution trend aggregation module, and a remote inference feedback control module. By constructing read and write access sequences under multiple access frequencies, the system collects response behaviors, generates map structures, identifies and models defect areas, and realizes remote adaptive control.
It achieves accurate identification and control of defects in memory chips, reduces the probability of video frame corruption and chip damage, has high responsiveness and low intrusion, adapts to the actual operating environment of edge devices, and improves the early defect identification rate and system stability.
Smart Images

Figure CN120973578A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of defect detection, in particular to a storage chip defect detection method and system based on deep learning. BACKGROUND
[0002] In electronic systems, storage chips, as key data storage and carrier devices, are widely used in various terminal devices. Traditional storage chip defect detection methods mainly focus on the factory test stage, relying on electrical signal scanning, BIST self-checking structure or appearance detection means in the die stage. Although this method can exclude dominant failure units in the late stage of chip manufacturing, for storage chips deployed in long-term running devices, their defects often have dynamic, hidden and environmental coupling characteristics. Especially in high data-intensive terminals such as security monitoring and edge video devices, chips are more likely to have delayed, metastable or aggregated evolution defects due to frequent write operations, high temperature environments and sudden power failures. Such defects are difficult to be exposed in time under standard static detection means, and often show up as video frame loss, file read / write failure or algorithm model abnormality, which seriously affects the stability of the system. Therefore, there is an urgent need for an online defect detection mechanism that can be deployed on the device side and adapt to the actual running environment, especially for the application characteristics of edge devices, with high responsiveness, low invasiveness and high explainable intelligent detection system.
[0003] In recent years, with the development of deep learning technology, storage chip anomaly detection mechanisms based on model recognition capabilities have gradually attracted research attention. However, existing solutions focus more on data center type SSD running state evaluation and server level fault prediction, lacking the system capabilities of "lightweight deployment-real-time inference-trend tracking" in security video type devices. In the existing technology, there is no complete mechanism that can combine access behavior dynamic perception, abnormal atlas expression, defect micro-feature extraction and health evolution trend modeling for edge deployment solutions. In addition, most existing deep recognition methods rely on single static images and are difficult to capture the abnormal evolution trajectory in time series, lacking effective modeling of "abnormal transfer", "function degradation" and "spatial aggregation", resulting in lagging identification of early defects and difficulty in achieving active control and early warning response. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application provides a storage chip defect detection method and system based on deep learning, which solves the problems mentioned in the background art.
[0005] In order to achieve the above object, the present application is realized by the following technical scheme: a storage chip defect detection system based on deep learning, comprising a frequency disturbance perception module, a feature coupling graph construction module, a micro-deviation anomaly extraction network module, a behavior evolution trend aggregation module, a remote reasoning feedback control module and a multi-task adaptive scheduling module;
[0006] The frequency disturbance perception module is used to construct read-write access sequences under multiple access frequencies during the operation of the edge video device to stimulate the response behavior of the storage chip, and to collect response delays, control path dislocations and error correction activation behaviors in the access process, and to establish chip dynamic interference indicators;
[0007] The feature coupling graph construction module is used to integrate the above-mentioned response behaviors in the form of multi-dimensional parameters, and to generate a graph structure covering access behavior, control path and error correction response, and to form a coupling feature representation suitable for deep model input;
[0008] The micro-deviation anomaly extraction network module is used to receive the above-mentioned coupling feature graph, identify weak image features, regional abnormal structures or control response spots in the graph, output the position, type and confidence mark of the defect area, and generate anomaly records in an encoded manner;
[0009] The behavior evolution trend aggregation module is used to construct anomaly records at multiple time points into a time series, and to model the evolution trend based on indicators such as anomaly activity, position change trend and activation frequency, and to output a trend function for health level assessment;
[0010] The remote reasoning feedback control module is used to receive the health level result and match it with the control strategy table, perform chip protection, working state adjustment or fault reporting operation, and realize remote adaptive control of the chip running state;
[0011] The multi-task adaptive scheduling module is used to monitor the device running load, dynamically adjust the execution frequency, resource allocation and reasoning accuracy of the system, so that the collaborative operation between the chip detection task and the video main task does not conflict.
[0012] Preferably, the frequency disturbance perception module comprises a timing disturbance generation unit, a response behavior acquisition unit and a dynamic interference evaluation unit;
[0013] The timing disturbance generation unit is used to simulate the storage chip access situation of the edge video device under multiple load environments, dynamically generates a set of preset disturbance access sequences according to the device running scene, including high-definition video recording, night static monitoring or intelligent review, the sequence includes different read-write combination modes, address jump rules and access interval time difference contents;
[0014] The response behavior collection unit is connected to the execution interface back end of the timing disturbance sequence, and is used for real-time monitoring and recording of various response information generated by the chip during access, including time difference of storage controller return confirmation signal, address resolution abnormality ratio, error correction module enable frequency, cache hit state and reply misplacement situation, and archiving in a structured timestamp and behavior label manner;
[0015] The dynamic interference evaluation unit is used for attribution and disturbance intensity normalization evaluation of the collected response behavior sequence, and the processing mode is: respectively counting response fluctuation amplitude, instruction delay range, and control abnormal point aggregation density according to different access types, and converting the evaluation results into a composite disturbance level identifier representing chip stability, as transmission information transmitted into the next module to judge whether the chip is in a potential metastable state.
[0016] Preferably, the feature coupling graph construction module includes a behavior parameter integration unit, a graph mapping generation unit, and a coupling feature fusion unit.
[0017] The behavior parameter integration unit is used for receiving the composite disturbance level identifier and the multiple access response behavior indicators, and standardizing and rearranging the data items of different sources and different scales, including: unifying the time scale, converting the behavior frequency into the distribution density, and mapping the response abnormality rate to the equal-quantile interval, for graph processing.
[0018] The graph mapping generation unit is used for spatial mapping operation of various parameters through the integration output of the behavior parameter integration unit, and the mapping form includes: two-dimensional density matrix formed by access response delay, heat overlay generated by error correction trigger position, and channel drift graph formed by address jump structure. The graph adopts a fixed size coding structure, which is suitable for the input format of the subsequent deep model.
[0019] The coupling feature fusion unit is used for multi-channel merging operation of various graphs according to data homology and functional coupling, and the merging operation is: weighted superposition of different graph layers in the same region, establishment of correlation index between behavior responses of physically adjacent regions, and assignment of region markers to significant structural features in the graph. The result forms a unified feature coupling graph.
[0020] Preferably, the micro-deviation anomaly extraction network includes a local recognition preprocessing unit, a deep micro-deviation recognition unit, and an anomaly coding output unit.
[0021] The local recognition preprocessing unit is used for preprocessing the input graph through a perception region division method, including image noise elimination, scale unification, graph block cutting, and edge enhancement. The output graph is encoded in a region block sequence manner, which is convenient for local anomaly feature extraction.
[0022] The deep differential recognition unit is used for recognizing the abnormality of the subtle gray difference, edge distortion mode and the area of control path hot spot in the input atlas by deploying a deep convolutional network structure. The recognition strategy is as follows: local activation features are extracted under multiple convolutional receptive windows, feature comparison is performed on each image block, and an abnormal confidence level is generated. The final output includes the position index, type label and activation score of the suspected defect area.
[0023] The abnormal coding output unit is used for receiving the recognition result and generating a standardized abnormal coding, which includes the defect area number, center coordinate information, abnormal level, coverage page range and detection timestamp.
[0024] Preferably, the behavior evolution trend aggregation module includes a time series construction unit, a trend analysis modeling unit and a health level evaluation unit.
[0025] The time series construction unit is used for arranging the abnormal coding output by the abnormal coding output unit in chronological order to form an abnormal evolution record sequence with frame-level granularity. The sequence includes the defect type, spatial position, activity level and response variation information of each time slice, which is used to construct a defect trajectory path.
[0026] The trend analysis modeling unit is used to identify whether the defect behavior presents repeatability, persistence or expansibility. The identification method includes detecting the activation density, position offset rate and evolution linkage relationship between adjacent types of defects in the time dimension, and outputting the trend label: stable, transient and evolution.
[0027] The health level evaluation unit is used to divide the current running state of the chip into three levels of “healthy”, “warning” and “failure risk” according to the trend label result of the trend analysis modeling unit. The evaluation standard includes the logic comparison of the defect activation frequency threshold, evolution rate threshold and coverage area proportion index. The evaluation result is used to judge whether the intervention mechanism needs to be started, and is used as the input basis of the control module.
[0028] Preferably, the trend analysis modeling unit is used to receive and process the abnormal behavior coding sequence output by the time series construction unit. The trend attribution modeling is performed according to multiple dimensions such as defect activation frequency, position evolution behavior, defect category similarity and sustained response amplitude, so as to identify whether there is a potential defect behavior pattern with persistence, evolution or aggregation characteristics in the current storage chip, including abnormal activity calculation, spatial position offset trajectory analysis and abnormal type variation analysis.
[0029] The abnormal activity calculation is used to judge the activity degree of abnormal events in consecutive time slices, and judge whether it has a frequent activation trend. The number of triggers of the same type of defect and the spatial coincidence degree in a unit time window are counted to realize the following judgment logic:
[0030] If a certain type of anomaly repeatedly occurs in multiple frames and the confidence remains above the set threshold, it is considered an "active anomaly";
[0031] If the anomaly is intermittently triggered between frames or the number of activations is significantly reduced, it is marked as a "transient anomaly";
[0032] Spatial position offset trajectory analysis identifies whether the anomaly exhibits stability or migration in the physical address space by analyzing the spatial coordinates of the defect area at multiple times:
[0033] If the anomaly position remains relatively stable in the time series and the coordinate offset is less than the system error tolerance, it is marked as a "position fixed type anomaly";
[0034] If the defect exhibits continuous or jumping offset and the direction shows the same trend, increasing along the on-chip address line, it is judged as an "evolutionary expansion anomaly";
[0035] If multiple defects gradually converge around a certain central area, it is considered an "aggregative evolution";
[0036] Anomaly type variation analysis is used to analyze whether the anomaly category changes in the time series, including the gradual transition from "response delay type" to "error correction frequent type". This type of variation indicates that the defect has evolved from the performance degradation stage to the functional error stage. The system compares and judges according to the set anomaly category conversion mapping table. If it meets the specific conversion path, it outputs the "category evolution label".
[0037] Preferably, the health level evaluation unit mainly includes three evaluation processes, which respectively judge the anomaly activity, evolution trend rate, and spatial coverage breadth, and adopts a step-by-step rule fusion method to finally output the health level label:
[0038] Anomaly activity evaluation: used to count the active frequency of defect behavior per unit time to determine whether it reaches the preset threshold:
[0039] If the anomaly activation frequency in the last 3 time windows is lower than the minimum activity threshold, and the activation frequency per second is less than 2 times, it is rated as "normal";
[0040] If the activation frequency is in the medium interval, 2-8 times per second, it is rated as "warning";
[0041] If the activation frequency is continuously higher than the maximum activity threshold, the activation frequency is greater than 8 times per second, or a burst of dense anomaly clusters appears, it enters "high risk"
[0042] This stage mainly represents whether the short-term behavior of the chip is in a state of continuous instability;
[0043] Evolution trend rate assessment: used to measure the "worsening speed" of the anomaly in the time dimension, i.e. the growth rate of the defect confidence, distribution range or impact level per unit time:
[0044] If there is no significant increase in defect intensity, or a continuous decrease, it is defined as "tending to be stable", supporting the "normal" judgment;
[0045] If there is intermittent increase, but the amplitude is limited, it is defined as "fluctuating evolution", supporting "early warning";
[0046] If there is a sustained increasing trend in multiple cycles, and the rising amplitude is significantly more than the growth threshold of 30% or more, it is defined as "rapid evolution", directly classified as "high risk";
[0047] This stage reflects whether the chip is currently in a potential "worsening" situation;
[0048] Spatial coverage breadth assessment: used to assess the distribution range and density of the identified abnormal area in the address space of the storage chip:
[0049] If the abnormal block covers less than 10% of the address range, there is no cross-aggregation phenomenon, it is considered "locally controllable", supporting "normal";
[0050] If the coverage range is between 10% and 30%, or there is a skip distribution of address segments, it is considered "potential spread", supporting "early warning";
[0051] If the coverage ratio exceeds 30%, or multiple address clusters overlap to form a defect cluster phenomenon, it is defined as "multi-point instability", included in "high risk";
[0052] This stage assesses the "spread risk" and "page-level impact level" of the anomaly.
[0053] Preferably, the remote reasoning feedback control module includes a control strategy matching unit, an edge device response execution unit, and a remote reporting and recording unit;
[0054] The control strategy matching unit is used to select a matching item from the built-in control strategy table according to the chip health level, defect type and spatial distribution characteristics output by the health level assessment unit, and the control strategy covers: chip access mode adjustment, data channel switching, write rate limitation and video cache transfer;
[0055] The edge device response execution unit is used to convert the selected control strategy into bottom-level command instructions, and send them to the device system layer or chip control layer, to complete the actual action execution of chip working parameter adjustment, resource reallocation, and power supply limitation adjustment through SPI, I2C or GPIO interface, to ensure that the defect risk is responded at the hardware level;
[0056] The remote reporting and recording unit is used to generate an execution log and a state change record of a current control action, and to summarize a current health state of the chip and a response control result into structured reporting content and upload to an operation and maintenance management end or a remote monitoring platform.
[0057] Preferably, the multi-task adaptive scheduling module comprises a running resource monitoring unit and a detection scheduling decision unit.
[0058] The running resource monitoring unit is used to collect current system load information of the device in real time, including a video stream channel number, an AI model running state and a chip temperature power consumption, and to judge available resources of the chip and a calculation pressure caused by the detection task as a reference input for scheduling strategy adjustment.
[0059] The detection scheduling decision unit is used to dynamically determine strategy parameters of a running frequency, a detection accuracy and a spectrum update interval of each unit in the system according to a resource state judgment result of the running resource monitoring unit, to realize system-level task collaborative management, to ensure stability of recording and identification as a priority in a non-conflict premise, to execute a chip detection behavior, and to automatically optimize a scheduling strategy, to realize adaptive frequency adjustment, self-recovery and low-interference operation of the detection task.
[0060] A storage chip defect detection method based on deep learning comprises the following steps:
[0061] Step one: in the running process of the edge video device, a plurality of read-write access sequences under different access frequencies are constructed to stimulate response behaviors of the storage chip, and response delays, control path dislocations and error correction activation behaviors in the access process are collected to establish dynamic interference indicators of the chip.
[0062] Step two: the above response behaviors are integrated in the form of multi-dimensional parameters, a spectrum structure covering access behaviors, control paths and error correction responses is generated, and a coupled feature representation suitable for deep model input is formed by fusion;
[0063] Step three: the above coupled feature spectrum is received, weak image features, regional abnormal structures or control response spots in the spectrum are identified, positions, types and confidence markers of defect regions are output, and abnormal records are generated in an encoding manner;
[0064] Step four: the abnormal records at a plurality of time points are constructed into a time sequence, and an evolution trend model is established based on indicators of abnormal activity, position change trend and activation frequency, and a trend function for health level evaluation is output;
[0065] Step five: the health level result is received and matched with a control strategy table, a chip protection, a working state adjustment or a fault reporting operation is executed, and remote adaptive control of the chip running state is realized.
[0066] Step six: monitor the device running load, dynamically adjust the execution frequency, resource allocation and inference accuracy of the system, so that the cooperative operation between the chip detection task and the video main task does not conflict.
[0067] The application provides a storage chip defect detection method and system based on deep learning, which has the following beneficial effects:
[0068] (1) When the system is running, the interference behavior in the access process is collected, the chip dynamic interference index is established, the above-mentioned response behavior is integrated in the form of multi-dimensional parameters, the atlas structure is generated, and the coupled feature representation suitable for the input of the deep model is formed. The weak image features, regional abnormal structures or control response spots in the atlas are identified, the position, type and confidence mark of the defect area are output, the abnormal records are generated in the form of codes, the abnormal records at multiple time points are constructed as time series, and the evolution trend modeling is carried out based on the indexes of abnormal activity, position change trend and activation frequency. The trend function for health level evaluation is output, the health level result is received and matched with the control strategy table, and the chip protection, working state adjustment or fault reporting operation is performed.
[0069] (2) The application provides a storage chip defect detection system based on deep learning, which constructs a full-process detection closed loop with access behavior excitation, abnormality identification, trend modeling, remote control and resource scheduling capabilities around the actual operation needs of security monitoring and edge video equipment. The system relies on the hetero-frequency disturbance perception module to actively construct a variety of high-low frequency alternating access sequences to stimulate the response behavior of the storage chip, and then structure and code the multi-dimensional response parameter through the feature coupling atlas construction module and perform atlas fusion to ensure that the defect behavior has identifiable in spatial dimension. Through the micro-deviation abnormality extraction network module, the system can realize deep identification of fine-grained abnormal features generated by the chip during the access process and output structured results, and then the trend aggregation module completes the evolution trajectory tracking of the abnormal behavior in the time dimension, realizes the closed-loop processing of "excitation-detection-modeling-feedback", and fully covers the chip running state perception task.
[0070] (3) Compared with the traditional storage chip test which only relies on static electrical detection or function verification at the factory stage, the application innovatively introduces the defect detection path combining access behavior disturbance and evolution trend modeling, and realizes the active adjustment and dynamic correction of the chip running state through the remote reasoning feedback control module, effectively reducing the probability of problems such as video frame damage, data writing error or chip damage caused by potential defects. At the same time, the multi-task adaptive scheduling module can dynamically control the detection rhythm and model inference accuracy according to the actual calculation resource occupation of the edge device, so that the system can be embedded in resource-constrained devices for long-term operation without affecting the video acquisition, transmission and identification functions of the main task, realizing the deployment ability of cooperative coexistence of detection task and core business.
[0071] (4) The application is different from the model recognition mode of the prior art in the technical path, and the model recognition mode is mainly offline training + static input. Through the combination of behavior-driven input and trend evolution analysis, the accurate recognition and control response of the on-target failure type, hidden type and dynamic aggregation type defects are realized for the first time in the edge chip detection scene. Experimental results show that the system has significant improvement in chip defect early identification rate, abnormal false alarm suppression rate, system response speed and the like compared with the traditional scheme, especially in the unattended, high-load task edge security scene, has strong deployment adaptability and running stability, and provides a complete structure, advanced mechanism and engineering implementable technical path for intelligent sensing and pre-protection of the storage chip running state. BRIEF DESCRIPTION OF DRAWINGS
[0072] Figure 1 A storage chip defect detection system based on deep learning is provided.
[0073] Figure 2 A storage chip defect detection method based on deep learning is provided.
[0074] Figure 3 A chip defect detection trend line graph of the storage chip defect detection system based on deep learning is provided. DETAILED DESCRIPTION
[0075] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0076] Embodiment 1
[0077] The application provides a storage chip defect detection system based on deep learning, please refer to Figure 1 , which comprises a frequency disturbance perception module, a feature coupling graph construction module, a micro-deviation anomaly extraction network module, a behavior evolution trend aggregation module, a remote reasoning feedback control module and a multi-task adaptive scheduling module.
[0078] The frequency disturbance perception module is used to construct read-write access sequences under multiple access frequencies to stimulate the response behavior of the storage chip during the operation of the edge video device, and to collect the response delay, control path dislocation and error correction activation behavior in the access process, and to establish a chip dynamic interference index.
[0079] The feature coupling graph construction module is configured to integrate the above response behaviors in the form of multi-dimensional parameters, and generate a graph structure covering access behaviors, control paths and error correction responses, and fuse to form a coupling feature representation suitable for input of a deep model.
[0080] The micro-deviation anomaly extraction network module is configured to receive the above coupling feature graph, identify weak image features, regional anomaly structures or control response spots in the graph, output the position, type and confidence mark of the defect region, and generate an anomaly record in an encoded manner.
[0081] The behavior evolution trend aggregation module is configured to construct anomaly records at multiple time points into a time sequence, and model evolution trends based on indexes of anomaly activity, position change trend and activation frequency, and output a trend function for health level assessment.
[0082] The remote reasoning feedback control module is configured to receive the health level result and match it with a control strategy table, perform chip protection, working state adjustment or fault reporting operation, and realize remote adaptive control of the chip running state.
[0083] The multi-task adaptive scheduling module is configured to monitor device running load, dynamically adjust execution frequency, resource allocation and reasoning accuracy of the system, so that the collaborative operation between the chip detection task and the video main task does not conflict.
[0084] In the running process of the edge video device in this embodiment, read-write access sequences under multiple access frequencies are constructed to stimulate the response behavior of the storage chip, and response delays, control path dislocations and error correction activation behaviors in the access process are collected to establish chip dynamic interference indexes. The above response behaviors are integrated in the form of multi-dimensional parameters, and a graph structure covering access behaviors, control paths and error correction responses is generated, and a coupling feature representation suitable for input of a deep model is fused. The above coupling feature graph is received, weak image features, regional anomaly structures or control response spots in the graph are identified, the position, type and confidence mark of the defect region are output, and an anomaly record is generated in an encoded manner. Anomaly records at multiple time points are constructed into a time sequence, and evolution trends are modeled based on indexes of anomaly activity, position change trend and activation frequency, and a trend function for health level assessment is output. The health level result is received and matched with a control strategy table, chip protection, working state adjustment or fault reporting operation is performed, and remote adaptive control of the chip running state is realized. The device running load is monitored, and the execution frequency, resource allocation and reasoning accuracy of the system are dynamically adjusted, so that the collaborative operation between the chip detection task and the video main task does not conflict.
[0085] Embodiment 2
[0086] This embodiment is an explanation and description in Embodiment 1. Please refer to Figure 1Specifically, the inter-frequency disturbance perception module includes a timing disturbance generation unit, a response behavior collection unit, and a dynamic disturbance evaluation unit.
[0087] The timing disturbance generation unit is used to simulate the storage chip access situation of edge video equipment in various load environments, dynamically generates a set of preset disturbance access sequences according to the device running scene, including high-definition video recording, night static monitoring or intelligent review, and the sequence includes different read-write combination modes, address jump rules and access interval time difference contents.
[0088] The response behavior collection unit is used to connect to the execution interface of the timing disturbance sequence, and real-time monitors and records various response information generated by the chip during access, including the time difference of the storage controller return confirmation signal, the address resolution abnormality rate, the error correction module enable frequency, the cache hit state and the receipt error position condition, and archives them in a structured timestamp and behavior label manner.
[0089] The dynamic disturbance evaluation unit is used to attribute and normalize the disturbance intensity of the collected response behavior sequence. The processing method is: according to different access types, respectively, the response fluctuation amplitude, the instruction delay range, and the aggregation density of control abnormal points are counted, and the evaluation results are converted into a composite disturbance level identifier representing the stability of the chip, which is transmitted into the next module as transmission information to judge whether the chip is in a potential metastable state.
[0090] The feature coupling graph construction module includes a behavior parameter integration unit, a graph mapping generation unit, and a coupling feature fusion unit.
[0091] The behavior parameter integration unit is used to receive the composite disturbance level identifier and the multi-type access response behavior index, and to standardize and rearrange data items of different sources and different scales, including: unifying the time scale, converting the behavior frequency into the distribution density, and mapping the response abnormality rate to the equal-quantile interval, for graph processing.
[0092] The graph mapping generation unit is used to perform spatial mapping operation on various parameters through the integration output of the behavior parameter integration unit. The mapping form includes: the access response delay forms a two-dimensional density matrix, the error correction trigger position generates a heat map, and the address jump structure forms a channel drift graph. The graph adopts a fixed size coding structure to adapt to the input format of the subsequent deep model.
[0093] The coupling feature fusion unit is used to perform multi-channel merging operation on various graphs according to data homology and functional coupling. The merging operation is: weighting and superimposing different graph layers in the same area, establishing correlation index between behavior responses of physically adjacent areas, and assigning region labels to significant structural features in the graph. The result forms a unified feature coupling graph.
[0094] The micro-deviation anomaly extraction network comprises a local recognition preprocessing unit, a deep micro-deviation recognition unit and an anomaly coding output unit;
[0095] The local recognition preprocessing unit is configured to perform preprocessing operation on the input atlas by a perception region division manner, including image noise elimination, scale unification, image block cutting and edge enhancement, and the output atlas is encoded in a region block sequence manner, facilitating local anomaly feature extraction;
[0096] The deep micro-deviation recognition unit is configured to perform anomaly recognition on the subtle gray difference, edge distortion mode and region of control path hot spot in the input atlas by deploying a deep convolutional network structure, and the recognition strategy is as follows: local activation features are extracted under multiple convolution receptive windows, feature comparison is performed on each image block, and an anomaly confidence level is generated, and finally, the output is as follows: position index, type label and activation score of the suspected defect region;
[0097] The anomaly coding output unit is configured to receive the recognition result and generate a standardized anomaly code, and the content includes defect region number, center coordinate information, anomaly level, coverage page segment range and detection timestamp.
[0098] In this embodiment, by constructing the hetero-frequency disturbance perception module, the system can actively simulate the access behavior characteristics of edge video devices in various actual working states, and form a set of composite disturbance sequences including read-write mode combination, address jump rhythm and access timing difference, thereby breaking the limitations of traditional static testing methods in abnormal excitation coverage. By dynamically collecting the response behavior of the chip under disturbance, and conducting attribution analysis of fluctuation amplitude, control anomaly and aggregation distribution, the system can establish a chip stability evaluation mechanism that is more consistent with the actual running state, and timely capture potential sub-stable operation risks, providing a disturbance level identification with resolution for the subsequent deep recognition module. Secondly, combined with the feature coupling graph construction module, the system breaks through the traditional defect detection method which only relies on single-dimensional electrical characteristics. By performing density mapping, abnormal position heat map generation and jump structure channel mapping on multiple behavior response parameters in the spatial structure, a coupling graph is formed, which is visualized for defects, has clear structure and logic, and has the ability to express physical meaning. Further, with the help of homologous fusion and neighborhood coupling index mechanism, the saliency of abnormal features and the distinctiveness of recognition targets in the graph are effectively improved, enhancing the overall discrimination basis and data input quality of the system. Finally, through the three-layer recognition architecture of the micro-deviation anomaly extraction network, the system can identify the slight but representative defect behavior in the chip under high interference background. It improves the abnormal feature response degree in the graph through local block decomposition and edge enhancement preprocessing; through multi-window convolution perception and fine-grained feature comparison mechanism, the spatial features and heat focusing degree of the suspected area are accurately extracted; finally, the recognition results are standardized and encoded, forming an abnormal report data with positioning, typing and scoring, providing detailed basic information for the next stage of trend evaluation and control of the system. The combination of the three modules enables the system to detect potential defects online in real time without traditional lengthy testing, realizes the front-end early warning response and active chip protection mechanism for deployment side, and has high engineering adaptability and real-time reliability.
[0099] Embodiment 3
[0100] This embodiment is an explanation and description in Embodiment 1. Please refer to Figure 1 , in particular: the behavior evolution trend aggregation module includes a time series construction unit, a trend analysis modeling unit and a health level evaluation unit;
[0101] The time series construction unit is used to arrange the abnormal codes output by the abnormal code output unit in time sequence to form an abnormal evolution record sequence with frame level as granularity. The sequence contains defect types, spatial positions, activity levels and response variation information of each time slice, which is used to construct a defect trajectory path;
[0102] The trend analysis modeling unit is configured to identify whether the defect behavior presents repetitiveness, persistence or expansion, and the identification includes detecting the activation density of the abnormality in the time dimension, the position shift rate, the evolution linkage relationship between adjacent types of defects, and outputting a trend label: stable, transient and evolution;
[0103] The health level evaluation unit is configured to divide the current running state of the chip into three levels of "healthy", "warning" and "failure risk" according to the trend label result of the trend analysis modeling unit, and the evaluation standard includes a logical comparison of a defect activation frequency threshold, an evolution rate threshold and a coverage area proportion index, and the evaluation result is used to judge whether an intervention mechanism needs to be started and serves as an input basis of the control module.
[0104] The trend analysis modeling unit is configured to receive and process the abnormal behavior encoding sequence output by the time series construction unit, and perform trend attribution modeling according to multiple dimensions of defect activation frequency, position evolution behavior, defect category similarity and sustained response amplitude, so as to identify whether there is a potential defect behavior mode with persistent, evolutionary or aggregative characteristics in the current storage chip, including abnormal activity calculation, spatial position shift trajectory analysis and abnormal type variation analysis:
[0105] The abnormal activity calculation is configured to judge the activity degree of the abnormal event in the continuous time slice, judge whether it has a frequent activation trend, and realize the following judgment logic by counting the number of triggers of the same type of defect in a unit time window and the spatial coincidence degree:
[0106] If a certain type of abnormality repeatedly occurs in multiple frames and the confidence remains above a set threshold, it is considered as an "active abnormality";
[0107] If the abnormality has intermittent triggering between frames or the number of activations is significantly reduced, it is marked as a "transient abnormality";
[0108] The spatial position shift trajectory analysis identifies whether the abnormality presents stability or migration in the physical address space through the spatial coordinates of the defect region at multiple times:
[0109] If the abnormal position remains relatively stable in the time sequence, and the coordinate shift is less than the system error tolerance, it is marked as a "position fixed type abnormality";
[0110] If the defect presents continuous or jumping shift, and the direction presents the same trend, along the on-chip address line, it is judged as an "evolutionary expansion abnormality";
[0111] If multiple defects converge around a certain center region, it is considered as "aggregative evolution";
[0112] Abnormal type variation analysis is used to analyze whether the abnormal category in the time series is transformed, including gradually transforming from "response delay type" to "error correction frequent type", which indicates that the defect has evolved from the performance degradation stage to the functional error stage. The system compares according to the set abnormal category transformation mapping table, and if it meets the specific transformation path, the "category evolution label" is output.
[0113] The health level evaluation unit mainly includes three evaluation processes, which are respectively for judging the abnormal activity, evolution trend rate and spatial coverage breadth, and finally outputs the health level label by using the step-by-step rule fusion method:
[0114] Abnormal activity evaluation: used to count the active frequency of defect behavior per unit time, and judge whether it reaches the preset threshold:
[0115] If the abnormal activation frequency in the last 3 time windows is lower than the minimum activity threshold, the activation frequency per second is less than 2 times, and the "normal" is evaluated;
[0116] If the activation frequency is in the middle interval, 2-8 times per second, the "warning" is evaluated;
[0117] If the activation frequency is continuously higher than the maximum activity threshold, the activation frequency is greater than 8 times per second, or a burst of dense abnormal clusters appears, it enters "high risk"
[0118] This stage mainly represents whether the short-time behavior of the chip is in a sustained unstable state;
[0119] Evolution trend rate evaluation: used to measure the "deterioration speed" of the abnormality in the time dimension, that is, the growth rate of the defect confidence, distribution range or impact level per unit time:
[0120] If the defect intensity does not significantly increase or continuously decreases, it is defined as "trend stable", supporting "normal" judgment;
[0121] If there is intermittent rise, but the amplitude is limited, it is defined as "fluctuation evolution", supporting "warning";
[0122] If there is a continuous enhancement trend in multiple periods, and the rising amplitude is significantly more than the growth threshold of 30% or more, it is defined as "rapid evolution", and directly classified as "high risk";
[0123] This stage reflects whether the chip is currently in a potential "deterioration" situation;
[0124] Spatial coverage breadth evaluation: used to evaluate the distribution range and density of the identified abnormal region in the address space of the memory chip:
[0125] If the abnormal block covers an address range of less than 10% without cross-aggregation phenomenon, it is considered "locally controllable" and supports "normal";
[0126] If the coverage range is between 10% and 30%, or there is a step distribution of address segments, it is considered as "potential spread", supporting "early warning";
[0127] If the coverage ratio exceeds 30%, or multiple address clusters overlap to form a defect cluster, it is defined as "multi-point instability" and listed as "high risk";
[0128] This stage evaluates the "spread risk" and "page-level impact level" of the anomaly.
[0129] In this embodiment, the application sets up a behavior evolution trend aggregation module, so that the system can not only realize static identification of defects, but also can identify the dynamic evolution behavior of defects in the storage chip based on time series construction and trend modeling method, solving the problem that it is difficult to judge abnormal trend in the prior art only by single detection. By reorganizing the anomaly encoding result into an evolution track according to frame-level granularity, the system can perform trend attribution modeling from multiple dimensions such as defect activation frequency, spatial displacement path and category variation, accurately determine whether the current chip has a metastable state that tends to evolve, aggregate or persistently fluctuate. Further, the system constructs a health level evaluation mechanism, sets three evaluation dimensions of abnormal activity, evolution rate and spatial coverage breadth, and outputs clear health level labels combined with logic fusion rules, divided into "healthy", "warning" and "failure risk" three levels. This multi-dimensional and multi-stage dynamic evolution identification and grading judgment mechanism not only enhances the defect trend perception ability of the system, but also realizes the prospective early warning and hierarchical management of the chip behavior state, and has strong online running adaptability.
[0130] Embodiment 4
[0131] This embodiment is an explanation and description in embodiment 1, please refer to Figure 1 , specifically: the remote reasoning feedback control module includes a control strategy matching unit, an edge device response execution unit and a remote reporting and recording unit;
[0132] The control strategy matching unit is used to select a matching item from the built-in control strategy table according to the chip health level, defect type and spatial distribution characteristics output by the health level evaluation unit, and the control strategy covers: chip access mode adjustment, data channel switching, write rate limitation and video cache transfer;
[0133] The edge device response execution unit is used to convert the selected control strategy into bottom command instruction, and send it to the device system layer or chip control layer, to complete the actual action execution of chip working parameter adjustment, resource reallocation and power supply limitation adjustment through SPI, I2C or GPIO interface, to ensure that the defect risk is responded at the hardware level;
[0134] The remote reporting and recording unit is used for generating an execution log and a state change record of a current control action, and collecting a current health status of the chip and a response control result into structured reporting content, and uploading the structured reporting content to an operation and maintenance management end or a remote monitoring platform.
[0135] The multi-task adaptive scheduling module comprises a running resource monitoring unit and a detection scheduling decision unit.
[0136] The running resource monitoring unit is used for collecting current system load information of the device in real time, including a number of video stream channels, an AI model running state and a chip temperature power consumption, and judging available resources of the chip and a calculation pressure caused by the detection task, as a reference input for adjusting a scheduling strategy.
[0137] The detection scheduling decision unit is used for dynamically deciding strategy parameters of a running frequency, a detection accuracy and a spectrum update interval of each unit in the system according to a resource state judgment result of the running resource monitoring unit, realizing system-level task collaborative management, and scheduling results preferentially guaranteeing stability of recording and identification, and executing a chip detection behavior under a non-conflict premise, and automatically optimizing a scheduling strategy, and realizing adaptive frequency adjustment, self-recovery and low-interference operation of the detection task.
[0138] In the embodiment, through the remote reasoning feedback control module, a closed-loop mapping of a chip health status and a defect level judgment result to an actual control response is realized, and a closed-loop control chain from intelligent identification to physical intervention is completed. The system intelligently matches a control strategy according to a health evaluation label, supports triggering of multiple types of strategies including access path switching, rate limiting, channel transfer and cache management, and directly acts on a chip control layer and a system resource allocation unit through a control instruction issuing interface, so that a risk defect is suppressed by hardware-level intervention. Meanwhile, cooperating with a remote reporting mechanism, the system can log archive each control behavior and synchronize a running state, effectively supporting generation of a subsequent maintenance strategy and iteration optimization of a model. In the current technical system, most chip health detection methods only stay in an abnormality discovery stage, lack a linkage mechanism with a control layer, and cause a feedback lag or a human intervention dependence between identification and response. The scheme solves a problem that defect identification cannot be implemented by introducing a remote control linkage mechanism, especially in scenes such as high-definition video recording and night AI identification, which have very high requirements for stability, and can realize a whole-process operation path of “online detection-real-time response-result reporting”. Compared with a traditional technical scheme, the present application has significant advantages in intelligentization of a control strategy matching, rapidity of a response path and accessibility of an execution level.
[0139] Embodiment 5
[0140] A storage chip defect detection method based on deep learning, please refer to Figure 2 , specifically: comprising the following steps:
[0141] Step one: During the operation of the edge video device, read-write access sequences at multiple access frequencies are constructed to stimulate the response behavior of the memory chip, and the response delay, control path misalignment and error correction activation behavior during the access process are collected to establish dynamic interference indicators of the chip;
[0142] Step two: The above response behavior is integrated in the form of multi-dimensional parameters, and a graph structure covering access behavior, control path and error correction response is generated, and a coupled feature representation suitable for deep model input is formed;
[0143] Step three: Receive the above coupled feature graph, identify weak image features, regional abnormal structures or control response spots in the graph, output the location, type and confidence label of the defect area, and generate an abnormal record in an encoded manner;
[0144] Step four: Build a time series of abnormal records at multiple time points, and model the evolution trend based on the indicators of abnormal activity, location change trend and activation frequency, and output a trend function for health level assessment;
[0145] Step five: Receive the health level result and match it with the control strategy table to perform chip protection, working state adjustment or fault reporting operation, and realize remote adaptive control of the chip running state;
[0146] Step six: Monitor the device running load, dynamically adjust the execution frequency, resource allocation and reasoning accuracy of the system, so that the cooperative operation between the chip detection task and the video main task does not conflict.
[0147] In this embodiment, the storage chip defect detection system based on deep learning proposed by the application actively constructs multiple access modes during the running process of the edge video device, accurately stimulates the chip bottom layer response behavior, combines the multi-dimensional collection of non-characteristic signals such as access delay, control path dislocation and error correction activation, and realizes the dynamic characteristic quantization of the chip running state. On this basis, the atlas structure representation is constructed, and the potential physical or behavioral defect area in the storage chip is identified through the deep model, forming an abnormal record with space, type and intensity labels. Through this sequence of actions, the system completes the structured modeling and quantitative expression from the original response behavior to the defect identification result, providing accurate basis for subsequent trend judgment and risk decision. Compared with the prior art, the system not only has finer granularity defect detection capability, but also realizes time series modeling and trend attribution analysis of defect behavior. Through multi-dimensional identification of abnormal activity, spatial migration behavior and evolution path, the chip "healthy-warning-high risk" three-level state judgment logic is realized, and is automatically matched with the control strategy table to realize the closed-loop control of chip state adjustment, access mode switching and fault reporting. At the same time, the system introduces a resource perception and scheduling mechanism to ensure that the detection task is executed efficiently without affecting the stability of the main business (such as video recording, AI identification) of the device, forming a low-interference, high-adaptability running environment.
[0148] Although embodiments of the application have been shown and described, it is to be understood that the application is not limited to these embodiments. Since modifications, variations, and alterations are possible using the principles of the application in the light of the above teachings, it is therefore understood that changes can be made in the detail within the scope and range of equivalents of the embodiments of the application without departing from the scope of the application.
Claims
1. A memory chip defect detection system based on deep learning, characterized in that: It includes a heterogeneous frequency disturbance perception module, a feature coupling map construction module, a micro-deviation anomaly extraction network module, a behavior evolution trend aggregation module, a remote inference feedback control module, and a multi-task adaptation and scheduling module; The inter-frequency disturbance sensing module is used to construct read and write access sequences under multiple access frequencies during the operation of edge video devices to stimulate the response behavior of memory chips, and to collect response delay, control path misalignment and error correction activation behavior during the access process to establish dynamic interference indicators of the chip. The feature coupling graph construction module is used to integrate the above response behaviors in the form of multi-dimensional parameters, generate a graph structure covering access behavior, control path and error correction response, and fuse them to form a coupled feature representation suitable for deep model input; The micro-deviation anomaly extraction network module is used to receive the above coupled feature map, identify weak image features, regional abnormal structures or control response spots in the map, output the location, type and confidence marker of the defect area, and generate anomaly records in an encoded manner. The behavior evolution trend aggregation module is used to construct a time series from abnormal records at multiple time points, and to model the evolution trend based on indicators such as abnormal activity, location change trend and activation frequency, and output a trend function for health level assessment. The remote inference feedback control module is used to receive health level results and match them with the control strategy table, and perform chip protection, working status adjustment or fault reporting operations to realize remote adaptive control of chip operating status. The multi-task adaptation and scheduling module is used to monitor the operating load of the equipment and dynamically adjust the system's execution frequency, resource allocation, and inference accuracy to ensure that the collaborative operation between the chip detection task and the main video task does not conflict.
2. The memory chip defect detection system based on deep learning according to claim 1, characterized in that: The inter-frequency disturbance sensing module includes a timing disturbance generation unit, a response behavior acquisition unit, and a dynamic interference evaluation unit; The timing disturbance generation unit is used to simulate the memory chip access scenarios of edge video devices under various load environments. Based on the device operation scenario, including high-definition video recording, nighttime static monitoring or intelligent playback, it dynamically generates a set of preset disturbance access sequences. The sequences include different read / write combination methods, address transition rules and access interval time differences. The response behavior acquisition unit is used to connect to the execution interface backend of the timing disturbance sequence, and to monitor and record various response information generated by the chip during access in real time, including the time difference of the memory controller returning the confirmation signal, the address resolution anomaly ratio, the frequency of error correction module activation, cache hit status and receipt misalignment, and archive them in a structured timestamp and behavior tag format. The dynamic interference evaluation unit is used to perform attribution and normalization evaluation of the collected response behavior sequence. The processing method is as follows: according to different access types, the response fluctuation amplitude, instruction delay range, and control anomaly point aggregation density are statistically analyzed, and the evaluation results are converted into a composite disturbance level identifier that characterizes the chip stability. This information is then passed to the next module to determine whether the chip is in a potentially metastable state.
3. The memory chip defect detection system based on deep learning according to claim 1, characterized in that: The feature coupling graph construction module includes a behavior parameter integration unit, a graph mapping generation unit, and a coupling feature fusion unit; The behavior parameter integration unit is used to receive composite interference level identifiers and multiple access response behavior indicators, and to standardize and rearrange data items from different sources and at different scales, including: unifying the time scale, converting behavior frequency into distribution density, and mapping response anomaly rate to equal quantile intervals for graph processing; The graph mapping generation unit is used to integrate the output through the behavior parameter integration unit and perform spatial mapping operations on various parameters. The mapping forms include: access response delay to form a two-dimensional density matrix, error correction trigger position to generate a heat overlay, address jump structure to form a channel drift graph. The graph adopts a fixed-size encoding structure to adapt to the input format of subsequent depth models. The coupling feature fusion unit is used to perform multi-channel merging operations on various maps according to data homogeneity and functional coupling. The merging operation is as follows: weighted overlay of different map layers in the same region, establishing correlation indexes between behavioral responses of physically adjacent regions, and assigning regional labels to significant structural features in the map, forming a unified feature coupling map.
4. The memory chip defect detection system based on deep learning according to claim 1, characterized in that: The micro-deviation anomaly extraction network includes a local identification preprocessing unit, a deep micro-deviation identification unit, and an anomaly coding output unit; The local recognition preprocessing unit is used to preprocess the input map by perceptual region division, including image noise removal, scale unification, patch cropping and edge enhancement. The output map is encoded in the form of a region block sequence, which facilitates the extraction of local anomaly features. The deep differential recognition unit is used to identify anomalies in the input map by deploying a deep convolutional network structure, such as subtle gray-level differences, edge distortion patterns, and hot spots in the control path. The recognition strategy is to extract local activation features under multiple convolutional receptive windows, perform feature comparison for each map patch and generate anomaly confidence level. The final output is: the location index, type label and activation score of the suspected defect area. The anomaly coding output unit is used to receive the identification results and generate standardized anomaly codes, which include: defect area number, center coordinate information, anomaly level, coverage page range, and detection timestamp.
5. The memory chip defect detection system based on deep learning according to claim 1, characterized in that: The behavioral evolution trend aggregation module includes a time series construction unit, a trend analysis modeling unit, and a health level assessment unit; The time series construction unit is used to organize the abnormal codes output by the abnormal coding output unit in chronological order to form an abnormal evolution record sequence with frame-level granularity. The sequence contains information on the defect type, spatial location, activity level and response variation of each time slice, which is used to construct the defect trajectory path. The trend analysis modeling unit is used to identify whether defect behavior is repetitive, persistent, or expansive. The identification methods include: detecting the activation density of anomalies in the time dimension, the location offset rate, and the evolutionary linkage between adjacent defect types, and outputting trend labels: stable, transient, and evolving. The health level assessment unit is used to classify the current operating status of the chip into three levels: "healthy", "warning" and "failure risk" based on the trend label results of the trend analysis modeling unit. The assessment criteria include: logical comparison of defect activation frequency threshold, evolution rate threshold and coverage area ratio index. The assessment results are used to determine whether an intervention mechanism needs to be initiated and serve as the input basis for the control module.
6. The memory chip defect detection system based on deep learning according to claim 5, characterized in that: The trend analysis modeling unit receives and processes the abnormal behavior encoding sequence output by the time series construction unit. It performs trend attribution modeling based on multiple dimensions, including defect activation frequency, location evolution behavior, defect category similarity, and sustained response amplitude. This identifies whether the current memory chip exhibits potential defect behavior patterns with persistent, evolving, or clustered characteristics. The modeling includes: abnormal activity calculation, spatial location offset trajectory analysis, and abnormal type variation analysis. Anomaly activity calculation is used to determine the activity level of anomaly events within consecutive time slices, and to determine whether they have a frequent activation trend. It achieves the following judgment logic by statistically analyzing the number of triggers of the same type of defect within a unit time window and their spatial overlap: If a certain type of anomaly occurs repeatedly in multiple frames and the confidence level remains above a set threshold, it is considered an "active anomaly". If the anomaly is triggered intermittently between frames, or the number of activations is significantly reduced, it is marked as a "transient anomaly"; Spatial location offset trajectory analysis identifies whether anomalies exhibit stability or mobility in the physical address space by using the spatial coordinates of the defect area at multiple time points. If the anomaly location remains relatively stable in the time series and the coordinate offset is less than the system error tolerance, it is marked as a "fixed location anomaly". If the defect exhibits continuous or abrupt shifts, and the direction shows the same trend, growing along the on-chip address lines, it is judged as an "evolutionary expansion anomaly". If multiple defects gradually converge around a central region, it is considered "aggregative evolution"; Anomaly type variation analysis is used to analyze whether the anomaly category in the time series has changed, including the gradual change from "response delay type" to "error correction frequency type". Such variation indicates that the defect has evolved from the performance degradation stage to the functional error stage. The system compares and judges according to the set anomaly category transformation mapping table. If it meets the specific transformation path, it outputs the "category evolution label".
7. The memory chip defect detection system based on deep learning according to claim 5, characterized in that: The health level assessment unit mainly comprises three assessment processes, which respectively determine abnormal activity level, evolutionary trend rate, and spatial coverage. A hierarchical rule fusion method is used to ultimately output a health level label. Abnormal activity assessment: Used to statistically analyze the frequency of defective behavior within a unit of time and determine whether it reaches a preset threshold. If the frequency of abnormal activations is lower than the minimum activity threshold within three consecutive time windows, with less than 2 activations per second, it is rated as "normal". If the activation frequency is in the medium range, 2 to 8 times / second, it is rated as "warning". If the activation frequency consistently exceeds the maximum activity threshold (>8 times / second), or if a sudden surge of dense abnormal clusters occurs, the system enters a "high-risk" state. This stage primarily characterizes whether the chip's short-term behavior is in a state of continuous instability. Evolutionary trend rate assessment: Used to measure the "deterioration rate" of an anomaly over time, i.e., the rate at which the defect confidence, distribution range, or impact level increases per unit time. If the defect strength does not increase significantly or continues to decrease, it is defined as "stabilizing" and supports the "normal" judgment. If there are intermittent increases, but the magnitude is limited, it is defined as "fluctuation evolution" and supports "early warning"; If there is a continuous strengthening trend in multiple consecutive cycles, and the increase significantly exceeds the growth threshold by more than 30%, it is defined as "rapid evolution" and directly classified as "high risk". This stage reflects whether the chip is currently in a potentially "deteriorating" state; Spatial coverage breadth assessment: Used to evaluate the distribution range and density of identified anomalous regions within the memory chip address space. If the address range covered by the abnormal block is less than 10% and there is no cross-clustering phenomenon, it is considered "locally controllable" and supports "normal". If the coverage area is between 10% and 30%, or if there is a jumpy distribution of address ranges, it is considered "potential diffusion" and supports "early warning". If the coverage ratio exceeds 30%, or if multiple address clusters overlap and defective clusters appear, it is defined as "multi-point instability" and classified as "high risk". This stage assesses the "diffusion risk" and "page-level impact level" of the anomaly.
8. The memory chip defect detection system based on deep learning according to claim 1, characterized in that: The remote inference feedback control module includes a control strategy matching unit, an edge device response execution unit, and a remote reporting and recording unit; The control strategy matching unit is used to select matching items from the built-in control strategy table based on the chip health level, defect type and its spatial distribution characteristics output by the health level assessment unit. The control strategies cover: chip access mode adjustment, data channel switching, write rate limit and video buffer transfer. The edge device response execution unit is used to convert the selected control strategy into low-level command instructions and send them to the device system layer or chip control layer. It performs the actual actions of adjusting chip operating parameters, reallocating resources, and adjusting power supply limits through SPI, I2C or GPIO interfaces to ensure that defects and risks are responded to at the hardware level. The remote reporting and recording unit is used to generate execution logs and status change records of the current control actions, and summarizes the chip's current health status and response control results into structured reporting content, which is then uploaded to the operation and maintenance management terminal or remote monitoring platform.
9. The memory chip defect detection system based on deep learning according to claim 1, characterized in that: The multi-task adaptation and scheduling module includes a runtime resource monitoring unit and a detection and scheduling decision unit; The operation resource monitoring unit is used to collect the current system load information of the device in real time, including the number of video stream channels, the running status of the AI model and the chip temperature and power consumption, and to determine the available chip resources and the computing pressure brought by the detection task, as a reference input for adjusting the scheduling strategy; The detection scheduling decision unit is used to dynamically determine the operating frequency, detection accuracy, and map update interval of each unit in the system based on the resource status judgment results of the operating resource monitoring unit. This enables system-level task collaborative management, prioritizes the stability of recording and recognition, executes chip detection behavior without conflict, and automatically optimizes the scheduling strategy to achieve adaptive frequency adjustment, self-recovery, and low-interference operation of the detection task.
10. A method for detecting defects in memory chips based on deep learning, applied to the memory chip defect detection system based on deep learning as described in any one of claims 1 to 9, characterized in that: Includes the following steps: Step 1: During the operation of the edge video device, construct read and write access sequences under various access frequencies to stimulate the response behavior of the memory chip, and collect response delay, control path misalignment and error correction activation behavior during the access process to establish chip dynamic interference indicators. Step 2: Integrate the above response behaviors in the form of multi-dimensional parameters, generate a graph structure covering access behavior, control path and error correction response, and fuse them to form a coupled feature representation suitable for deep model input; Step 3: Receive the above coupled feature map, identify weak image features, regional abnormal structures or control response spots in the map, output the location, type and confidence marker of the defect area, and generate anomaly records in an encoded manner; Step 4: Construct anomaly records from multiple time points into a time series, and model the evolutionary trend based on indicators such as abnormal activity, location change trend, and activation frequency, outputting a trend function for health level assessment; Step 5: Receive the health level results and match them with the control strategy table, then perform chip protection, working status adjustment or fault reporting operations to achieve remote adaptive control of the chip's operating status. Step Six: Monitor the equipment's operating load and dynamically adjust the system's execution frequency, resource allocation, and inference accuracy to ensure that the chip detection task and the main video task do not conflict with each other.
Citation Information
Cited By
Cooperative control system for intelligent production line of non-woven fabric packaging bag
CN121455104A
Elevator safety dual-prevention mechanism operation effect evaluation method
CN121672298A