Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

448results about "Reliability/availability analysis" patented technology

Systems and methods for detailed cloud posture remediation recommendations utilizing custom Large Language Models (LLMs)

Systems and methods for detailed cloud posture remediation recommendations utilizing custom Large Language Models (LLMs). The present systems and methods are configured to perform the steps of scanning a cloud environment for posture control data; generating one or more alerts related to any of risky configurations and risky activities associated with the cloud environment; generating one or more remediation recommendations based on the one or more alerts; and providing the one or more alerts and the one or more remediation recommendations to administrators of the cloud environment.
Owner:ZSCALER INC

Method and system for predicting server hardware and server hardware component failures

A system for predicting system component failures within a computer system that comprises a plurality of hardware components and a plurality of software components. The system may comprise memory storing instructions that, when executed, cause a processor to: obtain performance metrics by monitoring a network interface of the computer system; generate component failure probabilities by processing the performance metrics; determine that a first component failure probability among the component failure probabilities exceeds a risk threshold; determine remedial actions that mitigate a first component failure probability; and mitigate the first component failure probability by initiating an execution of the remedial actions.
Owner:JPMORGAN CHASE BANK NA

Storage chip high and low temperature aging test chamber fault self-diagnosis system and method and computer equipment

The invention relates to the technical field of fault diagnosis, in particular to a fault self-diagnosis system and method for a high and low temperature aging test chamber of a storage chip and computer equipment. The temperature control abnormal index is obtained through the data of the temperature control system sensor, early warning can be performed when the temperature control system has abnormal fluctuation, the performance evaluation basis of the air circulation system is provided by the air system efficiency coefficient, the heating efficiency and the refrigeration coefficient are obtained through the data of the heating and refrigerating system sensor, and the performance of the heating and refrigerating system is evaluated. An electrical stability index is obtained through data of a test system sensor, the operation stability of an electrical system is evaluated, the stability of the electrical system is detected in time, and a fault probability index is obtained through a system entropy increase rate, the electrical stability index, a thermal coupling factor, a wind system efficiency coefficient and a temperature control abnormal index. Therefore, the stability and the reliability of the equipment can be improved, the maintenance cost can be reduced, the use efficiency of the equipment is improved, and the accuracy and the safety of test results are ensured.
Owner:SHANGHAI QITAI FENHUA SEMICON TECH CO LTD

Apparatus and method for data fault detection and repair

An apparatus for data fault detection and repair is disclosed. The apparatus comprises at least a processor and a memory communicatively connected to the at least a processor. The memory instructs the processor to receive a user profile relating to a user, wherein the user profile comprises at least provider data of a user. The memory instructs the processor to generate practitioner data as a function of the user profile. The memory additionally instructs the processor to retrieve remittance data as a function of the practitioner data. The memory then instructs the processor to identify a data fault in at least one of the practitioner data and the remittance data. The memory instructs the processor to initiate a data correction action based on the identified data fault.
Owner:EMERGIP LLC

Interactive data processing system failure management using hidden knowledge from predictive models

Methods and systems for managing data processing systems are disclosed. A data processing system may include and depend on the operation of hardware and / or software components. Inference models may be implemented to predict future system infrastructure outcomes (e.g., component failures) using information recorded in logs that reflect the operation of the components. However, the models may be complex “black boxes” and may generate critical outcome predictions for downstream consumers without explanations of how the predictions are determined, resulting in downstream consumers having low confidence in the predictions. Therefore, hidden knowledge (e.g., structured knowledge attributes) of the models may be extracted and / or used to understand the underlying processes that the models use to predict the system infrastructure outcomes. The hidden knowledge may be provided for interactively managing data processing system(s) failures in order to increase the likelihood of preventing and / or mitigating future data processing system failures.
Owner:DELL PROD LP

Industrial equipment health state analysis method

The invention discloses an industrial equipment health state analysis method, and relates to the technical field of industrial equipment health monitoring, and the method comprises the steps: obtaining vibration data, lubricating oil pollution image data and real-time motor current data of an automobile factory punching machine; generating a dynamic weight coefficient according to the ratio of the real-time motor current data to the rated current of the equipment; performing weighted summation to calculate a real-time health score; constructing a wear-fatigue coupling model to calculate and output an alarm prediction L value; and when the alarm prediction L value exceeds a set threshold value, generating an equipment maintenance instruction and adjusting equipment operation parameters. According to the method, through multi-physics field coupling modeling, dynamic weight distribution and closed-loop self-optimization mechanisms, the precision and real-time performance problems of equipment health state analysis in an automobile manufacturing scene are solved, the method is particularly suitable for high-load equipment such as a punching machine tool and a welding robot, and compared with a traditional method, the comprehensive operation and maintenance cost is reduced, and the non-planed downtime is shortened.
Owner:CHONGQING CREATION VOCATIONAL COLLEGE

Cluster reliability test method and system based on fault simulation

The invention provides a cluster reliability test method based on fault simulation, which comprises the following steps: acquiring injection parameter information, and performing fault simulation in a cluster based on the injection parameter information and a fault transfer model to obtain a complex fault scene; key performance index records are obtained in real time based on the complex fault scene; executing a reliability detection step based on a preset reliability detection model and the key performance index record, and obtaining a data analysis report; and obtaining a reliability score, an influence analysis result and a risk prediction result based on a data analysis report, a weighted scoring algorithm, an anomaly detection algorithm and linear regression analysis, and completing reliability detection of the cluster. According to the cluster reliability test method and system based on fault simulation provided by the invention, the test efficiency and the reliability detection level are greatly improved, the fault simulation is carried out in the cluster and the reliability detection model is combined, so that the accurate quantification of the reliability detection result is realized, and the reliability detection efficiency and the detection capability are improved.
Owner:CHINA SOUTHERN POWER GRID DIGITAL GRID GRP CO LTD

Fault early-warning method and apparatus for heterogeneous hard disk system

PCT designated stage expiredWO2025129877A1Reliability/availability analysisEnergy efficient computing
The present application relates to the technical field of device detection. Provided are a fault early-warning method and apparatus for a heterogeneous hard disk system. The method comprises: obtaining hard disk state attribute data of hard disks of different model numbers in a heterogeneous hard disk system, and performing cluster grouping on the hard disk state attribute data on the basis of the model numbers of the hard disks and a distribution difference of the hard disk state attribute data, so as to determine a hard disk cluster to which the hard disk state attribute data belongs (101); inputting data corresponding to each hard disk cluster into a hard disk fault prediction model of a multi-tower structure for abnormality detection processing, so as to obtain hard disk health indicator information, wherein the hard disk fault prediction model is obtained by means of performing training on the basis of sample hard disk state attribute data and hard disk health label information corresponding to the sample hard disk state attribute data (102); and performing fault early-warning on the heterogeneous hard disk system on the basis of the hard disk health indicator information (103).
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Extension of network control system into public cloud

Some embodiments provide a method for a first data compute node (DCN) operating in a public datacenter. The method receives an encryption rule from a centralized network controller. The method determines that the network encryption rule requires encryption of packets between second and third DCNs operating in the public datacenter. The method requests a first key from a secure key storage. Upon receipt of the first key, the method uses the first key and additional parameters to generate second and third keys. The method distributes the second key to the second DCN and the third key to the third DCN in the public datacenter.
Owner:VMWARE INC

Artificial-intelligence-assisted error prediction in integration processes

Conventional error detection for integration processes in an integration platform are inefficient and require significant expertise. Accordingly, an error prediction model is disclosed. The error prediction model may be operated to produce error predictions, based on the current design (e.g., lineage) of an integration process, during construction of that integration process (e.g., on a virtual canvas). A generative language model may also be used to provide the error predictions in natural language. This enables the efficient troubleshooting and resolution of errors in an integration process, prior to that integration process being deployed and executed, and without requiring significant expertise.
Owner:BOOMI LP

Communication data verification method and device, electronic equipment and storage medium

The invention discloses a communication data verification method and device, electronic equipment and a storage medium, and relates to the technical field of communication, and the method comprises the steps: after executing a write-in operation on a target terminal, immediately initiating a read-in operation, and through the design of writing first and reading second, immediately checking whether the written-in communication data is correctly received and processed after the data is sent, thereby improving the verification efficiency. According to the method and the device, the communication verification format is improved, and the write operation is seamlessly switched to the read operation under the condition of keeping the communication continuity, so that the reliability of data transmission is improved, and the response speed and the performance of the server are optimized. Therefore, the problems that a verification mechanism is incomplete during data transmission, and the reliability and accuracy of data transmission are poor in the prior art are solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Training and using a memory failure prediction model

The disclosure herein describes training and using an uncorrectable error (UE) state prediction model based on telemetry error data. Sets of UE state labels and non-UE state labels are generated from a first set of collected telemetry data, wherein the UE state labels each reference a UE and telemetry data of an interval prior to the referenced UE. Statistical features are extracted from telemetry data of the sets of UE state labels and non-UE state labels, and the extracted statistical features are used to train a UE state prediction model. A second set of collected telemetry data is obtained, and a UE event is predicted based on the second set of collected telemetry data using the trained UE state prediction model. A preventative operation is performed on a memory page of the system based on the predicted UE event, whereby the predicted UE event is prevented from occurring.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Solid state disk fault intelligent prediction system

The invention discloses an intelligent fault prediction system for a solid state disk, relates to the technical field of solid state disks, solves the problems that only the test condition of a single chip can be identified, and the test effect of comprehensive radiation cannot be achieved, and is based on the number of threads corresponding to the solid state disk, and different processing processes are executed. The method comprises the following steps: processing the solid state disk in different combination states, confirming numerical value characteristics associated with different processing processes, analyzing reading speed fluctuation states of the solid state disk in different combination states according to the confirmed numerical value characteristics, carrying out multidirectional characteristic testing according to a specific analysis process, and identifying the use performance of the solid state disk. According to the multi-azimuth and multi-chip testing method, the multi-azimuth and multi-chip testing method and the multi-azimuth and multi-chip testing system, the characteristic data associated in the use performance is subjected to interval range confirmation, whether the fault hidden danger exists or not is confirmed according to the change amplitude of the corresponding interval range, the comprehensiveness in the fault hidden danger testing process can be effectively guaranteed by adopting the multi-azimuth and multi-chip testing mode, and a better testing processing effect is achieved.
Owner:DONGGUAN LIJING TECH CO LTD

Context based health visualization system and method for a computing environment

Embodiments of the present disclosure provide a context based health visualization system and method for a computing environment that indicates context-based health information for the computing resources of a computing environment. According to one embodiment, an Information Handling System (IHS) includes computer-executable instructions to identify a maintenance task that needs to be performed on a computing resource, classify the maintenance task according to one of a plurality of health contexts, and generate, using information associated with the maintenance task, a health context score for the one health context. The maintenance task may be one that impacts an overall health score of the computing resource. The instructions may then cause the IHS to display the health context score for view by a user.
Owner:DELL PROD LP

Server aging test method and device, electronic equipment and storage medium

The invention relates to the technical field of server aging testing, in particular to a server aging testing method and device, electronic equipment and a storage medium. Then convolution processing is carried out on each aging data queue in a sliding mode through a difference operator, the obtained multiple first convolution queues are segmented one by one, queue segments obtained through segmentation are normalized, and multiple queue segments are obtained; performing dot product calculation on each queue segment, each same-time-segment queue segment and each same-part queue segment to obtain a plurality of dot product results, and adding representative values determined according to the plurality of dot product results into a test analysis matrix; and finally, according to the plurality of row vectors of the test analysis matrix, determining the abnormal probability of the plurality of tested parts of the server. According to the method, the state of the detected component is determined based on the cross point product and probability estimation, so that the problem component can be found early.
Owner:XIONGAN BAIXIN INFORMATION TECHNOLOGY CO LTD

Predictively Addressing Hardware Component Failures

The present invention extends to methods, systems, and computer program products for predictively addressing hardware component failures. Network packets can be received over time at a platform. Metrics derived from platform hardware components and derived from one or more workloads utilizing the platform hardware components can be monitored. Model training data can be formulated from the metrics. A health check model can be trained using the model training data. The health check model can be executed to compute a probability that a monitored platform hardware component is on a path to failure. It can be determined that the probability exceeds a threshold. A workload can be relocated from a pod containing the monitored platform hardware component to another pod. Additional network packets can be received over time at the platform. The workload can process data contained in the additional network packets at the other pod.
Owner:RAKUTEN SYMPHONY INC

System and method for predicting processing errors in a computing system

After generating a first customized recovery plan to resolve a first error associated with a first error message generated by a first software application, a jobs manager obtains identities of an input system, an output system or a combination thereof associated with the first software application. In response to determining that a second software application is associated with the same or similar input system, the same or similar output system, or the combination thereof as the first software application, the jobs manager determines that the first error associated with the first software application is predicted to occur relating to the second software application. Thereafter, the jobs manager generates a second customized recovery plan for the second software application based at least in part upon the first customized recovery plan generated for the first software application.
Owner:BANK OF AMERICA CORP

Systems and methods for detailed cloud posture remediation recommendations utilizing custom large language models (LLMs)

Systems and methods for detailed cloud posture remediation recommendations utilizing custom Large Language Models (LLMs). The present systems and methods are configured to perform the steps of scanning a cloud environment for posture control data; generating one or more alerts related to any of risky configurations and risky activities associated with the cloud environment; generating one or more remediation recommendations based on the one or more alerts; and providing the one or more alerts and the one or more remediation recommendations to administrators of the cloud environment.
Owner:ZSCALER INC

Fan life prediction method and system, fan, storage medium and program product

The embodiment of the invention provides a fan service life prediction method and system, a fan, a storage medium and a program product. The fan service life prediction method comprises the steps that the environment temperature of the fan is obtained; if the environment temperature of the fan meets the first condition, the mechanical factor service life and the environment factor service life of the fan are calculated; the first condition is that the environment temperature is smaller than or equal to first preset temperature or larger than or equal to second preset temperature, and the first preset temperature is smaller than the second preset temperature; and outputting the residual life of the fan after fusing the mechanical factor life and the environmental factor life. By fusing the mechanical factor life and the environmental factor life of the fan, the prediction precision of the fan life is improved, the cost does not need to be increased, and the residual life of the fan can be output in real time; and secondly, by monitoring the residual service life of the fan, the effects of good thermal management performance of the equipment and prolonging of the service life of the equipment can be achieved.
Owner:EMERSON NETWORK POWER CO LTD

Composite risk score for cloud software deployments

The techniques described herein provide a risk assessment framework that enhances the functionality of software deployment systems in cloud-based platforms. Generally described, the present techniques evaluate and consolidate various risk scores to classify a given computing cluster within a software deployment strategy. In various examples, a deployment system collects node-level feature data from the computing cluster to generate a dataset to train a prediction model to calculate constituent risk scores. In another aspect, the deployment system aggregates constituent risk scores to determine an overall risk of software failure. Likewise, the deployment system considers diverse criteria such as virtual machine size and virtual machine density to determine an overall impact of software deployment failure. The deployment system then calculates a composite risk score for the computing cluster as a function of the risk of software deployment failure and the impact of software deployment failure.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Automated large-scale failure mode effects analysis system

A computer-implemented method for conducting failure mode effects analysis (FMEA) at a large scale includes generating a set of failure modes. The method includes determining an impact score for one of the set of failure modes. The method includes determining a probability score for one of the set of failure modes. The method includes determining a detectability score for one of the set of failure modes. The method includes calculating risk priority scores for the failure modes based on the impact score, the probability score, and the detectability score. The method includes ranking the failure modes according to the calculated risk priority scores. The method includes generating a report including the ranked failure modes.
Owner:EXPRESS SCRIPTS STRATEGIC DEVELOPMENT INC

Guaranteeing online services based on predicting failures of storage devices

The present disclosure describes techniques for guaranteeing online services based on predicting failures of storage devices. Statistical data may be extracted on a regular basis by each of a plurality of storage devices. Each of the plurality of storage devices may comprise a set of NAND dies. Each of the set of NAND dies may be configured to measure and track a set of metrics indicating characteristics of each NAND die. Prediction data indicating potential failures of the plurality of storage devices may be generated. The prediction data may be shared with a host on a periodic basis. A strategy of decommissioning an aged storage device and adding a new storage device based on the prediction data may be created by the host. The data migration to the new storage device may be implemented.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD +1

Compatibility and reliability testing method for solid state disk

The invention relates to the technical field of data storage, and discloses a solid state disk compatibility and reliability testing method which comprises the steps of initializing a testing environment, installing an operating system, compiling and executing a testing script and configuring a compiling environment. Basic information of the solid state disk is recognized and classified, an automatic script is written, the classified basic information of the solid state disk is created into file clusters of different sizes according to the automatic script, continuous / random reading and writing are conducted, and file integrity is verified; and performing performance test on the classified basic information of the solid state disk by using a Fio tool and recording performance indexes of the solid state disk, wherein the performance indexes comprise sequential read-write speed, random IOPS and response delay, compiling an automatic script, and simulating a database load to perform small block random read-write operation. According to the invention, a plurality of modules for basic function and compatibility test, performance test, power management and reliability test, safety characteristic test and the like are integrated into a complete and automatic test scheme.
Owner:SHENZHEN JINGCUN TECH CO LTD

PCBA aging test method and system

The invention discloses a PCBA aging test method and system, and belongs to the technical field of aging test. The method specifically comprises the following steps: S1, historical data acquisition: acquiring PCBA aging related historical data, including but not limited to voltage parameters, current parameters, core area temperature, operation duration, environment humidity, fault types and fault occurrence time during PCBA operation; s2, preprocessing the historical data: preprocessing the historical data, including removing abnormal values, filling missing values and normalizing; 24-72-hour long-time power-on loading of a traditional aging test is not needed, the test period of a single PCBA is shortened to the minute level, the production takt is greatly improved, the large-scale batch production requirement is met, and aging fault recognition is more comprehensive; and recessive aging precursors such as capacitance attenuation and welding spot microcracks can be captured, and the sudden failure risk after the product leaves the factory is reduced.
Owner:ZHUHAI QILI ELECTRONICS CO LTD

Internet-based performance prediction system for software development

The invention relates to the technical field of computer system performance evaluation and fault prediction, and particularly discloses an internet-based performance prediction system for software development. The system comprises a data acquisition module, a feature extraction module, a performance prediction model module, a feedback optimization module and a visual analysis module. The system constructs a high-dimensional feature vector set fusing code complexity, a calling relation, an abnormal mode and the like by collecting source codes, logs and defect data, and outputs response time, a resource occupancy rate and a reliability index based on a multi-target prediction model. And further optimizing the model performance through an error feedback mechanism, and displaying a prediction result and a risk score in a graphic mode. The system supports cloud deployment, has cross-project sharing capability and data security guarantee, and is suitable for large software project performance bottleneck diagnosis and risk early warning.
Owner:HUAIAN ZHENGYI TECHNOLOGY CO LTD