A Backup Sampling Anticipation Production Data Loss Model and Method
By introducing a backup sampling and prediction production data loss model in data backup, the problem that traditional backup methods are difficult to adapt to abnormal conditions of production equipment and lack of personalized adjustments is solved, and more accurate data loss prediction and operation and maintenance resource optimization is achieved, and the reliability of the system is enhanced.
Patent Information
- Application Number
- CN202411826355.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-12
AI Technical Summary
Traditional data backup methods are difficult to adapt to the abnormal situations that may occur in the production process of production equipment, resulting in incomplete or loss of backup data, and lack the ability to dynamically adjust the personalized needs of different production equipment.
It provides a backup sampling and prejudice production data loss model, including a backup data acquisition module, a data loss behavior representation module, an associated sampling processing module and an operation and maintenance center module. It generates a prejudice list by predividing the backup time segments, portraying the backup behavior simulation rule curve, performing associated sampling of data loss behavior and evaluating the predicted probability.
It improves the accuracy of data loss prediction, helps operation and maintenance personnel to identify production equipment that may be at risk of data loss in advance, optimizes operation and maintenance resource allocation, reduces the impact of data loss on the production system, and enhances the reliability of the system.
Smart Images

Figure CN119759655B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data backup, and specifically provides a model and method for predicting production data loss in backup sampling applications. Background Art
[0002] In modern industrial production, data backup is an important means to ensure the security of production data and business continuity. However, with the continuous expansion of production scale and the sharp increase in the amount of production data, the problem of data loss during the data backup process has become increasingly prominent, posing a significant challenge to the production and operation of enterprises.
[0003] Traditional data backup methods often rely on fixed backup strategies and cycles. Due to performance differences between production devices, there is a lack of dynamic adjustment ability for the actual operating status of production devices. Such static backup strategies are difficult to adapt to various abnormal situations that may occur during the production process of production devices, such as equipment failures and network interruptions, resulting in incomplete or lost backup data. Moreover, due to differences in the amount of data, data change rate, and data importance generated by different production devices during the production process, using a unified backup strategy often fails to meet the personalized needs of different devices, exacerbating the risk of data loss.
[0004] To address the above problems, existing technologies generally attempt to establish an association model between data backup and data loss by collecting and analyzing various data generated by production devices during operation, so as to achieve the prediction and early warning of data loss risks. However, in actual applications, only the backup behavior of a single production device is considered, ignoring the possible mutual influences and correlation relationships between production devices, which is not conducive to effectively predicting or warning of the data loss risks of production devices. Summary of the Invention
[0005] The purpose of the present invention is to provide a model and method for predicting production data loss in backup sampling applications to solve the problems raised in the above background art.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] A model for predicting production data loss in backup sampling, which includes a backup data collection module, a data loss behavior characterization module, an associated sampling processing module, and an operation and maintenance center module connected in sequence.
[0008] The backup data collection module is used to pre-divide backup time segments and obtain the data loss ratio generated by each production device during backup production data within the backup time segments.
[0009] The data loss behavior characterization module characterizes the backup behavior simulation regular curve of the production equipment based on the backup time segment and the data loss ratio, and converts the backup behavior simulation regular curve into a backup behavior digital regular straight line;
[0010] The associated sampling processing module is used to arbitrarily select a production equipment as a prediction reference object, arbitrarily select another production equipment as a prediction associated object except the prediction reference object, and perform associated sampling of the data loss behavior between the prediction reference object and the prediction associated object;
[0011] The operation and maintenance center module evaluates the prediction probability between the prediction reference object and the prediction associated object based on the result of the associated sampling to generate a prediction list and output it to the operation and maintenance port.
[0012] Further, the backup data acquisition module includes a backup encoding unit and a data acquisition unit connected in sequence;
[0013] The backup encoding unit is used to uniformly divide the backup time segments for the time within a day and perform unified encoding on each backup time segment; it is also used to perform unified encoding on the production equipment;
[0014] The data acquisition unit is used to record the backup time nodes when the production equipment backs up the production data, and identify the backup time segments to which the backup time nodes belong. The data loss ratio is the ratio of the capacity value of the lost data packets to the total capacity value of the data packets when backing up the production data.
[0015] Further, the data loss behavior characterization module includes a punctuation unit and a feature transformation unit connected in sequence;
[0016] The punctuation unit is used to establish a two-dimensional coordinate system for the backup behavior, use the backup time segment as the independent variable of the abscissa of the two-dimensional coordinate system for the backup behavior, use the data loss ratio as the dependent variable of the ordinate of the two-dimensional coordinate system for the backup behavior, and then perform punctuation on the two-dimensional coordinate system for the backup behavior;
[0017] The feature transformation unit is used to smoothly connect in sequence each punctuation formed by the production equipment during backup to obtain the backup behavior simulation regular curve of the production equipment; it is also used to capture the punctuation of each peak and valley in the backup behavior simulation regular curve, and convert the backup behavior simulation regular curve into a backup behavior digital regular straight line.
[0018] Further, the associated sampling processing module includes an object selection unit and an associated sampling analysis unit connected in sequence;
[0019] The object selection unit is used to select a production device as the pre-judgment reference object, arbitrarily select a production device other than the pre-judgment reference object as the pre-judgment associated object, and associate the digital law straight line of the backup behavior corresponding to the pre-judgment associated object; it is also used to uniformly number each horizontal straight line in the digital law straight line of the backup behavior, and the horizontal straight line includes a high-level straight line or a low-level straight line.
[0020] The associated sampling and analysis unit is used to set a logical judgment function. If the horizontal straight line is a high-level straight line, the logical judgment function is set to 1. If the horizontal straight line is a low-level straight line, the logical judgment function is set to 0. Based on the logical judgment function, the backup time segment interval for associated sampling is analyzed and obtained.
[0021] Further, the operation and maintenance center module includes an associated sampling and statistics unit and a pre-judgment list generation unit that are sequentially connected in series.
[0022] The associated sampling and statistics unit is used to count all the abscissa intervals obtained from the associated sampling analysis between the pre-judgment reference object and the pre-judgment associated object, and generate an associated sampling set.
[0023] The pre-judgment list generation unit is used to evaluate the pre-judgment probability between the pre-judgment reference object and the pre-judgment associated object, generate a pre-judgment list, and send it to the operation and maintenance port.
[0024] A method for predicting the loss of production data in backup sampling is applied. This method includes the following steps:
[0025] Step S1: Pre-divide the backup time segment, and obtain the data loss ratio generated by each production device during the backup of production data within the backup time segment.
[0026] Step S2: Based on the backup time segment and the data loss ratio, depict the simulation law curve of the backup behavior of the production device, and convert the simulation law curve of the backup behavior into a digital law straight line of the backup behavior.
[0027] Step S3: Arbitrarily select a production device as the pre-judgment reference object, arbitrarily select a production device other than the pre-judgment reference object as the pre-judgment associated object, and conduct associated sampling of the data loss behavior between the pre-judgment reference object and the pre-judgment associated object.
[0028] Step S4: Based on the results of the associated sampling, evaluate the pre-judgment probability between the pre-judgment reference object and the pre-judgment associated object, generate a pre-judgment list, and output it to the operation and maintenance port.
[0029] Further, the specific implementation process of step S1 includes:
[0030] Divide the backup time segments evenly within a day, uniformly encode each backup time segment, and denote the \(i\)-th backup time segment as \(t\). i Uniformly encode the production equipment, and denote \(e\) production equipment as \(P\). e ;
[0031] Record the backup time nodes when the production equipment \(P\) e backs up production data, identify the backup time segment \(t\) to which the backup time node belongs i , and denote the data loss ratio when the production equipment \(P\) e backs up production data within the backup time segment \(t\) as \(CV(P\) i , \(t\) e , \(t\) i ). The data loss ratio is the ratio of the capacity value of the lost data packets to the total capacity value of the data packets when backing up production data.
[0032] Furthermore, the specific implementation process of step S2 includes:
[0033] Establish a two-dimensional coordinate system for backup behavior, use the backup time segment as the independent variable of the abscissa of the two-dimensional coordinate system for backup behavior, and use the data loss ratio as the dependent variable of the ordinate of the two-dimensional coordinate system for backup behavior. Then, punctuate on the two-dimensional coordinate system for backup behavior, and denote the punctuation corresponding to the data loss ratio \(CV(P\) e , \(t\) i ) as \([t\) i , \(CV(P\) e , \(t\) i )];
[0034] Sequentially and smoothly connect the punctuations formed by the production equipment \(P\) e during backup to obtain the backup behavior simulation regular curve of the production equipment \(P\) e , denoted as \(AL(P\) e );
[0035] Capture the punctuations of each peak and valley in the backup behavior simulation regular curve \(AL(P\) e ), and transform the backup behavior simulation regular curve \(AL(P\) e ) into a backup behavior digital regular straight line \(DL(P\) e ). The transformation process is as follows:
[0036] If the next punctuation of the peak punctuation is a valley punctuation, then fit the segment of the backup behavior simulation regular curve between the peak punctuation and the valley punctuation into a low-level straight line. The ordinate value of the low-level straight line is the ordinate value corresponding to the peak punctuation, the starting point of the abscissa interval of the low-level straight line is the abscissa value corresponding to the peak punctuation, and the ending point of the abscissa interval of the low-level straight line is the abscissa value corresponding to the valley punctuation;
[0037] If the next punctuation after the valley punctuation is a peak punctuation, then the segment of the backup behavior simulation regular curve between the valley punctuation and the peak punctuation is fitted to a high-level straight line. The ordinate value of the high-level straight line is the ordinate value corresponding to the valley punctuation, the starting point of the abscissa interval of the high-level straight line is the abscissa value corresponding to the valley punctuation, and the ending point of the abscissa interval of the high-level straight line is the abscissa value corresponding to the peak punctuation;
[0038] In the backup behavior simulation regular curve AL(P e ), alternately fit high or low-level straight lines in the alternating order of peak punctuation and valley punctuation to obtain the backup behavior digital regular straight line DL(P e );
[0039] In the above method, different from the design principle of analog-to-digital conversion, the low-level straight line in the present invention represents the growth law of the degree of data loss behavior, while the high-level straight line represents the decreasing law of the degree of data loss behavior.
[0040] Further, the specific implementation process of step S3 includes:
[0041] Taking the production equipment P e as the prediction reference object, arbitrarily select the r-th production equipment P e except the production equipment P r as the prediction correlation object, and correlate the backup behavior digital regular straight line DL(P r ) generated by the corresponding production equipment P r ;
[0042] Unify the numbers of each horizontal straight line in the backup behavior digital regular straight line DL(P e ). The horizontal straight lines include high-level straight lines or low-level straight lines. Denote any x-th horizontal straight line as dl x (P e );
[0043] Capture each horizontal straight line in the backup behavior digital regular straight line DL(P x ) within the abscissa interval of the horizontal straight line dl e (P r ), and generate an associated sample set, denoted as S[dl x (P e )] = {dl y (P r )|y ∈ [1, Y]}, where dl y (P r ) represents the y-th horizontal straight line in the backup behavior digital regular straight line DL(P r ), and Y represents the backup behavior digital regular straight line DL(Pr ) The total number of horizontal lines in
[0044] Set up a logical judgment function f(α), where α represents a horizontal line. If the horizontal line α is a high-level horizontal line, then let f(α) = 1; if the horizontal line α is a low-level horizontal line, then let f(α) = 0.
[0045] Conduct associated sampling for data loss behaviors:
[0046] If f[dl x (P e )] × f[dl y (P r )] = 1, then extract the abscissa interval of the horizontal line dl x (P e ), denoted as T x = [t a , t b , where t a represents the a-th backup time segment, and t b represents the b-th backup time segment.
[0047] If f[dl x (P e )] × f[dl y (P r )] = 0, then do not extract the abscissa interval of the horizontal line dl x (P e ).
[0048] In the above method, a logical judgment function is set up to quickly capture the backup time segments with the same data loss behavior pattern, that is, the backup time segments with the data loss behavior pattern of the high-level line state.
[0049] Furthermore, the specific implementation process of step S4 includes:
[0050] Statistically analyze all the abscissa intervals obtained from the associated sampling analysis between the backup behavior digital rule line DL(P e ) and the backup behavior digital rule line DL(P r ), and generate an associated sampling set, denoted as as(P e , P r ) = {T x |x ∈ [1, β]}, where β represents the total number of horizontal lines in the backup behavior digital rule line DL(P e ).
[0051] Taking the production equipment P e as the pre-judgment reference object, evaluate the pre-judgment probability when the production equipment P e is associated with the production equipment P r Among them, NUM[as(P e , P r )] represents the total number of abscissa intervals included in the associated sampling set as(P e , P r ), represents the total number of abscissa intervals included in the set obtained after performing the union operation between the associated sampling sets, and R represents the total number of production devices;
[0052] Sort the production devices as the predicted associated objects in descending order of the predicted probability, and generate a prediction list of production device P e denoted as list(P e ). When a data loss behavior occurs in production device P e , send the prediction list list(P e ) to the operation and maintenance port.
[0053] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: In an application for a backup sampling prediction production data loss model and method provided by the present invention, by pre-dividing the backup time segments and obtaining the data loss ratios of each production device within the backup time segments, the backup behavior characteristics of the production devices can be accurately captured, and the backup behavior simulation regular curve can be transformed into a digital regular straight line, making the characterization of the data loss behavior more intuitive and quantitative, so as to improve the accuracy of data loss prediction;
[0054] Perform associated sampling on the data loss behaviors between the predicted reference object and the predicted associated object, and evaluate the predicted probability based on the results of the associated sampling, which helps the operation and maintenance personnel to identify in advance the production devices that may have data loss risks, thereby optimizing the allocation of operation and maintenance resources, concentrating the limited operation and maintenance resources on high-risk devices, and improving the operation and maintenance efficiency;
[0055] When a data loss behavior occurs in a production device, the operation and maintenance personnel can quickly refer to the prediction list and take corresponding preventive or remedial measures, thereby effectively avoiding or reducing the impact of data loss on the production system and enhancing the reliability of the system;
[0056] The method and model of the present invention are not only applicable to predicting production data loss, but also can provide a scientific basis for formulating data backup strategies; through in-depth analysis of the backup behaviors of production devices, the operation and maintenance personnel can formulate data backup plans more scientifically, reasonably arrange the backup time and frequency, and ensure the security and integrity of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.
[0058] Figure 1 It is a schematic diagram of the steps of a method for predicting production data loss by backup sampling in the present invention. Specific implementation manner
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] In the first embodiment: A model for predicting production data loss by backup sampling is provided. The model includes a backup data acquisition module, a data loss behavior characterization module, an associated sampling processing module, and an operation and maintenance center module that are sequentially connected in series;
[0061] The backup data acquisition module is used to pre-divide backup time segments and obtain the data loss ratio generated when each production device backs up production data within the backup time segments;
[0062] Among them, the backup data acquisition module includes a backup encoding unit and a data acquisition unit that are sequentially connected in series;
[0063] The backup encoding unit is used to uniformly divide the time within a day into backup time segments and uniformly encode each backup time segment; it is also used to uniformly encode production devices;
[0064] The data acquisition unit is used to record the backup time nodes when production devices back up production data and identify the backup time segments to which the backup time nodes belong. The data loss ratio is the ratio of the capacity value of the lost data packets to the total capacity value of the data packets when backing up production data;
[0065] The data loss behavior characterization module, based on the backup time segments and the data loss ratio, depicts the simulation law curve of the backup behavior of the production device and converts the backup behavior simulation law curve into a digital law straight line of the backup behavior;
[0066] Among them, the data loss behavior characterization module includes a punctuation unit and a feature conversion unit that are sequentially connected in series;
[0067] The punctuation unit is used to establish a two-dimensional coordinate system for the backup behavior, use the backup time segment as the independent variable of the abscissa of the two-dimensional coordinate system for the backup behavior, use the data loss ratio as the dependent variable of the ordinate of the two-dimensional coordinate system for the backup behavior, and then punctuate on the two-dimensional coordinate system for the backup behavior;
[0068] The feature transformation unit is used to sequentially and smoothly connect each punctuation mark formed by the production equipment during backup to obtain the backup behavior simulation law curve of the production equipment; it is also used to capture the punctuation marks of each peak and valley in the backup behavior simulation law curve, and transform the backup behavior simulation law curve into a backup behavior digital law straight line;
[0069] The associated sampling processing module is used to arbitrarily select a production equipment as a pre-judgment reference object, and arbitrarily select a production equipment as a pre-judgment associated object except the pre-judgment reference object, and perform associated sampling of data loss behavior between the pre-judgment reference object and the pre-judgment associated object;
[0070] Among them, the associated sampling processing module includes an object selection unit and an associated sampling analysis unit that are sequentially connected;
[0071] The object selection unit is used to select a production equipment as a pre-judgment reference object, arbitrarily select a production equipment as a pre-judgment associated object except the pre-judgment reference object, and associate the backup behavior digital law straight line corresponding to the pre-judgment associated object; it is also used to uniformly number each horizontal straight line in the backup behavior digital law straight line, and the horizontal straight line includes a high-level straight line or a low-level straight line;
[0072] The associated sampling analysis unit is used to set a logical judgment function. If the horizontal straight line is a high-level straight line, the logical judgment function is set to 1. If the horizontal straight line is a low-level straight line, the logical judgment function is set to 0. Based on the logical judgment function, the backup time segment interval for associated sampling is analyzed;
[0073] The operation and maintenance center module, based on the results of associated sampling, evaluates the pre-judgment probability between the pre-judgment reference object and the pre-judgment associated object to generate a pre-judgment list and output it to the operation and maintenance port;
[0074] Among them, the operation and maintenance center module includes an associated sampling statistics unit and a pre-judgment list generation unit that are sequentially connected;
[0075] The associated sampling statistics unit is used to count all the abscissa intervals obtained from the associated sampling analysis between the pre-judgment reference object and the pre-judgment associated object, and generate an associated sampling set;
[0076] The pre-judgment list generation unit is used to evaluate the pre-judgment probability between the pre-judgment reference object and the pre-judgment associated object, and generate a pre-judgment list and send it to the operation and maintenance port.
[0077] Please refer to Figure 1 , in the second embodiment: A method for predicting data loss during backup sampling of production data is provided. The method includes the following steps:
[0078] Step S1: Pre-divide the backup time segments, and obtain the data loss ratio generated when each production device backs up production data within the backup time segments;
[0079] Exemplarily, evenly divide the time within a day into backup time segments, and uniformly encode each backup time segment. Denote the i-th backup time segment as t i ; Uniformly encode the production devices, and denote e production devices as P e ;
[0080] Record the backup time nodes when the production device P e backs up production data, identify the backup time segment t i to which the backup time node belongs, and denote the data loss ratio generated when the production device P e backs up production data within the backup time segment t i as CV(P e , t i ). The data loss ratio is the ratio of the capacity value of the lost data packets to the total capacity value of the data packets when backing up production data;
[0081] Step S2: Based on the backup time segments and the data loss ratio, characterize the backup behavior simulation law curve of the production device, and convert the backup behavior simulation law curve into a backup behavior digital law straight line;
[0082] Exemplarily, establish a two-dimensional coordinate system for backup behavior, use the backup time segment as the independent variable of the abscissa of the two-dimensional coordinate system for backup behavior, and use the data loss ratio as the dependent variable of the ordinate of the two-dimensional coordinate system for backup behavior. Then, punctuate on the two-dimensional coordinate system for backup behavior, and denote the punctuation corresponding to the data loss ratio CV(P e , t i ) as [t i , CV(P e , t i )];
[0083] Smoothly connect in sequence the punctuations formed by the production device P e during backup to obtain the backup behavior simulation law curve of the production device P e , denoted as AL(P e );
[0084] Capture the punctuations of the peaks and valleys in the backup behavior simulation law curve AL(P e ), and convert the backup behavior simulation law curve AL(P e ) into a backup behavior digital law straight line DL(P e ). The conversion process is as follows:
[0085] If the next punctuation after the peak punctuation is a trough punctuation, then the segment of the backup behavior simulation regular curve between the peak punctuation and the trough punctuation is fitted to a low-level straight line, the vertical coordinate value of the low-level straight line is the vertical coordinate value corresponding to the peak punctuation, the starting point of the horizontal coordinate interval of the low-level straight line is the horizontal coordinate value corresponding to the peak punctuation, and the ending point of the horizontal coordinate interval of the low-level straight line is the horizontal coordinate value corresponding to the trough punctuation;
[0086] If the next punctuation after the trough punctuation is a peak punctuation, then the segment of the backup behavior simulation regular curve between the trough punctuation and the peak punctuation is fitted to a high-level straight line, the vertical coordinate value of the high-level straight line is the vertical coordinate value corresponding to the trough punctuation, the starting point of the horizontal coordinate interval of the high-level straight line is the horizontal coordinate value corresponding to the trough punctuation, and the ending point of the horizontal coordinate interval of the high-level straight line is the horizontal coordinate value corresponding to the peak punctuation;
[0087] In the backup behavior simulation regular curve AL(P e ), the high or low-level straight lines are alternately fitted in the alternating order of the peak punctuation and the trough punctuation to obtain the backup behavior digital regular straight line DL(P e );
[0088] Step S3: Arbitrarily select a production device as the prediction reference object, and arbitrarily select a production device as the prediction associated object except the prediction reference object, and perform associated sampling of the data loss behavior between the prediction reference object and the prediction associated object;
[0089] Exemplarily, taking the production device P e as the prediction reference object, and arbitrarily selecting the r-th production device P e except the production device P r as the prediction associated object, and associating the backup behavior digital regular straight line DL(P r ) generated by the production device P r ;
[0090] Unify the numbering of each horizontal straight line in the backup behavior digital regular straight line DL(P e ), the horizontal straight lines include high-level straight lines or low-level straight lines, and denote any x-th horizontal straight line as dl x (P e );
[0091] Capture each horizontal straight line in the backup behavior digital regular straight line DL(P x ) within the horizontal coordinate interval of the horizontal straight line dl e (P r ), and generate an associated sample set, denoted as S[dl x (P e )] = {dl y (P r)|y ∈ [1, Y]}, where dl y (P r ) represents the y-th horizontal line in the backup behavior digital rule straight line DL(P r ), and Y represents the total number of horizontal lines in the backup behavior digital rule straight line DL(P r );
[0092] Set the logical judgment function f(α), where α represents the horizontal line. If the horizontal line α is a high-level horizontal line, then let f(α) = 1; if the horizontal line α is a low-level horizontal line, then let f(α) = 0;
[0093] Perform associated sampling for data loss behavior:
[0094] If f[dl x (P e )] × f[dl y (P r )] = 1, then extract the abscissa interval of the horizontal line dl x (P e ), denoted as T x = [t a , t b , where t a represents the a-th backup time segment, and t b represents the b-th backup time segment;
[0095] If f[dl x (P e )] × f[dl y (P r )] = 0, then do not extract the abscissa interval of the horizontal line dl x (P e );
[0096] Step S4: Based on the results of associated sampling, evaluate the prediction probability between the prediction reference object and the prediction associated object to generate a prediction list and output it to the operation and maintenance port;
[0097] Exemplarily, count all the abscissa intervals obtained from the associated sampling analysis between the backup behavior digital rule straight line DL(P e ) and the backup behavior digital rule straight line DL(P r ), and generate an associated sampling set, denoted as as(P e , P r ) = {T x |x ∈ [1, β]}, where β represents the total number of horizontal lines in the backup behavior digital rule straight line DL(P e );
[0098] For the production equipment P eAs a pre-judgment reference object, evaluate production equipment P e Associated production equipment P r The pre-judgment probability when Among them, NUM[as(P e , P r )] represents the total number of abscissa intervals included in the associated sampling set as(P e , P r ), Represents the set obtained by performing the union operation between associated sampling sets The total number of abscissa intervals included in it, and R represents the total number of production equipment;
[0099] Sort the production equipment as the pre-judgment associated object in descending order of the pre-judgment probability, and generate the pre-judgment list of production equipment P e Denoted as list(P e ). When data loss behavior occurs in production equipment P e , send the pre-judgment list list(P e ) to the operation and maintenance port.
[0100] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0101] Finally, it should be noted that: the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting production data loss applied to backup sampling, characterized in that: The method comprises the following steps: Step S1: pre-dividing the backup time segment, and obtaining the data loss ratio generated when each production device backs up production data within the backup time segment; Step S2: based on the backup time segment and the data loss ratio, characterize the backup behavior simulation regularity curve of the production equipment, and convert the backup behavior simulation regularity curve into a backup behavior digital regularity straight line; Step S3: randomly selecting a production device as a prediction reference object, randomly selecting a production device other than the prediction reference object as a prediction associated object, and performing association sampling of data loss behaviors between the prediction reference object and the prediction associated object; Step S4: Based on the result of the association sampling, the prediction probability between the prediction reference object and the prediction associated object is evaluated to generate a prediction list, and output it to the operation and maintenance port; The specific implementation process of step S1 includes: The time in a day is evenly divided into backup time segments, and each backup time segment is uniformly coded. The i-th backup time segment is recorded as t i ; Unify the production equipment and record e production equipment as P e ; Record production equipment P e Back up the backup time node of the production data and identify the backup time segment t to which the backup time node belongs i , the production equipment P e In the backup time segment t i The data loss ratio generated when backing up production data is recorded as CV (P e , t i ), the data loss ratio is the ratio of the data packet capacity value lost when backing up production data to the total data packet capacity value; In step S2: A two-dimensional coordinate system of backup behavior is established, and the backup time segment is used as the abscissa independent variable of the two-dimensional coordinate system of backup behavior, and the data loss ratio is used as the ordinate dependent variable of the two-dimensional coordinate system of backup behavior. Then, points are placed on the two-dimensional coordinate system of backup behavior, and the data loss ratio CV(P e , t i ) The corresponding punctuation is denoted as [t i , CV(P e , t i )]; Smoothly connect the production equipment P in sequence e The various punctuations formed during the backup process obtain the production equipment P e The backup behavior simulation regularity curve is denoted as AL(P e ); Capture the backup behavior simulation regularity curve AL(P e ) and the backup behavior simulation regular curve AL(P e ) is transformed into the backup behavior digital law straight line DL(P e ).
2. According to claim 1, a method for predicting production data loss applied to backup sampling is characterized in that: The backup behavior simulation law curve AL(P e ) is transformed into the backup behavior digital law straight line DL(P e ), the conversion process is as follows: If the next punctuation point of the peak punctuation point is a trough punctuation point, the backup behavior simulation regular curve segment between the peak punctuation point and the trough punctuation point is fitted into a low horizontal straight line, the ordinate value of the low horizontal straight line is the ordinate value corresponding to the peak punctuation point, the starting point of the abscissa interval of the low horizontal straight line is the abscissa value corresponding to the peak punctuation point, and the end point of the abscissa interval of the low horizontal straight line is the abscissa value corresponding to the trough punctuation point; If the next punctuation point of the trough punctuation point is a peak punctuation point, the backup behavior simulation regular curve segment between the trough punctuation point and the peak punctuation point is fitted into a high horizontal straight line, the ordinate value of the high horizontal straight line is the ordinate value corresponding to the trough punctuation point, the starting point of the abscissa interval of the high horizontal straight line is the abscissa value corresponding to the trough punctuation point, and the end point of the abscissa interval of the high horizontal straight line is the abscissa value corresponding to the peak punctuation point; In the backup behavior simulation law curve AL(P e ) in the alternating order of the peak punctuation points and the trough punctuation points, and the backup behavior digital regularity straight line DL (P e ).
3. The method for predicting production data loss by backup sampling according to claim 2 is characterized in that: The specific implementation process of step S3 includes: Production equipment P e In order to predict the reference object, in addition to the production equipment P e Randomly select the rth production equipment P r As the predicted associated object, the associated production equipment P r The corresponding backup behavior digital regularity straight line DL (P r ); The digital law straight line DL(P e ) are uniformly numbered, and the horizontal straight lines include high horizontal straight lines or low horizontal straight lines, and any x-th horizontal straight line is denoted as dl x (P e ); Snap horizontal line dl x (P e ) Backup behavior digital regularity straight line DL(P r ) and generate the associated sample set, denoted as S[dl x (P e )]={dl y (P r )|y∈[1,Y]}, where dl y (P r ) represents the digital law straight line DL(P r ) is the y-th horizontal line in the image, where Y represents the digital regularity line of backup behavior DL(P r ) The total number of horizontal lines in Set a logic judgment function f(α), where α represents a horizontal straight line. If the horizontal straight line α is a high-level straight line, f(α)=1; if the horizontal straight line α is a low-level straight line, f(α)=0; Perform correlation sampling of data loss behavior: If f[dl x (P e )]×f[dl y (P r )]=1, then extract the horizontal line dl x (P e ), denoted as T x =[t a , t b ], where t a represents the ath backup time segment, t b Indicates the bth backup time segment; If f[dl x (P e )]×f[dl y (P r )]=0, then the horizontal line dl is not extracted x (P e )’s horizontal axis interval.
4. The method for predicting production data loss by backup sampling according to claim 3 is characterized in that: The specific implementation process of step S4 includes: Statistical backup behavior digital law straight line DL (P e ) and the backup behavior digital law straight line DL(P r ) to obtain all the horizontal axis intervals obtained by performing associated sampling analysis, and generate an associated sampling set, recorded as (P e , P r )={T x |x∈[1,β]}, β represents the digital law straight line DL(P e ) The total number of horizontal lines in Production equipment P e To predict the reference object, evaluate the production equipment P e Related production equipment r The predicted probability Among them, NUM[as(P e , P r )] represents the associated sampling set as(P e , P r ), Represents the set obtained by performing a union operation between associated sampling sets The total number of horizontal axis intervals included in , R represents the total number of production equipment; The production equipment as prediction related objects are sorted in descending order of prediction probability, and the production equipment P is generated. e The prediction list is recorded as list(P e ), when production equipment P e When data loss occurs, send a prediction list (P e ) to the operation and maintenance port.
5. A model for predicting production data loss applied to backup sampling, executing a method for predicting production data loss applied to backup sampling as claimed in any one of claims 1 to 4, characterized in that: The model includes a backup data acquisition module, a data loss behavior characterization module, a correlation sampling processing module and an operation and maintenance center module which are sequentially connected; The backup data acquisition module is used to pre-divide the backup time segments and obtain the data loss ratio generated when each production device backs up production data within the backup time segment; The data loss behavior characterization module describes the backup behavior simulation regularity curve of the production equipment based on the backup time segment and the data loss ratio, and converts the backup behavior simulation regularity curve into a backup behavior digital regularity straight line; The associated sampling processing module is used to arbitrarily select a production device as a prejudgment reference object, arbitrarily select a production device other than the prejudgment reference object as a prejudgment associated object, and perform associated sampling of data loss behavior between the prejudgment reference object and the prejudgment associated object; The operation and maintenance center module evaluates the prediction probability between the prediction reference object and the prediction associated object based on the result of the associated sampling to generate a prediction list and output it to the operation and maintenance port.
6. The model for predicting production data loss for backup sampling according to claim 5 is characterized in that: The backup data acquisition module includes a backup encoding unit and a data acquisition unit connected in sequence; The backup coding unit is used to evenly divide the time in a day into backup time segments and uniformly code each backup time segment; it is also used to uniformly code production equipment; The data acquisition unit is used to record the backup time node of the production equipment to back up the production data, and identify the backup time segment to which the backup time node belongs. The data loss ratio is the ratio of the data packet capacity value lost when backing up the production data to the total data packet capacity value.
7. The model for predicting production data loss for backup sampling according to claim 6 is characterized in that: The data loss behavior characterization module includes a punctuation unit and a feature conversion unit connected in sequence; The punctuation unit is used to establish a backup behavior two-dimensional coordinate system, take the backup time segment as the abscissa independent variable of the backup behavior two-dimensional coordinate system, take the data loss ratio as the ordinate dependent variable of the backup behavior two-dimensional coordinate system, and then perform punctuation on the backup behavior two-dimensional coordinate system; The feature conversion unit is used to smoothly connect the various punctuation points formed by the production equipment during backup in sequence to obtain the backup behavior simulation regular curve of the production equipment; it is also used to capture the punctuation points of each peak and trough in the backup behavior simulation regular curve, and convert the backup behavior simulation regular curve into a backup behavior digital regular straight line.
8. The model for predicting production data loss for backup sampling according to claim 7 is characterized in that: The associated sampling processing module includes an object selection unit and an associated sampling analysis unit connected in sequence; The object selection unit is used to select a production device as a prejudgment reference object, select any production device other than the prejudgment reference object as a prejudgment associated object, and associate the backup behavior digital regular straight line generated corresponding to the prejudgment associated object; and is also used to uniformly number each horizontal straight line in the backup behavior digital regular straight line, wherein the horizontal straight line includes a high horizontal straight line or a low horizontal straight line; The associated sampling analysis unit is used to set a logical judgment function. If the horizontal line is a high horizontal line, the logical judgment function is set to 1; if the horizontal line is a low horizontal line, the logical judgment function is set to 0. Based on the logical judgment function, the backup time segment interval used for associated sampling is analyzed.
9. The model for predicting production data loss for backup sampling according to claim 8 is characterized in that: The operation and maintenance center module includes an associated sampling statistics unit and a prediction list generation unit connected in sequence; The associated sampling statistical unit is used to count all horizontal coordinate intervals obtained by performing associated sampling analysis between the predicted reference object and the predicted associated object, and generate an associated sampling set; The prediction list generating unit is used to evaluate the prediction probability between the prediction reference object and the prediction associated object, and generate a prediction list, and send it to the operation and maintenance port.
Citation Information
Patent Citations
Game data processing method and device, electronic equipment and readable storage medium
CN115248749A
Cloud disaster recovery backup method and system for big data
CN117271222A