Inspection index sequence feature extraction method and device, storage medium and equipment
By normalizing and performing interval distance analysis on multiple test result sequences during a patient's medical visit, feature vectors are extracted, solving the feature extraction problem of multiple types of test results, improving the prediction accuracy of machine learning, and making it applicable to tumor diagnosis and treatment and other scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to effectively utilize the sequential relationships of multiple test results generated during a patient's visit, lacking the ability to extract information-rich feature vectors, which affects the accuracy of subsequent machine learning modeling.
This paper provides a feature extraction method for a test index sequence. Through normalization, interval distance analysis, and eigenvalue calculation, a feature vector of length λ is extracted. The method includes setting the dimension of the feature vector and determining the number of index categories. Different methods are used to calculate the eigenvalues for different dimensions.
It enables the extraction of feature vectors from multiple test index sequences, improving the accuracy of machine learning prediction tasks. It has good versatility and robustness, and is applicable to tumor diagnosis and treatment as well as other scenarios.
Smart Images

Figure CN115859076B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of laboratory testing technology, and more specifically, to a method, apparatus, storage medium, and device for feature extraction of a sequence of testing indicators. Background Technology
[0002] Laboratory testing determines the contents, properties, concentration, quantity, and other characteristics of a submitted substance through physical or chemical examinations performed in a laboratory. Because some laboratory tests, such as tumor marker testing, are conducted throughout a patient's medical history, from initial consultation and diagnosis to treatment, monitoring, and eventual discharge, multiple and varied test results are generated, and these results exhibit a sequential relationship. Extracting feature vectors rich in useful information from these test results would be of great help in subsequent machine learning modeling based on laboratory test results. Summary of the Invention
[0003] The purpose of this disclosure is to provide a method, apparatus, storage medium, and device for feature extraction of test index sequences, for extracting feature vectors from test numerical results that have a sequential relationship.
[0004] To achieve the above objectives, this disclosure provides a method for feature extraction of a test index sequence, including:
[0005] Obtain the original test index sequence, which includes multiple index values of different categories generated based on multiple tests;
[0006] Each index value in the original test index sequence is normalized to obtain a normalized index value, and a normalized index sequence is obtained based on the normalized index value.
[0007] Determine the set dimension λ of the feature vector and the number of index categories n in the normalized index sequence;
[0008] Based on different values of the interval distance, the index difference feature values are extracted from the normalized index values that satisfy the interval distance in the normalized index sequence to obtain the overall index difference feature value.
[0009] For any dimension from the first to the nth dimension in the feature vector, the feature value of the dimension is determined based on the index distribution frequency of the corresponding category in the normalized index sequence and the overall index difference feature value.
[0010] For any dimension from the (n+1)th to the λth dimension in the feature vector, determine the target value of the interval distance corresponding to the dimension, and determine the feature value of the dimension based on the indicator difference feature value corresponding to the interval distance under the target value and the overall indicator difference feature value.
[0011] The eigenvector is obtained based on the eigenvalues of dimensions 1 to λ.
[0012] Optionally, the step of extracting index difference feature values from the normalized index values that satisfy the interval distance in the normalized index sequence based on different values of the interval distance to obtain the overall index difference feature value includes:
[0013] Take values for the interval distance j from 1 to λ;
[0014] Based on the value of the interval distance j, i is taken from 1 to Lj respectively. For the i-th normalized index value in the normalized index sequence, the difference information between the i-th normalized index value and the (i+j)-th normalized index value is calculated to obtain Lj difference information, where L is the number of normalized index values in the normalized index sequence.
[0015] Based on the Lj difference information, determine the indicator difference characteristic value corresponding to the interval distance j under the given value;
[0016] The overall index difference characteristic value is obtained by summing the index difference characteristic values corresponding to interval distances ranging from 1 to λ.
[0017] Optionally, determining the index difference feature value corresponding to the interval distance j under the given value based on Lj difference information includes:
[0018] The indicator difference characteristic value corresponding to the interval distance j under the given value is calculated according to the following formula:
[0019]
[0020] Where, θ j N represents the characteristic value of the index difference corresponding to the interval distance j under the given value. i N represents the value of the i-th normalized index. i+j Represents the (i+j)th normalized index value, (N i -N i+j ) 2 This represents the difference between the i-th normalized index value and the (i+j)-th normalized index value.
[0021] Optionally, determining the feature value of any dimension from the first to the nth dimension in the feature vector, based on the index distribution frequency of the corresponding category in the normalized index sequence and the overall index difference feature value, includes:
[0022] For any dimension from the 1st to the nth dimension of the feature vector, the feature value of that dimension is determined by the following formula:
[0023]
[0024] Among them, X ε Let f represent the eigenvalue of the ε-th dimension from the 1st to the nth dimension, w represent the weight, and f ε Π represents the index distribution frequency of the ε-th dimension corresponding to the category in the normalized index sequence, and Π represents the overall index difference characteristic value.
[0025] Optionally, determining the target value of the interval distance corresponding to the dimension includes:
[0026] The difference between the dimension and the number of index categories n is used as the target value of the interval distance corresponding to the dimension.
[0027] Optionally, for any dimension from n+1 to λ in the feature vector, determining the target value of the interval distance corresponding to the dimension, and determining the feature value of the dimension based on the indicator difference feature value corresponding to the interval distance under the target value and the overall indicator difference feature value, includes:
[0028] For any dimension from the (n+1)th to the λth dimension in the feature vector, the feature value of that dimension is determined by the following formula:
[0029]
[0030] Among them, X ε Let w represent the eigenvalue of the ε-th dimension from the (n+1)-λ-th dimension, w represent the weight, ε-n represent the target value of the interval distance corresponding to the ε-th dimension, and θ represent the eigenvalue of the ε-th dimension. ε-n The interval distance represents the index difference characteristic value corresponding to the target value ε-n, and Π represents the overall index difference characteristic value.
[0031] Optionally, the normalization process for each index value in the original test index sequence to obtain a normalized index value includes:
[0032] Calculate the mean and standard deviation of the indicators in the original test indicator sequence;
[0033] For each index value in the original test index sequence, the index value is normalized according to the index mean and standard deviation to obtain the corresponding normalized index value.
[0034] This disclosure also provides a feature extraction device for a test index sequence, comprising:
[0035] The original sequence acquisition module is used to acquire the original test index sequence, which includes multiple index values of different categories generated based on multiple tests.
[0036] The normalized sequence acquisition module is used to normalize each index value in the original test index sequence to obtain a normalized index value, and obtain a normalized index sequence based on the normalized index value.
[0037] The information determination module is used to determine the set dimension λ of the feature vector and the number of index categories n in the normalized index sequence;
[0038] The difference feature extraction module is used to extract the difference feature values of the indicators from the normalized indicator sequence that satisfy the interval distance based on different values of the interval distance, so as to obtain the overall indicator difference feature value.
[0039] The first feature value determination module is used to determine the feature value of any dimension from the first to the nth dimension in the feature vector based on the index distribution frequency of the corresponding category in the normalized index sequence and the overall index difference feature value.
[0040] The second feature value determination module is used to determine the target value of the interval distance corresponding to any dimension from the (n+1)th to the λth dimension in the feature vector, and to determine the feature value of the dimension based on the indicator difference feature value corresponding to the interval distance under the target value and the overall indicator difference feature value.
[0041] The feature vector acquisition module is used to obtain the feature vector based on the feature values of the first to λth dimensions.
[0042] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the feature extraction method for the test index sequence provided in this disclosure.
[0043] This disclosure also provides an electronic device, including:
[0044] A memory on which computer programs are stored;
[0045] A processor is configured to execute the computer program in the memory to implement the feature extraction method for the test index sequence provided in this disclosure.
[0046] This technical solution addresses the feature extraction problem of multiple types and multiple tests generated during medical visits. It achieves the extraction of λ-dimensional feature vectors from the test indicator sequence, providing crucial feature vectors for subsequent machine learning prediction tasks or patient similarity calculations, thereby improving prediction accuracy. This technical solution has good versatility and robustness, and can be used, but is not limited to, tumor diagnosis and treatment scenarios. It is also applicable to test indicator sequences obtained in other scenarios.
[0047] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0048] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings:
[0049] Figure 1 A flowchart of a feature extraction method for a test index sequence in an exemplary embodiment is shown;
[0050] Figure 2 A flowchart illustrating a specific implementation of step S104 in an exemplary embodiment is shown;
[0051] Figure 3 A block diagram of a feature extraction apparatus for a test index sequence is shown in an exemplary embodiment;
[0052] Figure 4 A block diagram of an electronic device in an exemplary embodiment is shown. Detailed Implementation
[0053] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0054] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0055] In medicine, numerous indicators are used to measure various human body parameters, thereby assisting clinical research and treatment of diseases. For example, tumor markers exist in a patient's blood, body fluids, cells, or tissues. Due to their high specificity, high sensitivity, organ specificity, ability to reflect dynamic changes in tumors, detection of treatment effectiveness, recurrence, and metastasis, and ease of disease assessment, as well as the ability to monitor treatment outcomes, they are of great value in the auxiliary diagnosis, differential diagnosis, observation of treatment efficacy, monitoring of disease recurrence, and prognostic evaluation of tumors.
[0056] The testing of tumor markers is conducted throughout the entire process, from the patient's initial consultation, diagnosis, treatment, and monitoring until their final discharge from the hospital. During this period, multiple tests generate various marker values of different categories. These marker values have a sequential relationship, which is known as the test marker sequence.
[0057] This disclosure provides a method for feature extraction from a sequence of test indicators, used to extract feature vectors from the sequence of test indicators.
[0058] Figure 1 A flowchart of a feature extraction method for a test index sequence in an exemplary embodiment is shown, such as... Figure 1 As shown, the feature extraction method includes:
[0059] S101, Obtain the original test index sequence, which includes multiple index values of different categories generated based on multiple tests.
[0060] For example, the original test indicator sequence is a sequence of indicator values for different categories of tumor markers, and these indicator values are obtained based on multiple tests of the tumor marker indicators. Table 1 shows several common categories of tumor marker indicators, and the original test indicator sequence may include indicator values for the following categories of tumor marker indicators.
[0061] Table 1
[0062]
[0063]
[0064] S102, normalize each index value in the original test index sequence to obtain the normalized index value, and obtain the normalized index sequence based on the normalized index value.
[0065] As an example, the mean and standard deviation of the indicators in the original test indicator sequence are calculated. For each indicator value in the original test indicator sequence, the indicator value is normalized according to the mean and standard deviation to obtain the corresponding normalized indicator value. After obtaining the normalized indicator value for each indicator value, the normalized indicator sequence is obtained based on the normalized indicator value. It should be understood that the normalized indicator sequence, compared to the original test indicator sequence, only normalizes the numerical values of the indicator values; the order, number, and number of indicator categories of the indicator values remain unchanged.
[0066] S103, determine the set dimension λ of the feature vector and the number of index categories n in the normalized index sequence.
[0067] The dimension of the feature vector is preset. For example, if the value of λ is set to 10, it means that the final feature vector will have a length of 10 dimensions.
[0068] It should be noted that, Figure 1The order shown is not intended to limit the execution order of the steps in this embodiment. In an exemplary embodiment, step S103 can directly determine the number of index categories n from the original test index sequence, and therefore can be executed before step S102.
[0069] S104. Based on different values of the interval distance, extract the indicator difference feature values from the normalized indicator values that satisfy the interval distance in the normalized indicator sequence to obtain the overall indicator difference feature values.
[0070] As an example, the interval distance is taken multiple times. After each interval distance is taken, for every two normalized index values in the normalized index sequence that satisfy the interval distance, the corresponding index difference feature value is extracted. This index difference feature value is used to characterize the difference between the normalized index values that satisfy the interval distance. The above process is performed for different interval distance values to obtain the index difference feature values corresponding to the interval distance under different values. The index difference feature values corresponding to the interval distance under different values are summed to obtain the overall index difference feature value.
[0071] It is worth noting that in this technical solution, based on the normalized index sequence, the index difference feature value is extracted from two normalized index values that satisfy a set interval distance. Then, the size of the interval distance is changed by taking different values. Under some values of the interval distance, two normalized index values that satisfy the interval distance in the normalized index sequence may represent the numerical results of the same category of index at different test time points. Therefore, the index difference feature values obtained under different values are summed, and the final overall index difference feature value can be used to reflect the difference characteristics of the index between different test time points as a whole. In subsequent steps, based on the difference characteristics of the index between different test time points, i.e., the overall difference feature value, the feature values of each dimension in the feature vector are calculated, and thus the feature vector is obtained.
[0072] S105, for any dimension from the first to the nth dimension in the feature vector, determine the feature value of that dimension based on the index distribution frequency of the corresponding category in the normalized index sequence and the overall index difference feature value.
[0073] S106. For any dimension from the (n+1)th to the λth dimension in the feature vector, determine the target value of the interval distance corresponding to that dimension, and determine the feature value of that dimension based on the indicator difference feature value and the overall indicator difference feature value corresponding to the interval distance under the target value.
[0074] Understandably, the feature vector to be extracted consists of feature values of length λ. Specifically, the feature values for dimensions 1 to n and dimensions (n+1 to λ) are calculated using different methods. For dimensions 1 to n, the feature values are calculated using the frequency distribution of the corresponding category's indicators and the overall indicator difference feature value. Since the normalized indicator sequence contains only n categories of indicators, for the remaining dimensions (n+1 to λ), the feature values are calculated using the corresponding indicator difference feature values and the overall indicator difference feature value.
[0075] Understandably, the first to nth dimensions of the feature vector correspond to the n categories of the index values in the normalized index sequence.
[0076] Specifically, for any dimension from 1 to n in the feature vector, the feature value of that dimension is determined based on the indicator distribution frequency of the corresponding category in the normalized indicator sequence and the overall indicator difference feature value. The indicator distribution frequency represents the proportion of the number of indicators of the corresponding category in the normalized indicator sequence to the total number of indicators.
[0077] For example, the normalized index sequence includes 50 normalized index values, including 10 normalized index values for alpha-fetoprotein, 15 normalized index values for ferritin, 20 normalized index values for embryonic antigens, and 5 index values for cancer antigen 125, that is, the number of index categories n is 4.
[0078] For the first dimension of the feature vector, the index distribution frequency of the alpha-fetoprotein category corresponding to the first dimension is determined to be 1 / 5. Based on the index distribution frequency and the overall index difference feature value obtained earlier, the feature value of the first dimension is determined.
[0079] For the second dimension of the feature vector, the index distribution frequency of the ferritin category corresponding to the second dimension is determined to be 3 / 10. Based on this index distribution frequency and the overall index difference feature value obtained earlier, the feature value of the second dimension is determined.
[0080] For the third dimension in the feature vector, the index distribution frequency of the embryonic antigen category corresponding to the third dimension is determined to be 2 / 5. Based on the index distribution frequency and the overall index difference feature value obtained earlier, the feature value of the third dimension is determined.
[0081] For the fourth dimension in the feature vector, based on the indicator distribution frequency of the cancer antigen 125 category corresponding to the fourth dimension being 1 / 10, the feature value of the fourth dimension is determined according to the indicator distribution frequency and the overall indicator difference feature value obtained earlier.
[0082] Specifically, for any dimension from the (n+1)th to the λth dimension in the feature vector, it is necessary to determine the target value of the interval distance corresponding to that dimension, and determine the feature value of that dimension based on the indicator difference feature value and the overall indicator difference feature value corresponding to the interval distance under the target value.
[0083] As an example, the difference between this dimension and the number of indicator categories n can be used as the target value of the interval distance corresponding to this dimension.
[0084] For example, for the fifth dimension in the feature vector, the target value of the interval distance corresponding to this dimension is determined to be 1. Based on the indicator difference feature value corresponding to the interval distance when it is 1 and the overall indicator difference feature value obtained earlier, the feature value of the fifth dimension is determined.
[0085] For the 6th dimension in the feature vector, the target value of the interval distance corresponding to this dimension is determined to be 2. Based on the indicator difference feature value corresponding to the interval distance with a value of 2 and the overall indicator difference feature value obtained earlier, the feature value of the 6th dimension is determined.
[0086] For the 7th to 10th dimensions of the feature vector, the process of determining the feature values of the corresponding dimensions is described above and will not be repeated here.
[0087] S107, based on the eigenvalues of dimensions 1 to λ, obtain the eigenvectors.
[0088] Since the aforementioned steps have determined the eigenvalues of dimensions 1 to n and dimensions n+1 to λ in the eigenvector, the eigenvector can be obtained based on the eigenvalues of dimensions 1 to λ.
[0089] The above technical solution addresses the feature extraction problem of multiple types and multiple tests generated during medical visits. It achieves the extraction of λ-dimensional feature vectors from the test indicator sequence, providing crucial feature vectors for subsequent machine learning prediction tasks or patient similarity calculations, thereby improving prediction accuracy. This technical solution has good versatility and robustness, and can be used, but is not limited to, tumor diagnosis and treatment scenarios. It is also applicable to test indicator sequences obtained in other scenarios.
[0090] Figure 2 A flowchart illustrating a specific implementation of step S104 in an exemplary embodiment is shown, as follows: Figure 2 As shown, step S104 can obtain the overall index difference characteristic value through the following steps:
[0091] S201, where j takes values for intervals from 1 to λ.
[0092] Here, λ is the defined dimension of the feature vector. For example, if the value of λ is set to 10, then the interval distance j is determined to take values from 1 to 10.
[0093] S202, based on the value of the interval distance j, i is taken as 1 to Lj respectively. For the i-th normalized index value in the normalized index sequence, the difference information between the i-th normalized index value and the (i+j)-th normalized index value is calculated to obtain Lj difference information.
[0094] After assigning values to the interval distance j, based on this assigned value, i is then set from 1 to Lj. For the i-th normalized index value in the normalized index sequence, the difference between the i-th and (i+j)-th normalized index values is calculated. Since i ranges from 1 to Lj, a total of Lj difference values are obtained. Here, L represents the number of normalized index values in the normalized index sequence.
[0095] For example, the square of the difference between the i-th normalized index value and the (i+j)-th normalized index value can be used as the difference information between the two.
[0096] S203, based on the Lj difference information, obtain the indicator difference characteristic value corresponding to the interval distance j under this value.
[0097] For example, the indicator difference characteristic value corresponding to the interval distance j for this value can be calculated according to the following formula:
[0098]
[0099] Where, θ j N represents the characteristic value of the index difference corresponding to the interval distance j at this value. i N represents the value of the i-th normalized index. i+j Represents the (i+j)th normalized index value, (N i -N i+j ) 2 This represents the difference between the i-th normalized index value and the (i+j)-th normalized index value.
[0100] S204: Sum the indicator difference characteristic values corresponding to interval distances ranging from 1 to λ to obtain the overall indicator difference characteristic value.
[0101] After each interval distance j is assigned a value, the corresponding index difference characteristic value of the interval distance j under that value is obtained through steps S202 to S203. Then, the index difference characteristic values corresponding to the interval distances from 1 to λ are summed to obtain the overall index difference characteristic value.
[0102] For example, the overall index difference value is denoted as Π.
[0103] For example, step S105, for any dimension from the first to the nth dimension of the feature vector, can determine the feature value of that dimension using the following formula:
[0104]
[0105] Among them, X ε Let f represent the eigenvalue of the ε-th dimension from the 1st to the nth dimension, w represent the weight, and f ε Let represent the frequency distribution of the index corresponding to the ε-th dimension in the normalized index sequence, and Π represent the overall index difference characteristic value. It should be noted that the 1 in the denominator of the above formula is related to... Equivalent, f k This represents the frequency distribution of the index for the k-th category out of n categories.
[0106] For example, step S106, for any dimension from the (n+1)th to the λth dimension of the feature vector, can determine the feature value of that dimension using the following formula:
[0107]
[0108] Among them, X ε Let w represent the eigenvalue of the ε-th dimension from the (n+1)-λ-th dimension, w represent the weight, ε-n represent the target value of the interval distance corresponding to the ε-th dimension, and θ represent the eigenvalue of the ε-th dimension. ε-n This represents the characteristic value of the index difference corresponding to the interval distance under the target value ε-n, and Π represents the characteristic value of the overall index difference. It should be noted that the 1 in the denominator of the above formula is related to... Equivalent, f k This represents the frequency distribution of the index for the k-th category out of n categories.
[0109] In this embodiment, λ and w are adjustable variables. λ represents the dimension of the final feature vector, and w represents the influence of the overall index difference feature value on the feature values of each dimension in the feature vector. By default, they can be set to 10 and 0.1, respectively.
[0110] The feature extraction method for test indicator sequences provided in this disclosure is applicable to various numerical test indicator sequences. This method addresses the problem of ineffective utilization of indicator values from multiple tests conducted throughout the entire course of a patient's diagnosis and treatment. This method can serve as input for subsequent machine learning prediction tasks, identifying patients with similar characteristics among multiple patients and thus recommending personalized treatment plans.
[0111] Figure 3 A block diagram of a feature extraction apparatus for a test index sequence is shown in an exemplary embodiment, such as Figure 3 As shown, the feature extraction device 300 for the test index sequence includes:
[0112] The original sequence acquisition module 301 is used to acquire the original test index sequence, which includes multiple index values of different categories generated based on multiple tests.
[0113] The normalized sequence acquisition module 302 is used to normalize each index value in the original test index sequence to obtain a normalized index value, and obtain a normalized index sequence based on the normalized index value.
[0114] The information determination module 303 is used to determine the set dimension λ of the feature vector and the number of index categories n in the normalized index sequence;
[0115] The difference feature extraction module 304 is used to extract the difference feature values of the index from the normalized index sequence that satisfy the interval distance based on different values of the interval distance, so as to obtain the overall index difference feature value.
[0116] The first feature value determination module 305 is used to determine the feature value of any dimension from the first to the nth dimension in the feature vector based on the index distribution frequency of the corresponding category of the dimension in the normalized index sequence and the overall index difference feature value.
[0117] The second feature value determination module 306 is used to determine the target value of the interval distance corresponding to any dimension from the (n+1)th to the λth dimension in the feature vector, and to determine the feature value of the dimension based on the indicator difference feature value corresponding to the interval distance under the target value and the overall indicator difference feature value.
[0118] The feature vector acquisition module 307 is used to obtain the feature vector based on the feature values of the first to λth dimensions.
[0119] Optionally, the differential feature extraction module 304 includes:
[0120] The interval distance value module is used to select values for interval distance j from 1 to λ.
[0121] The difference information calculation module is used to calculate the difference information between the i-th normalized index value and the (i+j)-th normalized index value in the normalized index sequence based on the value of the interval distance j, where i is from 1 to Lj respectively, to obtain Lj difference information, where L is the number of normalized index values in the normalized index sequence.
[0122] The first difference feature determination module is used to determine the index difference feature value corresponding to the interval distance j under the given value based on Lj difference information;
[0123] The second difference feature determination module is used to sum the indicator difference feature values corresponding to interval distances ranging from 1 to λ to obtain the overall indicator difference feature value.
[0124] Optionally, the first difference feature determination module is used to calculate the index difference feature value corresponding to the interval distance j under the given value according to the following formula:
[0125]
[0126] Where, θ j N represents the characteristic value of the index difference corresponding to the interval distance j under the given value. i N represents the value of the i-th normalized index. i+j Represents the (i+j)th normalized index value, (N i -N i+j ) 2 This represents the difference between the i-th normalized index value and the (i+j)-th normalized index value.
[0127] Optionally, the first eigenvalue determination module 305 is used to determine the eigenvalue of any dimension from the first to the nth dimension of the eigenvector using the following formula:
[0128]
[0129] Among them, X ε Let f represent the eigenvalue of the ε-th dimension from the 1st to the nth dimension, w represent the weight, and f ε Π represents the index distribution frequency of the ε-th dimension corresponding to the category in the normalized index sequence, and Π represents the overall index difference characteristic value.
[0130] Optionally, the second feature value determination module 306 is used to take the difference between the dimension and the number of index categories n as the target value of the interval distance corresponding to the dimension.
[0131] Optionally, the second eigenvalue determination module 306 is used to determine the eigenvalue of any dimension from the (n+1)th to the λth dimension of the eigenvector using the following formula:
[0132]
[0133] Among them, X ε Let w represent the eigenvalue of the ε-th dimension from the (n+1)-λ-th dimension, w represent the weight, ε-n represent the target value of the interval distance corresponding to the ε-th dimension, and θ represent the eigenvalue of the ε-th dimension. ε-n The interval distance represents the index difference characteristic value corresponding to the target value ε-n, and Π represents the overall index difference characteristic value.
[0134] Optionally, the normalized sequence acquisition module 302 is used to calculate the mean and standard deviation of the indicators in the original test indicator sequence; and for each indicator value in the original test indicator sequence, normalize the indicator value according to the mean and standard deviation to obtain the corresponding normalized indicator value.
[0135] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0136] Figure 4 This is a block diagram illustrating an electronic device 400 according to an exemplary embodiment. (Refer to...) Figure 4 The electronic device 400 includes a processor 422, which may be one or more, and a memory 432 for storing computer programs executable by the processor 422. The computer programs stored in the memory 432 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processor 422 may be configured to execute the computer program to perform the feature extraction method for the aforementioned test index sequence.
[0137] Additionally, the electronic device 400 may also include a power supply component 426 and a communication component 450. The power supply component 426 may be configured to perform power management of the electronic device 400, and the communication component 450 may be configured to enable communication of the electronic device 400, such as wired or wireless communication. Furthermore, the electronic device 400 may also include an input / output (I / O) interface 458. The electronic device 400 can operate on an operating system stored in the memory 432.
[0138] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the feature extraction method for the test index sequence described above. For example, the non-transitory computer-readable storage medium may be the memory 432 including the program instructions described above, which may be executed by the processor 422 of the electronic device 400 to complete the feature extraction method for the test index sequence described above.
[0139] In another exemplary embodiment, a computer program product is also provided, comprising a computer program executable by a programmable device, the computer program having a code portion for performing the feature extraction method of the above-described test index sequence when executed by the programmable device.
[0140] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0141] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0142] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. A method for feature extraction of a test index sequence, characterized in that, include: Obtain the original test index sequence, which includes multiple index values of different categories generated based on multiple tests; Each index value in the original test index sequence is normalized to obtain a normalized index value, and a normalized index sequence is obtained based on the normalized index value. Determine the set dimension λ of the feature vector and the number of index categories n in the normalized index sequence; Based on different values of the interval distance, for each pair of normalized index values that satisfy the interval distance in the normalized index sequence, the index difference feature value corresponding to the interval distance is extracted, and the index difference feature values corresponding to the interval distance under different values are summed to obtain the overall index difference feature value. For any dimension from the first to the nth dimension in the feature vector, the feature value of the dimension is determined based on the index distribution frequency of the corresponding category in the normalized index sequence and the overall index difference feature value. For any dimension from the (n+1)th to the λth dimension in the feature vector, determine the target value of the interval distance corresponding to the dimension, and determine the feature value of the dimension based on the indicator difference feature value corresponding to the interval distance under the target value and the overall indicator difference feature value. The eigenvector is obtained based on the eigenvalues of dimensions 1 to λ.
2. The method according to claim 1, characterized in that, Based on different values of the interval distance, for each pair of normalized index values in the normalized index sequence that satisfy the interval distance, the index difference feature value corresponding to the interval distance is extracted, and the index difference feature values corresponding to the interval distance under different values are summed to obtain the overall index difference feature value, including: Take values for the interval distance j from 1 to λ; Based on the value of the interval distance j, i is taken from 1 to Lj respectively. For the i-th normalized index value in the normalized index sequence, the difference information between the i-th normalized index value and the (i+j)-th normalized index value is calculated to obtain Lj difference information, where L is the number of normalized index values in the normalized index sequence. Based on the Lj difference information, determine the indicator difference characteristic value corresponding to the interval distance j under the given value; The overall index difference characteristic value is obtained by summing the index difference characteristic values corresponding to interval distances ranging from 1 to λ.
3. The method according to claim 2, characterized in that, The step of determining the indicator difference characteristic value corresponding to the interval distance j under the given value based on Lj difference information includes: The indicator difference characteristic value corresponding to the interval distance j under the given value is calculated according to the following formula: in, This represents the characteristic value of the index difference corresponding to the interval distance j under the given value. This represents the value of the i-th normalized index. This represents the (i+j)th normalized index value. This represents the difference between the i-th normalized index value and the (i+j)-th normalized index value.
4. The method according to claim 1, characterized in that, For any dimension from the 1st to the nth dimension of the feature vector, the feature value of the dimension is determined based on the index distribution frequency of the corresponding category in the normalized index sequence and the overall index difference feature value, including: For any dimension from the 1st to the nth dimension of the feature vector, the feature value of that dimension is determined by the following formula: in, Represents the first to nth dimension eigenvalues of dimension Indicates weight, Indicates the first normalized index in the sequence. The frequency distribution of indicators corresponding to the category is denoted by Π, which represents the overall indicator difference characteristic value.
5. The method according to claim 1, characterized in that, The determination of the target value of the interval distance corresponding to the dimension includes: The difference between the dimension and the number of index categories n is used as the target value of the interval distance corresponding to the dimension.
6. The method according to claim 5, characterized in that, For any dimension from n+1 to λ in the feature vector, determine the target value of the interval distance corresponding to that dimension, and determine the feature value of that dimension based on the indicator difference feature value corresponding to the interval distance under the target value and the overall indicator difference feature value, including: For any dimension from the (n+1)th to the λth dimension in the feature vector, the feature value of that dimension is determined by the following formula: in, Represents the n+1 to λth dimension eigenvalues of dimension Indicates weight, Indicates the first The target value corresponding to the interval distance of the dimension. This indicates the interval distance within the target value. The corresponding indicator difference characteristic value is Π, which represents the overall indicator difference characteristic value.
7. The method according to claim 1, characterized in that, The normalization process for each index value in the original test index sequence to obtain a normalized index value includes: Calculate the mean and standard deviation of the indicators in the original test indicator sequence; For each index value in the original test index sequence, the index value is normalized according to the index mean and standard deviation to obtain the corresponding normalized index value.
8. A feature extraction device for a test index sequence, characterized in that, include: The original sequence acquisition module is used to acquire the original test index sequence, which includes multiple index values of different categories generated based on multiple tests. The normalized sequence acquisition module is used to normalize each index value in the original test index sequence to obtain a normalized index value, and obtain a normalized index sequence based on the normalized index value. The information determination module is used to determine the set dimension λ of the feature vector and the number of index categories n in the normalized index sequence; The difference feature extraction module is used to extract the index difference feature value corresponding to the interval distance based on different values of the interval distance, according to each pair of normalized index values that satisfy the interval distance in the normalized index sequence, and to sum the index difference feature values corresponding to the interval distance under different values to obtain the overall index difference feature value. The first feature value determination module is used to determine the feature value of any dimension from the first to the nth dimension in the feature vector based on the index distribution frequency of the corresponding category in the normalized index sequence and the overall index difference feature value. The second feature value determination module is used to determine the target value of the interval distance corresponding to any dimension from the (n+1)th to the λth dimension in the feature vector, and to determine the feature value of the dimension based on the indicator difference feature value corresponding to the interval distance under the target value and the overall indicator difference feature value. The feature vector acquisition module is used to obtain the feature vector based on the feature values of the first to λth dimensions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.
10. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Regional electric energy quality monitoring and analyzing system
CN113240261A
Abnormality detection method and device, electronic equipment and storage medium
CN113377568A