Target number statistical method
By grouping and analyzing the dataset and building an adaptive model, the problems of generalization ability and computational complexity of existing methods in multidimensional data processing are solved, and efficient and interpretable target quantity statistics are achieved.
Patent Information
- Application Number
- CN202511476248.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-01-23
AI Technical Summary
Existing methods for statistical analysis of target quantities have poor generalization ability, high processing complexity, narrow applicability, and cannot effectively process multi-dimensional data and extract potential features. Furthermore, they are difficult to meet real-time requirements in resource-constrained scenarios.
The data is divided into target dataset A and target dataset B, and one-dimensional and multi-dimensional data analysis are performed respectively. One-dimensional mathematical models and linear or nonlinear regression models are constructed. The calculation process is optimized by difference time series and dimensionality reduction, and statistical models that are suitable for various data types are constructed.
It improves multi-dimensional data processing capabilities, enhances model interpretability and adaptability, reduces computational complexity, and is suitable for resource-constrained scenarios such as embedded devices, meeting real-time requirements.
Smart Images

Figure CN121387995A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for statistical analysis of target quantities, and pertains to the field of mathematical statistics. Background Technology
[0002] Existing statistical methods for determining the number of targets have the following shortcomings:
[0003] Poor generalization ability: Most existing statistical methods rely on density plot statistical methods, which are highly dependent on the distribution of data. When the distribution of test data and training data differs greatly, the statistical accuracy will decrease significantly; it is unable to extract the latent features of a certain type of data, and the interpretability is not strong.
[0004] High processing complexity: Existing statistical methods still use deep convolutional neural networks, resulting in high model complexity and high computational resource consumption; in resource-constrained scenarios such as embedded devices, deployment is difficult; in dense small target scenarios, generating high-resolution density maps requires high computational costs and is difficult to meet real-time requirements.
[0005] Narrow applicability: Existing statistical methods can only perform statistical analysis on certain types or categories of data and have extremely high requirements for data quality. When processing multi-dimensional data, they cannot identify core features. Summary of the Invention
[0006] In view of the shortcomings of existing technologies, the purpose of this invention is to provide a method for statistical analysis of target quantities, which aims to solve the problem of low efficiency in statistical analysis of multiple types of data.
[0007] To achieve the above objectives, the present invention provides a target quantity statistics method comprising:
[0008] Obtain the raw data provided by the data source, and divide the raw data into target dataset A and target dataset B according to the data type;
[0009] Analyze the target dataset A; determine whether the data in dataset A is one-dimensional; if so, perform time series difference analysis on the one-dimensional data and construct a one-dimensional mathematical model; if not, determine whether there is a main attribute in the multi-dimensional data; if there is a main attribute, analyze the influence weight of non-main attributes on the main attribute in the multi-dimensional data, identify important attributes, and construct a linear regression model or a nonlinear regression model based on the main attribute; if there is no main attribute, perform dimensionality reduction on the attributes in the multi-dimensional data and then construct a linear mathematical model.
[0010] Analyze the target dataset B; in dataset B, count the number of times the target image and target audio appear per unit time, then perform time series difference analysis to construct a mathematical model of the number of times the target image and target audio appear; summarize the mathematical model and feed it back to the data source; obtain the identity information of the data source, create a storage structure based on the identity information and save the original data.
[0011] Furthermore, the specific steps for constructing the one-dimensional mathematical model are as follows:
[0012] One-dimensional data is denoted as od (1) ~od (oi) ;oi represents the number of one-dimensional data;
[0013] Using ADF to detect od (1) ~od (oi) Perform differential detection to obtain the difference order d; for od (1) ~od (oi) Perform d-order differencing and process with autocorrelation function and partial autocorrelation function to obtain autoregressive parameter p and moving average parameter q;
[0014] Calculate od (1) ~od (oi) Mean AOD, variance SOD, white noise ε (1) ~ε (oi) ;
[0015] Calculate the autocorrelation coefficients φ from order 1 to order p. (1) ~φ (p) ;
[0016] {od (1) ~od (oi) Given sequence T, calculate the differences from order 1 to p of sequence T to obtain sequence T. (1) ~T (p) ;
[0017] Calculate the sequence T with respect to the sequence T (1) ~T (p) The autocovariance is obtained by γ. (1) ~γ (p) ;
[0018] Constructed matrix B (φ) Matrix B (1) Sum matrix B (2) ; where matrix B (2) All elements on the diagonal are sod; Transform sequence T about sequence T (1) ~T (p-1) autocovariance γ (1) ~γ (p-1) symmetrically arranged in matrix B (2) Both sides of the diagonal;
[0019] Calculate φ (1) ~φ (p) Value: B (φ) = (B (2) ) -1 *B (1) ;
[0020] Compare the magnitudes of p and q, and calculate the moving average coefficients θ from the 1st to the qth order. (1) ~θ (q) ;
[0021] If p ≥ q, then according to sequence T with respect to sequence T (1) ~T (q) autocovariance to γ (1) ~γ (q) Construct matrix Bφ (q) Matrix B (θ) And matrix Bq;
[0022] Calculate θ (1) ~θ (q) Value: B( θ )=(Bq) -1 *Bφ( q );
[0023] If p < q, calculate the differences from order 1 to q of sequence T to obtain sequence Tq. (1) ~Tq (q) .
[0024] 3. The target quantity statistics method according to claim 2, characterized in that the step of constructing a one-dimensional mathematical model further includes:
[0025] Calculate the sequence T with respect to the sequence Tq (1) ~Tq (q) autocovariance γq (1) ~γq (q) And construct the matrix equation to calculate φ (1) ~φ (p) The value;
[0026] {ε (1) ~ε (oi) Let T be the sequence T after m transformations, given the sequence Et. Then the one-dimensional mathematical model is:
[0027]
[0028] Among them, Ii (oi×1) Let φ be a (oi×1) matrix of all 1s. (n) θ represents the autocorrelation coefficient of the nth order. (n)Et represents the moving average coefficient of the nth order. (n) This represents the nth-order difference of the sequence Et.
[0029] Furthermore, the specific steps for determining whether a subject attribute exists are as follows:
[0030] Let the dimension of multidimensional data be (w+1) and the quantity be d;
[0031] Let ya be an attribute in multidimensional data. Use the pandas library to determine whether the remaining attributes in the multidimensional data have a unique mathematical dependency on ya.
[0032] If it exists, then ya is the main attribute; analyze the influence weight of non-main attributes on main attributes in multidimensional data, determine important attributes, and construct a linear regression model or a nonlinear regression model based on the main attributes;
[0033] If it does not exist, then each attribute in the multidimensional data is independent of each other. Dimensionality reduction is performed on the attributes in the multidimensional data to construct a linear mathematical model.
[0034] Furthermore, the specific steps for determining the primary attributes are as follows:
[0035] The numerical values of the main attributes of multidimensional data are denoted as yt. (1) ~yt (d) The numerical value of a non-subject attribute is denoted as xu. (1,1) ~xuA (d,w) ;
[0036] Construct matrix C (y) And construct matrix C (x) ;
[0037] Calculate matrix C (x) The average value wx corresponding to the elements in columns 1 to w. (1) ~wx (w) ;
[0038] Calculate yt (1) ~yt (d) The average value of ayt;
[0039] matrix C (x) By dividing by columns, we obtain the first to the wth non-primary attribute matrices Cx. (1) ~Matrix Cx (w) ;
[0040] Calculate matrix Cx (1) ~Matrix Cx (w) The decentralized matrix is obtained as matrix Dx. (1) ~ Matrix Dx (w) ; Calculate matrix C (y) Decentralized matrix D(y) ;
[0041] Define formula s-1:
[0042]
[0043] Among them, Cx (v) Dx represents the v-th non-subject attribute matrix; (v) Represents matrix Cx (v) The decentralization matrix; α (v) This represents the weight coefficient of the v-th non-subject attribute to the subject attribute; the value of v ranges from 1 to w.
[0044] Calculate the weight coefficients of the first to wth non-subject attributes on the subject attributes to obtain α. (1) ~α(w);
[0045] Extract α (1) ~α (w) The maximum value α in (max) Minimum value α (min) Calculate the combined weight Aα;
[0046] Compare α (1) ~α (w) The non-primary attributes with a weight coefficient greater than Aα are considered as primary attributes, depending on the magnitude of Aα.
[0047] Furthermore, the specific steps for constructing the linear regression model are as follows:
[0048] Calculate and determine whether the absolute value of the Pearson correlation coefficient of the multidimensional data and 1 is less than or equal to the error determination coefficient ζ; if it is less than or equal to ζ, construct a linear regression model; if it is greater than ζ, construct a nonlinear regression model.
[0049] Constructing a linear regression model for multidimensional data;
[0050] The linear regression coefficient of multidimensional data is denoted as β. (0) ~β (w) ;β (1) ~β (w) Represents the linear regression coefficients of the 1st to wth non-subject attributes;
[0051] Construct matrix Bx, matrix B (β) Calculate β (0) ~β (w) Value:
[0052] The mathematical expression for a linear regression model for multidimensional data is:
[0053] yt (a) =β (0) +(β(1) ×xu (a,1) )+(β (2) ×xu (a,2) )+......+(β (w) ×xu (a,w) );
[0054] Among them, xu (a,1) ~xu (a,w) This represents the values of the first to the wth non-subject attributes of the a-th multidimensional data.
[0055] Furthermore, the specific steps for constructing the nonlinear regression model are as follows:
[0056] Define the decision steps for a random forest decision tree:
[0057] Define formula T-1:
[0058] Among them, ix (b) wx represents the exponent of the main attribute value relative to the b-th non-main attribute value. (b) Representing matrix C (x) The average value corresponding to the elements in column b;
[0059] Calculate the exponent ix of the main attribute value relative to the values of the 1st to wth non-main attributes according to formula T-1. (1) ~ix (w) ;
[0060] Let the nonlinear function of the principal attribute and each non-principal attribute be fu(xu) (a,b) ), function fu(xu (a,b) The coefficients of the 1st to wth non-subject attributes in the equation are μ. (1) ~μ (w) ;
[0061] Let μ be the coefficient of the b-th non-subject attribute. (b) ; function fu(xu (a,b) The initial mathematical form of ) is:
[0062]
[0063] Define the residual sum of squares (RSS):
[0064]
[0065] Using RSS as the loss function, μ is calculated on the RSS. (1) ~μ (w) The partial derivative;
[0066] Calculate μ for RSS (b) The partial derivative is
[0067] Let the learning rate be η; let μ (b) The value after k updates is (μ) (b) (k) ), μ (b) The value after (k+1) updates is (μ) (b) (k+1) ), ;
[0068] Define formula T-2:
[0069]
[0070] Based on the decision-making steps defined above, a nonlinear regression model is constructed using the random forest algorithm.
[0071] Furthermore, the specific steps for dimensionality reduction of attributes in multidimensional data are as follows:
[0072] Multidimensional data where each attribute is independent is treated as independent multidimensional data.
[0073] Defined matrix F:
[0074]
[0075] Among them, ru (1,1) ~ru (d,(w+1)) This represents the value of the first to (w+1)th attribute in the first to dth independent multidimensional data;
[0076] Using the elements of each column of matrix F as clustering objects, a clustering algorithm is used to calculate the cluster centers from the 1st to the (w+1th)th column in matrix F, resulting in cu. (1) ~cu (w+1) Define matrix Fc;
[0077] Dimensionality reduction is performed on matrix F to obtain matrix Fw: Fw = F * Fc; the elements in Fw are treated as one-dimensional data, and the processing steps of difference time series analysis are used to construct a linear mathematical model of independent multidimensional data.
[0078] Compared with the prior art, the beneficial effects of the present invention are:
[0079] Highly applicable: This invention can process multi-dimensional structured data; common multi-dimensional data exists in the form of two-dimensional tables (such as Excel spreadsheets and database tables). This invention can uncover the potential mathematical patterns between the attributes of multi-dimensional data and construct mathematical models to demonstrate the influence of various non-subject attributes of multi-dimensional data on the subject attributes; furthermore, the statistical model constructed by this invention has strong interpretability, can clarify the relationships between structured data variables, and provide a basis for decision-making by relevant departments.
[0080] High adaptability: This invention provides different mathematical models for various data types. It offers corresponding solutions for processing single-dimensional, multi-dimensional, or other non-numerical data. Furthermore, the statistical models constructed by this invention are highly interpretable, clearly defining the relationships between structured data variables and providing a basis for decision-making by relevant departments.
[0081] The calculation is simple; this invention optimizes the numerical calculation in the analysis and construction process of difference time series and regression models. By using matrix operations, it transforms complex algebraic operations into matrix operations that are easier for computers to process, thereby further improving the speed of data analysis and processing. Attached Figure Description
[0082] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0083] Figure 1 This is a schematic diagram of the method of the present invention;
[0084] Figure 2 This is a schematic diagram of the variable types in this invention;
[0085] Figure 3 This is a schematic diagram of the data in this invention. Detailed Implementation
[0086] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0087] Please see Figure 1 and Figure 3 One method for statistical analysis of target quantities includes:
[0088] Step S1: Obtain the raw data provided by the data source, and divide the raw data into target dataset A (structured data) and target dataset B (unstructured data) according to the data type;
[0089] It should be noted that the data source in this invention refers to the individual, unit, or department that uses this invention (a method for statistical analysis of target quantities) to perform mathematical statistics;
[0090] Step S2: Analyze the target dataset A; determine whether the data in dataset A is one-dimensional data; if so, perform time series difference analysis on the one-dimensional data and construct a one-dimensional mathematical model; if not, determine whether there is a main attribute in the multi-dimensional data; if there is a main attribute, analyze the influence weight of non-main attributes on the main attribute in the multi-dimensional data, determine the important attributes, and construct a linear regression model or a nonlinear regression model based on the main attribute; if there is no main attribute, perform dimensionality reduction on the attributes in the multi-dimensional data, and then construct a linear mathematical model.
[0091] The specific steps of step S2 are as follows:
[0092] Step S21: If the data in the target dataset A is one-dimensional data, then perform time series difference on the one-dimensional data to construct a one-dimensional mathematical model;
[0093] One-dimensional data is denoted as od (1) ~od (oi) ;oi represents the number of one-dimensional data;
[0094] Using ADF to detect od (1) ~od (oi) Perform differential detection to obtain the difference order d;
[0095] For od (1) ~od (oi) Perform d-order differencing and process with autocorrelation function (ACF) and partial autocorrelation function (PACF) to obtain autoregressive parameter p and moving average parameter q (the range of values for p and q is [1, oi) and the values of (oi-p) and (oi-q) are both greater than or equal to 1);
[0096] Calculate od (1) ~od (oi) The mean value is AOD, and the variance is SOD.
[0097] Calculate od (1) ~od (oi) White noise ε (1) ~ε (oi) ; where ε (1) =od (1) -aod; ε (2) =od (2) -aod; and so on, ε (oi) =od (oi) -aod;
[0098] Calculate the autocorrelation coefficients φ from order 1 to order p. (1) ~φ (p) ;
[0099] {od (1) ~od(oi) Given sequence T, calculate the differences from order 1 to p of sequence T to obtain sequence T. (1) ~T (p) ;
[0100] Calculate the sequence T with respect to the sequence T (1) ~T (p) The autocovariance is obtained by γ. (1) ~γ (p) ;
[0101] Construct a (1×p) matrix B (φ) :
[0102] Construct a (1×p) matrix B (1) :
[0103] Construct a (p×p) matrix B (2) : Wherein, matrix B (2) All elements on the diagonal are sod; Transform sequence T about sequence T (1) ~T (p-1) autocovariance γ (1) ~γ (p-1) symmetrically arranged in matrix B (2) Both sides of the diagonal;
[0104] Calculate φ (1) ~φ (p) Value: B (φ) = (B (2) ) -1 *B (1) Where -1 represents the inverse of the matrix, and * represents matrix multiplication;
[0105] Compare the magnitudes of p and q, and calculate the moving average coefficients θ from the 1st to the qth order. (1) ~θ (q) ;
[0106] If p ≥ q, then according to sequence T with respect to sequence T (1) ~T (q) autocovariance to γ (1) ~γ (q) Construct a (1×q) matrix Bφ (q) :
[0107] Construct a (1×q) matrix B (θ) :
[0108] Construct a (q×q) matrix Bq: In this matrix, the elements on the diagonal of matrix Bq are all sod; the sequence T is expressed about the sequence T.(1) ~T (q-1) autocovariance γ (1) ~γ (q-1) symmetrically arranged in matrix B (2) Both sides of the diagonal;
[0109] Calculate θ (1) ~θ (q) Value: B( θ )=(Bq) -1 *Bφ( q );
[0110] If p < q, calculate the differences from order 1 to q of sequence T to obtain sequence Tq. (1) ~Tq (q) ;
[0111] Calculate the sequence T with respect to the sequence Tq (1) ~Tq (q) The autocovariance is obtained as γq. (1) ~γq (q) ;
[0112] Define the function ss(i):
[0113] Where, γq (j) Represents sequence T with respect to sequence Tq (j) The autocovariance; θ (i) θ (j) and θ (j+i) , represent the moving average coefficients of the i-th, i-th and (j+i)-th orders respectively, where i and j both range from 1 to (q-1);
[0114] Construct the matrix equation:
[0115] Where, θ (j+1) This represents the moving average coefficient of the (j+1)th order;
[0116] γq (1) ~γq (q) Substitute into the matrix equation and calculate φ (1) ~φ (p) The value;
[0117] {ε (1) ~ε (oi) Let T be the sequence T after m transformations, given the sequence Et. Then the one-dimensional mathematical model is:
[0118]
[0119] Among them, Ii (oi×1) Let φ be a (oi×1) matrix of all 1s.(n) θ represents the autocorrelation coefficient of the nth order. (n) Et represents the moving average coefficient of the nth order. (n) Let m represent the nth order difference of the sequence Et; the range of m is: 1≤m≤[(p,q)]. min The value of n ranges from 1 to m;
[0120] Step S22: Please refer to Figure 2 If the data in the target dataset A is multidimensional data, then remove the categorical and Boolean variables from the multidimensional data, and denote the dimension of the multidimensional data (after removing the categorical and Boolean variables) as (w+1) and the number as d;
[0121] Step S221: Let ya be an attribute in the multidimensional data. Use the pandas library to determine whether the remaining attributes in the multidimensional data (i.e., the other w attributes other than ya) have a unique mathematical dependency relationship with ya.
[0122] If it exists, then ya is the main attribute; construct a linear regression model or a nonlinear regression model based on the main attribute, and analyze the influence weight of non-main attributes on the main attribute in the multidimensional data to determine the important attributes;
[0123] The numerical values of the main attributes of multidimensional data are denoted as yt. (1) ~yt (d) The numerical value of a non-subject attribute is denoted as xu. (1,1) ~xuA (d,w) ;
[0124] Construct matrix C (y) :
[0125]
[0126] Construct matrix C (x) :
[0127] Among them, (yt) (1) xu (1,1) ~xuA (1,w) (yt) represents the first multidimensional data (after removing categorical and Boolean variables); (2) xu (2,1) ~xuA (2,w) () represents the second (after removing categorical and Boolean variables) multidimensional data; and so on, (yt) (d) xu (d,1) ~xuA (d,w) ) represents the d-th multidimensional data point (after removing categorical and Boolean variables);
[0128] Calculate matrix C (x)In the expression, the average value of the elements in rows 1 to d is used to obtain dx. (1) ~dx (d) The average value of the elements in columns 1 to w is used to obtain wx. (1) ~wx (w) ;
[0129] Calculate yt (1) ~yt (d) The average value of ayt;
[0130] matrix C (x) By dividing by columns, we obtain the first to the wth non-primary attribute matrices Cx. (1) ~Matrix Cx (w) ;
[0131] Matrix Cx (1) : Matrix Cx (2) : And so on, matrix Cx (w) :
[0132] Calculate matrix Cx (1) ~Matrix Cx (w) The decentralized matrix is obtained as matrix Dx. (1) ~ Matrix Dx (w) ;(where Dx (1) The formula for calculating it is: Dx (1) =Cx (1) -wx (1) *Ii (d×1) ;Dx (2) The formula for calculating it is: Dx (2) =Cx (2) -wx (2) *Ii (d×1) And so on, Dx (w) The formula for calculating it is: Dx (w) =Cx (w) -wx (w) *Ii (d×1) Among them, Ii (d×1) (represents a (d×1) matrix of all 1s)
[0133] Calculate matrix C (y) Decentralized matrix D (y) (Matrix D) (y) The formula for calculating D is: (y) =C (y) -ayt*Ii (d×1) );
[0134] Define formula s-1:
[0135]
[0136] Among them, Cx (v) Dx represents the v-th non-subject attribute matrix; (v) Represents matrix Cx (v) The decentralization matrix; α (v) This represents the weight coefficient of the v-th non-subject attribute to the subject attribute; the value of v ranges from 1 to w.
[0137] Calculate the weight coefficients of the first to wth non-subject attributes on the subject attributes to obtain α. (1) ~α(w);
[0138] Extract α (1) ~α (w) The maximum value α in (max) Minimum value α (min) ;
[0139] Define formula s-2:
[0140] Where Aα represents the combined weight;
[0141] Compare α (1) ~α (w) Based on the magnitude of Aα, non-primary attributes with weight coefficients greater than Aα are considered as primary attributes;
[0142] It should be noted that if the multidimensional data is two-dimensional data, step S221 is skipped and step S222 is executed directly.
[0143] Step S222: Construct a linear or nonlinear regression model for the multidimensional data; construct the relation s-3:
[0144] (Pearson correlation coefficient); where ζ represents the error determination coefficient (ζ is 0.01; users or relevant technical personnel can adjust the value of ζ according to actual needs); xu (a,b) Representing matrix C (x) The element in row a, column b (i.e., the value of the b-th non-subject attribute in the a-th multidimensional data (after removing categorical and Boolean variables); wx (b) Representing matrix C (x) The average value of the elements in column b; yt (a) Representing matrix C (y) The element in row a (i.e., the value of the a-th multidimensional data subject attribute (after removing categorical and Boolean variables); axx represents the transition parameter, where the value of a ranges from 1 to d, and the value of b ranges from 1 to w;
[0145] Determine whether the relation s-3 holds true (i.e., whether the absolute value of the Pearson correlation coefficient of the multidimensional data (after removing categorical and Boolean variables) and 1 is less than or equal to the error determination coefficient ζ);
[0146] If relation s-3 holds, it means that the main attributes and various non-main attributes in the multidimensional data satisfy a linear relationship, and a linear regression model can be constructed.
[0147] Construct a (d×(w+1)) matrix Bx:
[0148]
[0149] The linear regression coefficient of multidimensional data is denoted as β. (0) ~β (w) ; where β (1) ~β ( w ) Represents the linear regression coefficients of the 1st to wth non-subject attributes;
[0150] Construct a matrix B of ((w+1)×1) (β) :
[0151]
[0152] Calculate β (0) ~β (w) Value: Where T represents the transpose of a matrix, -1 represents the inverse of a matrix, and * represents matrix multiplication;
[0153] The mathematical expression for a linear regression model for multidimensional data is:
[0154]
[0155] Among them, xu (a,1) ~xu (a,w) This represents the values of the 1st to wth non-subject attributes of the a-th multidimensional data;
[0156] If relation s-3 does not hold, it means that the main attributes and various non-main attributes in the multidimensional data satisfy a non-linear relationship, and a non-linear regression model is constructed.
[0157] Define the decision steps for a random forest decision tree:
[0158] Define formula T-1:
[0159] Among them, ix (b) wx represents the exponent of the main attribute value relative to the b-th non-main attribute value. (b) Representing matrix C (x) The average value corresponding to the elements in column b;
[0160] Calculate the exponent ix of the main attribute value relative to the values of the 1st to wth non-main attributes according to formula T-1. (1) ~ix (w) ;
[0161] Let the nonlinear function of the principal attribute and each non-principal attribute be fu(xu) (a,b) ), function fu(xu (a,b) The coefficients of the 1st to wth non-subject attributes in the equation are μ. (1) ~μ (w) ;
[0162] Let μ be the coefficient of the b-th non-subject attribute. (b) ; function fu(xu (a,b) The initial mathematical form of ) is:
[0163]
[0164] Define the residual sum of squares (RSS):
[0165]
[0166] Using RSS as the loss function, μ is calculated on the RSS. (1) ~μ (w) The partial derivative;
[0167] Calculate μ for RSS (b) The partial derivative is
[0168] Let the learning rate be η; (η takes the value of 0.001; users or relevant technical personnel can adjust the value of η according to actual needs) Let μ (b) The value after k updates is (μ) (b) (k) ), μ ( b ) The value after (k+1) updates is (μ) (b) (k+1) ), ;
[0169] Define formula T-2:
[0170]
[0171] Based on the decision-making steps defined above, a nonlinear regression model is constructed using the random forest algorithm;
[0172] It should be noted that, due to the uncertainty of nonlinear regression, the nonlinear regression models constructed for different datasets may be different. Therefore, this invention only provides a computer processing method for "constructing a nonlinear regression model" and the initial mathematical form of the nonlinear model.
[0173] If it does not exist, then each attribute in the multidimensional data is independent of each other. Dimensionality reduction is performed on the attributes in the multidimensional data to construct a linear mathematical model.
[0174] Multidimensional data where "each attribute is independent" is treated as independent multidimensional data.
[0175] Define a matrix F of (d×(w+1)) to represent multidimensional data where "each attribute is independent". Matrix F:
[0176]
[0177] Among them, ru (1,1) ~ru (d,(w+1)) This represents the value of the first to (w+1)th attribute in the first to dth independent multidimensional data;
[0178] Using the elements of each column of matrix F as clustering objects, a clustering algorithm is used to calculate the cluster centers from the 1st to the (w+1th)th column in matrix F, resulting in cu. (1) ~cu (w+1) (If a column in matrix F has multiple cluster centers, then the total number of cluster centers in that column is the average of the multiple cluster centers.)
[0179] Define a matrix Fc of ((w+1)×1):
[0180]
[0181] Dimensionality reduction of matrix F yields matrix Fw: Fw = F * Fc; elements in Fw are treated as one-dimensional data, and the processing steps of step S1 above are used to construct a linear mathematical model of independent multidimensional data.
[0182] Step S3: Analyze the target dataset B; in dataset B, (using Natural Language Processing (NLP) and Computer Vision (CV) to extract dynamic features from dataset B), count the number of occurrences of target images and target audio per unit time, then perform time series difference analysis to construct a mathematical model of the occurrence frequency of target images and target audio; summarize the mathematical models constructed in (steps S2 and S3) (one-dimensional mathematical model (one-dimensional data); linear regression model, non-linear regression model, linear mathematical model (multi-dimensional data); mathematical model of the occurrence frequency of target images and target audio) and feed them back to the data source; obtain the identity information of the data source, create a storage structure based on the identity information, and save the original data;
[0183] It should be noted that in this invention, "target image" and "target audio" refer to the images and audio that need to be counted.
[0184] The above formulas are all dimensionless calculations. The formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation. For example, there are weighting coefficients and proportional coefficients. The values set are to quantify each parameter to obtain a specific value, which is convenient for subsequent comparison. The values of the weighting coefficients and proportional coefficients are only required to not affect the proportional relationship between the parameters and the quantified values.
[0185] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for statistically analyzing target quantities, characterized in that, The method includes: Obtain the raw data provided by the data source, and divide the raw data into target dataset A and target dataset B according to the data type; Analyze the target dataset A; determine whether the data in dataset A is one-dimensional; if so, perform time series difference analysis on the one-dimensional data and construct a one-dimensional mathematical model; if not, determine whether there is a main attribute in the multi-dimensional data; if there is a main attribute, analyze the influence weight of non-main attributes on the main attribute in the multi-dimensional data, identify important attributes, and construct a linear regression model or a nonlinear regression model based on the main attribute; if there is no main attribute, perform dimensionality reduction on the attributes in the multi-dimensional data and then construct a linear mathematical model. Analyze the target dataset B; in dataset B, count the number of times the target image and target audio appear per unit time, then perform time series difference analysis to construct a mathematical model of the number of times the target image and target audio appear; summarize the mathematical model and feed it back to the data source; obtain the identity information of the data source, create a storage structure based on the identity information and save the original data.
2. The target quantity statistics method according to claim 1, characterized in that, The specific steps for constructing a one-dimensional mathematical model are as follows: One-dimensional data is denoted as od (1) ~od (oi) ;oi represents the number of one-dimensional data; Using ADF to detect od (1) ~od (oi) Perform differential detection to obtain the difference order d; for od (1) ~od (oi) Perform d-order differencing and process with autocorrelation function and partial autocorrelation function to obtain autoregressive parameter p and moving average parameter q; Calculate od (1) ~od (oi) Mean AOD, variance SOD, white noise ε (1) ~ε (oi) ; Calculate the autocorrelation coefficients φ from order 1 to order p. (1) ~φ (p) ; {od (1) ~od (oi) Given sequence T, calculate the differences from order 1 to p of sequence T to obtain sequence T. (1) ~T (p) ; Calculate the sequence T with respect to the sequence T (1) ~T (p) The autocovariance is obtained by γ. (1) ~γ (p) ; Constructed matrix B (φ) Matrix B (1) Sum matrix B (2) ; where matrix B (2) All elements on the diagonal are sod; Transform sequence T about sequence T (1) ~T (p-1) autocovariance γ (1) ~γ (p-1) symmetrically arranged in matrix B (2) Both sides of the diagonal; Calculate φ (1) ~φ (p) Value: B (φ) = (B (2) ) -1 *B (1) ; Compare the magnitudes of p and q, and calculate the moving average coefficients θ from the 1st to the qth order. (1) ~θ (q) ; If p ≥ q, then according to sequence T with respect to sequence T (1) ~T (q) autocovariance to γ (1) ~γ (q) Construct matrix Bφ (q) Matrix B (θ) And matrix Bq; Calculate θ (1) ~θ (q) Value: B (θ) =(Bq) -1 *Bφ (q) ; If p < q, calculate the differences from order 1 to q of sequence T to obtain sequence Tq. (1) ~Tq (q) .
3. The target quantity statistics method according to claim 2, characterized in that, The steps for constructing a one-dimensional mathematical model also include: Calculate the sequence T with respect to the sequence Tq (1) ~Tq (q) autocovariance γq (1) ~γq (q) And construct the matrix equation to calculate φ (1) ~φ (p) The value; {ε (1) ~ε (oi) Let T be the sequence T after m transformations, given the sequence Et. Then the one-dimensional mathematical model is: Among them, Ii (oi×1) Let φ be a (oi×1) matrix of all 1s. (n) θ represents the autocorrelation coefficient of the nth order. (n) Et represents the moving average coefficient of the nth order. (n) This represents the nth-order difference of the sequence Et.
4. The target quantity statistics method according to claim 3, characterized in that, The specific steps to determine whether a subject attribute exists are as follows: Let the dimension of multidimensional data be (w+1) and the quantity be d; Let ya be an attribute in multidimensional data. Use the pandas library to determine whether the remaining attributes in the multidimensional data have a unique mathematical dependency on ya. If it exists, then ya is the main attribute; Analyze the influence weights of non-subject attributes on subject attributes in multidimensional data, identify important attributes, and construct linear or nonlinear regression models based on subject attributes; If it does not exist, then each attribute in the multidimensional data is independent of each other. Dimensionality reduction is performed on the attributes in the multidimensional data to construct a linear mathematical model.
5. The target quantity statistics method according to claim 4, characterized in that, The specific steps for determining the main attributes are as follows: The numerical values of the main attributes of multidimensional data are denoted as yt. (1) ~yt (d) The numerical value of a non-subject attribute is denoted as xu. (1,1) ~xuA (d,w) ; Construct matrix C (y) And construct matrix C (x) ; Calculate matrix C (x) The average value wx corresponding to the elements in columns 1 to w. (1) ~wx (w) ; Calculate yt (1) ~yt (d) The average value of ayt; matrix C (x) By dividing by columns, we obtain the first to the wth non-primary attribute matrices Cx. (1) ~Matrix Cx (w) ; Calculate matrix Cx (1) ~Matrix Cx (w) The decentralized matrix is obtained as matrix Dx. (1) ~ Matrix Dx (w) ; Calculate matrix C (y) Decentralized matrix D (y) ; Define formula s-1: Among them, Cx (v) Dx represents the v-th non-subject attribute matrix; (v) Represents matrix Cx (v) The decentralization matrix; α (v) This represents the weight coefficient of the v-th non-subject attribute to the subject attribute; Calculate the weight coefficients of the first to wth non-subject attributes on the subject attributes to obtain α. (1) ~α(w); Extract α (1) ~α (w) The maximum value α in (max) Minimum value α (min) Calculate the combined weight Aα; Compare α (1) ~α (w) The non-primary attributes with a weight coefficient greater than Aα are considered as primary attributes, depending on the magnitude of Aα.
6. The target quantity statistics method according to claim 4, characterized in that, The specific steps for constructing a linear regression model are as follows: Calculate and determine whether the absolute value of the Pearson correlation coefficient of the multidimensional data and 1 is less than or equal to the error determination coefficient ζ; if it is less than or equal to ζ, construct a linear regression model; if it is greater than ζ, construct a nonlinear regression model. Constructing a linear regression model for multidimensional data; The linear regression coefficient of multidimensional data is denoted as β. (0) ~β (w) ;β (1) ~β (w) Represents the linear regression coefficients of the 1st to wth non-subject attributes; Construct matrix Bx, matrix B (β) Calculate β (0) ~β (w) Value: The mathematical expression for a linear regression model for multidimensional data is: yt (a) =b (0) +(β (1) ×xu (a,1) )+(β (2) ×xu (a,2) )+......+(b (w) ×xu (a,w) ); Among them, xu (a,1) ~xu (a,w) This represents the values of the first to the wth non-subject attributes of the a-th multidimensional data.
7. The target quantity statistics method according to claim 6, characterized in that, The specific steps for constructing a nonlinear regression model are as follows: Define the decision steps for a random forest decision tree: Define formula T-1: Among them, ix (b) wx represents the exponent of the main attribute value relative to the b-th non-main attribute value. (b) Representing matrix C (x) The average value corresponding to the elements in column b; Calculate the exponent ix of the main attribute value relative to the values of the 1st to wth non-main attributes according to formula T-1. (1) ~ix (w) ; Let the nonlinear function of the principal attribute and each non-principal attribute be fu(xu) (a,b) ), function fu(xu (a,b) The coefficients of the 1st to wth non-subject attributes in the equation are μ. (1) ~μ (w) ; Let μ be the coefficient of the b-th non-subject attribute. (b) ; function fu(xu (a,b) The initial mathematical form of ) is: Define the residual sum of squares (RSS) as the loss function, and calculate μ for the RSS. (1) ~μ (w) The partial derivative; Calculate μ for RSS (b) The partial derivative is Let the learning rate be η; let μ (b) The value after k updates is (μ) (b) (k) ), μ (b) The value after (k+1) updates is (μ) (b) (k+1) ), ; Define formula T-2: Based on the decision-making steps defined above, a nonlinear regression model is constructed using the random forest algorithm.
8. The target quantity statistics method according to claim 4, characterized in that, The specific steps for dimensionality reduction of attributes in multidimensional data are as follows: Multidimensional data where each attribute is independent is treated as independent multidimensional data. Defined matrix F: Among them, ru (1,1) ~ru (d,(w+1)) This represents the value of the first to (w+1)th attribute in the first to dth independent multidimensional data; Using the elements of each column of matrix F as clustering objects, a clustering algorithm is used to calculate the cluster centers from the 1st to the (w+1th)th column in matrix F, resulting in cu. (1) ~cu (w+1) Define matrix Fc; Dimensionality reduction is performed on matrix F to obtain matrix Fw: Fw = F * Fc; the elements in Fw are treated as one-dimensional data, and the processing steps of difference time series analysis are used to construct a linear mathematical model of independent multidimensional data.