Visual representation global modeling method and system based on feature extraction

By employing visual bag-of-vocabulary feature encoding, spatiotemporal continuous visual autoencoder, and pixel-level semantic feature mapping algorithm, combined with a high-dimensional visual representation intelligent analysis platform, the problem of insufficient local details and spatiotemporal continuity in visual representation methods is solved, achieving high-precision global feature association and improving visual processing capabilities in complex scenes.

CN121904531APending Publication Date: 2026-04-21ZHENJIANG ZHIGU HIGH END EQUIP RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHENJIANG ZHIGU HIGH END EQUIP RES INST CO LTD
Filing Date
2026-01-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing visual representation methods lack coordination in feature encoding and spatiotemporal feature learning, resulting in insufficient accuracy of local details and spatiotemporal continuity, as well as insufficient utilization of pixel-level semantic information, which affects the modeling accuracy and adaptability in complex scenes.

Method used

Feature vectors are generated using a visual bag-of-vocabulary feature encoding algorithm. Dimensionality reduction and feature reconstruction are performed using a spatiotemporal continuous visual autoencoder. A pixel-level semantic feature mapping algorithm is combined to establish the mapping relationship between pixels and semantic categories. Finally, a high-dimensional visual representation intelligent analysis platform is used for feature selection and global modeling to construct a global feature association matrix.

Benefits of technology

It improves the accuracy of local details and spatiotemporal continuity in visual data processing, reduces feature fragmentation and detail loss, enhances adaptability and analysis accuracy in complex scenarios, and provides high-precision visual processing support for fields such as autonomous driving and intelligent monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904531A_ABST
    Figure CN121904531A_ABST
Patent Text Reader

Abstract

The invention discloses a visual representation global modeling method and system based on feature extraction, and the method comprises the steps: processing input visual data through employing a visual vocabulary bag feature coding algorithm, generating a feature vector, and obtaining space-time continuous feature representation through the dimension reduction reconstruction of a space-time continuous visual autoencoder; pixel and semantic category mapping is established by using a pixel-level semantic feature mapping algorithm, a pixel-level semantic feature map is generated, and the step comprises the sub-steps of dimension analysis, model construction and the like; inputting the feature map into a high-dimensional visual representation intelligent analysis platform, and screening through sub-steps of segmentation, standardization, correlation analysis and the like to obtain high-dimensional screening features; and on the basis of the screening features, constructing and optimizing a global feature incidence matrix through a visual representation global modeling technology, and outputting a global model. The system comprises corresponding function units, the visual data global association rule is accurately captured, and adaptability and analysis accuracy in a complex scene are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual information processing technology, and in particular to a method and system for global modeling of visual representations based on feature extraction. Background Technology

[0002] In current visual information processing, with the emergence of large amounts of high-definition visual data and dynamic visual sequences, the requirements for the modeling accuracy and global correlation of visual representations are constantly increasing. Traditional visual representation methods often focus on local feature extraction, making it difficult to effectively integrate spatiotemporal information and pixel-level semantic associations, resulting in modeling limitations in tasks such as complex scene analysis and dynamic target tracking. The feature extraction-based global modeling method and system for visual representations addresses this need. Relying on technologies such as visual bag-of-vocabulary feature encoding, spatiotemporal continuous visual autoencoders, and pixel-level semantic feature mapping, combined with a high-dimensional visual representation intelligent analysis platform, it constructs a visual representation model that covers both local and global aspects, taking into account spatiotemporal and semantic dimensions. This aims to solve problems such as insufficient fusion of multi-dimensional visual information and weak global feature associations, and is applicable to multiple fields such as autonomous driving visual perception, intelligent surveillance image analysis, and medical image diagnosis, providing technical support for high-precision visual information processing.

[0003] Existing technologies have two significant drawbacks in visual representation modeling: First, existing methods lack effective coordination between feature encoding and spatiotemporal feature learning, often performing local feature encoding or spatiotemporal feature extraction separately without establishing a deep correlation mechanism between the two. This makes it difficult for the generated feature representations to simultaneously achieve accuracy in local details and spatiotemporal continuity. When processing dynamically changing visual data, feature fragmentation or loss of details is prone to occur, affecting the accuracy of subsequent modeling. Second, in the high-dimensional feature analysis and global modeling stages, existing technologies do not fully utilize pixel-level semantic information, fail to construct a complete pixel-semantic-global feature association system, and lack flexible feature selection and weight adjustment mechanisms. This makes the generated global model prone to including redundant features, making it difficult to accurately capture the global correlation patterns of visual data, resulting in reduced model adaptability and analytical accuracy in complex scenarios. Summary of the Invention

[0004] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a method and system for global modeling of visual representations based on feature extraction.

[0005] The technical solution adopted in this invention is a global modeling method for visual representation based on feature extraction, comprising the following steps: S1, using a visual bag-of-vocabularies feature encoding algorithm to encode the input visual data, and by constructing a visual vocabulary dictionary, mapping local features in the visual data to corresponding visual words to generate a visual bag-of-vocabularies feature vector; S2, using a spatiotemporal continuous visual autoencoder to learn the spatiotemporal dimension features of the visual bag-of-vocabularies feature vector obtained in S1, performing dimensionality reduction transformation on the feature vector through the encoder, and then reconstructing the features through the decoder to obtain a spatiotemporal continuous feature representation; S3, using a pixel-level semantic feature mapping algorithm to process the spatiotemporal continuous feature table output in S2. S3. Perform pixel-level semantic association processing to establish a mapping relationship between pixels and semantic categories, generating a pixel-level semantic feature map; S4. Input the pixel-level semantic feature map obtained in S3 into the high-dimensional visual representation intelligent analysis platform, and perform association analysis and feature filtering on the high-dimensional semantic features through the platform's built-in feature analysis module to obtain high-dimensional filtered features; S5. Based on the high-dimensional filtered features output in S4, use visual representation global modeling technology to perform global feature fusion, construct a global feature association matrix, and perform global modeling of visual representation; S6. Perform feature optimization processing on the global feature association matrix constructed in S5, and output the final global model of visual representation by adjusting the feature weight coefficients.

[0006] Furthermore, the feature vector generation expression of the visual bag-of-vocabularies feature encoding algorithm is as follows: ,in, For visual bag-of-vocabulary feature vectors, For the number of local features, For the first The weight coefficients of each local feature. For the first Local features, Hist is a visual vocabulary dictionary. This is a histogram statistical function used to count the frequency of words corresponding to each local feature in the visual vocabulary dictionary.

[0007] Furthermore, the feature reconstruction expression of the spatiotemporal continuous visual autoencoder is: ,in, To represent the spatiotemporal continuity features after reconstruction, The input is the visual bag-of-vocabulary feature vector. For encoder functions, Here is the encoder weight matrix. For encoder bias vector, For decoder functions, This is the decoder weight matrix. This is the decoder bias vector.

[0008] Furthermore, the semantic feature map generation expression of the pixel-level semantic feature mapping algorithm is as follows: ,in, pixel coordinates The pixel-level semantic feature map values ​​at that location. For the number of semantic categories, For activation function, For convolution operations, coordinates Spatiotemporal continuous eigenvalues ​​at that location For the first Convolution kernels corresponding to class semantics, For the first Convolution bias corresponding to class semantics For the first The identifier vector of class semantics.

[0009] Furthermore, the feature selection expression of the high-dimensional visual representation intelligent analysis platform is as follows: ,in, For high-dimensional feature selection, The high-dimensional semantic features are the input. For feature selection function, This is a correlation calculation function used to compute the autocorrelation matrix of high-dimensional semantic features. This is the relevance weighting coefficient. This is a variance calculation function used to calculate the variance of high-dimensional semantic features. This is the variance weighting coefficient.

[0010] Furthermore, the expression for constructing the global feature association matrix for the global modeling of visual representation is as follows: ,in, This is the global feature correlation matrix. The number of rows for high-dimensional feature selection. The number of columns for high-dimensional filtering features. The first feature for high-dimensional screening row vectors For high-dimensional feature selection Transpose of a column vector To model the weight matrix globally, Model the bias matrix for the global model.

[0011] Further, step S3 includes the following sub-steps: S31, obtaining the spatiotemporal continuous feature representation output in step S2, performing feature dimension analysis on the feature representation, and determining the number of spatial and channel dimensions of the feature representation to provide a dimensional basis for subsequent pixel-level semantic mapping; S32, based on the analyzed feature dimension information, constructing an initial association model for pixel-level semantic mapping, and setting the initial mapping parameters of the model, including the number of semantic categories and pixel association thresholds; S33, inputting the spatiotemporal continuous feature representation into the initial association model, and performing matching calculations on the features corresponding to each pixel and semantic category features through the feature matching module inside the model to obtain preliminary pixel-semantic matching results; S34, adjusting the mapping parameters of the association model according to the preliminary matching results, optimizing the matching accuracy between pixels and semantic categories, and finally generating a pixel-level semantic feature map.

[0012] Further, step S4 includes the following sub-steps: S41, converting the pixel-level semantic feature map obtained in S3 into a high-dimensional feature matrix, and dividing the high-dimensional feature matrix into multiple sub-feature matrices according to preset feature segmentation rules; S42, inputting the divided sub-feature matrices into the feature preprocessing module of the high-dimensional visual representation intelligent analysis platform, performing feature standardization processing on each sub-feature matrix to eliminate feature scale differences between different sub-matrices; S43, starting the platform's feature analysis module, performing feature correlation analysis on the standardized sub-feature matrices, and calculating the correlation coefficients of features within each sub-matrix and features between sub-matrices; S44, based on the correlation coefficient results, selecting feature combinations with high correlation, integrating the selected feature combinations into high-dimensional selected features, and transmitting them to step S5.

[0013] Further, S5 includes the following sub-steps: S51, receiving the high-dimensional filtered features output by S4, verifying the feature dimensions to confirm whether the feature dimensions meet the input requirements of global modeling of visual representation; if not, adjusting the dimensions; S52, constructing a feature fusion framework for global modeling of visual representation, setting the feature fusion levels and fusion parameters for each level within the framework, including fusion weights and fusion window size; S53, inputting the high-dimensional filtered features into the feature fusion framework, performing layer-by-layer fusion processing on the features through the fusion modules at each level, and adjusting the fusion parameters in real time during the fusion process; S54, constructing a global feature association matrix based on the layer-by-layer fused features, quantifying the association relationships of global features through matrix operations, and completing the global modeling of visual representation.

[0014] A global modeling system for visual representations based on feature extraction is proposed. This system, applied to a global modeling method for visual representations based on feature extraction, includes: a visual bag-of-vocabulary feature encoding unit (VBOV) for performing VBOV feature encoding on input visual data, constructing a visual vocabulary dictionary and generating VBOV feature vectors; the output of this unit is connected to a spatiotemporal continuous feature learning unit (SPLM); a SPLM receiving the feature vectors output by the VBOV feature encoding unit, performing spatiotemporal dimension feature learning through a spatiotemporal continuous visual autoencoder, and outputting spatiotemporal continuous feature representations; its output is connected to a pixel-level semantic mapping unit (PSLM); and a pixel-level semantic mapping unit receiving the spatiotemporal continuous feature representations and using pixel-level semantic feature mapping algorithms. The system establishes a mapping relationship between pixels and semantic categories and generates pixel-level semantic feature maps. The output end is connected to the high-dimensional feature analysis unit. The high-dimensional feature analysis unit receives the pixel-level semantic feature maps, performs feature association analysis and filtering through the high-dimensional visual representation intelligent analysis platform, and outputs high-dimensional filtered features. Its output end is connected to the global feature modeling unit. The global feature modeling unit receives the high-dimensional filtered features, constructs a global feature association matrix using visual representation global modeling technology, performs global visual representation modeling, and its output end is connected to the feature optimization output unit. The feature optimization output unit receives the global feature association matrix, adjusts the feature weight coefficients, and outputs the final global visual representation model. This unit is the final output module of the system.

[0015] Beneficial Effects: This invention proposes a global modeling method and system for visual representation based on feature extraction. After generating feature vectors through a visual bag-of-vocabularies feature encoding algorithm, the vectors are directly input into a spatiotemporally continuous visual autoencoder for dimensionality reduction and feature reconstruction. This establishes a close connection between local feature encoding and spatiotemporal feature learning, ensuring that feature representation possesses both local detail accuracy and spatiotemporal continuity. This effectively avoids feature fragmentation and detail loss in dynamic visual data processing, improving subsequent modeling accuracy. Addressing the issues of insufficient utilization of pixel-level semantic information and lack of flexible feature selection mechanisms in existing high-dimensional feature analysis and global modeling technologies, this method first establishes a mapping relationship between pixels and semantic categories through a pixel-level semantic feature mapping algorithm, generating a pixel-level semantic feature map. This map is then input into a high-dimensional visual representation intelligent analysis platform for association analysis and feature selection. Subsequently, it combines visual representation global modeling technology to construct a global feature association matrix and adjust feature weights, building a complete pixel-semantic-global feature association system. This reduces redundant features, accurately captures global association patterns in visual data, enhances adaptability and analysis accuracy in complex scenarios, and provides reliable support for high-precision visual processing in fields such as autonomous driving and intelligent monitoring. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method steps of the present invention; Figure 2This is a diagram showing the system unit composition of the present invention. Detailed Implementation

[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] like Figure 1 As shown, the visual representation global modeling method based on feature extraction includes the following steps: S1. The visual vocabulary bag feature encoding algorithm is used to encode the input visual data. By constructing a visual vocabulary dictionary, the local features in the visual data are mapped to the corresponding visual words to generate the visual vocabulary bag feature vector. Specifically, step S1 is the core step of processing the input visual data using the visual bag-of-vocabulary feature encoding algorithm to generate feature vectors. During implementation, the basic parameters of the input visual data must first be clearly defined. Typically, an RGB color image with a resolution of 1920×1080 pixels or a dynamic video sequence with a frame rate of 30 frames per second is selected to ensure the data has sufficient visual detail and information density. Next, a visual vocabulary dictionary is constructed. This requires extracting local features from 100,000 to 500,000 sets of sample visual data. Extraction is done using a sliding method with a fixed 16×16 pixel window. Within each window, 64-dimensional or 128-dimensional feature descriptors are extracted. Then, the K-means clustering algorithm is used to cluster the descriptors, with the number of clusters set to 1000 to 5000. Each cluster center is a visual vocabulary, ultimately forming a dictionary containing 1000 to 5000 words. Subsequently, local features are extracted from the input visual data using the same parameters. Each feature is matched with words in the dictionary using the nearest neighbor matching algorithm. The number of matches for each word is counted, and a feature vector with the same dimension as the number of words in the dictionary is generated. This vector can accurately represent the local feature distribution of the input data, providing reliable basic data for subsequent spatiotemporal feature learning and ensuring the accuracy of feature processing in subsequent steps.

[0019] S2. Spatiotemporal continuous visual autoencoder is used to learn the spatiotemporal dimension features of the visual bag-of-vocabularies feature vector obtained in S1. The feature vector is dimensionality reduced by the encoder and then reconstructed by the decoder to obtain spatiotemporal continuous feature representation. Specifically, step S2 uses a spatiotemporal continuous visual autoencoder to learn the spatiotemporal dimension features of the feature vectors generated in S1. During implementation, the autoencoder network structure parameters are first determined: the encoder is a 3-layer fully connected network with the same number of neurons in the input layer as the feature vector dimension of S1 (1000 to 5000), the hidden layer has 500 to 2000 neurons and 100 to 500 neurons respectively, and the output layer (bottleneck layer) has 50 to 200 neurons, achieving feature dimensionality reduction. The decoder is also a 3-layer fully connected network with the same number of neurons in the input and bottleneck layers, the hidden layers have 100 to 500 neurons and 500 to 2000 neurons respectively, and the output layer is the same as the encoder input layer, used for feature reconstruction. During training, the mean squared error is used as the loss function, the learning rate is set to 0.001 to 0.01, and the iterations are 1000 to 5000 times. 32 to 128 sets of feature vectors are processed in batches. The training data uses 10,000 to 100,000 sets of feature vectors of the same type as the input data of S1 to ensure that the model grasps the general spatiotemporal feature patterns. After training, the S1 feature vector is input into the encoder and processed by a linear transformation and activation function to obtain a low-dimensional vector. Then, it is input into the decoder and reconstructed by inverse transformation to output a spatiotemporally continuous feature representation. This representation retains local feature information and incorporates spatiotemporal correlation rules, effectively improving the global feature expression ability and providing high-quality feature support for subsequent semantic mapping.

[0020] S3. Utilize the pixel-level semantic feature mapping algorithm to perform pixel-level semantic association processing on the spatiotemporal continuous feature representation output by S2, establish the mapping relationship between pixels and semantic categories, and generate a pixel-level semantic feature map. Specifically, step S3 uses a pixel-level semantic feature mapping algorithm to establish a mapping between pixels and semantic categories for the spatiotemporal continuous feature representation output by S2. During implementation, the dimension of the feature representation in S2 is first analyzed. If the input data is a 1920×1080 pixel image, the feature representation space dimension is downsampled to 240×270 pixels, and the channel dimension is consistent with the output of the S2 decoder (50 to 200 channels), clarifying the processing range and scale. Next, an initial association model is constructed, setting 10 to 50 semantic categories (including target objects, background regions, etc.), with a pixel association threshold of 0.6 to 0.8. The model uses a convolutional neural network with 3 convolutional layers (3×3 pixel convolutional kernels, stride 1, same padding) and 2 pooling layers (2×2 pixel pooling kernels, stride 2) to extract semantic features. The feature representation is input into the model, where multi-scale semantic features are extracted by convolutional layers, and dimensionality is reduced by pooling layers. Finally, a fully connected layer calculates the matching probability value for each semantic category corresponding to each pixel. The semantic attribution of a pixel is determined based on its probability value and association threshold. If the probability value is higher than the threshold, the corresponding category is labeled. At the same time, the model parameters (such as convolution kernel weights and association thresholds) are adjusted based on the labeling results. The matching accuracy is improved by iterative optimization 500 to 1000 times. Finally, a pixel-level semantic feature map is generated. Each pixel in the map carries clear semantic information, which lays the semantic foundation for subsequent high-dimensional feature analysis and improves the targeting of feature analysis.

[0021] S4. Input the pixel-level semantic feature map obtained in S3 into the high-dimensional visual representation intelligent analysis platform. Through the feature analysis module built into the platform, perform correlation analysis and feature filtering on the high-dimensional semantic features to obtain high-dimensional filtered features. Specifically, in step S4, the pixel-level semantic feature map generated in S3 is input into the high-dimensional visual representation intelligent analysis platform for feature selection. During implementation, the feature map format is first converted to a high-dimensional feature matrix. If the feature map is 240×270 pixels and contains 10 to 50 semantic classes, the matrix dimension is set to (240×270)×(10 to 50), i.e., 64800×50, maintaining the integrity and accuracy of pixel semantic information during the conversion. The matrix is ​​divided into 36 to 100 sub-feature matrices according to a uniform segmentation rule, with each sub-matrix having a dimension of 720×50 to 1740×50, reducing the data size for a single analysis and improving processing efficiency. The sub-matrices are input into the platform's feature preprocessing module, where standardization converts the elements into normally distributed data with a mean of 0 and a standard deviation of 1, eliminating scale differences and ensuring a unified analysis benchmark. The feature analysis module is then activated, using the Pearson correlation coefficient algorithm to calculate the correlation coefficients (ranging from -1 to 1) of features within and between sub-matrices, setting a threshold of 0.7 to 0.9 to select highly correlated feature combinations and avoid redundant feature interference. Finally, the filtered features are integrated according to their original pixel positions to form high-dimensional filtered features ranging from 32400×50 to 58320×50. This maintains spatial correlation and semantic coherence, providing high-quality feature data for subsequent global modeling and ensuring the accuracy of the global model.

[0022] S5. Based on the high-dimensional filtering features output by S4, global feature fusion is performed using visual representation global modeling technology to construct a global feature association matrix and perform global modeling of visual representation. Specifically, step S5, based on the high-dimensional filtering features of S4, uses visual representation global modeling technology to construct a global feature association matrix. During implementation, the dimension of the filtered features is first verified. If the dimension is between 32400×50 and 58320×50, it needs to be confirmed that the row dimension (number of pixels) is 1 / 4 to 1 / 2 of the number of pixels in the input data, and the column dimension (semantic feature dimension) is consistent with the number of semantic categories in S3. If they do not match, interpolation or cropping methods are used to adjust them to meet the modeling input standards. A graph neural network structure fusion framework containing a feature input layer, a fusion layer, and a matrix output layer is constructed. The fusion layer has 3 to 5 levels, with each level having a fusion window of 5×5 to 9×9 pixels to capture feature associations in different ranges. The initial fusion weights are 0.1 to 0.3 (to be optimized during subsequent training). The validated selected features are input into the framework and sequentially passed to each fusion level. Local region feature associations are calculated via a window and weighted summation (weights being the fusion weights) to obtain local fusion features. During fusion, a gradient descent algorithm (learning rate 0.001 to 0.01, iterations 300 to 800) is used to adjust the weights in real time to ensure accurate reflection of local association patterns. Based on the local fusion features, a global feature association matrix of (32400 to 58320) × (32400 to 58320) is constructed through matrix operations. The element values ​​(0 to 1) represent the association strength between corresponding two pixel features, achieving global association quantification of visual representation, completing global modeling, providing a global feature foundation for subsequent optimization, and improving the overall association of the model.

[0023] S6. Perform feature optimization processing on the global feature association matrix constructed in S5. By adjusting the feature weight coefficients, output the final global visual representation model.

[0024] Specifically, step S6 optimizes the global feature association matrix from S5 and outputs the final model. During implementation, the optimization objective and parameters are first determined. The objective is to reduce matrix redundancy and noise, and increase the proportion of effective associations. The parameters include a feature weight adjustment threshold of 0.2 to 0.4 (to filter effective associations), a regularization coefficient of 0.001 to 0.01 (to prevent overfitting), and 200 to 500 iterations (to ensure effectiveness). A threshold screening method is used for initial optimization. Matrix elements are iterated; those below the threshold are set to 0 to eliminate weakly associated features and reduce redundancy, while those above the threshold retain effective associations. L2 regularization is introduced, adding a regularization term (coefficient multiplied by the sum of squares of matrix elements) to the loss function to constrain the range of element values, avoiding overfitting and adapting to different visual data. Cross-validation is used for evaluation during optimization. 1000 to 5000 sets of test data of the same type as the input data are selected. The optimized matrix is ​​used for modeling, and the accuracy of target semantic recognition and feature association is calculated. The optimization stops when it reaches 90% or higher; otherwise, the threshold and coefficients are adjusted and the optimization is repeated. Finally, the optimized matrix is ​​integrated with key parameters from the early stages (such as the visual vocabulary dictionary, autoencoder parameters, and the number of semantic categories) to form a complete global visual representation model. This model contains global feature association information and parameters for each stage, and can be directly used for visual data modeling and analysis, providing efficient and accurate technical support for fields such as autonomous driving and intelligent monitoring, and meeting practical application needs.

[0025] Preferably, the feature vector generation expression of the visual bag-of-vocabularies feature encoding algorithm is: ,in, For visual bag-of-vocabulary feature vectors, For the number of local features, For the first The weight coefficients of each local feature. For the first Local features, Hist is a visual vocabulary dictionary. This is a histogram statistical function used to count the frequency of words corresponding to each local feature in the visual vocabulary dictionary.

[0026] Specifically, the feature vector generation process of the visual bag-of-vocabularies feature encoding algorithm requires first clarifying the core parameters and computational logic of the algorithm to ensure that the feature vector accurately reflects the local feature distribution of the input visual data. First, determine the number of local features. Based on the resolution and complexity of the input visual data, set the number of local features to 500 to 2000. The higher the resolution and the more complex the scene, the larger the number of local features should be to cover more detailed information. The weight coefficient of each local feature needs to be set in conjunction with its importance. By statistically analyzing the frequency and discriminative power of local features in the sample data, adjust the weight coefficient to 0.5 to 1.2. Features with high frequency and strong discriminative power have larger weight coefficients to enhance their contribution to the final feature vector. The construction of the visual vocabulary dictionary should maintain a scale of 1000 to 5000 words to ensure that the dictionary includes common visual feature types. In the computation process, each local feature is first extracted. Then, a histogram statistical function is used to count the frequency of the corresponding word in the visual vocabulary dictionary for that feature. The weight coefficient of each local feature is multiplied by its corresponding frequency, and the results are summed to obtain the final visual vocabulary feature vector. This process highlights key local features by adjusting the weight coefficients and ensures the integrity of the feature distribution through histogram statistics. The generated feature vector provides an accurate local feature foundation for subsequent spatiotemporal continuous feature learning, effectively improving the efficiency and accuracy of feature processing in subsequent steps.

[0027] Preferably, the feature reconstruction expression of the spatiotemporal continuous visual autoencoder is: ,in, To represent the spatiotemporal continuity features after reconstruction, The input is the visual bag-of-vocabulary feature vector. For encoder functions, Here is the encoder weight matrix. For encoder bias vector, For decoder functions, This is the decoder weight matrix. This is the decoder bias vector.

[0028] Specifically, in the feature reconstruction process of a spatiotemporally continuous visual autoencoder, it is crucial to clearly define the parameter settings and reconstruction logic of the encoder and decoder to ensure the spatiotemporal continuity of the reconstructed feature representation. The dimension of the encoder's weight matrix must match the dimension of the input feature vector and the number of neurons in the hidden layer. If the dimension of the input feature vector is 1000 to 5000, and the number of neurons in the first hidden layer is 500 to 2000, then the dimension of the encoder's first-layer weight matrix should be set to (1000 to 5000) × (500 to 2000), the dimension of the second-layer weight matrix should be set to (500 to 2000) × (100 to 500), and the dimension of the bottleneck layer weight matrix should be set to (100 to 500) × (50 to 200). The dimension of the bias vector should be consistent with the number of neurons in the corresponding layer, set to 1 × (500 to 2000), 1 × (100 to 500), and 1 × (50 to 200). The dimensions of the decoder weight matrix are inversely related to those of the encoder. The first layer is set to (50 to 200) × (100 to 500), the second layer to (100 to 500) × (500 to 2000), and the third layer to (500 to 2000) × (1000 to 5000). The dimension of the bias vector is consistent with the number of neurons in each layer of the decoder. During computation, the generated visual bag-of-vocabulary feature vector is first input into the encoder, and after linear transformation of the weight matrix of each layer and processing by the activation function, a low-dimensional feature vector is obtained. Then, the low-dimensional vector is input into the decoder, and the features are reconstructed through inverse transformation. This process ensures the accuracy of feature dimensionality reduction and reconstruction through the matching weight matrix and bias vector. The generated spatiotemporally continuous feature representation can retain local features and incorporate spatiotemporal correlations, providing high-quality feature input for pixel-level semantic feature mapping.

[0029] Preferably, the semantic feature map generation expression of the pixel-level semantic feature mapping algorithm is: ,in, pixel coordinates The pixel-level semantic feature map values ​​at that location. For the number of semantic categories, For activation function, For convolution operations, coordinates Spatiotemporal continuous eigenvalues ​​at that location For the first Convolution kernels corresponding to class semantics, For the first Convolution bias corresponding to class semantics For the first The identifier vector of class semantics.

[0030] Specifically, the semantic feature map generation process of the pixel-level semantic feature mapping algorithm requires clearly defining the number of semantic categories, convolution kernel parameters, and computational logic to ensure that each pixel accurately matches a semantic category. The number of semantic categories is set to 10 to 50 categories based on the application scenario, including core categories such as target objects and background regions. The number of categories needs to be adapted to the semantic complexity of the input visual data; the more complex the scenario, the larger the number of categories. The convolution kernel size is uniformly set to 3×3 pixels, and the number is consistent with the number of semantic categories (10 to 50). Each convolution kernel corresponds to one type of semantic feature. The convolution kernel parameters are adjusted through pre-training so that each convolution kernel can extract feature information corresponding to its semantic category. The convolution bias vector dimension is set to 1×(10 to 50), with each element corresponding to a semantic category bias value, ranging from -0.5 to 0.5, used to adjust the offset of the convolution operation result. The activation function uses the Sigmoid or ReLU function to map the convolution operation result to the range of 0 to 1, facilitating subsequent probability judgment. During computation, convolution operations are first performed on each pixel location represented by spatiotemporally continuous features. After superimposing the bias value corresponding to the semantic category, the result is processed by an activation function and then multiplied by the label vector of that semantic category to obtain the feature value corresponding to that semantic category for that pixel. The calculation results for all semantic categories are then aggregated to form a pixel-level semantic feature map. This process ensures the targeted nature of semantic feature extraction through dedicated convolution kernels and biases, while activation functions and label vectors improve semantic matching accuracy. The generated feature map provides clear semantic support for high-dimensional feature analysis.

[0031] Preferably, the feature selection expression of the high-dimensional visual representation intelligent analysis platform is: ,in, For high-dimensional feature selection, The high-dimensional semantic features are the input. For feature selection function, This is a correlation calculation function used to compute the autocorrelation matrix of high-dimensional semantic features. This is the relevance weighting coefficient. This is a variance calculation function used to calculate the variance of high-dimensional semantic features. This is the variance weighting coefficient.

[0032] Specifically, the feature selection process of the high-dimensional visual representation intelligent analysis platform requires determining the relevance weight coefficient, variance weight coefficient, and selection logic to ensure that the selected features possess high value and low redundancy. The relevance weight coefficient is set to 0.4 to 0.7 based on the importance of feature relevance; if greater emphasis is placed on the correlation patterns between features, the coefficient value is set higher. The variance weight coefficient is set to 0.3 to 0.6, and its sum with the relevance weight coefficient is 1, used to balance the impact of feature dispersion on the selection results. The dimension of the high-dimensional semantic feature matrix needs to match the pixel-level semantic feature map. If the feature map is 240×270 pixels with 10 to 50 semantic classes, then the dimension of the high-dimensional semantic feature matrix is ​​64800×50. The calculation process begins by multiplying the high-dimensional semantic feature matrix by its transpose using a correlation calculation function to obtain the feature autocorrelation matrix, reflecting the strength of the association between features. Next, a variance calculation function is used to calculate the variance of each feature in the high-dimensional semantic feature matrix, reflecting the feature's dispersion. The autocorrelation matrix is ​​then multiplied by a correlation weighting coefficient, and the variance is multiplied by a variance weighting coefficient; the sum of these two values ​​yields a comprehensive evaluation value for feature selection. Finally, a feature selection function selects features with comprehensive evaluation values ​​higher than a preset threshold (0.6 to 0.8), forming the high-dimensional selected features. This process balances the influence of correlation and variance through weighting coefficients, ensuring the rationality of feature selection through comprehensive evaluation. The generated high-dimensional selected features provide high-quality feature data for global modeling of visual representations, reducing redundant interference.

[0033] Preferably, the expression for constructing the global feature association matrix for the global modeling of visual representation is: ,in, This is the global feature correlation matrix. The number of rows for high-dimensional feature selection. The number of columns for high-dimensional filtering features. The first feature for high-dimensional screening row vectors For high-dimensional feature selection Transpose of a column vector To model the weight matrix globally, Model the bias matrix for the global model.

[0034] Specifically, the process of constructing the global feature association matrix for global modeling of visual representation requires clearly defining the number of rows and columns of the high-dimensional selected features, the weight matrix, and the bias matrix parameters to ensure that the matrix can accurately quantify the global feature associations. The number of rows for the high-dimensional selected features is set to 32,400 to 58,320 rows based on the number of pixels, and the number of columns is consistent with the number of semantic categories, ranging from 10 to 50 columns. The number of rows and columns must match the dimension of the high-dimensional selected features. The global modeling weight matrix is ​​set to a dimension of (10 to 50) × (10 to 50), where each element represents the association weight between different semantic features. Through prior training and adjustment, elements corresponding to closely associated features have larger values ​​(0.5 to 1.2), while elements with weak associations have smaller values ​​(0.1 to 0.4). The global modeling bias matrix has the same dimension as the weight matrix, with element values ​​ranging from -0.3 to 0.3, used to adjust the baseline values ​​of the matrix operation results. During calculation, the transpose of each row vector and column vector of the high-dimensional selected features is first multiplied to obtain the local feature correlation matrix. The sum of all local feature correlation matrices is then divided by the product of the number of rows and columns to obtain the average local correlation matrix. Finally, the average local correlation matrix is ​​multiplied by the global modeling weight matrix, and the global modeling bias matrix is ​​superimposed to obtain the global feature correlation matrix. This process balances local correlation differences through averaging, optimizes correlation quantization accuracy through weight and bias matrices, and generates a global feature correlation matrix that provides a global correlation foundation for feature optimization, improving the accuracy and completeness of the final visual representation of the global model.

[0035] Preferably, step S3 includes the following sub-steps: S31, obtaining the spatiotemporal continuous feature representation output in step S2, performing feature dimension analysis on the feature representation, and determining the number of spatial and channel dimensions of the feature representation to provide a dimensional basis for subsequent pixel-level semantic mapping; S32, based on the analyzed feature dimension information, constructing an initial association model for pixel-level semantic mapping, and setting the initial mapping parameters of the model, including the number of semantic categories and pixel association thresholds; S33, inputting the spatiotemporal continuous feature representation into the initial association model, and performing matching calculations on the features corresponding to each pixel and semantic category features through the feature matching module inside the model to obtain preliminary pixel-semantic matching results; S34, adjusting the mapping parameters of the association model according to the preliminary matching results, optimizing the matching accuracy between pixels and semantic categories, and finally generating a pixel-level semantic feature map.

[0036] Specifically, step S3 includes the implementation details of four sub-steps, S31 to S34, to ensure accurate generation of pixel-level semantic feature maps. In S31, the spatiotemporal continuous feature representation output from S2 is first acquired. Its spatial and channel dimensions are determined using a feature dimension analysis tool. If the input visual data is 1920×1080 pixels, the spatial dimension of the spatiotemporal continuous feature representation is downsampled to 240×270 pixels. The channel dimension is consistent with the output of the S2 decoder, ranging from 50 to 200, laying the dimensional foundation for subsequent mapping. In S32, when constructing the initial association model, the number of semantic categories is set to 10 to 50, and the pixel association threshold is set to 0.6 to 0.8. The model uses three convolutional layers (3×3 pixel convolutional kernels, stride 1, same padding) and two pooling layers (2×2 pixel pooling kernels, stride 2) to complete the initial parameter configuration. After inputting the spatiotemporal continuous feature representation into the model, S33 calculates the matching probability between each pixel feature and the semantic category feature through the feature matching module. The matching calculation uses the cosine similarity algorithm to ensure the reliability of the initial matching results. S34 adjusts the model parameters based on the initial results, setting the convolution kernel weight adjustment range to ±0.1 and the association threshold fine-tuning step size to 0.05. Iterates and optimizes 500 to 1000 times until the pixel semantic matching accuracy reaches more than 90%, finally generating a pixel-level semantic feature map, providing semantically clear feature input for S4.

[0037] Preferably, step S4 includes the following sub-steps: S41, converting the pixel-level semantic feature map obtained in S3 into a high-dimensional feature matrix, and dividing the high-dimensional feature matrix into multiple sub-feature matrices according to preset feature segmentation rules; S42, inputting the divided sub-feature matrices into the feature preprocessing module of the high-dimensional visual representation intelligent analysis platform, performing feature standardization processing on each sub-feature matrix to eliminate feature scale differences between different sub-matrices; S43, starting the platform's feature analysis module, performing feature correlation analysis on the standardized sub-feature matrices, and calculating the correlation coefficients of features within each sub-matrix and features between sub-matrices; S44, based on the correlation coefficient results, selecting feature combinations with high correlation, integrating the selected feature combinations into high-dimensional selected features, and transmitting them to step S5.

[0038] Specifically, step S4 achieves high-dimensional feature filtering through four sub-steps, S41 to S44, ensuring high-quality features. S41 converts the pixel-level semantic feature map generated in S3 into a high-dimensional feature matrix. If the feature map is 240×270 pixels and contains 10 to 50 semantic classes, the matrix dimension is set to 64800×50. A uniform segmentation method is used to divide it into 36 to 100 sub-feature matrices, each with dimensions ranging from 720×50 to 1740×50, reducing the data processing scale. S42 inputs the sub-matrices into the preprocessing module of the high-dimensional visual representation intelligent analysis platform. Standardization is used to convert the elements into normally distributed data with a mean of 0 and a standard deviation of 1, eliminating feature scale differences between sub-matrices and controlling the processing error within ±0.01. S43 activates the feature analysis module, using the Pearson correlation coefficient algorithm to calculate the correlation coefficients of features within and between sub-matrices. The calculation precision is retained to four decimal places, and the correlation coefficient ranges from -1 to 1, accurately reflecting the strength of feature association. S44 sets the correlation coefficient threshold to 0.7 to 0.9, filters out feature combinations that are higher than the threshold, and integrates them into high-dimensional filtered features of 32400×50 to 58320×50 according to the original pixel position. During the integration process, the feature position deviation does not exceed 1 pixel to ensure the correlation of feature space and provide high-quality feature data for global modeling of S5.

[0039] Preferably, step S5 includes the following sub-steps: S51, receiving the high-dimensional filtered features output from S4, verifying the feature dimensions to confirm whether the feature dimensions meet the input requirements for global modeling of visual representation; if not, adjusting the dimensions; S52, constructing a feature fusion framework for global modeling of visual representation, setting the feature fusion levels and fusion parameters for each level within the framework, including fusion weights and fusion window sizes; S53, inputting the high-dimensional filtered features into the feature fusion framework, performing layer-by-layer fusion processing on the features through fusion modules at each level, and adjusting the fusion parameters in real time during the fusion process; S54, constructing a global feature association matrix based on the layer-by-layer fused features, quantifying the association relationships of global features through matrix operations, and completing the global modeling of visual representation.

[0040] Specifically, step S5 completes the global modeling of visual representation through four sub-steps, S51 to S54, to ensure the accuracy of the global feature association matrix. S51 receives the high-dimensional filtered features output from S4 and verifies whether their dimensions meet the modeling requirements. If the feature dimensions are between 32400×50 and 58320×50, it needs to be confirmed that the row dimension is 1 / 4 to 1 / 2 of the number of pixels in the input data, and the column dimension is consistent with the number of semantic categories in S3. If they do not match, linear interpolation is used for adjustment, with the dimension adjustment error not exceeding 5%. S52 constructs a feature fusion framework, setting 3 to 5 fusion levels. The fusion window size for each level is 5×5 to 9×9 pixels, and the initial fusion weights are 0.1 to 0.3, with weight precision retained to 3 decimal places, providing a parameter basis for feature fusion. S53 inputs the validated filtering features into the framework. Each level calculates the local region feature associations through a fusion window and performs a weighted summation. A gradient descent algorithm is used to adjust the fusion weights in real time, with a learning rate of 0.001 to 0.01 and 300 to 800 iterations to ensure that the fused features accurately reflect the local association patterns. S54 constructs a global feature association matrix based on the local fused features. The matrix dimension is (32400 to 58320) × (32400 to 58320), and the element values ​​represent the pixel feature association strength (0 to 1). The matrix calculation error is controlled within ±0.005, completing the global modeling of visual representation and providing a reliable global feature foundation for the feature optimization in S6.

[0041] The visual bag-of-words feature encoding algorithm is the core technology in this invention used to extract local features from input visual data and generate standardized feature vectors. Essentially, it simulates the "bag of words" concept in text processing, transforming local features in visual data into quantifiable combinations of "visual words." In implementation, local features are first extracted from 100,000 to 500,000 sets of sample visual data. A fixed 16×16 pixel sliding window is used, with each window extracting 64-dimensional or 128-dimensional feature descriptors. Then, K-means clustering (1000 to 5000 clusters) is used to generate a visual vocabulary dictionary. Next, local features are extracted from the target visual data using the same parameters. Nearest neighbor matching maps each feature to a visual word in the dictionary, and the number of matches for each word is counted. Combined with feature weight coefficients of 0.5 to 1.2 (set according to feature importance), a feature vector with the same dimensions as the dictionary is generated. This algorithm transforms complex visual data into structured feature vectors, eliminating interference caused by differences in visual data formats. It provides a unified and accurate local feature foundation for subsequent spatiotemporal continuous feature learning, avoiding subsequent modeling deviations caused by non-standard local feature representations. At the same time, it highlights key features through weight adjustment, improving the efficiency and relevance of overall feature processing, and laying a reliable initial feature support for the entire global modeling process of visual representation.

[0042] The spatiotemporal continuous visual autoencoder is a technique used in this invention for learning and reconstructing spatiotemporal features from local feature vectors. Essentially, it's a neural network structure containing an encoder and a decoder, capable of preserving the spatiotemporal correlation information of features during dimensionality reduction. In implementation, a 3-layer fully connected encoder and a 3-layer fully connected decoder are constructed: the number of neurons in the encoder's input layer matches the dimension of the visual bag-of-vocabularies feature vectors (1000 to 5000), the hidden layers have 500 to 2000 neurons, the bottleneck layer has 100 to 500 neurons, and the decoder has 50 to 200 neurons. The decoder structure is symmetrical to the encoder, with the input layer matching the bottleneck layer dimension and the output layer matching the encoder's input layer dimension. During training, the mean squared error is used as the loss function, the learning rate is 0.001 to 0.01, and the iterations are 1000 to 5000 times. 32 to 128 sets of feature vectors are processed in batches, and the model is trained using 10,000 to 100,000 sets of sample data. In use, the local feature vectors are input to the encoder for dimensionality reduction, then reconstructed by the decoder, outputting a spatiotemporally continuous feature representation. This autoencoder integrates local features with spatiotemporal correlation information, solving the problem of lack of temporal and spatial continuity of local features. It provides feature inputs with both details and spatiotemporal regularity for subsequent pixel-level semantic feature mapping, avoids feature fragmentation in dynamic visual data processing, and improves the global expressive power of features. It is a key technical bridge connecting local features and semantic features.

[0043] The pixel-level semantic feature mapping algorithm is a technique used in this invention to establish associations between pixels and semantic categories and generate feature maps with semantic information. Essentially, it transforms abstract features into pixel labels with clear semantic meanings through neural networks and matching rules. In implementation, the dimensions of the spatiotemporally continuous feature representation are first analyzed (e.g., 240×270 pixel spatial dimension, 50 to 200 channel dimensions); then, an initial association model is constructed containing 3 convolutional layers (3×3 convolutional kernels, stride 1, same padding) and 2 pooling layers (2×2 pooling kernels, stride 2), setting 10 to 50 semantic categories and a pixel association threshold of 0.6 to 0.8; the feature representation is input into the model, and after convolution to extract semantic features and pooling to reduce dimensionality, the matching probability of each pixel corresponding to each semantic category is calculated through a fully connected layer; finally, the semantic attribution of the pixel is determined based on the probability and threshold, and the model parameters are iteratively optimized (convolutional kernel weights adjusted by ±0.1, threshold fine-tuning stride 0.05, iterated 500 to 1000 times) to generate a pixel-level semantic feature map. This algorithm assigns explicit semantic labels to pixels, establishes a direct association between features and semantics, solves the problem of traditional features lacking semantic information, enables subsequent high-dimensional feature analysis to be targeted based on semantics, reduces the interference of meaningless features, provides semantic-level feature support for global modeling, and significantly improves the accuracy of the final model's semantic understanding of visual data.

[0044] The High-Dimensional Visual Representation Intelligent Analysis Platform, as described in this invention, is an integrated tool for high-dimensional feature processing and filtering of pixel-level semantic feature maps. Essentially, it's a comprehensive system including preprocessing, analysis, and filtering modules, focusing on extracting high-value information from high-dimensional semantic features. In implementation, the pixel-level semantic feature map (e.g., 240×270 pixels, 10 to 50 semantic categories) is first converted into a 64800×50 high-dimensional feature matrix, then divided into 36 to 100 sub-feature matrices (each ranging from 720×50 to 1740×50) using a uniform segmentation method. The preprocessing module standardizes the sub-matrix elements to a normal distribution with a mean of 0 and a standard deviation of 1 (error within ±0.01). The analysis module calculates the correlation coefficient between features using the Pearson correlation coefficient algorithm (precision retained to 4 decimal places). Finally, highly correlated feature combinations are filtered using a threshold of 0.7 to 0.9, and integrated into high-dimensional filtered features ranging from 32400×50 to 58320×50 (positional deviation ≤ 1 pixel) based on the original pixel position. This platform reduces the redundancy of high-dimensional features, extracts core features valuable for global modeling, and solves the problems of large data volume, computational complexity and redundancy of high-dimensional features. It ensures the uniformity of analysis benchmarks through standardization and retains key features through correlation screening, providing high-quality and low-redundancy feature inputs for subsequent global modeling of visual representation, reducing the computational burden of global modeling, improving the efficiency of the model in utilizing key features, and ensuring the accuracy and practicality of the final global model.

[0045] like Figure 2As shown, a visual representation global modeling system based on feature extraction is applied to a feature extraction-based visual representation global modeling method. The system includes: a visual bag-of-vocabulary feature encoding unit, used to perform visual bag-of-vocabulary feature encoding on the input visual data, constructing a visual vocabulary dictionary and generating visual bag-of-vocabulary feature vectors; the output of this unit is connected to a spatiotemporal continuous feature learning unit; a spatiotemporal continuous feature learning unit, used to receive the feature vectors output by the visual bag-of-vocabulary feature encoding unit, learn spatiotemporal dimension features through a spatiotemporal continuous visual autoencoder, and output spatiotemporal continuous feature representations; its output is connected to a pixel-level semantic mapping unit; and a pixel-level semantic mapping unit, used to receive spatiotemporal continuous feature representations and utilize pixel-level semantic feature mapping... The algorithm establishes a mapping relationship between pixels and semantic categories and generates pixel-level semantic feature maps. The output end is connected to the high-dimensional feature analysis unit. The high-dimensional feature analysis unit receives the pixel-level semantic feature maps, performs feature association analysis and filtering through the high-dimensional visual representation intelligent analysis platform, and outputs high-dimensional filtered features. Its output end is connected to the global feature modeling unit. The global feature modeling unit receives the high-dimensional filtered features, constructs a global feature association matrix using visual representation global modeling technology, performs global visual representation modeling, and its output end is connected to the feature optimization output unit. The feature optimization output unit receives the global feature association matrix, adjusts the feature weight coefficients, and outputs the final global visual representation model. This unit is the final output module of the system.

[0046] A global modeling method and system for visual representation based on feature extraction directly feeds the feature vectors generated by the visual bag-of-vocabulary feature encoding into a spatiotemporally continuous visual autoencoder for dimensionality reduction and reconstruction through a fixed process, forming a coherent "encoding-learning-reconstruction" chain, rather than processing local features and spatiotemporal information in isolation. This design allows feature representation to retain the local details captured by the visual bag-of-vocabulary algorithm while incorporating the spatiotemporally continuous attributes learned by the autoencoder, completely solving the problems of feature fragmentation and detail loss that easily occur in dynamic visual data processing. This lays a high-precision feature foundation for subsequent modeling stages and significantly improves the reliability of the overall modeling.

[0047] This method and system possess significant advantages in semantic information utilization and global modeling mechanisms, precisely addressing the shortcomings of existing technologies in semantic utilization and rigid filtering mechanisms. It first establishes a direct association between pixels and semantic categories through a pixel-level semantic feature mapping algorithm, generating feature maps rich in semantic information. Then, relying on a high-dimensional visual representation intelligent analysis platform, it performs association analysis and dynamic filtering on the feature maps, rather than simply using a fixed feature set. Subsequently, it combines visual representation global modeling technology to construct an association matrix and adjust feature weights, forming a complete system of "pixel semantic mapping - high-dimensional feature filtering - global association modeling." This fully leverages the semantic value of pixels while reducing redundant features through flexible filtering, accurately capturing global patterns in visual data, significantly enhancing adaptability and analytical accuracy in complex scenarios, and meeting the high-precision visual processing needs of multiple fields.

[0048] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A global modeling method for visual representation based on feature extraction, characterized in that, Includes the following steps: S1. The visual vocabulary bag feature encoding algorithm is used to encode the input visual data. By constructing a visual vocabulary dictionary, the local features in the visual data are mapped to the corresponding visual words to generate the visual vocabulary bag feature vector. S2. Using a spatiotemporal continuous visual autoencoder, the visual bag-of-vocabularies feature vector obtained in S1 is subjected to spatiotemporal dimension feature learning. The feature vector is then subjected to dimensionality reduction transformation by the encoder and feature reconstruction by the decoder to obtain spatiotemporal continuous feature representation. S3. The spatiotemporal continuous feature representation output by S2 is subjected to pixel-level semantic association processing using a pixel-level semantic feature mapping algorithm to establish the mapping relationship between pixels and semantic categories and generate a pixel-level semantic feature map. S4. The pixel-level semantic feature map obtained in S3 is input into a high-dimensional visual representation intelligent analysis platform. The high-dimensional semantic features are subjected to association analysis and feature filtering by the built-in feature analysis module of the platform to obtain high-dimensional filtered features. S5. Based on the high-dimensional filtering features output by S4, global feature fusion is performed using visual representation global modeling technology to construct a global feature association matrix and perform global modeling of visual representation; S6. Feature optimization processing is performed on the global feature association matrix constructed by S5. By adjusting the feature weight coefficients, the final global model of visual representation is output.

2. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, The feature vector generation expression of the visual bag-of-vocabularies feature encoding algorithm is: ,in, For visual bag-of-vocabulary feature vectors, For the number of local features, For the first The weight coefficients of each local feature. For the first Local features, Hist is a visual vocabulary dictionary. This is a histogram statistical function used to count the frequency of words corresponding to each local feature in the visual vocabulary dictionary.

3. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, The feature reconstruction expression of the spatiotemporal continuous visual autoencoder is: ,in, To represent the spatiotemporal continuity features after reconstruction, The input is the visual bag-of-vocabularies feature vector. For encoder functions, Here is the encoder weight matrix. For encoder bias vector, For decoder functions, This is the decoder weight matrix. This is the decoder bias vector.

4. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, The semantic feature map generation expression of the pixel-level semantic feature mapping algorithm is: ,in, pixel coordinates The pixel-level semantic feature map values ​​at that location. For the number of semantic categories, For activation function, For convolution operations, coordinates Spatiotemporal continuous eigenvalues ​​at that location For the first Convolution kernels corresponding to class semantics, For the first Convolution bias corresponding to class semantics For the first The identifier vector of class semantics.

5. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, The feature selection expression of the high-dimensional visual representation intelligent analysis platform is: ,in, For high-dimensional feature selection, The high-dimensional semantic features are the input. For feature selection function, This is a correlation calculation function used to compute the autocorrelation matrix of high-dimensional semantic features. This is the relevance weighting coefficient. This is a variance calculation function used to calculate the variance of high-dimensional semantic features. This is the variance weighting coefficient.

6. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, The expression for constructing the global feature association matrix of the global modeling of visual representation is as follows: ,in, This is the global feature correlation matrix. The number of rows for high-dimensional feature selection. The number of columns for high-dimensional filtering features. The first feature for high-dimensional screening row vectors For high-dimensional feature selection Transpose of a column vector To model a weight matrix for the whole system, Model the bias matrix for the global model.

7. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, Step S3 includes the following sub-steps: S31, obtaining the spatiotemporal continuous feature representation output in step S2, performing feature dimension analysis on the feature representation, and determining the number of spatial and channel dimensions of the feature representation to provide a dimensional basis for subsequent pixel-level semantic mapping; S32, based on the analyzed feature dimension information, constructing an initial association model for pixel-level semantic mapping, and setting the initial mapping parameters of the model, including the number of semantic categories and pixel association thresholds; S33, inputting the spatiotemporal continuous feature representation into the initial association model, and performing matching calculations on the features corresponding to each pixel and semantic category features through the feature matching module inside the model to obtain preliminary pixel-semantic matching results; S34, adjusting the mapping parameters of the association model according to the preliminary matching results, optimizing the matching accuracy between pixels and semantic categories, and finally generating a pixel-level semantic feature map.

8. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, S4 includes the following steps: S41, converting the pixel-level semantic feature map obtained in S3 into a high-dimensional feature matrix, and dividing the high-dimensional feature matrix into multiple sub-feature matrices according to preset feature segmentation rules; S42, inputting the divided sub-feature matrices into the feature preprocessing module of the high-dimensional visual representation intelligent analysis platform, performing feature standardization processing on each sub-feature matrix to eliminate feature scale differences between different sub-matrices; S43, starting the platform's feature analysis module, performing feature correlation analysis on the standardized sub-feature matrices, and calculating the correlation coefficients of features within each sub-matrix and features between sub-matrices; S44, based on the correlation coefficient results, selecting feature combinations with high correlation, integrating the selected feature combinations into high-dimensional selected features, and transmitting them to step S5.

9. The method for global modeling of visual representations based on feature extraction according to claim 1, characterized in that, S5 includes the following sub-steps: S51, receiving the high-dimensional filtered features output from S4, verifying the feature dimensions to confirm whether the feature dimensions meet the input requirements of global modeling of visual representation; if not, adjusting the dimensions; S52, constructing a feature fusion framework for global modeling of visual representation, setting the feature fusion levels and fusion parameters for each level within the framework, including fusion weights and fusion window sizes; S53, inputting the high-dimensional filtered features into the feature fusion framework, performing layer-by-layer fusion processing on the features through fusion modules at each level, and adjusting the fusion parameters in real time during the fusion process; S54, constructing a global feature association matrix based on the layer-by-layer fused features, quantifying the association relationships of global features through matrix operations, and completing the global modeling of visual representation.

10. A global modeling system for visual representation based on feature extraction, characterized in that, The system is applied to the feature extraction-based global modeling method for visual representations as described in claim 1, comprising: a visual bag-of-vocabulary feature encoding unit, used to perform visual bag-of-vocabulary feature encoding processing on input visual data, construct a visual vocabulary dictionary and generate visual bag-of-vocabulary feature vectors, the output of which is connected to a spatiotemporal continuous feature learning unit; a spatiotemporal continuous feature learning unit, used to receive the feature vectors output by the visual bag-of-vocabulary feature encoding unit, perform spatiotemporal dimension feature learning through a spatiotemporal continuous visual autoencoder and output spatiotemporal continuous feature representations, the output of which is connected to a pixel-level semantic mapping unit; and a pixel-level semantic mapping unit, used to receive spatiotemporal continuous feature representations and establish pixel-level semantic feature mapping using a pixel-level semantic feature mapping algorithm. The system maps semantic categories and generates pixel-level semantic feature maps. The output is connected to a high-dimensional feature analysis unit. The high-dimensional feature analysis unit receives the pixel-level semantic feature maps, performs feature association analysis and filtering through a high-dimensional visual representation intelligent analysis platform, and outputs high-dimensional filtered features. Its output is connected to a global feature modeling unit. The global feature modeling unit receives the high-dimensional filtered features, constructs a global feature association matrix using global visual representation modeling technology, and performs global visual representation modeling. Its output is connected to a feature optimization output unit. The feature optimization output unit receives the global feature association matrix, adjusts the feature weight coefficients, and outputs the final global visual representation model. This unit is the system's final output module.