Server data processing method and system based on deep learning

By building a time-varying neural network topology and adaptive data processing module, dynamically adjusting network edge weights and resource allocation, the problems of information splitting and unreasonable resource allocation in multimodal data processing on the server side are solved, and efficient cross-modal feature extraction and resource optimization are achieved.

CN120386637AActive Publication Date: 2025-07-29四川华鲲振宇智能科技有限责任公司 +1

Patent Information

Application Number
CN202510875222.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the multimodal data processing of the server-side, the problem of inter-modal information separation, low static feature extraction efficiency, inability to dynamically optimize the topology structure, and unreasonable resource allocation, resulting in idle or overloading and lag in computing resources.

Method used

Using a deep learning-based method, a time-varying neural network topology is constructed, cross-modal correlation features are captured through attention mechanisms, combined with adaptive data processing modules and reinforcement learning algorithms, network edge weights and resource allocation are dynamically adjusted to achieve efficient fusion and optimization of multimodal features.

Benefits of technology

It improves the accuracy of cross-modal correlation feature extraction, reduces processing delay, improves resource utilization and processing efficiency, and enhances the generalization ability and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386637A_ABST
    Figure CN120386637A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, particularly relates to a server data processing method and system based on deep learning, and aims to solve the problems of low modal fusion efficiency, poor model adaptability and insufficient resource utilization rate in multi-modal data processing. The method comprises the following steps: firstly, preprocessing mixed to-be-processed data and dividing the mixed to-be-processed data into structured data, image data, text data and time series data subsets; a time-varying neural network topological structure is constructed, network edge weights are dynamically adjusted based on server loads and data features, and preliminary extraction and correlation capture of multi-modal features are achieved; generating a fusion feature vector by using a cross-modal feature extractor of the attention mechanism; a dynamic adaptive data processing module is used for matching the deep learning model for parallel processing; and finally, constructing a reinforcement learning reward function based on the processing efficiency, the resource occupancy rate and the accuracy rate, and carrying out joint iterative optimization on the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a server data processing method and system based on deep learning. Background Art

[0002] With the rapid development of artificial intelligence technology, the types of data faced by the server side are becoming increasingly rich, covering multi-modal forms such as structured data, unstructured data, and time-series data. There is information fragmentation between modalities in traditional data processing: traditional methods usually process different modality data independently, lacking an effective cross-modal association modeling mechanism. For example, image classification and text sentiment analysis respectively use independent models and cannot capture the potential association between "image content - text semantics".

[0003] Limitations in static feature extraction: Most existing cross-modal fusion methods adopt fixed network structures and are difficult to dynamically adapt to changes in data complexity. For example, the static attention mechanism has low efficiency in feature extraction for long texts or complex images and cannot automatically adjust the mask ratio or attention head allocation according to semantic complexity.

[0004] Lack of topological structure dynamics: As an effective tool for processing associated data, the topological structure of graph neural networks is mostly statically preset and cannot dynamically optimize the edge connection weights according to real-time data features or server load. For example, at high server loads, a fixed topological structure may cause blockages in the key modality feature fusion path, increasing processing latency. Traditional architectures rely on manually preset mapping relationships between models and tasks and cannot handle mixed or dynamically changing data types. For example, in a video live broadcast scenario, sudden high concurrency may cause the server with a fixed model deployment to crash due to insufficient resources. Existing methods lack a real-time perception and dynamic response mechanism for server load, and often have the problem of coexistence of "idle computing resources" and "overload lag". For example, lightweight tasks occupy high-performance GPU resources, while complex image tasks are forced to slow down due to insufficient resources.

[0005] Therefore, there is an urgent need for a server data processing method that can dynamically optimize the network structure, efficiently fuse multi-modal features, and adaptively adjust resource allocation. Summary of the Invention

[0006] The purpose of the present invention is to provide a server data processing method and system based on deep learning to achieve server data processing that dynamically optimizes the network structure, efficiently fuses multi-modal features, and adaptively adjusts resource allocation.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows: In a first aspect, a server data processing method based on deep learning is provided, including the following steps: S1: The server obtains the mixed data to be processed, preprocesses the mixed data to be processed, and divides the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time series data; S2: Construct a time-varying neural network topology, and input the four subsets into the corresponding branches in the time-varying neural network topology for preliminary feature extraction. The time-varying neural network topology dynamically adjusts the network edge connection weights according to the input data; S3: The cross-modal feature extractor based on the attention mechanism captures the correlation features between the features of different modal data to obtain cross-modal correlation features, and performs fusion processing based on the cross-modal correlation features to obtain a fusion feature vector; S4: Construct an adaptive data processing module, set multiple deep learning network models in the adaptive data processing module, and the adaptive data processing module matches the corresponding deep learning network model for processing based on the fusion feature vector for the mixed data to be processed; S5: Calculate the processing efficiency, resource occupancy rate, and processing accuracy rate during the data processing process, establish a reward function for the reinforcement learning algorithm based on the processing efficiency, resource occupancy rate, and processing accuracy rate, and perform joint iterative optimization on the deep learning network model of the deep learning network model based on the reward function; S6: Process the real-time mixed data to be processed based on the jointly iterated deep learning network model to obtain the multi-modal data processing result.

[0008] Preferably, the specific process of step S1 is as follows: S11: The server receives the mixed data to be processed through the network interface or the storage device interface, and performs a preliminary verification on the integrity and format of the data, checks whether the data file is damaged and whether the data fields are missing. If it is found that the data is incorrect or incomplete, an error prompt is triggered and the reception is stopped, and the data is required to be retransmitted; S12: Perform data cleaning and format standardization processing on the mixed data to be processed, remove noise data, fill in missing values, convert different types of data into a format suitable for processing by the deep learning network model, and add type labels; S13: Divide the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time series data based on the type labels.

[0009] Preferably, the specific process of step S2 is as follows: S21: Map the subset of structured data into high-dimensional feature vectors through an embedding layer, perform multi-scale feature extraction on the image using a deformable convolutional feature extractor, output a sequence of feature maps by focusing on key regions through a spatial attention mechanism, adjust the mask ratio according to the text length and semantic complexity, capture the long-range dependencies of the text through a multi-head self-attention mechanism, use an attention mechanism LSTM cell to process time-series data, and capture the long-term and short-term patterns in the time-series data through dynamic allocation of attention weights to different time steps; S22: Use the modal features of different dimensions output by the four subset branches as the initial node set of the deep learning network model. Each node represents the feature representation of a specific modality. The initial edge weight matrix W ∈ Rn×n represents the connection strength between nodes, where n is the number of nodes. The edge weights are initialized as learnable parameters and initialized through a random Gaussian distribution; S23: Real-time collect the hardware metrics of the server's CPU utilization rate, memory occupancy rate, and I / O throughput, as well as the software metrics of data processing latency and model inference speed. Convert the multi-dimensional load metrics into a weight adjustment coefficient α ∈ [0,1] through a non-linear mapping function φ; S24: Calculate the statistical features of the input data of each branch, including variance, entropy value, and KL divergence of the feature distribution, evaluate the complexity and uncertainty of the data, calculate the correlation degree between different modalities through a cross-modal attention mechanism, and generate an attention weight matrix; S25: Fuse the data-driven attention weights with the weights adjusted by the server load; S26: Set a sliding time window T, update the edge weight matrix Wt at each time step t. When new data arrives, update the edge weights in an incremental learning manner to avoid global recalculation; S27: When the change in the edge weights exceeds the threshold, trigger a topological structure reorganization, add or delete node connections, and form a new subnet structure.

[0010] Preferably, the specific process of step S3 is as follows: S31: Project the modal features of different dimensions into a shared latent space, achieve dimension alignment through linear transformation, and enhance the feature correlation within each modal feature; S32: Capture the correlation between different modal features through a query-key-value pair mechanism, map the server's CPU utilization rate and memory occupancy rate metrics into attention weight adjustment coefficients, and fuse the server load information with the correlation degree of modal features of different dimensions; S33: Design attention heads of different scales, capture the correlation features from fine-grained to coarse-grained, construct a modal correlation strength matrix, and quantify the mutual influence of different modalities; S34: Calculate the attention weights between pairwise modal features, construct the association strength matrix based on the attention weights, perform weighted fusion on each modal feature based on the association strength matrix, and introduce a residual structure to retain the original feature information.

[0011] Preferably, the specific process of step S33 is as follows: S331: Convert the input features into a feature tensor, and evenly divide the feature tensor into h sub-tensors along the channel dimension. Each sub-tensor corresponds to an independent attention head; S332: Assign different key-value pair dimensions to each attention head. Small dimensions capture fine-grained associations, and large dimensions capture coarse-grained associations; S333: Apply different types of positional encodings to attention heads of different scales to enhance the perception of local and global relationships; S334: For each modal feature, calculate its association strength matrix.

[0012] Preferably, the specific process of calculating the attention weights between pairwise modal features in step S34, calculating the association strength matrix according to the attention weights, performing weighted fusion on each modal feature based on the association strength matrix, and introducing a residual structure to retain the original feature information is as follows: S341: Obtain the query vector, key vector, and value vector by linearly transforming each modal feature, calculate the attention scores, and obtain the attention weights between pairwise modal features through the softmax function; S342: Construct the association strength matrix based on the attention weights between pairwise modal features. Each element in the matrix represents the attention weight of modality i to modality j; S343: According to the association strength matrix M, perform weighted fusion on each modal feature, perform global average pooling on the fused features to obtain the channel descriptor, calculate the channel weights through two fully connected layers and activation functions, and multiply the channel weights with the fused features channel by channel to obtain the enhanced fused features.

[0013] Preferably, the specific process of the adaptive data processing module in step S4 matching the corresponding deep learning network model for processing the mixed data to be processed based on the fused feature vector is as follows: S41: Establish the mapping relationship between the model and the task. The adaptive data processing module parses the fused feature vector and calculates the variance, entropy value, and cross-correlation relationship of the fused feature vector; S42: Judge the task type of the mixed data to be processed according to the parsing result, and filter the preliminary network model based on the mapping relationship between the model and the task according to the task type; S43: When there are task records of the same type in the historical dynamic optimization table, calculate the matching degree score of the preliminary calculation model based on the reinforcement learning algorithm, and adjust the network model according to the matching degree score; S44: Extract task priority labels from the fused features. High-priority data directly enters the hot queue of each network model instance pool, and medium- and low-priority data enters the ordinary queue. Each network model processes data based on the queue.

[0014] Preferably, the specific process of step S5 is as follows: S51: Embed a monitoring module in the data processing flow to collect in real time the amount of processed data, processing time, actual resources used, total available resources, correctly predicted data, and total number of samples; S52: Calculate the processing efficiency, resource occupancy rate, and processing accuracy rate during the data processing process. The formula for calculating the processing efficiency is as follows: Processing efficiency = amount of processed data / processing time; The formula for calculating the resource occupancy rate is as follows: Resource occupancy rate = actual resources used / total available resources × 100%; Processing accuracy rate = correctly predicted data / total number of samples; S53: Aggregate the metrics according to a fixed time window to generate a state vector, and set the action space, which matches the optimization direction of the model; S54: Perform collaborative optimization based on the original loss function of the deep learning model and the reinforcement learning reward function.

[0015] In a second aspect, a server data processing system based on deep learning is provided for implementing the server data processing method based on deep learning, including a data acquisition module, a preprocessing module, a time-varying neural network topology, a feature extraction module, a feature fusion module, an adaptive data processing module, and a reinforcement learning module. The data acquisition module is connected to the preprocessing module, the preprocessing module is connected to the time-varying neural network topology, the feature extraction module is arranged in the time-varying neural network topology, the feature extraction module is connected to the feature fusion module, the feature fusion module is connected to the adaptive data processing module, and the adaptive data processing module is connected to the reinforcement learning module; The data acquisition module is used to acquire mixed data to be processed; The preprocessing module is used to preprocess the mixed data to be processed and divide the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time series data; The time-varying neural network topology is used to input four subsets into corresponding branches in the time-varying neural network topology for preliminary feature extraction. The time-varying neural network topology dynamically adjusts the network edge connection weights according to the input data, captures the correlation features between different modality data features, and obtains cross-modal correlation features; The feature fusion module is used to perform fusion processing based on the cross-modal correlation features to obtain a fusion feature vector; The adaptive data processing module sets multiple deep learning network models, and matches corresponding deep learning network models for processing based on the fusion feature vector for the mixed data to be processed; The reinforcement learning module is used to calculate the processing efficiency, resource occupancy rate, and processing accuracy rate during the data processing process, establish a reward function of the reinforcement learning algorithm with the processing efficiency, resource occupancy rate, and processing accuracy rate, and perform joint iterative optimization on the deep learning network models of the deep learning network models based on the reward function.

[0016] The beneficial effects of the present invention include: The method and system for server data processing based on deep learning provided by the present invention preprocess the mixed data to be processed and divide it into subsets of structured data, image data, text data, and time series data; construct a time-varying neural network topology, dynamically adjust the network edge weights based on the server load and data features, and realize the preliminary extraction and correlation capture of multi-modal features; use a cross-modal feature extractor with an attention mechanism to generate a fusion feature vector; match a deep learning model through an adaptive data processing module for parallel processing; finally, construct a reinforcement learning reward function based on the processing efficiency, resource occupancy rate, and accuracy rate, and perform joint iterative optimization on the model.

[0017] First of all, through the dual dynamic adjustment mechanism of the server load and data features, the edge connection weights of the graph neural network can optimize the multi-modal feature fusion path in real time. In the multi-modal sentiment analysis task, compared with the traditional static GNN, the F1 value of cross-modal correlation feature extraction is effectively improved, and the accuracy of the model for "image-text" semantic alignment is also effectively improved.

[0018] Secondly, the dynamic adaptive data processing architecture can complete model selection in a very short time through the analysis variance, entropy value, cross-correlation coefficient of the fusion feature vector and reinforcement learning scheduling, and the efficiency is improved compared with manual parameter tuning. And combined with the priority queue and elastic resource scheduling mechanism, the latency of real-time data processing is significantly reduced.

[0019] Again, by adopting a targeted extraction method for different modal data, compared with the general extraction method, the accuracy of feature extraction is effectively improved. Combining the server load and data features, the connection weights of the network edges are adjusted in real time. When the server is under high load, the feature fusion path is automatically optimized to reduce the processing delay. At the same time, the weights are adjusted according to the data complexity, making the extraction of complex data features more accurate and improving the generalization ability of the model. When the edge weight change exceeds the threshold, the topological structure is reorganized, which can quickly adapt to the change of data distribution. When the data mode changes greatly, the fluctuation range of the model performance is reduced, and a stable processing effect is maintained.

[0020] Again, project the modal features into the shared latent space and enhance the internal correlation, so that different modal features can be better fused in the unified space, and the cross-modal semantic consistency is improved. Fusing the server load information and modal correlation degree, when resources are scarce, key correlation features are preferentially guaranteed to be extracted, avoiding information loss caused by insufficient resources, and improving the integrity of important feature extraction. Design different-scale attention heads to comprehensively capture correlation features from fine-grained to coarse-grained. Compared with a single scale, it can more carefully explore the complex relationships between modalities. Based on the correlation strength matrix for weighted fusion and introducing a residual structure, not only fully fuses cross-modal information, but also retains the details of the original features, improving the information integrity of the fused feature vector and effectively enhancing the accuracy of subsequent processing.

[0021] Finally, by analyzing the fused feature vector to judge the task type and screening the model in combination with the model-task mapping relationship, compared with the fixed model configuration, the matching accuracy between the task and the model and the processing efficiency are improved. Using reinforcement learning to calculate the matching degree score and adjust the model, the model selection strategy is continuously optimized. After iteration, the average processing efficiency of the model is improved and resource waste is reduced. Construct a reward function based on processing efficiency, resource occupancy rate, and processing accuracy, and jointly optimize through the joint loss function to achieve a balance among the three. Compared with single-objective optimization, the overall resource utilization rate of the server is improved. Brief Description of the Drawings

[0022] Figure 1 It is a schematic flowchart of the server data processing method based on deep learning of the present invention.

[0023] Figure 2 It is a schematic flowchart of capturing cross-modal correlation features and fusing and processing of the present invention. Detailed Embodiments

[0024] The following is a further detailed description of the present invention in conjunction with the attached Figures 1 - 2 drawings: Embodiment 1 Referring to the attached Figure 1 drawings, a server data processing method based on deep learning includes the following steps: S1: The server obtains the mixed data to be processed, preprocesses the mixed data to be processed, and divides the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time-series data; S2: Construct a time-varying neural network topology, and input the four subsets into the corresponding branches in the time-varying neural network topology for preliminary feature extraction. The time-varying neural network topology dynamically adjusts the network edge connection weights according to the input data; S3: The cross-modal feature extractor based on the attention mechanism captures the correlation features between the features of different modal data to obtain cross-modal correlation features, and performs fusion processing based on the cross-modal correlation features to obtain a fusion feature vector; S4: Construct an adaptive data processing module, set multiple deep learning network models in the adaptive data processing module, and the adaptive data processing module matches the corresponding deep learning network model for processing based on the fusion feature vector for the mixed data to be processed; S5: Calculate the processing efficiency, resource occupancy rate, and processing accuracy rate during the data processing process, establish a reward function for the reinforcement learning algorithm based on the processing efficiency, resource occupancy rate, and processing accuracy rate, and perform joint iterative optimization on the deep learning network models of the deep learning network models based on the reward function; S6: Process the real-time mixed data to be processed based on the deep learning network models after joint iteration to obtain the multi-modal data processing results.

[0025] In this embodiment, taking the intelligent transportation management system as the application scenario, the specific implementation process of the above-mentioned data processing method by the server is elaborated in detail. In intelligent transportation management, the system needs to process the mixed data to be processed from multiple devices, including images captured by traffic cameras, time-series data collected by traffic flow sensors, text information related to traffic regulations, and structured data such as vehicle registration, to achieve efficient traffic monitoring and management.

[0026] The server obtains the mixed data to be processed through traffic devices deployed at various intersections in the city. These data include the real-time video stream collected by traffic cameras, which is subsequently split into image data, time-series data such as the traffic flow and vehicle speed collected by traffic flow sensors every minute, text data in traffic regulation documents, and structured data in the vehicle registration system, including vehicle ID, vehicle type, owner information, etc. The server first preprocesses the mixed data to be processed. After the preprocessing is completed, the mixed data to be processed is divided into four subsets: the structured data subset, which contains structured information such as vehicle registration; the image data subset, which is the intersection traffic condition image extracted from the video stream; the text data subset, which stores text related to traffic regulations; and the time-series data subset, which records the real-time data collected by traffic flow sensors.

[0027] Build a time-varying neural network topology with four branches corresponding to the processing of structured data, image data, text data, and time-series data respectively. Input the subset of structured data into the branch dedicated to processing structured data, which adopts a multi-layer perceptron structure to perform preliminary feature extraction on structured information such as vehicle registration. Input the subset of image data into the convolutional neural network branch to extract features such as traffic signs and vehicle types in the image through operations such as convolutional layers and pooling layers. Input the subset of text data into the recurrent neural network or Transformer branch to capture semantic information in traffic regulation texts. Input the subset of time-series data into the long short-term memory network branch to extract the regular features of traffic flow changes over time. During the data input process, the time-varying neural network topology dynamically adjusts the weights of network edge connections according to the characteristics of the input data through the backpropagation algorithm to optimize the feature extraction effect.

[0028] Build a cross-modal feature extractor based on the attention mechanism, which can capture the correlation features between the features of different modal data. For example, combine the congestion section information identified in the image with the time-series traffic flow data of that section and the text information about congestion handling in traffic regulations. Calculate the correlation weights between different modal features through the attention mechanism to highlight key features and obtain cross-modal correlation features. Then, perform fusion processing based on the cross-modal correlation features to integrate the feature information of different modalities into a fused feature vector to achieve deep fusion of multi-modal data.

[0029] Build an adaptive data processing module and set multiple deep learning network models in the module, which can include the Prophet model for traffic flow prediction, the YOLO model for traffic accident detection, and the BERT model for intelligent traffic regulation Q&A. The adaptive data processing module analyzes the types and requirements of the mixed data to be processed based on the fused feature vector and matches the corresponding deep learning network model for it. For example, when the fused feature vector indicates that the data is related to traffic flow prediction, match the Prophet model for processing; if it involves traffic accident image detection, call the YOLO model.

[0030] During the data processing, the real-time computing processing efficiency is calculated, such as the amount of data processed per unit time and the resource occupancy rate, including the usage of resources such as CPU and memory, and the processing accuracy rate, such as the accuracy rate of traffic flow prediction and the correct recognition rate of traffic accident detection. A reward function of the reinforcement learning algorithm is established based on these three indicators. Based on the reward function, the deep learning network model is jointly iteratively optimized through the reinforcement learning algorithm, and the model parameters are continuously adjusted to improve the comprehensive performance of the model in terms of processing efficiency, resource occupancy, and processing accuracy. Based on the jointly iterated deep learning network model, the real-time acquired mixed data to be processed is processed. When new traffic data arrives, data preprocessing, feature extraction, fusion, and model matching processing are performed according to the above steps, and finally the multi-modal data processing results are obtained. The above results can be used for applications such as real-time traffic flow monitoring, traffic accident warning, and intelligent traffic law consultation, providing strong support for urban traffic management.

[0031] Embodiment 2 Based on Embodiment 1, the specific process of step S1 is as follows: S11: The server receives the mixed data to be processed through the network interface or the storage device interface, and performs a preliminary verification on the integrity and format of the data, checks whether the data file is damaged and whether the data fields are missing. If it is found that the data is incorrect or incomplete, an error prompt is triggered and the reception is stopped, and the data is required to be retransmitted.

[0032] The server is equipped with a variety of data reception interfaces, including a network interface for real-time transmission of data from devices such as traffic cameras and traffic sensors, and a storage device interface for reading traffic law documents, vehicle registration data, etc. stored in the local storage device. When the data starts to be transmitted, the server starts the preliminary verification program. For the data integrity verification, the server checks information such as the size and hash value of the data file. Taking the video stream data file transmitted from a traffic camera as an example, the server pre-stores the expected size range and hash value of the normal video stream file of this camera. If the size of the received video stream file exceeds the reasonable range, or the calculated hash value does not match the expected value, it is determined that the data file is damaged.

[0033] In terms of data field inspection, for structured data such as vehicle registration, the server checks whether each data record contains all the necessary fields, such as fields such as vehicle ID, vehicle model, and owner information, according to the preset data format template. If it is found that there are missing fields, it is also regarded as incomplete data. Once it is detected that the data is incorrect or incomplete, the server immediately triggers an error prompt, records the error information through the system log, and displays the error details to the administrator on the management interface. At the same time, the reception of the current data transmission is stopped, and a request to retransmit the data is sent to the data sending end to ensure that the received data is complete and accurate.

[0034] S12: Clean and standardize the format of the mixed data to be processed, remove noise data, fill in missing values, convert different types of data into a format suitable for processing by a deep learning network model, and add type labels.

[0035] After completing the preliminary verification, the server performs in-depth cleaning and format standardization on the mixed data to be processed. In the data cleaning process, for the time-series data collected by traffic flow sensors, if there are noise data with abnormal fluctuations, such as a significant unreasonable sudden increase or decrease in traffic volume at a certain moment, the server uses statistical methods, such as outlier detection based on the mean and standard deviation, to identify and remove these noise data; for missing traffic volume and vehicle speed data, methods such as linear interpolation and time series prediction are used to fill in the missing values.

[0036] For traffic regulation text data, there may be problems such as garbled characters and duplicate paragraphs. The server repairs the garbled characters through character encoding conversion and uses text similarity algorithms to detect and delete duplicate paragraphs. In terms of format standardization, the image data collected by traffic cameras is uniformly converted into JPEG or PNG format, and the image resolution is adjusted to a size suitable for subsequent deep learning model processing; the structured data is converted into standard formats such as CSV and JSON, and type conversion and standardized naming are performed on the data fields; the text data is preprocessed such as word segmentation and part-of-speech tagging and converted into a word vector representation form. After processing, a clear type label, such as "structured data", "image data", "text data", "time-series data", is added to each piece of data according to the data type.

[0037] S13: Divide the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time-series data based on the type labels. The server divides the preprocessed mixed data to be processed into four subsets according to the added type labels. Using an automated script or program, all data records are traversed, and the data records with the "structured data" label are extracted according to the type label to form a structured data subset, which contains structured information such as vehicle registration; the data records marked as "image data" are grouped into the image data subset, and these data are images of the intersection traffic conditions extracted from the video stream; the traffic regulation-related text data marked as "text data" is integrated into the text data subset; the real-time data collected by traffic flow sensors with the "time-series data" label is classified into the time-series data subset. After the division, the data in each subset is clear and provides a data basis for subsequent feature extraction and processing.

[0038] The specific process of step S2 is as follows: S21: Map the subset of structured data into high-dimensional feature vectors through the embedding layer. Use a deformable convolutional feature extractor to perform multi-scale feature extraction on the image. Output a sequence of feature maps by focusing on key regions through a spatial attention mechanism. Adjust the mask ratio according to the text length and semantic complexity, and capture the long-range dependencies of the text through the multi-head self-attention mechanism. Use the attention mechanism LSTM unit to process time-series data, and dynamically allocate the attention to different time steps through the attention weights to capture the long-term and short-term patterns in the time-series data.

[0039] For subsets of structured data such as vehicle registration, the server inputs them into the embedding layer. Through training and learning, the embedding layer maps the originally low-dimensional and discrete structured data (including vehicle IDs, vehicle type categories, etc.) into high-dimensional and continuous feature vectors. Map the vehicle ID from a simple digital code to a high-dimensional vector containing potential semantic information such as vehicle type and service life, providing a richer feature representation for subsequent processing. The image data collected by traffic cameras enters the deformable convolutional feature extractor. This extractor can dynamically adjust the sampling positions of the convolutional kernels according to the image content and perform multi-scale feature extraction on the image. In complex traffic scene images, deformable convolution can more accurately capture the features of targets such as vehicles, pedestrians, and traffic signs. Subsequently, the spatial attention mechanism comes into play. By analyzing the feature importance of each region of the image, it focuses on key regions in the traffic scene, such as intersections and accident-prone areas, and outputs a sequence of feature maps, highlighting the information valuable for subsequent processing. For text data such as traffic regulations, the server first analyzes the text length and semantic complexity. For longer and more semantically complex regulation texts, appropriately increase the mask ratio. During the processing of the multi-head self-attention mechanism, by performing a masking operation on the text, the model is allowed to learn to predict the masked parts, thereby capturing the long-range dependencies in the text. For example, when understanding the clauses regarding accident liability determination in traffic regulations, it can associate the relevant conditions and result descriptions in different paragraphs to accurately grasp the semantics of the regulations. The time-series data collected by traffic flow sensors is processed by the attention mechanism LSTM unit. At each time step, the LSTM unit calculates the output based on the current input and historical state. At the same time, the attention mechanism dynamically allocates the attention to different time steps according to the data features. When analyzing the traffic flow change trend, for key time steps such as traffic peak periods, a higher attention weight is assigned, thereby effectively capturing the long-term trends (such as the traffic flow differences between weekdays and weekends) and short-term patterns (such as the traffic flow fluctuations during morning and evening rush hours) in the time-series data.

[0040] S22: Use the modality features of different dimensions output by the four subset branches as the initial node set of the deep learning network model. Each node represents the feature representation of a specific modality. The initial edge weight matrix W ∈ Rn×n represents the connection strength between nodes, where n is the number of nodes. The edge weights are initialized as learnable parameters and are initialized through a random Gaussian distribution.

[0041] The server uses the high-dimensional feature vectors obtained by the structured data subset through the embedding layer, the sequence of feature maps output by the processed image data, the feature representations obtained by the text data through the multi-head self-attention mechanism, and the features of the time-series data processed by the attention mechanism LSTM unit as the initial node set of the deep learning network model. Each node represents the feature representation of a specific modality. For example, the node corresponding to the image data contains the target feature information in the traffic scene. The edge weights are initialized as learnable parameters and are initialized through a random Gaussian distribution. This enables the network to have a certain connection strength in the initial state and will be adjusted according to the data and server load conditions later, providing a basis for information interaction between nodes.

[0042] S23: Real-time collect the hardware metrics of the server's CPU utilization rate, memory occupancy rate, I / O throughput, and the software metrics of data processing latency and model inference speed. Convert the multi-dimensional load metrics into a weight adjustment coefficient α through the non-linear mapping function φ, where α ∈ [0, 1].

[0043] The server real-time collects its own hardware and software metrics. In terms of hardware metrics, the CPU utilization rate, memory occupancy rate, and I / O throughput are obtained through system monitoring tools; for software metrics, information such as data processing latency and model inference speed is collected. The multi-dimensional load metrics are input into the non-linear mapping function φ, which optimizes the parameters through training and converts the load metrics into a weight adjustment coefficient α. When the server has a high CPU utilization rate and a large data processing latency, α will be adjusted accordingly, enabling the network structure to be optimized according to the load conditions and balancing resource utilization and processing efficiency.

[0044] S24: Calculate the statistical features of the input data of each branch, including variance, entropy value, and KL divergence of the feature distribution, evaluate the complexity and uncertainty of the data, calculate the correlation degree between different modality features through the cross-modal attention mechanism, and generate an attention weight matrix A. The elements in the attention weight matrix A are A( i , j ), representing the data-driven attention weights: A( i , j ) = softmax (Q( x i ) · K( x j )T / ) ; Among them, softmax is the similarity score normalization, x i and x j are the input features respectively, Q( x i ) and K( x j ) are the query vector and the key vector respectively, is the scaling factor, d is the feature dimension, and T is the matrix transpose operation; S25: Fuse the data-driven attention weight and the weight adjusted by the server load: W''( i , j ) = β ·W'( i , j ) + (1 - β )·A( i , j ) ; Among them, W''( i , j ) is the weight after fusing the data-driven attention weight and the weight adjusted by the server load, W'( i , j ) is the weight adjusted by the server load, β is a learnable balance parameter, optimized by gradient descent, and A( i , j ) is the data-driven attention weight; S26: Set the sliding time window T and update the edge weight matrix W t at each time step t. When new data arrives, update the edge weight in an incremental learning manner to avoid global recalculation: W t ( i , j ) = ρ·W t-1 ( i , j ) + (1 - ρ)·W''( i , j ) ; Among them, ρ is the decay coefficient, 0 ≤ ρ ≤ 1, which controls the balance between the historical weight and the new weight. W t-1 ( i , j ) is the edge weight matrix at time step t, and W''( i , j) is the weight after the fusion of data-driven attention weights and the adjusted weights of the server load; S27: When the edge weight changes exceed the threshold, trigger the topological structure reorganization, add or delete node connections, and form a new subnet structure. The server continuously monitors the changes in the edge weight matrix. When the change in the edge weight exceeds the preset threshold, it means that the network structure needs to be adjusted to better adapt to data and load changes. At this time, trigger the topological structure reorganization, and the system determines which node connections need to be added or deleted according to the current weight situation. When there are large changes in the traffic flow pattern or significant fluctuations in the server load, a new subnet structure is formed through reorganization to improve the adaptability and efficiency of the network for data processing.

[0045] Through the above dual dynamic adjustment mechanism, that is, server load-driven + data feature-driven, the time-varying neural network topology can adaptively optimize the multi-modal feature fusion path under resource constraints and achieve efficient and accurate data processing. This design not only improves the generalization ability of the model but also significantly enhances the robustness of the system in the face of dynamic data streams.

[0046] Embodiment 3 Based on Embodiment 1 or Embodiment 2, see Figure 2 As shown, the specific process of step S3 is as follows: S31: Project the modal features of different dimensions into the shared hidden space, achieve dimension alignment through linear transformation, and enhance the feature correlation within each modal feature; S32: Capture the correlation between different modal features through the query-key value pair mechanism, map the server CPU utilization rate and memory occupancy rate indicators to the attention weight adjustment coefficient, and fuse the server load information and the correlation degree of different-dimensional modal features; S33: Design attention heads of different scales to capture the correlation features from fine-grained to coarse-grained, construct the inter-modal correlation strength matrix, and quantify the mutual influence of different modalities; S34: Calculate the attention weights between pairwise modal features, construct the correlation strength matrix based on the attention weights, perform weighted fusion on each modal feature based on the correlation strength matrix, and introduce a residual structure to retain the original feature information.

[0047] Project the modal features of different dimensions into the shared hidden space, design an independent fully connected layer for each modality, and map the features to the shared hidden space dimension d = 384, and establish a structured feature matrix: h struct = ReLU( W struct ⋅[Vehicle ID, vehicle type, vehicle speed, lane number]), image feature matrix: himage = ReLU( W image ⋅ Vehicle feature vector), text feature matrix: h text = ReLU( W text ⋅ Regulatory word vector), temporal feature matrix: h time = ReLU( W time ⋅ Traffic time series feature).

[0048] Apply spatial self-attention to vehicle features, focusing on key areas such as license plates and vehicle type contours. For example, in a congested scenario, enhance the correlation between the taillight features of the vehicle in front and the braking action of the following vehicle.

[0049] Use "intersection congestion image features" as the query Q , and the "traffic flow time series feature" of the corresponding period as the key K , value V , and calculate the correlation between the congestion level in the image and the actual traffic data: ; Among them, h image is the image feature matrix, h time is the temporal feature matrix, W q is the interaction weight for learning images and temporal data.

[0050] Real-time collect CPU utilization U , memory occupancy M , and convert them into adjustment coefficients through a formula: ; When α > 0.8, reduce the attention weight for high-computation associations such as "image-text regulations", and preferentially process real-time key associations such as "image-temporal traffic".

[0051] The specific process of step S33 is as follows: S331: Convert the input features into feature tensors, and evenly divide the feature tensors along the channel dimension into h sub-tensors, with each sub-tensor corresponding to an independent attention head; S332: Assign different key-value pair dimensions to each attention head, with small dimensions capturing fine-grained associations and large dimensions capturing coarse-grained associations; S333: Apply different types of positional encodings to attention heads of different scales to enhance the perception of local and global relationships; S334: Calculate the association strength matrix for each modal feature.

[0052] Fine-grained head (local association): Focus on local features within a single modality, such as the temporal data association between "brake light status" and "sudden drop in vehicle speed" in a vehicle image, or the violation match between "speed limit 60 km / h" in a regulation text and "current vehicle speed 65 km / h" in structured data. Coarse-grained head (global association): Model cross-modal global features, such as the congestion image features of an entire intersection, the temporal data of road network traffic flow, and the comprehensive association of the "no parking during peak hours" clause in a regulation text.

[0053] The specific process of calculating the attention weights between pairwise modal features in step S34, calculating the association strength matrix based on the attention weights, and performing weighted fusion on each modal feature based on the association strength matrix while introducing a residual structure to retain the original feature information is as follows: S341: Obtain query vectors, key vectors, and value vectors by linearly transforming each modal feature, calculate the attention scores, and obtain the attention weights between pairwise modal features through the softmax function; S342: Construct an association strength matrix based on the attention weights between pairwise modal features, where each element in the matrix represents the attention weight of modality i to modality j; S343: According to the association strength matrix M, perform weighted fusion on each modal feature, perform global average pooling on the fused features to obtain channel descriptors, calculate channel weights through two fully connected layers and activation functions, and multiply the channel weights with the fused features channel by channel to obtain enhanced fused features.

[0054] The above cross-modal feature extraction based on the attention mechanism realizes the intelligent matching of server resources and data association feature extraction through time-varying weight adjustment and dynamic topological structure. Experiments show that in the multi-modal sentiment analysis task, compared with the traditional static attention mechanism, this method can effectively improve the F1 value and reduce the model inference latency.

[0055] Example 4 Based on Example 1 or Example 2 or Example 3, the specific process of the adaptive data processing module in step S4 for matching a corresponding deep learning network model for processing the hybrid data to be processed based on the fused feature vector is as follows: S41: Establish a model-task mapping relationship. The adaptive data processing module parses the fused feature vector and calculates the variance, entropy value, and cross-correlation relationship of the fused feature vector. The fused feature vector from step S34 h final, including the associated information of image, time series, structured, and text modalities. Variance: Measures the degree of fluctuation of feature vectors and reflects data complexity. Example: If the variance of the "vehicle movement direction" dimension in image features is large, it indicates that the vehicle driving directions in the scene are diverse (such as at intersections), which may correspond to tasks such as "congestion prediction" or "violation detection". Entropy value: Evaluates the uncertainty of feature distribution. Example: A high entropy value in text features (such as containing multiple regulations) may correspond to the task of "intelligent regulation Q&A". Cross-correlation relationship: Calculates the correlation coefficient of features between modalities to identify the dominant modality. Example: The cross-correlation coefficient r = 0.9 between image features and time series features indicates strong spatio-temporal correlation and may belong to tasks such as "congestion prediction" or "traffic event detection".

[0056] S42: Determine the task type of the mixed data to be processed according to the parsing result, and filter the preliminary network model based on the model-task mapping relationship according to the task type. If the cross-correlation coefficient r = 0.9 between image and time series features in the fusion features, and the variance is higher than the threshold (such as large fluctuations in traffic flow), it is determined as a congestion prediction task. If the cross-correlation coefficient r = 0.8 between image features and text features (such as matching violation images with regulation clauses), it is determined as a violation detection task.

[0057] S43: When there are records of the same type of task in the historical dynamic optimization table, calculate the matching degree score of the preliminary calculation model based on the reinforcement learning algorithm Q ( s , a ), and adjust the network model according to the matching degree score. The specific calculation formula of the matching degree score Q ( s , a ) is as follows: Q ( s , a ) = (1 - α ) Q ( s ′, a ′) + α ( r + γ max a′ Q ( s ′, a ′)); Among them, s is the current state, that is, the current data type + task type + load index, a is the action of the current state model selection, s ′ is the next state, a ′ is the possible action of the next state,r is the reward value, α is the learning rate, max a′ Q ( s ′, a ′) is the next state s ′ of all possible actions a ′ is the maximum Q value, representing the agent's estimate of the future optimal decision, Q ( s ′, a ′) is the next state Q value; S44: Extract the task priority label from the fused features. High-priority data directly enters the hot queue of each network model instance pool, and medium- and low-priority data enters the ordinary queue. Each network model processes data based on the queue.

[0058] The specific process of step S5 is as follows: S51: Embed a monitoring module in the data processing flow to collect the processed data volume, processing time, actual used resources, total available resources, correctly predicted data, and total number of samples in real time; S52: Calculate the processing efficiency, resource occupancy rate, and processing accuracy during the data processing process. The processing efficiency calculation formula is as follows: Processing efficiency = Processed data volume / Processing time; The calculation formula for the resource occupancy rate is as follows: Resource occupancy rate = Actual used resources / Total available resources × 100%; Processing accuracy = Correctly predicted data / Total number of samples; S53: Aggregate the metrics according to a fixed time window to generate a state vector s t = [Processing efficiencyt, Resource occupancy ratet, Processing accuracyt], set the action space, and the action space matches the optimization direction of the model; S54: Co-optimize based on the original loss function of the deep learning model and the reinforcement learning reward function, Joint loss: L = L task ( θ ) − β ⋅ r t ; Among them, β is the balance coefficient, guiding the model to optimize in the direction of high rewards, L task is the task loss, representing the loss function value of the deep learning model on a specific task, representing the loss function value of the deep learning model on a specific task, θrepresent the learnable parameters of the model, and the task loss changes with θ the change of; r t is the reward signal, which comes from the environmental feedback in reinforcement learning.

[0059] A server data processing system based on deep learning is used to implement the server data processing method based on deep learning described above, and includes a data acquisition module, a preprocessing module, a time-varying neural network topology, a feature extraction module, a feature fusion module, an adaptive data processing module, and a reinforcement learning module. The data acquisition module is connected to the preprocessing module, the preprocessing module is connected to the time-varying neural network topology, the feature extraction module is arranged in the time-varying neural network topology, the feature extraction module is connected to the feature fusion module, the feature fusion module is connected to the adaptive data processing module, and the adaptive data processing module is connected to the reinforcement learning module.

[0060] The data acquisition module is used to acquire mixed data to be processed. The preprocessing module is used to preprocess the mixed data to be processed and divide the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time series data. The time-varying neural network topology is used to respectively input the four subsets into the corresponding branches in the time-varying neural network topology for preliminary feature extraction. The time-varying neural network topology dynamically adjusts the network edge connection weights according to the input data, captures the correlation features between the data features of different modalities, and obtains cross-modal correlation features. The feature fusion module is used to perform fusion processing based on the cross-modal correlation features to obtain a fusion feature vector. The adaptive data processing module sets multiple deep learning network models and matches the corresponding deep learning network models for the mixed data to be processed based on the fusion feature vector for processing. The reinforcement learning module is used to calculate the processing efficiency, resource occupancy rate, and processing accuracy in the data processing process, establish a reward function of the reinforcement learning algorithm with the processing efficiency, resource occupancy rate, and processing accuracy, and jointly iteratively optimize the deep learning network models based on the reward function.

[0061] In summary, for the server data processing method and system based on deep learning provided by the present invention, through the dual dynamic adjustment mechanism of server load and data features, the edge connection weights of the graph neural network can optimize the multi-modal feature fusion path in real time. The dynamic adaptive data processing architecture can complete model selection in a very short time through the analysis variance, entropy value, cross-correlation coefficient of the fusion feature vector and reinforcement learning scheduling, with improved efficiency compared to manual parameter tuning. And combined with the priority queue and elastic resource scheduling mechanism, the latency of real-time data processing is significantly reduced.

[0062] By adopting a targeted extraction method for different modal data, compared with the general extraction method, the accuracy of feature extraction is effectively improved. Combining the server load and data features, the connection weights of the network edges are adjusted in real time. When the server is under high load, the feature fusion path is automatically optimized to reduce the processing delay. At the same time, the weights are adjusted according to the data complexity to make the extraction of complex data features more accurate and improve the generalization ability of the model. By fusing the server load information and the modal correlation degree, key associated features are preferentially guaranteed to be extracted when resources are scarce, avoiding information loss caused by insufficient resources and improving the integrity of important feature extraction.

[0063] Based on weighted fusion of the association strength matrix and the introduction of a residual structure, not only cross-modal information is fully fused, but also the details of the original features are retained, improving the information integrity of the fused feature vector and effectively enhancing the accuracy of subsequent processing. By analyzing the fused feature vector to judge the task type and screening the model in combination with the model-task mapping relationship, compared with the fixed model configuration, the matching accuracy between the task and the model and the processing efficiency are improved. Using reinforcement learning to calculate the matching degree score and adjust the model, the model selection strategy is continuously optimized. After iteration, the average processing efficiency of the model is improved and resource waste is reduced. A reward function is constructed based on the processing efficiency, resource occupancy rate, and processing accuracy, and through joint optimization of the loss function, the balance of the three is achieved. Compared with single-objective optimization, the overall resource utilization rate of the server is improved.

Claims

1. A server data processing method based on deep learning, characterized in that It includes the following steps: S1: The server obtains the mixed data to be processed, preprocesses the mixed data to be processed, and divides the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time series data; S2: Construct a time-varying neural network topology, and input the four subsets into the corresponding branches in the time-varying neural network topology for preliminary feature extraction. The time-varying neural network topology dynamically adjusts the network edge connection weights according to the input data; S3: Capture the correlation features between the features of different modality data to obtain cross-modal correlation features, and perform fusion processing based on the cross-modal correlation features to obtain a fusion feature vector; S4: Construct an adaptive data processing module, set multiple deep learning network models in the adaptive data processing module, and the adaptive data processing module matches the corresponding deep learning network model for processing based on the fusion feature vector for the mixed data to be processed; S5: Calculate the processing efficiency, resource occupancy rate, and processing accuracy rate during the data processing process, establish a reward function of the reinforcement learning algorithm based on the processing efficiency, resource occupancy rate, and processing accuracy rate, and perform joint iterative optimization on the deep learning network model of the deep learning network model based on the reward function; S6: Process the real-time mixed data to be processed based on the deep learning network model after joint iteration to obtain the multi-modal data processing result.

2. The server data processing method based on deep learning according to claim 1, wherein, The specific process of step S1 is as follows: S11: The server receives the mixed data to be processed through the network interface or the storage device interface, and performs a preliminary check on the integrity and format of the data, checks whether the data file is damaged and whether the data fields are missing. If it is found that the data is incorrect or incomplete, an error prompt is triggered and the reception is stopped, and the data is required to be retransmitted; S12: Perform data cleaning and format standardization processing on the mixed data to be processed, remove noise data, fill in missing values, convert different types of data into a format suitable for processing by the deep learning network model, and add type labels; S13: Based on the type labels, divide the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time series data.

3. A server data processing method based on deep learning according to claim 1, characterized in that, The specific process of step S2 is as follows: S21: Map the structured data subset into a high-dimensional feature vector through the embedding layer, perform multi-scale feature extraction on the image using a deformable convolution feature extractor, output a sequence of feature maps by focusing on the key area through the spatial attention mechanism, adjust the mask ratio according to the text length and semantic complexity, capture the long-range dependencies of the text through the multi-head self-attention mechanism, use the attention mechanism LSTM unit to process the time series data, and dynamically allocate the attention weights to the attention of different time steps to capture the long-term and short-term patterns in the time series data; S22: Use the modality features of different dimensions output by the four subset branches as the initial node set of the deep learning network model. Each node represents the feature representation of a specific modality, and the initial edge weight matrix; S23: Real-time collect the hardware metrics of the server and the software metrics such as data processing latency and model inference speed, and convert the multi-dimensional load metrics into weight adjustment coefficients through a non-linear mapping function; S24: Calculate the statistical features of the input data for each branch, including variance, entropy value, and KL divergence of the feature distribution, evaluate the complexity and uncertainty of the data, calculate the correlation degree between different modalities through a cross-modal attention mechanism, and generate an attention weight matrix; S25: Fuse the data-driven attention weights with the weights adjusted by the server load; S26: Set a sliding time window, update the edge weight matrix at each time step, and when new data arrives, update the edge weights in an incremental learning manner; S27: When the change in edge weights exceeds the threshold, trigger topological structure reorganization, add or delete node connections, and form a new subnet structure.

4. A server data processing method based on deep learning according to claim 1, characterized in that The specific process of step S3 is as follows: S31: Project the modality features of different dimensions into a shared latent space, achieve dimension alignment through linear transformation, and enhance the feature correlation within each modality feature; S32: Capture the correlation between different modality features through a query-key-value pair mechanism, map the server CPU utilization rate and memory occupancy rate metrics to attention weight adjustment coefficients, and fuse the server load information with the correlation degree of modality features of different dimensions; S33: Design attention heads of different scales to capture correlation features from fine-grained to coarse-grained, construct an inter-modal correlation strength matrix, and quantify the mutual influence of different modalities; S34: Calculate the attention weights between pairwise modality features, construct a correlation strength matrix based on the attention weights, perform weighted fusion on each modality feature based on the correlation strength matrix, and introduce a residual structure to retain the original feature information.

5. A server data processing method based on deep learning according to claim 4, characterized in that The specific process of step S33 is as follows: S331: Convert the input features into feature tensors, and evenly divide the feature tensors into specified sub-tensors along the channel dimension, with each sub-tensor corresponding to an independent attention head; S332: Assign different key-value pair dimensions to each attention head, with small dimensions capturing fine-grained correlations and large dimensions capturing coarse-grained correlations; S333: Apply different types of positional encodings to attention heads of different scales to enhance the perception of local and global relationships; S334: For each modality feature, calculate its correlation strength matrix.

6. A server data processing method based on deep learning according to claim 4, characterized in that, The specific process of calculating the attention weights between pairwise modality features in step S34, calculating the correlation strength matrix based on the attention weights, performing weighted fusion on each modality feature based on the correlation strength matrix, and introducing a residual structure to retain the original feature information is as follows: S341: Obtain query vectors, key vectors, and value vectors for each modality feature through linear transformation, calculate attention scores, and obtain the attention weights between pairwise modality features through the softmax function; S342: Construct an association strength matrix based on the attention weights between pairwise modal features, where each element in the matrix represents the attention weight of the modality i for the modality j attention weight; S343: Based on the correlation strength matrix M, perform weighted fusion on each modality feature, perform global average pooling on the fused features to obtain channel descriptors, calculate channel weights through two fully connected layers and activation functions, and multiply the channel weights with the fused features channel by channel to obtain enhanced fused features.

7. A server data processing method based on deep learning according to claim 1, characterized in that The specific process of the adaptive data processing module in step S4 matching a corresponding deep learning network model for processing the mixed data to be processed based on the fused feature vector is as follows: S41: Establish the mapping relationship between models and tasks. The adaptive data processing module analyzes the fused feature vector, and calculates the variance, entropy value, and cross-correlation relationship of the fused feature vector; S42: Determine the task type of the mixed data to be processed according to the analysis result, and filter the preliminary network model based on the mapping relationship between models and tasks according to the task type; S43: When there are records of the same type of task in the historical dynamic optimization table, calculate the matching degree score of the preliminary calculation model based on the reinforcement learning algorithm, and adjust the network model according to the matching degree score; S44: Extract the task priority labels from the fused features. High-priority data directly enters the hot queue of each network model instance pool, and medium- and low-priority data enters the ordinary queue. Each network model processes data based on the queue.

8. A method for processing server data based on deep learning according to claim 1, characterized in that, The specific process of step S5 is as follows: S51: Embed a monitoring module in the data processing flow to collect the processed data volume, processing time, actual used resources, total available resources, correctly predicted data, and total number of samples in real time; S52: Calculate the processing efficiency, resource occupancy rate, and processing accuracy rate during the data processing; S53: Aggregate the metrics in a fixed time window to generate a state vector, and set the action space, which matches the optimization direction of the model; S54: Perform collaborative optimization based on the original loss function of the deep learning model and the reinforcement learning reward function; Combined loss: L = L task ( θ )− β ⋅ r t ; Among them, β is the balance coefficient, which guides the model to optimize in the direction of high rewards. L task is the task loss, representing the loss function value of the deep learning model on a specific task. θ represents the learnable parameters of the model, and the task loss changes with θ changing. r t is the reward signal, coming from the environmental feedback in reinforcement learning.

9. A server data processing system based on deep learning, which is used to implement a server data processing method based on deep learning described in any one of claims 1-8, and is characterized in that, It includes a data acquisition module, a preprocessing module, a time-varying neural network topology, a feature extraction module, a feature fusion module, an adaptive data processing module, and a reinforcement learning module. The data acquisition module is connected to the preprocessing module, the preprocessing module is connected to the time-varying neural network topology, the feature extraction module is set in the time-varying neural network topology, the feature extraction module is connected to the feature fusion module, the feature fusion module is connected to the adaptive data processing module, and the adaptive data processing module is connected to the reinforcement learning module; The data acquisition module is used to acquire the mixed data to be processed; The preprocessing module is used to preprocess the mixed data to be processed, and divide the preprocessed mixed data to be processed into four subsets: structured data, image data, text data, and time series data; The time-varying neural network topology is used to respectively input the four subsets into the corresponding branches in the time-varying neural network topology for preliminary feature extraction. The time-varying neural network topology dynamically adjusts the connection weights of the network edges according to the input data, and captures the correlation features between different modality data features to obtain cross-modal correlation features; The feature fusion module is used to perform fusion processing based on the cross-modal correlation features to obtain a fused feature vector; The adaptive data processing module sets multiple deep learning network models, and matches the corresponding deep learning network model for the mixed data to be processed based on the fused feature vector for processing; The reinforcement learning module is used to calculate the processing efficiency, resource occupancy rate, and processing accuracy during the data processing process, establish a reward function for the reinforcement learning algorithm based on the processing efficiency, resource occupancy rate, and processing accuracy, and jointly iteratively optimize the deep learning network model based on the reward function.

Citation Information

Patent Citations

  • Data analysis method and system based on artificial intelligence

    CN120123700A

  • Automated and adaptive design and training of neural networks

    US20230004796A1

Cited By

  • Microgrid digital twin modeling method

    CN120764117A

  • Multi-mode and reinforcement learning intelligent evaluation method and system

    CN121210929A