Multi-Agent Visualization Data Offloading Method Based on Information Sensing and Reinforcement Learning

By introducing multi-agent visual data unloading methods of information perception and reinforcement learning in the intelligent driving system, the vehicle independently decides on the image compression ratio, solving the problems of insufficient resource utilization and unstable service quality in visual data processing, achieving efficient visual data processing and recognition accuracy, and broadening the intelligent driving application of low-cost vehicles.

CN119917183BActive Publication Date: 2025-07-29BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510415258.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-29
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In the processing of visual data, existing intelligent driving systems have problems such as high computing resource requirements, huge amount of visual data, insufficient resource utilization and unstable service quality, especially in low-cost vehicles, which are difficult to achieve efficient processing.

Method used

Using a multi-agent visual data unloading method based on information perception and reinforcement learning, the vehicle status and task data are collected through edge computing devices, a DQN model is built, and the vehicle independently decides the image compression ratio, and uses image information entropy, original size and number of vehicles connected to the edge device as the decision basis to flexibly adjust the image compression level.

Benefits of technology

It significantly reduces the transmission cost of visual data, improves processing efficiency and recognition accuracy, and the system can quickly adapt to the dynamic traffic environment, improves the performance and stability of the intelligent driving system, and broadens the intelligent driving application market for low-cost vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917183B_ABST
    Figure CN119917183B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data processing, and discloses a multi-agent visual data offloading method based on information perception and reinforcement learning, including the following steps: collecting vehicle status and task data by an edge computing device set and processing them to form training set data; training a multi-agent reinforcement learning network based on DQN to construct a reinforcement learning decision model, that is, a DQN model; the edge computing device distributes the trained DQN model to each vehicle entering its communication range, and the vehicle makes an autonomous decision to determine the compression ratio of the picture sent to the edge computing device. The present invention adopts the above multi-agent visual data offloading method based on information perception and reinforcement learning, based on DQN, enabling each vehicle to flexibly adjust the compression level of the image according to the real-time computing resource load, aiming to balance the time consumption and processing accuracy of visual data processing in a vehicle environment with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and particularly to a multi-agent visual data offloading method based on information perception and reinforcement learning. Background Art

[0002] In recent years, with the rapid development of Internet technology, intelligent driving technology has gradually matured. Currently, intelligent vehicles mainly rely on advanced hardware sensing systems and artificial intelligence algorithms to perceive and analyze road conditions in real time. Visual data, due to its rich details, has become an important part of sensing data and can provide direct insights into road conditions. However, high-end models are difficult to be widely applied in the short term due to their high costs. In addition, visual data processing requires extremely high computing resources, especially in low-cost vehicles where the on-vehicle computing capabilities often cannot meet the requirements. To solve this problem, Vehicular Fog Computing (VFC) has gradually become a viable solution. VFC transforms public transportation vehicles or roadside units with sufficient space and power supply into fog nodes, enabling customer vehicles to transmit visual data to fog nodes through vehicle-to-vehicle (V2V) communication to utilize rich edge computing resources for data processing.

[0003] Although VFC shows potential in visual data processing, it still faces multiple challenges. First, the amount of visual data is huge and usually redundant, and directly transmitting all raw data will exhaust limited resources. Second, although preprocessing technologies such as compression and resolution adjustment can reduce the amount of data, improper compression methods may significantly reduce the recognition accuracy. Therefore, the trade-off between time consumption and processing accuracy becomes a major problem. In addition, the complexity of the vehicle environment increases the challenges, and the fluctuations in the number of fog nodes and customer vehicles result in unstable service supply and demand.

[0004] In summary, the existing technologies have obvious defects in visual data processing, resource utilization, and service quality guarantee, and there is an urgent need for improved solutions to enhance the overall performance of intelligent driving systems. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-agent visual data offloading method based on information perception and reinforcement learning, which is based on DQN reinforcement learning to determine how visual data is compressed before offloading, aiming to balance the time consumption and processing accuracy of visual data processing in a resource-constrained VFC environment.

[0006] To achieve the above purpose, the present invention provides a multi-agent visual data offloading method based on information perception and reinforcement learning, including the following steps:

[0007] Step S1, collect vehicle status and task data based on edge computing devices, and process the data set to construct training set data;

[0008] Step S2: Train the multi-agent reinforcement learning network based on the deep Q-network to construct a DQN model for decision-making;

[0009] Step S3: After the training of the deep Q-network is completed, the edge computing device distributes the trained DQN model to each vehicle entering its communication range. After the vehicle obtains the trained deep Q-network, it makes decisions autonomously to determine the compression ratio of the pictures sent to the edge computing device.

[0010] Preferably, in step S1, the specific process of training set data collection and processing is as follows:

[0011] Step S11: Each vehicle entering the communication range of the edge computing device is regarded as an agent. The vehicle first obtains a visual picture and records the time when the picture is obtained; based on the edge computing device, obtain the relevant information of the image processing task generated by the vehicle connected to it at the current moment;

[0012] Step S12: After the information collection is completed, the edge computing device processes each image at different compression ratios. The specific compression ratios are: 1.0, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, and 0.1;

[0013] Step S13: The edge computing device generates the state required for the training of the deep Q-network according to the picture information sent by different agents, and records the accuracy rate of each picture after being processed by a specific application program and the total time spent from the generation of the picture to the end of the processing, which is convenient for the DQN in the edge computing device to perform subsequent training.

[0014] Preferably, in step S11, the relevant information of the image processing task generated by the vehicle at the current moment includes: the two-dimensional information entropy of the image, the original image file, and the number of vehicles that the edge computing device can connect simultaneously;

[0015] Among them, the two-dimensional information entropy of the image is used to evaluate the complexity and information volume of the image; the original image file is used for subsequent processing.

[0016] Preferably, in step S12, for each compression ratio, the edge computing device respectively records the information entropy of each image, different compression ratios, and the corresponding compression time, the processing time of the YOLO model, and the image recall rate after the YOLO model processes the image;

[0017] Among them, the information entropy of each image is used to analyze the impact of different compressions on the image information; the processing time of the YOLO model is used to evaluate the processing speed of the model under different image qualities; the image recall rate after the YOLO model processes the image is used to evaluate the recognition accuracy of the image after being processed by the YOLO model.

[0018] Preferably, in step S2, the training process of the multi-agent reinforcement learning network based on DQN is as follows:

[0019] Step S21: First, define the state state as a triple (e, s, F); where e represents the information entropy of the image, reflecting the complexity and information content of the image; s represents the original size of the image, indicating the visual data size of the visual data captured by the vehicle set; F represents the computing resource load on the edge computing node, indicating the number of vehicles simultaneously connected to the edge computing device at the current moment;

[0020] Step S22: Define the action space action as follows:

[0021] ;

[0022] Among them, 1, 0.9, 0.8,..., 0.1 represent the picture compression ratio;

[0023] Step S23: With the goal of minimizing the time consumption T and maximizing the recognition accuracy G of all images, define the reward of the reinforcement learning as follows:

[0024] ;

[0025] Among them, represents the average recognition accuracy of all pictures processed by the edge computing node, represents the total average processing time of all pictures processed by the edge computing node, that is, the sum of the compression time and the model processing time; is a weight factor used to adjust the emphasis of the DQN model on and ;

[0026] Step S24: Therefore, the final optimization goal of the DQN reinforcement learning is as follows:

[0027] .

[0028] Preferably, in step S3, the trained DQN model is allocated to each vehicle entering its communication range, and the specific process of the vehicle's autonomous decision-making is as follows:

[0029] Step S31: When the vehicle has a new visual data processing task, the computing device on the vehicle calculates the information entropy of the image and the original size of the picture, and at the same time the vehicle sends information to the edge computing device to obtain the total number of connected vehicles;

[0030] Step S32: Then, take the information entropy of the image, the original size of the picture, and the total number of vehicles connected to the edge computing device as the state, input it into the trained DQN model, and then make corresponding decisions based on the state to give the optimal compression ratio for this image;

[0031] Step S33: Finally, the vehicle adjusts the resolution of the picture according to the optimal compression ratio given by the DQN model, and then sends the picture with adjusted resolution to the edge computing device for subsequent processing.

[0032] Preferably, due to the obvious difference in the traffic flow size in different times and regions, in different times of a day, when the traffic flow changes, the deep Q networks in the edge computing devices in different regions can achieve mutual transmission.

[0033] Preferably, the edge computing device trains deep Q network models under different traffic flow pressures for different traffic flows. When the traffic flow changes, switch the deep Q network model assigned to the vehicle, so as to ensure the efficiency of the compression scheme.

[0034] Therefore, the present invention adopts the above-mentioned multi-agent visual data offloading method based on information perception and reinforcement learning. By introducing the image information entropy, the original size of the image, and the number of vehicles connected to the edge device simultaneously as the decision basis, each vehicle can flexibly adjust the compression level of the image according to the real-time computing resource load, significantly reducing the transmission cost of redundant visual data and improving the overall efficiency of visual data processing; adopting experience-based multi-agent reinforcement learning (MARL), endowing each vehicle with the ability of autonomous decision-making, enabling it to flexibly cope with complex traffic conditions in a fully observable environment, not relying on traditional complex mathematical modeling, but learning through historical experience and automatically adjusting the offloading strategy, so that the system can quickly adapt to the dynamically changing traffic environment; through the simulation evaluation on real-world traffic data, the significant advantages of the present invention in terms of visual data offloading efficiency and recognition accuracy are verified, not only improving the performance of the existing intelligent driving system, but also opening up a new market for the intelligent driving application of low-cost vehicles.

[0035] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Brief Description of the Drawings

[0036] Figure 1 is the flowchart of the multi-agent visual data offloading method based on information perception and reinforcement learning of the present invention;

[0037] Figure 2 is the distributed framework diagram of the multi-agent reinforcement learning of the present invention;

[0038] Figure 3It is the traffic flow change diagram of a certain area in the embodiment of the present invention;

[0039] Figure 4 It is the experimental results of different strategies when the traffic flow is low in the embodiment of the present invention;

[0040] Figure 5 It is the experimental results of different strategies when the traffic flow is high in the embodiment of the present invention. Specific embodiments

[0041] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0042] As Figure 1 shown, the multi-agent visual data offloading method based on information perception and reinforcement learning includes the following steps:

[0043] Step S1: Collect vehicle status and task data based on edge computing devices, and process the data set to construct training set data.

[0044] Step S11: Each vehicle entering the communication range of the edge computing device is regarded as an agent. The vehicle first obtains a visual picture and records the time when the picture is obtained; based on the edge computing device, obtain the relevant information of the image processing task generated by the vehicle connected to it at the current moment.

[0045] Among them, the relevant information of the image processing task generated by the vehicle at the current moment includes: the two-dimensional information entropy of the image, the original image file, and the number of vehicles that the edge computing device can connect simultaneously. The two-dimensional information entropy of the image can be used to evaluate the complexity and information volume of the image; the original image file can be used for subsequent processing.

[0046] Step S12: After the information collection is completed, the edge computing device processes each image with different compression ratios. The specific compression ratios are: 1.0, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, and 0.1.

[0047] For each compression ratio, the edge computing device respectively records the information entropy of each image, different compression ratios, and the corresponding compression time, the processing time of the YOLO model, and the image recall rate after the YOLO model processes the image.

[0048] Among them, the information entropy of each image can be used to analyze the influence of different compressions on the image information; the processing time of the YOLO model can be used to evaluate the processing speed of the model under different image qualities; the image recall rate after the YOLO model processes the image can be used to evaluate the recognition accuracy of the image after being processed by the YOLO model.

[0049] Step S13: The edge computing device generates the state required for training the Deep Q-Network (DQN) based on the image information sent by different agents, and records the accuracy rate after each image is processed by a specific application program and the total time spent from the generation of the image to the end of processing, which facilitates the subsequent training of the DQN in the edge computing device.

[0050] Based on the systematically collected and processed data above, it provides comprehensive basic data for subsequent model training to optimize the overall performance of image compression and processing, and finally forms the training set data, whose data format is shown in Table 1.

[0051] Table 1 Data format of the training set

[0052] ;

[0053] Step S2: Train the multi-agent reinforcement learning network based on DQN to construct a DQN model for decision-making.

[0054] Given the dynamic characteristics of the vehicle network, multiple vehicles (agents) and an edge server interact to process visual data of different volumes, similar to the complex dynamic problem P that traditional dynamic optimization methods strive to solve. Based on this, a distributed framework is constructed for multi-agent reinforcement learning (MARL), as Figure 2 shown.

[0055] Each vehicle, as an individual agent, learns to make optimal decisions by interacting with the environment and other agents. Therefore, P is represented as a process of multi-agent stochastic game, and MARL is used to solve it. In addition, the design of the multi-agent system also allows new agents to be inserted into the system, thus having high scalability in the dynamic VFC environment.

[0056] Among them, the training process of the multi-agent reinforcement learning network based on DQN is as follows:

[0057] Step S21: First, define the state as a triple (e, s, F); where e represents the information entropy of the image, reflecting the complexity and information content of the image; s represents the original size of the image, indicating the visual data size of the visual data captured by the vehicle set; F represents the computing resource load on the edge computing node, indicating the number of vehicles connected to the edge computing device at the current moment.

[0058] Step S22: Define the action space action as follows:

[0059] ;

[0060] Among them, 1, 0.9, 0.8, ..., 0.1 represent the image compression ratios.

[0061] Step S23: With the goal of minimizing the time consumption T and maximizing the recognition accuracy G of all images, define the reward of the reinforcement learning as follows:

[0062] ;

[0063] Among them, represents the average recognition accuracy of all the images processed by the edge computing node, represents the total average processing time of all the images processed by the edge computing node, that is, the sum of the compression time and the model processing time; is a weight factor used to adjust the emphasis degree of the DQN model on and .

[0064] Step S24: Therefore, the final optimization goal of the DQN reinforcement learning is as follows:

[0065] ;

[0066] In the present invention, the state is continuous, and it is no longer suitable to store the state value or action value through a table using traditional reinforcement learning for today's problems. Therefore, in the present invention, the value function approximation based on DQN is used to calculate the action value, and the core of DQN is to use the Q function to make decisions. Among them, in order to accelerate the training process, the k-step average reward is incorporated into the normal DQN, and its training process is shown in Table 2.

[0067] Table 2 Training process of the deep Q-network

[0068] ;

[0069] The present invention introduces the image information entropy, the original size of the image, and the number of vehicles simultaneously connected to the edge device as the decision basis, enabling each vehicle to flexibly adjust the image compression level according to the real-time computing resource load. This innovation significantly reduces the transmission cost of redundant visual data and improves the overall efficiency of visual data processing.

[0070] When the vehicle is connected to the edge computing node, it can quickly select the optimal compression ratio, quickly offload important visual information, ensure that the key data reaches the processing node in the shortest time, not only reduce the consumption of the vehicle's transmission bandwidth, but also improve the service response speed, and at the same time ensure the service quality, solving the contradiction between the time consumption and recognition accuracy of the traditional method.

[0071] Step S3: After the DQN training is completed, the edge computing device distributes the trained DQN model to each vehicle that enters its communication range. After obtaining the trained DQN model, the vehicle makes its own decision to determine the compression ratio of the pictures sent to the edge computing device.

[0072] Step S31: When the vehicle has a new visual data processing task, the computing device on the vehicle calculates the information entropy of the image and the original size of the picture. At the same time, the vehicle sends information to the edge computing device to learn the total number of vehicles it is connected to.

[0073] Step S32: Then, take the information entropy of the image, the original size of the picture, and the total number of vehicles connected to the edge computing device as the state, and input them into the trained DQN model. The DQN makes corresponding decisions based on the state and gives the optimal compression ratio for this image.

[0074] Step S33: Finally, the vehicle adjusts the resolution of the picture according to the optimal compression ratio given by the DQN, and then sends the picture with adjusted resolution to the edge computing device for subsequent processing.

[0075] In addition, the traffic flow sizes in different regions at different times will have obvious differences. During different times of a day, when the traffic flow changes, the DQN networks in the edge computing devices of different regions can transmit to each other.

[0076] For example, Region A currently has a set of DQN networks for dealing with high traffic flow, and Region B currently has a set of DQN models for dealing with low traffic flow. When the traffic flows in A and B change, Region A and Region B can exchange their DQN models, thereby reducing the number of training times, accelerating the transformation of the compression schemes in both places, and more quickly and efficiently coping with the rapidly changing traffic flow.

[0077] In addition, the edge computing device can train DQN models under different traffic flow pressures for different traffic flows. When the traffic flow changes, it can switch the DQN model assigned to the vehicle, thereby ensuring the efficiency of the compression scheme.

[0078] In this process, the present invention adopts experience-based MARL, endows each vehicle with the ability of autonomous decision-making, enables it to flexibly cope with complex traffic conditions in a fully observable environment, does not rely on traditional complex mathematical modeling, but learns through historical experience, automatically adjusts the offloading strategy, so that the system can quickly adapt to the dynamically changing traffic environment. In case of dense traffic or emergencies, the data offloading strategy is optimized in a timely manner, improving the stability and reliability of the entire intelligent driving system. In addition, the information sharing mechanism among vehicles also enhances the overall collaborative combat ability, further improving traffic safety.

[0079] Embodiment

[0080] This embodiment conducts experiments based on the road traffic flow data in a certain area, and the changes in its traffic flow are as Figure 3 shown. Among them, the performance of five strategies including the strategy proposed in the present invention is compared, as specifically shown below:

[0081] (1) N-C: Do not compress the visual images captured by each vehicle.

[0082] (2) Random: Adopt a random compression ratio for the visual images captured by each vehicle.

[0083] (3) IOVE-T: In multi-agent reinforcement learning training, set = 0.2, which means that the IOVE strategy at this time focuses on the time part , that is, the IOVE strategy tends to reduce the overall time of the service.

[0084] (4) IOVE-B: In multi-agent reinforcement learning training, set = 0.7, which means that the IOVE strategy at this time focuses on the revenue part , that is, the IOVE strategy tends to make the overall recognition accuracy of the images higher.

[0085] (5) IOVE-A: In multi-agent reinforcement learning training, set = 0.9, which means that the IOVE strategy at this time focuses more on the revenue part than IOVE-B , that is, the IOVE strategy is more inclined to make the overall recognition accuracy of the images higher than IOVE-B.

[0086] As Figure 4 and Figure 5 shown, it shows the performance of different strategies in terms of time and accuracy in low-traffic and high-traffic scenarios. In both cases, N-C has the highest accuracy but the highest time cost. On the contrary, IOVE-T has the lowest average processing time in both cases, but the lowest accuracy.

[0087] Figure 3 and Figure 4 The results show that IOVE-T, IOVE-B, and IOVE-A show consistent trends in both cases. Specifically, as β increases, the accuracy improves, while the time cost also increases. However, the impact on time is more obvious in the case of high traffic.

[0088] In the low-traffic scenario, the average processing time of IOVE-A is only 0.3 seconds longer than that of IOVE-T; while in the high-traffic scenario, the difference increases to 0.64 seconds. This difference is due to the increase in the number of vehicles and multitasking under high traffic conditions, resulting in a more significant time cost.

[0089] In the case of low traffic, IOVE-B and IOVE-A achieved precisions of 0.48 and 0.53 respectively, which are higher than 0.38 of Randon. In the high-traffic scenario, the precisions of IOVE-B and IOVE-A are both better than Random. It is worth noting that in the high-traffic scenario, the precision of IOVE-A is close to the highest precision of N-C, with only a difference of 0.02, but the average processing time is saved by 0.13 s. This proves the effectiveness of the IOVE algorithm in making wise decisions based on image entropy and size, thus optimizing the compression ratio.

[0090] Through the simulation evaluation on real-world traffic data in this embodiment, the significant advantages of the IOVE algorithm in terms of visual data offloading efficiency and recognition accuracy are verified, which not only improves the performance of the existing intelligent driving system, but also opens up a new market for the intelligent driving applications of low-cost vehicles. Usually, traditional high-end models usually require expensive computing hardware, while the present invention reduces the dependence on these resources, meaning that more consumers can enjoy the convenience and safety brought by intelligent driving.

[0091] Therefore, the present invention adopts the above multi-agent visual data offloading method based on information perception and reinforcement learning, based on DQN reinforcement learning, enabling each vehicle to flexibly adjust the compression level of the image according to the real-time computing resource load and automatically adjust the offloading strategy, so that the system can quickly adapt to the dynamically changing traffic environment, balance the time consumption and processing accuracy of visual data processing in a resource-constrained vehicle environment, improve the performance of the existing intelligent driving system, and open up a new market for the intelligent driving applications of low-cost vehicles.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-agent visual data offloading method based on information perception and reinforcement learning, characterized in that Including the following steps: Step S1: Collect vehicle status and task data based on edge computing devices, process the collected data, and construct a training data set; Among them, the specific process of training set data collection and processing is as follows: Step S11: Each vehicle entering the communication range of the edge computing device is regarded as an agent. The vehicle first obtains a visual image and records the time when the image is obtained; based on the edge computing device, obtain relevant information about the image processing tasks generated by the vehicles connected to it at the current moment; Step S12: After completing the information collection, the edge computing device processes each image at different compression ratios. The specific compression ratios are: 1.0, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, and 0.1; Step S13: The edge computing device generates the state required for deep Q-network training according to the image information sent by different agents, and records the accuracy rate of each image after being processed by a specific application program and the total time spent from the generation of the image to the end of processing, which is convenient for the DQN in the edge computing device to perform subsequent training; Step S2: Train a multi-agent reinforcement learning network based on the deep Q-network to construct a DQN model for decision-making; The training process of the multi-agent reinforcement learning network based on DQN is as follows: Step S21: First, define the state as a triple (e, s, F); where e represents the information entropy of the image, reflecting the complexity and information volume of the image; s represents the original size of the image, indicating the visual data size of the visual data captured by the vehicle set; F represents the computing resource load on the edge computing node, indicating the number of vehicles connected to the edge computing device at the current moment; Step S22: Define the action space action as follows: ; Among them, 1, 0.9, 0.8,..., 0.1 represent the image compression ratio; Step S23: With the goal of minimizing the time consumption T and maximizing the recognition accuracy G of all images, define the reward of reinforcement learning as follows: ; Among them, represents the average recognition accuracy of all the pictures processed by the edge computing node, represents the total average processing time of all the pictures processed by the edge computing node, that is, the sum of the compression time and the model processing time; is a weight factor used to adjust the DQN model's and degree of emphasis; Step S24: Therefore, the final optimization goal of DQN reinforcement learning is as follows: ; Step S3: After the deep Q-network training is completed, the edge computing device distributes the trained DQN model to each vehicle entering its communication range. After the vehicle obtains the trained deep Q-network, it makes an autonomous decision on its own to determine the compression ratio of the image sent to the edge computing device; The specific process of distributing the trained DQN model to each vehicle entering its communication range and the vehicle making an autonomous decision on its own is as follows: Step S31: When the vehicle has a new visual data processing task, the computing device on the vehicle calculates the information entropy of the image and the original size of the image. At the same time, the vehicle sends information to the edge computing device to learn the total number of vehicles connected to it; Step S32: Then, take the information entropy of the image, the original size of the picture, and the total number of vehicles connected to the edge computing device as the state, input it into the trained DQN model, and then make corresponding decisions according to the state to give the optimal compression ratio for this image; Step S33: Finally, the vehicle adjusts the resolution of the picture according to the optimal compression ratio given by the DQN model, and then sends the picture with adjusted resolution to the edge computing device for subsequent processing.

2. The multi-agent visual data offloading method based on information perception and reinforcement learning according to claim 1, characterized in that: In step S11, the relevant information of the image processing task generated by the vehicle at the current moment includes the two-dimensional information entropy of the image, the original image file, and the number of vehicles that the edge computing device can connect simultaneously; Among them, the two-dimensional information entropy of the image is used to evaluate the complexity and information volume of the image; the original image file is used for subsequent processing.

3. The multi-agent visual data offloading method based on information perception and reinforcement learning according to claim 1, characterized in that: In step S12, for each compression ratio, the edge computing device respectively records the information entropy of each image, different compression ratios, and the corresponding compression time, the processing time of the YOLO model, and the image recall rate after the YOLO model processes the image; Among them, the information entropy of each image is used to analyze the impact of different compressions on the image information; the processing time of the YOLO model is used to evaluate the processing speed of the model under different image qualities; the image recall rate after the YOLO model processes the image is used to evaluate the recognition accuracy of the image after being processed by the YOLO model.

4. The multi-agent visual data offloading method based on information perception and reinforcement learning according to claim 1, characterized in that: Since there will be obvious differences in the traffic flow size in different regions at different times, during different times of a day, when the traffic flow changes, the deep Q networks in the edge computing devices in different regions can achieve mutual transmission.

5. The multi-agent visualization data offloading method based on information perception and reinforcement learning according to claim 4, characterized in that: The edge computing device trains the deep Q network model under different traffic flow pressures for different traffic flows, and when the traffic flow changes, switches the deep Q network model assigned to the vehicle, so as to ensure the efficiency of the compression scheme.

Citation Information

Patent Citations

  • Internet of vehicles computing unloading method and system based on distributed edge intelligence

    CN117369899A

  • Vehicle task unloading system and method based on deep Q network

    CN119938277A