Intelligent camera resource competition optimization method based on game learning
By introducing multi-agent resource optimization methods based on game learning in the intelligent camera system, combined with the zebra optimization algorithm, the problem of resource management pressure and dynamic changes in the intelligent camera system is solved, and efficient and stable resource allocation and system performance improvement are achieved.
Patent Information
- Application Number
- CN202510581904.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In smart camera systems, with the growth of the number of cameras and system scale, the pressure on resource management has intensified, resulting in resource waste, load imbalance and system bottlenecks. It is difficult for existing technologies to effectively deal with dynamic changes in resource demand and environmental uncertainty.
Using a game learning-based intelligent camera resource competition optimization method, combined with collaborative multi-agent reinforcement learning and improved zebra optimization algorithm, the agent realizes intelligent optimization of resource allocation by dynamically adjusting resource management strategies and global resource configuration.
It significantly improves resource utilization, system response speed, stability and ability to adapt to dynamic environments, avoids resource exclusivity and resource hunger, and improves the overall performance and rationality of resource allocation of the system.
Smart Images

Figure CN120224006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of geographic information technology, and in particular to a smart camera resource competition optimization method based on game learning. Background Art
[0002] With the rapid development of artificial intelligence, the Internet of Things, and edge computing technologies, smart camera systems have been widely used in various scenarios, including smart cities, smart transportation, smart security, industrial monitoring, and other fields. Smart cameras can effectively improve the automation level and decision-making efficiency of the system by collecting and analyzing video data in real time. However, with the rapid growth in the number of smart cameras and the continuous expansion of the system scale, the competition between cameras for limited computing resources, communication bandwidth, and storage resources has become increasingly prominent, resulting in increased pressure on resource management and a decline in the overall system performance, which has become a key issue that needs to be solved urgently.
[0003] In the prior art, research on resource management of intelligent camera systems mainly focuses on single-agent optimization or centralized scheduling schemes. Some methods allocate resources through preset fixed rules, such as static allocation based on camera priority, task urgency, or fixed resource ratios. Such methods can still meet basic needs in small-scale scenarios, but in practical applications with multiple cameras and large-scale dynamic changes, they often lack adaptability and flexibility, resulting in resource waste, load imbalance, and even system bottlenecks, and cannot effectively cope with the challenges brought about by dynamic changes in resource demand and environmental uncertainty.
[0004] On the other hand, some technologies have introduced resource scheduling methods based on traditional optimization algorithms (such as linear programming, integer programming, genetic algorithms, particle swarm optimization, etc.). These methods have improved the optimization level of resource allocation to a certain extent, but due to the dynamic interactions and complex dependencies between nodes in the smart camera system, and the time-series and high uncertainty of resource requirements, traditional optimization methods are difficult to respond to system state changes in real time, and have problems such as slow convergence, insufficient global search capabilities, and poor adaptability. At the same time, traditional centralized optimization models often rely on complete global information, which makes it difficult to apply them efficiently in distributed deployments or edge environments with limited communications.
[0005] With the development of technologies such as deep reinforcement learning and multi-agent systems (MAS), more and more research has begun to explore the application of reinforcement learning to resource management problems, leveraging the autonomous learning and decision-making capabilities of agents to achieve intelligent optimization of resource allocation. In particular, cooperative multi-agent reinforcement learning (CMARL) provides a new approach to solving resource coordination problems under distributed and multi-decision-making entities. In such methods, each agent can adaptively update its strategy based on environmental feedback and the behaviors of other agents, achieving local optimality and global cooperation. However, in existing technologies, most cooperative reinforcement learning methods mainly focus on improving the performance of individual agents, lacking systematic consideration of the resource competition conflicts brought about by the complex game relationships among agents, resulting in limited improvement in the overall system efficiency in resource-constrained or conflict-intensive environments.
[0006] Existing technologies generally overlook that resource competition is essentially a multi-agent game behavior. In a real intelligent camera system, each camera hopes to maximize its own resource benefits while having to coordinate and compete with other cameras in an environment with limited resources. Therefore, reinforcement learning relying solely on individual optimal strategies cannot fully reflect the game characteristics in the system, which may lead to the phenomenon of some cameras monopolizing resources and some cameras suffering from resource starvation, ultimately reducing the stability and resource utilization rate of the entire system.
[0007] In addition, there has been much technical research on integrating swarm intelligence optimization algorithms (such as particle swarm optimization, genetic algorithms, ant colony optimization) into the field of intelligent camera resource scheduling. However, the technical path of using the zebra optimization algorithm (ZOA) combined with game learning ideas for resource competition optimization has not been publicly reported. As an emerging swarm intelligence optimization algorithm, the zebra optimization algorithm is inspired by the behavior of zebra herds and has excellent global search capabilities and collaborative optimization characteristics, enabling effective exploration and exploitation in the resource allocation space. However, in existing technologies, there is no systematic method to deeply integrate the zebra optimization algorithm into a multi-agent game learning framework to achieve dynamic adjustment and global optimization of resource allocation. Existing zebra optimization applications mainly focus on static problems such as parameter optimization and path planning, lacking targeted mechanism design for dynamic and competitive resource management problems, resulting in limited optimization effects.
[0008] Therefore, how to provide an intelligent camera resource competition optimization method based on game learning is an urgent problem for those skilled in the art to solve. Summary of the Invention
[0009] An object of the present invention is to propose an intelligent camera resource competition optimization method based on game learning. The present invention fully combines cooperative multi-agent reinforcement learning and the zebra optimization algorithm, and details the technical solution in an intelligent camera system where agents dynamically adjust resource management strategies based on game learning and continuously optimize global resource allocation through an improved zebra optimization algorithm, having the advantages of high resource utilization rate, fast system response speed, strong stability, and good adaptability to dynamic environments.
[0010] An intelligent camera resource competition optimization method based on game learning according to an embodiment of the present invention includes the following steps:
[0011] S1. Initialize the intelligent camera system, collect resource demand data of each intelligent camera, and generate a preliminary resource allocation plan;
[0012] S2. Regard each intelligent camera as an agent, and use cooperative multi-agent reinforcement learning to initialize and optimize the resource management strategy of each intelligent camera in real time. Each agent can learn and optimize resource allocation behavior in the environment according to the operating state and resource requirements of the system;
[0013] S3. Through a shared reward mechanism, the agent optimizes the resource management strategy according to the current resource consumption and system performance feedback, and collaborates with other agents to generate local feedback information of each agent;
[0014] S4. Use the zebra optimization algorithm to perform global optimization adjustment according to the local feedback information of each agent, update the resource configuration, and adjust the preliminary resource allocation plan;
[0015] S5. Dynamically adjust the resource management strategy according to environmental changes and feedback to the agent, and each agent can adaptively respond to system changes;
[0016] S6. Through an iterative process, continuously update the resource management strategy of cooperative multi-agent reinforcement learning and the global resource configuration of the zebra optimization algorithm, and finally form a final resource allocation plan;
[0017] S7. Perform performance evaluation on the final resource allocation plan, verify the bandwidth utilization rate, latency, and resource utilization efficiency of the intelligent camera system, and further optimize the resource management strategy.
[0018] Optionally, the resource demand data of the intelligent camera specifically includes bandwidth, computing power, and storage space, which are used to initially generate a resource allocation plan.
[0019] Optionally, the preliminary resource allocation plan specifically refers to the preliminary resource configuration according to the resource demand data of the intelligent camera, and reasonably allocates the resources of each intelligent camera.
[0020] Optionally, S2 specifically includes:
[0021] S21. Regarding each intelligent camera as an agent, initialize the resource management strategy of each intelligent camera, where the resource management strategy is used to control the resource allocation behavior of the intelligent camera;
[0022] S22. According to the resource requirement data of each camera, enable each agent to continuously optimize the resource management strategy by interacting with other agents in the environment through collaborative multi-agent reinforcement learning;
[0023] S23. Set the state space S, action space A, and reward function R for each agent, where:
[0024] The state space S represents the current resource usage of the intelligent camera, including bandwidth, computing power, and storage space. Each state is represented by a three-dimensional vector: S = [B, C, D], where B is the bandwidth usage, C is the computing resource consumption, and D is the storage space occupancy;
[0025] The action space A represents the resource allocation decisions that the intelligent camera can take, including bandwidth allocation, computing task scheduling, and storage resource allocation. The action space is represented by a discrete multi-dimensional decision vector: A = [A b , A c , A d , where A b is the bandwidth allocation action, A c is the computing resource allocation action, and A d is the storage resource allocation action;
[0026] The reward function R calculates the reward value based on the overall system performance feedback and resource consumption. The goal is to maximize the resource utilization efficiency and system stability. Specifically:
[0027]
[0028] where α and β are weighting coefficients, B max , C max and D max are the maximum available resources for bandwidth, computing, and storage, L t is the latency, D miss is the packet loss rate, s t represents the state at time step t, a t represents the action taken by the agent at time step t, B t , C t and D t are the bandwidth, computing resources, and storage resources used at the current time step t respectively, L max is the maximum tolerance value of the latency, R(st , a t ) is an immediate reward;
[0029] S24. Define the state transition probability P(s, a, s'), which represents the probability that the agent transfers to state s' after taking action a in state s, and update the Q-value function by combining with the deep Q-network through the innovative reward function in reinforcement learning, which is used to evaluate the optimal action strategy in each state:
[0030]
[0031] Among them, Q(s t , a t ) is the Q-value of taking action a in the current state s t , γ is the discount factor, t is the maximum Q-value of the next state s', λ is the update speed parameter, a represents the action at time step t + 1; ′
[0032] S25. Enhance the adaptability of the agent in the dynamic environment through the local optimization mechanism. In each state update, the agent no longer only depends on the global optimization strategy, but further optimizes the decision-making strategy according to the feedback obtained in the local environment;
[0033] S26. Through the cooperation and competition among multiple agents, the agent can dynamically adjust the resource management strategy according to the environmental changes in the system.
[0034] Optionally, the S3 specifically includes:
[0035] S31. Through the shared reward mechanism, the agent generates a reward signal according to the current resource consumption situation and the system performance feedback. After each agent executes the resource allocation decision, it will obtain an immediate reward according to the impact of the agent's behavior on the system performance. The immediate reward not only reflects the efficiency of the agent itself, but also takes into account the cooperation effect with other agents;
[0036] S32. The agent evaluates the current resource usage situation, including bandwidth, computing resources, and storage resources, and determines the impact of resource consumption on the system performance in combination with the overall performance feedback of the system;
[0037] S33. After each agent executes the resource allocation decision, it will share the reward signal with other agents. The shared reward mechanism means that the agent not only pays attention to its own efficiency, but also adjusts the cooperation strategy with other agents through the shared reward signal;
[0038] S34. The agent optimizes the resource management strategy according to the shared reward information. After each agent receives feedback from other agents, it adjusts its resource allocation behavior to coordinate the resource allocation in the multi-agent system and avoid excessive consumption or competition of resources by a single agent.
[0039] S35. The agent obtains the system state information in real time through interaction with the environment and adjusts its behavior strategy according to the shared reward mechanism. The shared reward signal helps the agent identify which behaviors are most effective in cooperation and promotes the global optimization of system resources.
[0040] S36. The agent generates local feedback information and shares it with other agents. Through the shared reward signal and collaborative optimization strategy, the agents jointly adjust the resource allocation.
[0041] Optionally, the specific steps of S4 include:
[0042] S41. The zebra optimization algorithm is adopted to perform global optimization adjustment according to the local feedback information of each agent. Each agent adjusts the resource configuration through the shared reward mechanism and local contribution value, and collaborates through the group cooperation mechanism to jointly optimize the system resource allocation.
[0043] S42. The collaborative efficiency optimization mechanism is adopted. The agent adjusts the resource allocation strategy according to the current resource usage situation, system performance feedback, and the behaviors of other agents. The collaborative efficiency optimization mechanism dynamically evaluates the performance of individual resource configurations and combines the behavior adjustment of the agent with the global system performance. When each agent adjusts the resource allocation, it considers the local resource efficiency and group cooperation benefits:
[0044]
[0045] Where C collaborative (s t ) is the collaborative benefit score of the agent at time step t, γ1 and γ2 are the weighting coefficients for adjusting the local resource efficiency and group cooperation benefits, B t , C t , D t are the usage amounts of bandwidth, computing resources, and storage resources at the current time step t respectively, B max , C max , D max are the maximum values of bandwidth, computing resources, and storage resources respectively, R(s t , a t ) is the reward function, N i is the number of individuals collaborating with agent i, n is the number of agents participating in resource optimization and decision-making, s t represents the state at time step t, a tDenote the action taken by the agent at time step t;
[0046] S43. Through the individual movement mechanism of the zebra optimization algorithm, each agent adjusts its resource allocation according to the gap between its own local optimal solution and the global optimal solution:
[0047] r i (t + 1) = r i (t) + μ1·δ i ·(r best - r i (t)) + μ2·δ i ·(r global - r i (t));
[0048] Where r i (t) is the resource allocation of the i-th agent at time step t, r i (t + 1) is the resource allocation of the i-th agent at time step t + 1, μ1 and μ2 are coefficients controlling individual local search and global search, δ i is a random perturbation factor used to simulate the randomness in nature, r best and r global are the optimal resource allocation of the current agent itself and the global optimal resource allocation respectively;
[0049] S44. Through the calculation of the local contribution value, each agent evaluates its contribution to the system efficiency according to its own behavior and combines the collaborative benefit score C collaborative (s t ), and finally adjusts the resource allocation decision. The local contribution value C local (s t ) is:
[0050]
[0051] Where ε is an adjustment factor, R(s t , a t ) is a reward function;
[0052] S45. Through the group collaboration mechanism, agents share local feedback information and global optimization results among themselves. Based on the collaborative optimization strategy, agents adjust their resource allocation decisions, coordinate the resource allocation among multiple agents, avoid excessive individual competition, and the updated resource allocation is:
[0053]
[0054] Where r adjusted is the adjusted resource allocation, and η is a collaborative adjustment coefficient that controls the contribution of global optimization;
[0055] After the agent executes the resource allocation update, through the adaptive behavior adjustment mechanism, it automatically adjusts the search strategy, determines whether to increase the proportion of exploration or local exploitation based on the updated information of the global optimal solution. As the search progresses, the agent gradually reduces its dependence on global search and enhances the accuracy of local optimization;
[0056] S47. The zebra optimization algorithm updates the resource configuration through multiple iterations, gradually approaching the global optimal resource allocation scheme. After each iteration, the agent updates its own behavior strategy according to the new resource configuration, and finally forms an effective resource scheduling scheme;
[0057] S48. After the agent completes the optimization adjustment, it generates the final resource allocation scheme and coordinates with other agents in the system;
[0058] S49. Perform performance evaluation on the final resource allocation scheme, verify the bandwidth utilization, latency, and resource utilization efficiency of the intelligent camera system, and further adjust the resource management strategy according to the evaluation results.
[0059] Optionally, the S6 specifically includes:
[0060] S61. The agent evaluates the effectiveness of the existing resource management strategy and resource configuration according to the current resource usage situation and the global optimization goal;
[0061] S62. The agent dynamically adjusts the resource allocation behavior by interacting with the environment, and optimizes its own resource management strategy in real time according to the shared feedback information;
[0062] S63. Through collaborative multi-agent reinforcement learning, each agent continuously optimizes the decision-making process and adjusts the resource allocation strategy under the changing state of the intelligent camera system;
[0063] S64. After each update, the agent uses the zebra optimization algorithm to globally optimize the resource configuration, gradually guiding the entire intelligent camera system to converge to the optimal solution;
[0064] S65. After each optimization iteration, the agent will make fine-tuning according to the updated global resource configuration and personal strategy to adapt to the changing requirements of the intelligent camera system and environmental conditions;
[0065] S66. Through multiple iteration processes, a stable resource allocation scheme is finally formed, and the optimal utilization of the resources of the intelligent camera system is achieved through the resource allocation scheme, reducing the latency of the intelligent camera system and improving the overall stability and response speed of the intelligent camera system.
[0066] The beneficial effects of the present invention are:
[0067] Through the deep integration of collaborative multi-agent reinforcement learning and the zebra optimization algorithm, the present invention establishes a complete intelligent camera resource competition optimization mechanism, which can effectively overcome the problems in the prior art such as single resource management method, lack of dynamic adaptability, and insufficient modeling of cooperation and competition relationships. By modeling each intelligent camera as an agent and combining the idea of game learning, the agent not only updates its strategy based on its own resource consumption, but also can dynamically adjust its behavior according to the overall system efficiency and the cooperation effect with other agents, realizing the unity of local optimality and global coordination, and significantly improving the rationality of resource allocation and the overall system performance.
[0068] By designing a collaborative efficiency optimization mechanism and introducing a joint evaluation criterion of local resource efficiency and group cooperation benefit in the resource competition process, the agent considers both its own needs and the overall system interests in the resource allocation decision, avoiding the problems of resource monopolization and frequent conflicts caused by only pursuing individual optimality in traditional reinforcement learning methods. At the same time, through the improved application of the zebra optimization algorithm, global resource allocation optimization is carried out in the multi-agent system. The agent realizes a dynamic balance between local exploration and global exploitation, effectively improving the system convergence speed, reducing the risk of falling into local optimality, and ensuring the global optimality and stability of the resource allocation scheme.
[0069] In multiple iteration processes, the agent continuously optimizes the resource management strategy through environmental interaction feedback and continuously updates the global resource allocation under the guidance of the zebra optimization algorithm, enabling the system to maintain efficient operation in a dynamically changing load environment. Finally, the formed resource allocation scheme can adaptively adjust in a dynamic and complex environment, realizing the maximization of resource utilization rate, the minimization of system latency, the optimal load balance, and significantly improving the overall response speed and stability of the intelligent camera system.
[0070] Therefore, by deeply integrating the game learning mechanism into the intelligent camera resource competition optimization, the present invention establishes a resource management framework with both cooperation and competition, and realizes the global optimal resource scheduling for dynamic environments through the zebra optimization algorithm, breaking through the limitations of the prior art in the field of resource dynamic competition management. It has significant beneficial effects such as high resource utilization efficiency, good system stability, strong adaptability, and fast optimization convergence speed, and has broad application prospects and important engineering promotion value in wide application scenarios of intelligent cameras such as smart cities, intelligent transportation, and intelligent security. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0072] Figure 1Flowchart of an intelligent camera resource competition optimization method based on game learning proposed by the present invention;
[0073] Figure 2 Flowchart of optimizing and adjusting global resource allocation by an improved zebra optimization algorithm for an intelligent camera resource competition optimization method based on game learning proposed by the present invention. Specific implementation manner
[0074] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0075] Refer to Figure 1 and Figure 2 , an intelligent camera resource competition optimization method based on game learning, including the following steps:
[0076] S1. Initialize the intelligent camera system, collect the resource requirement data of each intelligent camera, and generate a preliminary resource allocation plan;
[0077] S2. Regard each intelligent camera as an agent, and use cooperative multi-agent reinforcement learning to initialize and optimize the resource management strategy of each intelligent camera in real time. Each agent can learn and optimize resource allocation behavior in the environment according to the operating state and resource requirements of the system;
[0078] S3. Through a shared reward mechanism, the agent optimizes the resource management strategy according to the current resource consumption and system performance feedback, and collaborates with other agents to generate local feedback information for each agent;
[0079] S4. Use the zebra optimization algorithm to perform global optimization and adjustment according to the local feedback information of each agent, update the resource configuration, and adjust the preliminary resource allocation plan;
[0080] S5. Dynamically adjust the resource management strategy according to environmental changes and feedback to the agent, and each agent can adaptively respond to system changes;
[0081] S6. Through an iterative process, continuously update the resource management strategy of cooperative multi-agent reinforcement learning and the global resource configuration of the zebra optimization algorithm, and finally form a final resource allocation plan;
[0082] S7. Evaluate the performance of the final resource allocation plan, verify the bandwidth utilization rate, latency, and resource utilization efficiency of the intelligent camera system, and further optimize the resource management strategy.
[0083] In this embodiment, the resource requirement data of the intelligent camera specifically includes bandwidth, computing power, and storage space, which are used to initially generate a resource allocation plan.
[0084] In this embodiment, the initial resource allocation plan specifically refers to the initial resource configuration based on the resource requirement data of the intelligent camera, and reasonably allocates the resources of each intelligent camera.
[0085] In this embodiment, S2 specifically includes:
[0086] S21. Treat each intelligent camera as an agent, and initialize the resource management strategy of each intelligent camera. The resource management strategy is used to control the resource allocation behavior of the intelligent camera;
[0087] S22. According to the resource requirement data of each camera, through collaborative multi-agent reinforcement learning, enable each agent to continuously optimize the resource management strategy by interacting with other agents in the environment;
[0088] S23. Set the state space S, action space A, and reward function R of each agent, where:
[0089] The state space S represents the current resource usage of the intelligent camera, including bandwidth, computing power, and storage space. Each state is represented by a three-dimensional vector: S = [B, C, D], where B is the bandwidth usage, C is the computing resource consumption, and D is the storage space occupancy;
[0090] The action space A represents the resource allocation decisions that the intelligent camera can take, including bandwidth allocation, computing task scheduling, and storage resource allocation. The action space is represented by a discrete multi-dimensional decision vector: A = [A b , A c , A d , where A b is the bandwidth allocation action, A c is the computing resource allocation action, and A d is the storage resource allocation action;
[0091] The reward function R calculates the reward value based on the overall system performance feedback and resource consumption. The goal is to maximize the resource utilization efficiency and system stability. Specifically:
[0092]
[0093] Among them, α and β are weighting coefficients, B max , C max and D max are the maximum available resources for bandwidth, computing, and storage, L t is the latency, D miss is the packet loss rate, st represents the state at time step t, a t represents the action taken by the agent at time step t, B t , C t and D t are the usage amounts of bandwidth, computing resources, and storage resources at the current time step t respectively, L max is the maximum tolerance value of latency, R(s t , a t ) is the immediate reward;
[0094] S24. Define the state transition probability P(s, a, s'), which represents the probability that the agent transfers to state s' after taking action a in state s, and update the Q-value function by combining it with the innovative reward function in reinforcement learning to evaluate the optimal action policy in each state:
[0095]
[0096] where Q(s t , a t ) is the Q-value of taking action a in the current state s t , γ is the discount factor, t is the maximum Q-value of the next state s', λ is the update speed parameter, a represents the action at time step t + 1; ′
[0097] S25. Enhance the adaptability of the agent in the dynamic environment through the local optimization mechanism. When updating the state each time, the agent no longer solely depends on the global optimization strategy, but further optimizes the decision-making strategy according to the feedback obtained in the local environment;
[0098] S26. Through the cooperation and competition among multiple agents, the agents can dynamically adjust the resource management strategy according to the environmental changes in the system.
[0099] In this embodiment, the S3 specifically includes:
[0100] S31. Through the shared reward mechanism, the agent generates a reward signal according to the current resource consumption situation and the system performance feedback. After each agent executes the resource allocation decision, it will obtain an immediate reward according to the impact of the agent's behavior on the system performance. The immediate reward not only reflects the efficiency of the agent itself but also takes into account the cooperation effect with other agents;
[0101] S32. The agent evaluates the current resource usage situation, including bandwidth, computing resources, and storage resources, and determines the impact of resource consumption on the system performance in combination with the overall system performance feedback;
[0102] S33. After each agent executes the resource allocation decision, it will share the reward signal with other agents. The shared reward mechanism means that the agent not only focuses on its own efficiency but also adjusts the cooperation strategy with other agents through the shared reward signal;
[0103] S34. The agent optimizes the resource management strategy according to the shared reward information. After each agent receives the feedback from other agents, it adjusts its own resource allocation behavior to coordinate the resource allocation in the multi-agent system and avoid excessive consumption or competition of resources by a single agent;
[0104] S35. The agent obtains the system state information in real time through interaction with the environment and adjusts its behavior strategy according to the shared reward mechanism. The shared reward signal helps the agent identify which behaviors are most effective in cooperation and promotes the global optimization of system resources;
[0105] S36. The agent generates local feedback information and shares it with other agents. Through the shared reward signal and collaborative optimization strategy, the agents jointly adjust the resource allocation.
[0106] In this embodiment, the specific content of S4 is as follows:
[0107] S41. The zebra optimization algorithm is adopted to perform global optimization adjustment according to the local feedback information of each agent. Each agent adjusts the resource configuration through the shared reward mechanism and local contribution value, and through the group cooperation mechanism, collaboratively optimizes the system resource allocation;
[0108] S42. The collaborative efficiency optimization mechanism is adopted. The agent adjusts the resource allocation strategy according to the current resource usage situation, system efficiency feedback, and the behaviors of other agents. The collaborative efficiency optimization mechanism dynamically evaluates the efficiency of individual resource configuration and combines the behavior adjustment of the agent with the global system efficiency. When each agent adjusts the resource allocation, it considers the local resource efficiency and group cooperation benefits:
[0109]
[0110] Among them, C collaborative (s t ) is the collaborative benefit score of the agent at time step t, γ1 and γ2 are the weighting coefficients for adjusting the local resource efficiency and group cooperation benefits, B t , C t , D t are the usage amounts of bandwidth, computing resources, and storage resources at the current time step t respectively, B max , C max , D max are the maximum values of bandwidth, computing resources, and storage resources respectively, R(s t, a t ) is the reward function, N i is the number of individuals collaborating with agent i, n is the number of agents participating in resource optimization and decision-making, s t represents the state at time step t, a t represents the action taken by the agent at time step t;
[0111] S43. Through the individual movement mechanism of the zebra optimization algorithm, each agent adjusts its resource allocation according to the gap between its own local optimal solution and the global optimal solution:
[0112] r i (t + 1) = r i (t) + μ1·δ i ·(r best -r i (t)) + μ2·δ i ·(r global -r i (t));
[0113] Among them, r i (t) is the resource allocation of the i-th agent at time step t, r i (t + 1) is the resource allocation of the i-th agent at time step t + 1, μ1 and μ2 are coefficients controlling individual local search and global search, δ i is a random perturbation factor used to simulate the randomness in nature, r best and r global are the optimal resource allocation of the current agent itself and the global optimal resource allocation respectively;
[0114] S44. Through local contribution value calculation by the agent, each agent evaluates the contribution of its own behavior to the system efficiency and combines the collaborative benefit score C collaborative (s t ), and finally adjusts the resource allocation decision. The local contribution value C local (s t ) is:
[0115]
[0116] Among them, ε is the adjustment factor, R(s t , a t ) is the reward function;
[0117] S45. Through the group collaboration mechanism, agents share local feedback information and global optimization results among themselves. Based on the collaborative optimization strategy, agents adjust their resource allocation decisions, coordinate the resource allocation among multiple agents, avoid excessive individual competition, and the updated resource allocation is:
[0118]
[0119] Among them, r adjusted is the adjusted resource allocation, η is the collaborative adjustment coefficient, controlling the contribution of global optimization;
[0120] S46. After the agent executes the resource allocation update, through the adaptive behavior adjustment mechanism, it automatically adjusts the search strategy, and decides whether to increase the proportion of exploration or local exploitation based on the updated information of the global optimal solution. As the search progresses, the agent gradually reduces its dependence on global search and enhances the accuracy of local optimization;
[0121] S47. The zebra optimization algorithm updates the resource allocation through multiple iterations, gradually approaching the global optimal resource allocation scheme. After each iteration, the agent updates its own behavior strategy according to the new resource allocation, and finally forms an effective resource scheduling scheme;
[0122] S48. After the agent completes the optimization adjustment, it generates the final resource allocation scheme and coordinates with other agents in the system;
[0123] S49. Perform performance evaluation on the final resource allocation scheme, verify the bandwidth utilization, latency and resource utilization efficiency of the intelligent camera system, and further adjust the resource management strategy according to the evaluation results.
[0124] In this embodiment, the S6 specifically includes:
[0125] S61. The agent evaluates the effectiveness of the existing resource management strategy and resource allocation according to the current resource usage situation and the global optimization goal;
[0126] S62. The agent dynamically adjusts the resource allocation behavior by interacting with the environment, and optimizes its own resource management strategy in real time according to the shared feedback information;
[0127] S63. Through collaborative multi-agent reinforcement learning, each agent continuously optimizes the decision-making process and adjusts the resource allocation strategy under the changing state of the intelligent camera system;
[0128] S64. After each update, the agent uses the zebra optimization algorithm to globally optimize the resource allocation, gradually guiding the entire intelligent camera system to converge to the optimal solution;
[0129] S65. After each optimization iteration, the agent will make fine-tuning according to the updated global resource allocation and personal strategy to adapt to the changing requirements of the intelligent camera system and environmental conditions;
[0130] S66. Through multiple iterative processes, a stable resource allocation scheme is finally formed. The optimal utilization of the resources of the intelligent camera system is achieved through the resource allocation scheme, the latency of the intelligent camera system is reduced, and the overall stability and response speed of the intelligent camera system are improved.
[0131] Example 1:
[0132] To verify the feasibility of the present invention in implementation, the present invention is applied to a certain urban traffic management center. The urban traffic management center has deployed more than 1200 intelligent cameras in the main urban area for real-time traffic flow monitoring, violation capture, and emergency warning. These camera systems have long faced serious problems such as resource contention, uneven node loads, high latency, and untimely data processing. Especially during the morning and evening rush hours, the bandwidth requirements of the camera nodes surge, and the system frequently experiences local congestion. Some cameras experience disconnection, image freezing, or even frame loss, seriously affecting traffic dispatching efficiency and safety warning capabilities.
[0133] In response to the above problems, the present invention proposes an intelligent camera resource competition optimization method based on game learning and conducts practical applications in the intelligent transportation management system. During the application process, each camera node is first regarded as an intelligent agent, and the demand characteristics of its bandwidth, computing resources, and storage resources in different time periods are collected. The intelligent agent initializes its resource management strategy according to the system operation status and resource competition situation through a collaborative multi-agent reinforcement learning mechanism and continuously makes dynamic adjustments according to the system feedback during real-time operation.
[0134] During the resource allocation process, each camera node feeds back its own resource utilization status and system performance changes in real time through a shared reward mechanism, and not only adjusts its strategy based on its own resource requirements but also dynamically evolves its behavior according to the cooperation benefits with other intelligent agents. There is a natural resource game relationship among intelligent agents. Each node needs to maximize its own resource benefits while coordinating with other nodes to ensure the overall system performance. The present invention introduces a collaborative efficiency optimization mechanism to comprehensively evaluate the local resource efficiency and group cooperation benefits, guiding each camera node to make decisions, so that the entire system can still maintain the balance and efficiency of resource allocation in a complex competition environment.
[0135] To further improve the global optimality of resource allocation, an improved zebra optimization algorithm is introduced in the application. At each stage of the system operation, the zebra optimization algorithm performs global optimization adjustment of resource allocation according to the local feedback information generated by each intelligent agent. Through the individual movement mechanism and group cooperation mechanism, the intelligent agent continuously updates its resource allocation behavior, gradually converges to the global optimal resource allocation scheme, and can make dynamic adaptive adjustments according to environmental changes.
[0136] In practical applications, the test was conducted for a total of 30 days from March 1, 2025 to March 30, 2025, covering ordinary weekdays, weekends and holidays, for continuous operation testing. The test area is the core urban area, with a total area of about 15 square kilometers, involving 1200 cameras, and the data collection granularity is 5 minutes. Three methods were compared and tested: the traditional static resource allocation method, the ordinary reinforcement learning optimization method, and the combined method of game learning and zebra optimization proposed in the present invention. The test content covers key performance indicators such as system bandwidth utilization rate, resource scheduling delay, node disconnection rate, and data processing success rate.
[0137] According to the statistical results of the actual operation data for 30 consecutive days, after applying the method of the present invention, the bandwidth utilization rate of the system has increased from 72.5% of the traditional method to 92.4%, and the average delay has decreased from 380 milliseconds to 188 milliseconds; the node disconnection rate has dropped from 1.8% to 0.3%, and the data processing success rate has increased to 97.8%. Especially during the traffic peak period, the number of resource conflicts has decreased by 58.7% compared with the traditional method, and the congestion mitigation time has been shortened by more than 62 seconds. The above data show that the present invention can significantly improve the resource allocation efficiency, reduce the system response delay, enhance the system stability and anti-interference ability, and show excellent adaptive optimization ability in complex dynamic load environments.
[0138] Table 1 Comparison table of optimization effects of intelligent camera resource competition
[0139]
[0140] It can be clearly seen from the comparison table of the optimization effects of intelligent camera resource competition that the method of the present invention is significantly superior to the traditional static resource allocation method and the ordinary reinforcement learning optimization method in all key performance indicators. In terms of the average bandwidth utilization rate, the traditional static resource allocation method only reaches 72.5%, the ordinary reinforcement learning optimization method has been improved to 81.2%, while the comprehensive method based on game learning and zebra optimization algorithm of the present invention is further improved to 92.4%. This result shows that through the dynamic game learning and resource collaborative optimization among agents, the present invention can more fully explore and utilize the available bandwidth resources of the system, effectively reduce resource waste, and improve the overall network resource utilization efficiency.
[0141] In terms of the system response speed, i.e., the system average latency metric, the latency of the traditional static method is as high as 380 milliseconds. Although the latency of the ordinary reinforcement learning method has decreased, it still reaches 290 milliseconds. The method of the present invention further reduces the latency to 188 milliseconds. By introducing collaborative multi-agent reinforcement learning to adjust the resource management strategy in real time and combining with the zebra optimization algorithm to dynamically optimize the global resource allocation, the present invention can quickly respond to changes in resource requirements, significantly reduce the queuing waiting time caused by resource competition among camera nodes, thereby significantly reducing the overall system latency and improving the real-time performance of data processing.
[0142] The node dropout rate is also an important indicator to measure the system stability. Data shows that under the traditional static configuration, the node dropout rate is 1.8%. The ordinary reinforcement learning optimization can reduce it to 0.9%. After applying the method of the present invention, the node dropout rate is further reduced to only 0.3%. This change fully demonstrates that through the effective cooperation among agents by the game learning mechanism and the reasonable resource allocation guided by the zebra optimization algorithm, the problem of node dropout caused by improper resource allocation is effectively alleviated, greatly enhancing the stability and availability of the system.
[0143] In terms of the data processing success rate, the present invention also shows great advantages. The data processing success rate of the traditional static configuration method is 85.4%. The ordinary reinforcement learning optimization method increases it to 91.6%. While the method of the present invention reaches a high level of 97.8%. This indicates that in a complex dynamic load environment, the present invention can ensure the integrity and processing success rate of camera data to the greatest extent through dynamic resource scheduling, thereby supporting the accurate processing and rapid response of a large amount of real-time video data by the traffic command center.
[0144] During the peak period when resource competition is intense, the method of the present invention also demonstrates excellent performance. In terms of the indicator of the number of resource conflicts per hour during the peak period, the average number of conflicts per hour of the traditional method is 9.2 times. The ordinary reinforcement learning optimization reduces it to 6.1 times. While the method of the present invention is effectively controlled at 3.8 times. This shows that through the agents' dynamic optimization decision-making based on game learning and combined with the global adjustment by the zebra optimization algorithm, the resource contention phenomenon among cameras can be significantly reduced, improving the coordination and efficiency of system resource allocation.
[0145] In addition, after congestion occurs during the traffic peak period, the congestion mitigation time of the system is also an important indicator. The traditional method needs about 135 seconds to relieve congestion once. The ordinary reinforcement learning optimization method shortens it to 92 seconds. While the present invention further shortens the average congestion mitigation time to 56 seconds. By reconfiguring resources and adjusting scheduling faster, the present invention can restore the smooth operation of the system in the shortest time, avoid long-term congestion caused by resource bottlenecks, and greatly improve the dynamic recovery ability of the system in a high-pressure environment.
[0146] Based on the comprehensive analysis of the above data, the conclusion can be drawn that by constructing a multi-agent resource optimization framework based on a game learning mechanism and continuously optimizing the global resource allocation in combination with an improved zebra optimization algorithm, the present invention effectively improves the resource utilization efficiency, response speed, stability, and dynamic adaptability of the intelligent camera system. Especially in the actual application environment with large fluctuations in resource demand and intense competition, the method provided by the present invention can significantly reduce latency, reduce conflicts, improve system throughput and stability, providing reliable technical support for large-scale intelligent camera systems such as intelligent transportation and urban surveillance, and having extremely high engineering application value and promotion prospects.
[0147] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A smart camera resource competition optimization method based on game learning, characterized in that: The steps include: S1, initialize the smart camera system, collect resource demand data of each smart camera, and generate a preliminary resource allocation plan; S2, each smart camera is regarded as an intelligent agent, and collaborative multi-agent reinforcement learning is used to initialize and optimize the resource management strategy of each smart camera in real time. Each intelligent agent can learn and optimize resource allocation behavior in the environment according to the system's operating status and resource requirements; S3. Through the shared reward mechanism, the agent optimizes the resource management strategy and collaborates with other agents based on the current resource consumption and system performance feedback to generate local feedback information for each agent; S4, using the zebra optimization algorithm to perform global optimization adjustments based on the local feedback information of each agent, update resource configuration, and adjust the preliminary resource allocation plan; S5. Dynamically adjust resource management strategies according to environmental changes and provide feedback to agents. Each agent can adaptively respond to system changes. S6. Through the iterative process, the resource management strategy of collaborative multi-agent reinforcement learning and the global resource configuration of the zebra optimization algorithm are continuously updated to finally form the final resource allocation plan; S7. Perform performance evaluation on the final resource allocation solution to verify the bandwidth utilization, latency, and resource utilization efficiency of the smart camera system and further optimize the resource management strategy.
2. According to the game learning-based intelligent camera resource competition optimization method of claim 1, it is characterized in that: The resource requirement data of the smart camera specifically includes bandwidth, computing power and storage space, and is used to preliminarily generate a resource allocation plan.
3. According to the game learning-based intelligent camera resource competition optimization method of claim 1, it is characterized in that: The preliminary resource allocation plan specifically refers to preliminary resource configuration based on resource demand data of smart cameras to reasonably allocate resources of each smart camera.
4. According to the game learning-based intelligent camera resource competition optimization method of claim 1, it is characterized in that: The S2 specifically includes: S21, treating each smart camera as an intelligent agent, initializing the resource management strategy of each smart camera, the resource management strategy is used to control the resource allocation behavior of the smart camera; S22. Based on the resource demand data of each camera, collaborative multi-agent reinforcement learning enables each agent to continuously optimize the resource management strategy by interacting with other agents in the environment; S23. Set the state space S, action space A and reward function R of each agent, where: The state space S represents the current resource usage of the smart camera, including bandwidth, computing power, and storage space. Each state is represented by a three-dimensional vector: S = [B, C, D], where B is the bandwidth usage, C is the computing resource consumption, and D is the storage space occupancy; The action space A represents the resource allocation decisions that the smart camera can take, including bandwidth allocation, computing task scheduling, and storage resource allocation. The action space is represented by a discrete multidimensional decision vector: A = [A b ,A c ,A d ], where A b is the bandwidth allocation action, A c Assign actions to computing resources, A d Allocate actions for storage resources; The reward function R calculates the reward value based on the overall system performance feedback and resource consumption. The goal is to maximize resource utilization efficiency and system stability. Specifically: Among them, α and β are weighting coefficients, B max , C max and D max is the maximum available resource of bandwidth, computing and storage, L t is the delay, D miss is the packet loss rate, s t represents the state at time step t, a t represents the action taken by the agent at time step t, B t , C t and D t are the bandwidth, computing resources and storage resource usage at the current time step t, L max is the maximum tolerable delay, R(s t ,a t ) is an immediate reward; S24. Define the state transition probability P(s,a,s'), which represents the probability that the agent will transfer to state s' after taking action a in state s, and update the Q-value function by combining the innovative reward function update in reinforcement learning with the deep Q network to evaluate the optimal action strategy in each state: Among them, Q(s t , a t ) is the current state s t Take action a t Q value, γ is the discount factor, is the maximum Q value of the next state s', λ is the update speed parameter, a ′ Represents the action at time step t+1; S25. Enhance the adaptability of the agent in a dynamic environment through a local optimization mechanism. At each state update, the agent no longer relies solely on the global optimization strategy, but further optimizes the decision-making strategy based on the feedback obtained in the local environment. S26. Through collaboration and competition among multiple agents, agents can dynamically adjust resource management strategies in the system according to environmental changes.
5. According to the game learning-based intelligent camera resource competition optimization method of claim 1, it is characterized in that: The S3 specifically includes: S31. Through the shared reward mechanism, the agent generates a reward signal based on the current resource consumption and system performance feedback. After executing the resource allocation decision, each agent will receive an immediate reward based on the impact of the agent's behavior on the system performance. The immediate reward not only reflects the agent's own performance, but also takes into account the collaborative effect with other agents. S32, the agent evaluates the current resource usage, including bandwidth, computing resources, and storage resources, and determines the impact of resource consumption on system performance in combination with the overall performance feedback of the system; S33. After executing the resource allocation decision, each agent will share the reward signal with other agents. The shared reward mechanism means that the agent not only pays attention to its own performance, but also adjusts the cooperation strategy with other agents through the shared reward signal. S34, the agents optimize resource management strategies based on the shared reward information. After receiving feedback from other agents, each agent adjusts its own resource allocation behavior and coordinates resource allocation in the multi-agent system to avoid excessive resource consumption or competition of a single agent. S35, the intelligent agent obtains system status information in real time through interaction with the environment, and adjusts its behavior strategy according to the shared reward mechanism. The shared reward signal helps the intelligent agent identify which behaviors are most effective in cooperation and promotes the optimization of the overall resources of the system; S36. The agent generates local feedback information and shares it with other agents. By sharing reward signals and collaborative optimization strategies, the agents jointly adjust resource allocation.
6. According to the game learning-based intelligent camera resource competition optimization method of claim 1, it is characterized in that: The S4 specifically includes: S41, using the zebra optimization algorithm, making global optimization adjustments based on the local feedback information of each agent, each agent adjusts resource allocation through a shared reward mechanism and local contribution value, and collaboratively optimizes system resource allocation through a group collaboration mechanism; S42. Adopting the collaborative efficiency optimization mechanism, the agent adjusts the resource allocation strategy according to the current resource usage, system performance feedback and the behavior with other agents. The collaborative efficiency optimization mechanism combines the agent's behavior adjustment with the global system efficiency by dynamically evaluating the efficiency of individual resource allocation. When adjusting resource allocation, each agent considers local resource efficiency and group collaboration benefits: Among them, C collaborative (s t ) is the collaborative benefit score of the agent at time step t, γ1 and γ2 are weighted coefficients for adjusting local resource efficiency and group collaboration benefits, B t ,C t ,D t are the bandwidth, computing resources and storage resource usage at the current time step t, respectively, max ,C max , D max are the maximum values of bandwidth, computing resources and storage resources respectively, R(s t ,a t ) is the reward function, N i is the number of individuals cooperating with agent i, n is the number of agents involved in resource optimization and decision-making, s t represents the state at time step t, a t represents the action taken by the agent at time step t; S43, Zebra optimization algorithm uses the individual movement mechanism, and each agent adjusts resource allocation according to the gap between its local optimal solution and the global optimal solution: r i (t+1)=r i (t)+μ1·δ i ·(r best -r i (t))+μ2·δ i ·(r global -r i (t)); Among them, r i (t) is the resource allocation of the ith agent at time step t, r i (t+1) is the resource configuration of the ith agent at time step t+1, μ1 and μ2 are the coefficients controlling the individual local search and global search, δ i is a random disturbance factor used to simulate the randomness in nature, r best and r global They are respectively the optimal resource configuration of the current agent itself and the global optimal resource configuration; S44, the agent calculates the local contribution value, and each agent evaluates the contribution of its own behavior to the system efficiency, and combines the synergy benefit score C collaborative (s t ), and finally adjust the resource allocation decision, the local contribution value C local (s t )for: Among them, ε is the adjustment factor, R(s t , a t ) is the reward function; S45. Through the group collaboration mechanism, the agents share local feedback information and global optimization results. Based on the collaborative optimization strategy, the agents adjust resource allocation decisions and coordinate resource allocation among multiple agents to avoid excessive competition among individuals. The updated resource configuration is: Among them, r adjusted is the adjusted resource allocation, η is the collaborative adjustment coefficient, which controls the contribution of global optimization; S46. After executing the resource allocation update, the agent automatically adjusts the search strategy through the adaptive behavior adjustment mechanism, and decides whether to increase the proportion of exploration or local development according to the updated information of the global optimal solution. As the search progresses, the agent gradually reduces the reliance on global search and enhances the accuracy of local optimization. S47, the zebra optimization algorithm updates resource configuration through multiple iterations, gradually approaching the global optimal resource allocation plan. After each iteration, the agent updates its own behavior strategy according to the new resource configuration, and finally forms an effective resource scheduling plan; S48, after completing the optimization adjustment, the agent generates a final resource allocation plan and coordinates with other agents in the system; S49. Perform performance evaluation on the final resource allocation solution to verify the bandwidth utilization, latency, and resource utilization efficiency of the smart camera system, and further adjust the resource management strategy based on the evaluation results.
7. According to the game learning-based intelligent camera resource competition optimization method of claim 1, it is characterized in that: The S6 specifically includes: S61, the agent evaluates the effectiveness of existing resource management strategies and resource allocation based on the current resource usage and global optimization goals; S62, the agent dynamically adjusts resource allocation behavior by interacting with the environment, and optimizes its resource management strategy in real time based on shared feedback information; S63, through collaborative multi-agent reinforcement learning, each agent continuously optimizes the decision-making process and adjusts the resource allocation strategy under the ever-changing state of the smart camera system; S64, after each update, the agent uses the zebra optimization algorithm to globally optimize the resource allocation, and gradually guides the entire intelligent camera system to converge to the optimal solution; S65. After each optimization iteration, the agent will fine-tune according to the updated global resource configuration and personal strategy to adapt to the changing smart camera system requirements and environmental conditions; S66. Through multiple iterations, a stable resource allocation scheme is finally formed. The resource allocation scheme is used to optimize the resource utilization of the smart camera system, reduce the delay of the smart camera system, and improve the overall stability and response speed of the smart camera system.