Automatic driving strategy intelligent switching method and system, electronic equipment and storage medium
By using the hierarchical selection reinforcement learning HOEAC method, combined with edge computing and cloud computing, and utilizing vibration frequency spectrum collection and analysis, the smart car can quickly and adaptively switch or expand the autonomous driving strategy, solving the problem of low strategy switching efficiency in existing technologies and achieving efficient strategy adjustment.
Patent Information
- Application Number
- CN202511077494.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-09-19
AI Technical Summary
Existing autonomous driving strategy switching technologies have difficulty in quickly adapting to scene changes in complex environments, especially those based on end-to-end deep reinforcement learning neural networks, which have high computing resource consumption and low training efficiency, making it difficult to achieve timely strategy switching.
The hierarchical selective reinforcement learning (HOEAC) method is adopted, combined with edge computing and cloud computing. Through vibration frequency spectrum collection and analysis, the target strategy is selected using a high-level decision network, and the deep reinforcement learning network parameters are updated in the edge computing terminal and cloud computing center to achieve rapid strategy switching and expansion of smart cars.
It enables smart cars to quickly and adaptively switch or expand autonomous driving strategies in complex environments, reduces computing resource consumption, improves the efficiency and adaptability of strategy switching, and avoids manual parameter selection and adjustment.
Smart Images

Figure CN120663945A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a method, system, electronic device, and storage medium for intelligent switching of autonomous driving strategies. Background Art
[0002] During autonomous driving, smart cars observe, understand, and predict their driving environment, making decisions and adjusting their speed and direction based on their driving strategies to ensure the safety of the vehicle, surrounding vehicles, and pedestrians, as well as smooth driving. The optimal driving strategy for smart cars varies in different scenarios, and in complex, real-world traffic environments, the types of scenarios can change dynamically. To ensure safe and smooth driving, smart cars should be able to intelligently adjust or switch their driving strategies based on the type of scenario.
[0003] Currently, autonomous driving policy switching technologies can be primarily categorized into two types: those based on given parameter rules and those based on deep reinforcement learning algorithms. Policy switching technologies based on given parameter rules require programmers to set different autonomous driving parameters or rules for different environmental scenarios, enabling the smart car to adjust its driving strategy for different scenarios. Intelligent policy switching technologies based on deep reinforcement learning can be further categorized into two types: end-to-end model architectures and hierarchical model architectures. The end-to-end model architecture directly inputs environmental states and outputs autonomous driving actions. Neural network parameters are updated through online training, enabling dynamic adjustment of behavioral strategies to environmental changes. However, the deep reinforcement learning neural networks in end-to-end models are large in size. Training these neural networks requires a large number of samples and long iteration times, resulting in high computational and time resource overhead, making it difficult to adapt to scenario changes and switch strategies in a timely manner. The hierarchical model architecture, by training multiple sub-deep reinforcement learning neural networks, enables rapid switching of behavioral strategies based on environmental changes, effectively addressing these issues with end-to-end models.
[0004] Therefore, the current intelligent switching technology of autonomous driving strategies needs to be improved. Summary of the Invention
[0005] In view of this, the purpose of this application is to propose an intelligent switching method, system, device and medium for autonomous driving strategies to solve the above technical problems.
[0006] Based on the above objectives, the first aspect of the present application provides a method for intelligent switching of autonomous driving strategies, which includes:
[0007] Collecting the vibration frequency spectrum of the smart car when it is driving, and converting the vibration frequency spectrum into a digital vibration frequency spectrum variable;
[0008] inputting the digitized vibration frequency spectrum variable into a high-level decision network, and utilizing the high-level decision network to select a target autonomous driving strategy from all stored autonomous driving strategies;
[0009] Switching the autonomous driving strategy of the smart car to the target autonomous driving strategy, and controlling the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy;
[0010] In the edge computing terminal of the smart car, updating the parameters of the underlying deep reinforcement learning network based on the autonomous driving training samples collected and stored by the smart car;
[0011] In the Internet of Vehicles cloud computing center, the parameters of the high-level decision network are updated based on the high-level decision network training samples collected and stored by each of the smart cars, and the autonomous driving strategy is expanded.
[0012] In one embodiment, the underlying deep reinforcement learning network includes a policy network, an evaluation network, a target policy network and a target evaluation network, wherein the policy network is used to output an autonomous driving policy, the evaluation network is used to output an evaluation of the pros and cons of the autonomous driving policy, the target policy network is used to store and update the policy network, and the target evaluation network is used to store and update the evaluation network.
[0013] In one embodiment, controlling the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy includes:
[0014] Controlling the smart car to take a road condition picture in real time, and inputting the road condition picture as the current state into the strategy network;
[0015] The policy network outputs a two-dimensional autonomous driving action; wherein the two-dimensional autonomous driving action includes a target speed and a target driving direction;
[0016] The speed and driving direction of the smart car are adjusted based on the two-dimensional autonomous driving action.
[0017] In one embodiment, in the edge computing terminal of the smart car, updating the parameters of the underlying deep reinforcement learning network based on the autonomous driving training samples collected and stored by the smart car includes:
[0018] In the edge computing terminal of the smart car, the current state, the corresponding two-dimensional autonomous driving action, the next moment state, and the current reward are stored as autonomous driving training samples in a local edge computing database, and the parameters of the underlying deep reinforcement learning network are updated; wherein, the current reward is the sum of the safety penalty item, the stability reward item, and the speed reward item of the smart car; the safety penalty item penalizes operations that cause any collision, the speed reward item rewards faster speeds, and the stability reward item penalizes sudden braking and acceleration with drastic speed changes.
[0019] In one embodiment, in the IoV cloud computing center, updating the parameters of the high-level decision network based on the high-level decision network training samples collected and stored by each of the smart vehicles and expanding the autonomous driving strategy include:
[0020] In the Internet of Vehicles cloud computing center, the digitized vibration frequency spectrum variables, the corresponding target autonomous driving strategy, and the accumulated rewards are used as training samples for a high-level decision network to update the parameters of the high-level decision network and expand the autonomous driving strategy; wherein the accumulated reward is the sum of the rewards within a preset time.
[0021] In one embodiment, collecting the vibration frequency spectrum of the smart car while it is traveling and converting the vibration frequency spectrum into a digital vibration frequency spectrum variable includes:
[0022] When the smart car is driving, a vibration sensor arranged on the wheel hub of the smart car is used to collect a vibration frequency spectrum within a preset continuous time, and the vibration frequency spectrum is converted into a digital vibration frequency spectrum variable.
[0023] In one embodiment, collecting the vibration frequency spectrum of the smart car while it is driving also includes:
[0024] The vibration energy harvesting device is used to convert the vibration energy of the wheel hub during driving into electrical energy for storage in real time, and to supply power to the vibration sensor.
[0025] Based on the same inventive concept, the second aspect of the present application provides an autonomous driving strategy intelligent switching system, which includes:
[0026] An acquisition module, configured to acquire a vibration frequency spectrum of the smart car while it is traveling, and convert the vibration frequency spectrum into a digital vibration frequency spectrum variable;
[0027] an autonomous driving strategy selection module, configured to input the digitized vibration frequency spectrum variable into a high-level decision network, and utilize the high-level decision network to select a target autonomous driving strategy from all stored autonomous driving strategies;
[0028] An autonomous driving strategy switching module, configured to switch the autonomous driving strategy of the smart car to the target autonomous driving strategy, and control the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy;
[0029] An underlying deep reinforcement learning network update module is configured to update parameters of the underlying deep reinforcement learning network in the edge computing terminal of the smart car based on autonomous driving training samples collected and stored by the smart car;
[0030] The high-level decision network update module is used to update the parameters of the high-level decision network based on the high-level decision network training samples collected and stored by each of the smart cars in the Internet of Vehicles cloud computing center, and to expand the autonomous driving strategy.
[0031] Based on the same inventive concept, the third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the processor executes the program, it implements the intelligent switching method of autonomous driving strategy as described in the first aspect above.
[0032] Based on the same inventive concept, the fourth aspect of this application provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the intelligent switching method of autonomous driving strategy described in the first aspect.
[0033] As can be seen from the above, the method for intelligently switching autonomous driving strategies provided in this application proposes a method based on hierarchical reinforcement learning with optional and extensible actor-critic (HOEAC) to solve the problem of intelligent switching of autonomous driving strategies. The intelligent vehicle can distinguish different road environments based on the intrinsic characteristics of the learned vibration frequency spectrum and specifically switch or expand the underlying deep reinforcement learning network representing different autonomous driving strategies. The HOEAC algorithm avoids manual parameter selection and adjustment and can quickly and adaptively switch or expand the underlying strategy.
[0034] This application designs a HOEAC neural network training method that combines edge computing and cloud computing. The smart car edge computing terminal uses autonomous driving strategy training samples to train the corresponding deep reinforcement learning network. The connected car cloud computing center can then use the high-level decision network training samples transmitted back by all smart cars to train the high-level decision network or expand the autonomous driving strategy. This collaborative edge computing and cloud computing training method solves the problem of imbalanced training samples between high-level and low-level neural networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0036] Figure 1 A flowchart of a method for intelligently switching autonomous driving strategies provided in one embodiment of the present application;
[0037] Figure 2 This is an architectural diagram of an intelligent switching method for autonomous driving strategies provided in one embodiment of the present application;
[0038] Figure 3 A schematic diagram of an intelligent switching method for autonomous driving strategies according to an embodiment of the present application;
[0039] Figure 4 A schematic diagram of the physical meaning of the three-dimensional variables of the vibration frequency spectrum provided in one embodiment of the present application;
[0040] Figure 5 A schematic diagram of ADC conversion of three-dimensional variables of a vibration frequency spectrum provided in one embodiment of the present application;
[0041] Figure 6 A flowchart of a method for intelligently switching autonomous driving strategies provided in another embodiment of the present application;
[0042] Figure 7 A schematic diagram of a high-level decision-making network policy expansion provided by another embodiment of the present application;
[0043] Figure 8 A schematic diagram of an intelligent switching system for autonomous driving strategies provided in another embodiment of the present application;
[0044] Figure 9 This is a schematic diagram of an electronic device according to another embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0046] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0047] Reference Figure 1-3 As shown, an embodiment of the present application provides a method for intelligent switching of autonomous driving strategies, the method comprising the following steps:
[0048] Step S10: collecting the vibration frequency spectrum of the smart car when it is driving, and converting the vibration frequency spectrum into a digital vibration frequency spectrum variable;
[0049] Step S20: Input the digitized vibration frequency spectrum variable into a high-level decision network, and use the high-level decision network to select a target autonomous driving strategy from all stored autonomous driving strategies;
[0050] Step S30: Switch the autonomous driving strategy of the smart car to the target autonomous driving strategy, and control the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy;
[0051] Step S40: In the edge computing terminal of the smart car, the parameters of the underlying deep reinforcement learning network are updated based on the autonomous driving training samples collected and stored by the smart car;
[0052] Step S50: In the Internet of Vehicles cloud computing center, the parameters of the high-level decision network are updated based on the high-level decision network training samples collected and stored by each smart car, and the autonomous driving strategy is expanded.
[0053] This application proposes a method for intelligently switching autonomous driving strategies, based on hierarchical selective reinforcement learning (HOEAC), to address the problem of intelligently switching autonomous driving strategies. The intelligent vehicle can distinguish different road environments based on the inherent characteristics of the learned vibration frequency spectrum and selectively switch or expand the underlying deep reinforcement learning network representing different autonomous driving strategies. The HOEAC algorithm avoids manual parameter selection and adjustment and can quickly and adaptively switch or expand the underlying strategy.
[0054] This application designs a HOEAC neural network training method that combines edge computing and cloud computing. The smart car edge computing terminal uses autonomous driving strategy training samples to train the corresponding deep reinforcement learning network. The connected car cloud computing center can then use the high-level decision network training samples transmitted back by all smart cars to train the high-level decision network or expand the autonomous driving strategy. This collaborative edge computing and cloud computing training method solves the problem of imbalanced training samples between high-level and low-level neural networks.
[0055] Step S10, collecting the vibration frequency spectrum of the smart car when it is driving, and converting the vibration frequency spectrum into a digital vibration frequency spectrum variable, including:
[0056] When the smart car is driving, a vibration sensor provided on the wheel hub of the smart car is used to collect a vibration frequency spectrum within a preset continuous time T, and the vibration frequency spectrum is converted into a digital vibration frequency spectrum variable.
[0057] Reference Figure 4 and 5 As shown in the figure, the vibration sensor detects and obtains a three-dimensional, simulated vibration frequency spectrum, which can be expressed by the vibration amplitude A, the vibration plane deflection angle φ, and the vibration vertical deflection angle The three variables represent the specific physical meanings as follows Figure 4 Furthermore, after sampling, holding, quantizing, and encoding the analog vibration frequency spectrum through ADC (analog-to-digital converter), the three-dimensional variable of the vibration frequency spectrum that changes continuously in time is converted into a digital quantity that changes discretely in time, φ D 、 Figure 5 The schematic diagram of ADC conversion of three-dimensional vibration frequency spectrum is shown. Each continuous detection obtains T seconds of digital vibration frequency spectrum variables. And transmitted to the edge computing terminal of the smart car, where F is the total spectrum length and Δf is the sampling frequency interval length.
[0058] In one embodiment, step S10, collecting the vibration frequency spectrum of the smart car while driving, also includes the following steps:
[0059] Step S11: Using a vibration energy harvesting device, the vibration energy of the wheel hub during driving is converted into electrical energy for storage in real time, and used to power the vibration sensor. This can effectively improve energy utilization efficiency and extend the range of the smart car.
[0060] Among them, different autonomous driving strategies are implemented by different underlying deep reinforcement learning networks. For the same environmental input, choosing different autonomous driving strategies may result in different adjustments to the autonomous driving speed and driving direction.
[0061] Specifically, step S20 includes: digitizing the vibration frequency spectrum variable fD Enter the high-level decision network where θ H Represents the parameters of the deep neural network. The high-level decision network outputs the value function H = {h1,h2,...,h C ,}, where C∈R is the total number of all selectable autonomous driving strategies. Based on the value function H, the autonomous driving strategy is sampled, and the probability of each autonomous driving strategy being sampled obeys the Softmax function, that is:
[0062]
[0063] Furthermore, the autonomous driving strategy is switched to the selected strategy, and the underlying deep reinforcement learning network corresponding to the strategy makes action decisions and outputs the target speed and driving direction of the autonomous driving based on the input information.
[0064] In one embodiment, the underlying deep reinforcement learning network includes a policy network, a critique network, a target-policy network, and a target-critique network. The policy network outputs the autonomous driving policy, the critique network outputs a rating of the policy's performance, the target-policy network stores and updates the policy network, and the target-critique network stores and updates the critique network. It's worth noting that when switching autonomous driving strategies, only the policy network corresponding to that strategy is called; the other three neural networks are used to update the parameters of the policy network corresponding to that strategy.
[0065] In one embodiment, in step S30, controlling the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy includes the following steps:
[0066] Step S31: Control the smart car to take a road condition picture in real time, and input the road condition picture as the current state into the strategy network;
[0067] Step S32: The policy network outputs a two-dimensional autonomous driving action; wherein the two-dimensional autonomous driving action includes the target speed and the target driving direction; specifically, the two-dimensional autonomous driving action can be expressed as a t ={v t ,ω t}, where v t is the target velocity, ω t is the target driving direction.
[0068] Step S33: Adjust the speed and direction of the smart car based on the two-dimensional autonomous driving action, thereby achieving operations such as acceleration, deceleration, and lane change of the smart car.
[0069] In one embodiment, step S40, in the edge computing terminal of the smart car, updating the parameters of the underlying deep reinforcement learning network based on the autonomous driving training samples collected and stored by the smart car, includes:
[0070] In the edge computing terminal of the smart car, the current state, the corresponding two-dimensional autonomous driving action, the next moment state, and the current reward are stored in the local edge computing database as autonomous driving strategy training samples, and the parameters of the underlying deep reinforcement learning network are updated; among them, the current reward is the sum of the safety penalty item, the stability reward item, and the speed reward item of the smart car; the safety penalty item penalizes any operation that causes any collision (including with pedestrians, vehicles, or obstacles, etc.), the speed reward item rewards faster speeds, and the stability reward item penalizes sudden braking and acceleration with drastic speed changes. The vehicle is accelerated when the driving situation is stable to save time, while reducing driving risks and improving passenger experience. Optionally, the reward function is as follows:
[0071] r t =I(collision)·ξ-||v t -v t-1 ||-||ω t -ω t-1 ||+v t
[0072] Where r t is the reward at time t, ξ is a large negative penalty term, and I(·) is the decision function: the value is 1 when the event occurs, otherwise it is 0; v t-1 is the target speed at time t-1, ω t-1 is the target driving direction at time t-1.
[0073] In one embodiment, step S50, in the IoV cloud computing center, updates the parameters of the high-level decision network based on the high-level decision network training samples collected and stored by each smart car, and expands the autonomous driving strategy, including:
[0074] In the Internet of Vehicles cloud computing center, the digitized vibration frequency spectrum variable f D , the corresponding target autonomous driving strategy, and the cumulative reward are used as training samples for the high-level decision network to update the parameters of the high-level decision network and expand the autonomous driving strategy; where the cumulative reward is the sum of the rewards within the preset time T. The cumulative reward R is:
[0075]
[0076] The training phase and decision-making phase of edge computing are carried out simultaneously, aiming to use the autonomous driving strategy training samples stored in the edge computing database to update the neural network parameters of the strategy network and optimize the autonomous driving action output strategy.
[0077] Ginseng Figure 6 As shown, in one embodiment, step S40 includes: first randomly selecting B1 groups of training samples of autonomous driving strategy i from the edge computing database (each group includes state s t 、Action a t , reward r t , the next moment state s t+1 ), represented by set B1. Since the optimization target of the autonomous driving strategy, that is, the optimization target of the strategy network parameters, needs to be output by the evaluation network, the evaluation network parameters need to be updated first. The optimization target of the evaluation network parameters is the time difference error of the action-state value function Q(s,a), which can be obtained by the autonomous driving strategy training sample set B1, the target strategy network Target Evaluation Network Calculated together:
[0078]
[0079] Among them, y t is the temporal difference estimate of the action-state value function Q(s,a); γ is the discount factor, ranging from γ∈[0,1), which is used to balance the rewards at the current moment and the future moment. The temporal difference target is calculated Then, the gradient descent method is used to update the evaluation network parameters
[0080]
[0081] where α Q is the update step size of the evaluation network.
[0082] Furthermore, the updated evaluation network is used to update the autonomous driving strategy, train the strategy network, and calculate the expected value of the action-state value function Q(s,a) as the optimization target of the autonomous driving strategy. Approximate calculation:
[0083]
[0084] Next, the policy network parameters are updated using the gradient descent method:
[0085]
[0086] Among them, α π is the update step size of the policy network.
[0087] Finally, soft-update the parameters of the target network:
[0088]
[0089] Where τ is the soft update step size.
[0090] To enable smart cars to intelligently switch autonomous driving strategies, the high-level decision network of the HOEAC algorithm must be trained. Given that the training samples for the high-level decision network consist of T seconds of vibration frequency spectra and accumulated rewards, a single smart car can only collect one set of training samples every T seconds, which is inefficient. Therefore, the cloud computing training phase aims to combine the high-level decision network training samples collected by all smart cars and collaboratively update the high-level decision network parameters.
[0091] Specifically, step S50 includes: the cloud computing center continuously receives high-level decision network training samples (including vibration frequency spectrum f D , the selected switching strategy c and the cumulative reward R in T seconds are stored in the cloud computing database. When executing training, a group B2 of samples is randomly selected from the cloud computing database for training. The sample set is represented by B2. First, the samples with different switching strategies in set B2 are classified and the variance of the cumulative reward of the samples with each autonomous driving strategy is calculated.
[0092]
[0093] Among them, N c Indicates the number of samples with scenario c in set B2. For the same strategy, if the effect is very different within a period of time after application, then it is necessary to split the strategy and treat it as two different strategy trainings. If the cumulative reward variance of strategy c is If the value of the last layer of the neural network parameter in the value function of the high-level neural network for calculating the autonomous driving strategy c is greater than the strategy expansion threshold Θ, the high-level decision network output dimension and the number of autonomous driving strategies C are increased by one dimension, such as Figure 7 As shown, the neural network parameters corresponding to the bold lines are copied and extended. If the strategy is extended, it is necessary to relabel the samples of the extended autonomous driving strategy c in the cloud computing database and the training sample set B2, and convert the accumulated rewards The switching strategy selected in the sample is changed to the newly added autonomous driving strategy c'.
[0094] Furthermore, the high-level value error is calculated to update the parameters of the high-level policy network. H ) can be expressed as:
[0095]
[0096] in, The cth one represents the high-level decision network t Dimensional value function output. Use gradient descent method to update high-level decision network parameters:
[0097]
[0098] Among them, α H is the update step size of the high-level decision network.
[0099] Finally, the cloud computing center needs to transmit the updated parameters back to each edge computing terminal to update the high-level policy network parameters of each terminal. It is worth noting that if the policy is extended, the underlying deep reinforcement learning neural network of the corresponding extended autonomous driving policy needs to be copied.
[0100] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0101] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0102] Based on the same inventive concept, corresponding to any of the above embodiments and methods, this application also provides an automatic driving strategy intelligent switching system, referring to Figure 8 As shown, the system includes the following modules:
[0103] An acquisition module is used to acquire the vibration frequency spectrum of the smart car when it is driving and convert the vibration frequency spectrum into a digital vibration frequency spectrum variable;
[0104] An autonomous driving strategy selection module is used to input the digitized vibration frequency spectrum variables into a high-level decision network and use the high-level decision network to select a target autonomous driving strategy from all stored autonomous driving strategies;
[0105] The autonomous driving strategy switching module is used to switch the autonomous driving strategy of the smart car to the target autonomous driving strategy and control the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy;
[0106] The underlying deep reinforcement learning network update module is used to update the parameters of the underlying deep reinforcement learning network in the edge computing terminal of the smart car based on the autonomous driving training samples collected and stored by the smart car;
[0107] The high-level decision network update module is used to update the parameters of the high-level decision network based on the high-level decision network training samples collected and stored by each smart car in the Internet of Vehicles cloud computing center, and expand the autonomous driving strategy.
[0108] This application provides an intelligent autonomous driving strategy switching system, which proposes a method based on hierarchical selective reinforcement learning (HOEAC) to solve the problem of intelligent switching of autonomous driving strategies. The intelligent vehicle can distinguish different road environments based on the inherent characteristics of the learned vibration frequency spectrum and specifically switch or expand the underlying deep reinforcement learning network representing different autonomous driving strategies. The HOEAC algorithm avoids manual parameter selection and adjustment and can quickly and adaptively switch or expand the underlying strategy.
[0109] This application designs a HOEAC neural network training method that combines edge computing and cloud computing. The smart car edge computing terminal uses autonomous driving strategy training samples to train the corresponding deep reinforcement learning network. The connected car cloud computing center can then use the high-level decision network training samples transmitted back by all smart cars to train the high-level decision network or expand the autonomous driving strategy. This collaborative edge computing and cloud computing training method solves the problem of imbalanced training samples between high-level and low-level neural networks.
[0110] The autonomous driving strategy intelligent switching system in this embodiment has the beneficial effects of the above-mentioned method embodiments, which will not be repeated here.
[0111] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the intelligent switching method of the autonomous driving strategy of any of the above embodiments is implemented.
[0112] Figure 9 1 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1101, a memory 1102, an input / output interface 1103, a communication interface 1104, and a bus 1105. The processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 are communicatively connected to each other within the device via the bus 1105.
[0113] The processor 1101 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0114] The memory 1102 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1102 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1102 and called and executed by the processor 1101.
[0115] The input / output interface 1103 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0116] The communication interface 1104 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0117] The bus 1105 comprises a path for transmitting information between the various components of the device (eg, the processor 1101 , the memory 1102 , the input / output interface 1103 , and the communication interface 1104 ).
[0118] It should be noted that although the above device only shows the processor 1101, the memory 1102, the input / output interface 1103, the communication interface 1104, and the bus 1105, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0119] The electronic device of the above embodiment is used to implement the corresponding automatic driving strategy intelligent switching method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0120] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the automatic driving strategy intelligent switching method described in any of the above embodiments.
[0121] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0122] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the intelligent switching method of autonomous driving strategy as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0123] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0124] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0125] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0126] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.
Claims
1. A method for intelligent switching of autonomous driving strategies, characterized in that: include: Collecting the vibration frequency spectrum of the smart car when it is driving, and converting the vibration frequency spectrum into a digital vibration frequency spectrum variable; inputting the digitized vibration frequency spectrum variable into a high-level decision network, and utilizing the high-level decision network to select a target autonomous driving strategy from all stored autonomous driving strategies; Switching the autonomous driving strategy of the smart car to the target autonomous driving strategy, and controlling the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy; In the edge computing terminal of the smart car, updating the parameters of the underlying deep reinforcement learning network based on the autonomous driving training samples collected and stored by the smart car; In the Internet of Vehicles cloud computing center, the parameters of the high-level decision network are updated based on the high-level decision network training samples collected and stored by each of the smart cars, and the autonomous driving strategy is expanded.
2. The method for intelligent switching of autonomous driving strategies according to claim 1, characterized in that: The underlying deep reinforcement learning network includes a policy network, an evaluation network, a target policy network and a target evaluation network, wherein the policy network is used to output the autonomous driving policy, the evaluation network is used to output the evaluation of the pros and cons of the autonomous driving policy, the target policy network is used to store and update the policy network, and the target evaluation network is used to store and update the evaluation network.
3. The method for intelligent switching of autonomous driving strategies according to claim 2, characterized in that: The controlling the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy includes: Controlling the smart car to take a road condition picture in real time, and inputting the road condition picture as the current state into the strategy network; The policy network outputs a two-dimensional autonomous driving action; wherein the two-dimensional autonomous driving action includes a target speed and a target driving direction; The speed and driving direction of the smart car are adjusted based on the two-dimensional autonomous driving action.
4. The method for intelligent switching of autonomous driving strategies according to claim 3, characterized in that: In the edge computing terminal of the smart car, updating the parameters of the underlying deep reinforcement learning network based on the autonomous driving training samples collected and stored by the smart car includes: In the edge computing terminal of the smart car, the current state, the corresponding two-dimensional autonomous driving action, the next moment state, and the current reward are stored as autonomous driving training samples in a local edge computing database, and the parameters of the underlying deep reinforcement learning network are updated; wherein, the current reward is the sum of the safety penalty item, the stability reward item, and the speed reward item of the smart car; the safety penalty item penalizes operations that cause any collision, the speed reward item rewards faster speeds, and the stability reward item penalizes sudden braking and acceleration with drastic speed changes.
5. The method for intelligent switching of autonomous driving strategies according to claim 4, characterized in that: The method of updating the parameters of the high-level decision network based on the high-level decision network training samples collected and stored by each smart car and expanding the autonomous driving strategy in the Internet of Vehicles cloud computing center includes: In the Internet of Vehicles cloud computing center, the digitized vibration frequency spectrum variables, the corresponding target autonomous driving strategy, and the accumulated rewards are used as training samples for a high-level decision network to update the parameters of the high-level decision network and expand the autonomous driving strategy; wherein the accumulated reward is the sum of the rewards within a preset time.
6. The method for intelligent switching of autonomous driving strategies according to claim 1, characterized in that: The collecting of the vibration frequency spectrum of the smart car when it is traveling and converting the vibration frequency spectrum into a digital vibration frequency spectrum variable includes: When the smart car is driving, a vibration sensor arranged on the wheel hub of the smart car is used to collect a vibration frequency spectrum within a preset continuous time, and the vibration frequency spectrum is converted into a digital vibration frequency spectrum variable.
7. The method for intelligent switching of autonomous driving strategies according to claim 1, characterized in that: The collecting of the vibration frequency spectrum of the smart car when it is driving also includes: The vibration energy harvesting device is used to convert the vibration energy of the wheel hub during driving into electrical energy for storage in real time, and to supply power to the vibration sensor.
8. An intelligent switching system for autonomous driving strategy, characterized in that: include: An acquisition module, configured to acquire a vibration frequency spectrum of the smart car while it is traveling, and convert the vibration frequency spectrum into a digital vibration frequency spectrum variable; an autonomous driving strategy selection module, configured to input the digitized vibration frequency spectrum variable into a high-level decision network, and utilize the high-level decision network to select a target autonomous driving strategy from all stored autonomous driving strategies; An autonomous driving strategy switching module, configured to switch the autonomous driving strategy of the smart car to the target autonomous driving strategy, and control the driving of the smart car through the underlying deep reinforcement learning network corresponding to the target autonomous driving strategy; An underlying deep reinforcement learning network update module is configured to update parameters of the underlying deep reinforcement learning network in the edge computing terminal of the smart car based on autonomous driving training samples collected and stored by the smart car; The high-level decision network update module is used to update the parameters of the high-level decision network based on the high-level decision network training samples collected and stored by each of the smart cars in the Internet of Vehicles cloud computing center, and to expand the autonomous driving strategy.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for intelligent switching of autonomous driving strategies as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the automatic driving strategy intelligent switching method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic driving control method and device
CN113219968A
Vehicle lane changing behavior decision-making method based on deep reinforcement learning and system thereof
CN114074680A
Automatic driving method and device and storage medium
CN114616158A
Vehicle-road collaborative automatic driving decision-making method based on hierarchical reinforcement learning
CN115100866A
Systems and methods for end-to-end learning of optimal driving policy
US20220388522A1