Model deployment method and system, edge access device, vehicle and storage medium
By splitting and distributing the target model through edge access, the high hardware requirements of traditional large model deployment methods are solved, and resource utilization is improved and dynamic adaptability is achieved.
Patent Information
- Application Number
- CN202410469452.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-18
- Publication Date
- 2025-10-24
AI Technical Summary
Traditional large-scale model deployment methods have high requirements on the hardware of deployment nodes, cannot be adjusted dynamically, and are not suitable for the ever-changing complex dynamic environment in intelligent transportation scenarios.
The target model is split through the edge access end to generate a policy file, which is then distributed and deployed on the edge access end and the vehicle end, utilizing the resources of multiple devices to achieve dynamic adjustment.
It improves resource utilization, reduces hardware resource requirements for a single device, and adapts to the dynamic changes in intelligent transportation scenarios.
Smart Images

Figure CN120835308A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent driving, in particular to a model deployment method and system, an edge access device, a vehicle and a storage medium. BACKGROUND
[0002] In recent years, the intelligent level of automobiles is continuously improved, and intelligent cockpits and automatic driving have become the main selling points of new automobiles. Artificial intelligence (AI) large models have strong feature extraction and generalization capabilities, and can explore deeper semantic information of things. However, the number of parameters of the order of hundreds of millions and the complex network structure of the AI large model increase the requirements for storage and computing resources of intelligent and connected automobiles, and automobile manufacturers demand lower vehicle manufacturing costs, which poses new challenges to the large-scale commercial use of vehicle-mounted large models. Under the premise of the high development of intelligent mobile devices and mobile communication technologies, the resources of edge servers, vehicle machines and in-vehicle mobile devices can be combined to distribute the large model, thereby reducing the hardware cost of intelligent and connected automobiles under the premise of ensuring the effectiveness of the time delay.
[0003] At present, there are mainly two methods for large model deployment and execution based on edge computing: one is to deploy and execute the model at the edge access point through edge computing and Internet of Things technology to reduce data transmission delay and bandwidth consumption, but this method deploys or executes the large model at a single node, and feeds back the model execution result through data return, which has high requirements for the hardware of the single node. The other is to deploy each network layer of the large model on multiple computing nodes, and obtain the data flow relationship between the nodes through a configuration file, so as to realize the rapid deployment and application of the large model on the device side, but this method only considers the static distributed deployment of the large model, and is not suitable for the complex dynamic environment that changes all the time in the intelligent transportation scenario. SUMMARY
[0004] Therefore, the present application provides a model deployment method, system, edge access device, vehicle and storage medium to solve the problem that the traditional large model deployment method has high requirements for the hardware of the deployment node and cannot be dynamically adjusted.
[0005] In a first aspect, the present application provides a model deployment method applied to an edge access end, which comprises:
[0006] obtaining target model information and resource occupation information sent by a vehicle end;
[0007] based on the target model information, finding and obtaining a target model requested to be deployed by the vehicle end;
[0008] split the target model based on the resource occupation information to obtain a first target sub-model with the deployment end being the edge access end and a second target sub-model with the deployment end being the vehicle end, and generate a policy file for representing cooperative execution information between the first target sub-model and the second target sub-model;
[0009] deploy the first target sub-model, and send the policy file and the second target sub-model to the vehicle end to enable the vehicle end to deploy the second target sub-model based on the policy file.
[0010] Thus, by using the edge access end to receive the target model information and the resource occupation information sent by the vehicle end, and splitting the target model corresponding to the target model information based on the resource occupation information, the edge access end obtains the first target sub-model with the deployment end being the edge access end, the second target sub-model with the deployment end being the vehicle end, and the policy file for representing cooperative execution information between the first target sub-model and the second target sub-model, deploys the first target sub-model at itself, and sends the policy file and the second target sub-model to the vehicle end to enable the vehicle end to complete the deployment of the second target sub-model, thereby realizing distributed dynamic deployment of a large model, fully utilizing the resources of each device, improving the resource utilization rate, and reducing the hardware resource requirement for a single device.
[0011] In an optional implementation, splitting the target model based on the resource occupation information to obtain the first target sub-model with the deployment end being the edge access end and the second target sub-model with the deployment end being the vehicle end includes:
[0012] model pruning and model quantization are performed on the target model to obtain an optimized target model;
[0013] The optimized target model is split into multiple sub-models based on storage resource occupation information, computing resource occupation information, and network resource occupation information in the resource occupation information.
[0014] The deployment end corresponding to each of the multiple sub-models is determined, the sub-model with the deployment end being the edge access end is taken as the first target sub-model, and the sub-model with the deployment end being the vehicle end is taken as the second target sub-model.
[0015] Thus, by performing model pruning and model quantization on the target model, the size of the target model itself is greatly reduced, the storage occupation space is reduced, the target model after model pruning and model quantization is split into multiple sub-models according to the resource occupation information of the vehicle end, and the deployment end of each sub-model is determined, so as to facilitate the deployment of each sub-model to a suitable device.
[0016] In an optional implementation, the method further includes:
[0017] Based on the policy file and the first target sub-model, the second target sub-model deployed cooperatively at the vehicle end performs model training;
[0018] And / or, based on the policy file and the first target sub-model, the second target sub-model deployed cooperatively at the vehicle end performs model inference, and transmits the obtained model inference result to the vehicle end.
[0019] Therefore, the resources of multiple devices are utilized for model training and / or model inference of the target model, improving resource utilization and efficiency of model execution.
[0020] In an optional implementation, the edge access end includes multiple edge access points, and the method further includes:
[0021] Receiving a node switching request sent by the vehicle end, and obtaining a candidate edge access point to be connected next by the vehicle end based on the node switching request;
[0022] Obtaining intermediate data generated when the current edge access point cooperates with the vehicle end to perform model training and / or model inference, the current edge access point being an edge access point currently connected with the vehicle end;
[0023] Controlling the current edge access point to disconnect with the vehicle end, and controlling the candidate edge access point to establish connection with the vehicle end;
[0024] Sending the policy file, the first target sub-model and the intermediate data to the candidate edge access point, so that the candidate edge access point synchronously stores the intermediate data, and deploys the first target sub-model according to the policy file.
[0025] Therefore, the candidate edge access point synchronously stores the intermediate data, and deploys the first target sub-model according to the policy file, realizes switching of the edge access point, so that the vehicle end continues to perform a calculation task of model training and / or model inference through the candidate edge access point, and ensures reliability and effectiveness of model execution.
[0026] In an optional implementation, after controlling the current edge access point to disconnect with the vehicle end, the method further includes:
[0027] Sending the target model, the policy file and the intermediate data to the cloud end, so that the cloud end deploys the target model according to the policy file, and performs model training and / or model inference based on the intermediate data and the target model.
[0028] Therefore, by sending the target model, the policy file and the intermediate data to the cloud end, the vehicle end can perform model training and / or model inference through the cloud end in necessary cases, meets use requirements of the vehicle end, and ensures reliability of the system.
[0029] In an optional implementation, the vehicle side includes a plurality of vehicles, each vehicle corresponding to a vehicle machine and at least one in-vehicle device; the method further includes:
[0030] monitoring the device state of the vehicle machine or the in-vehicle device in each vehicle in the vehicle side, and determining whether the vehicle side completes the deployment of the target model;
[0031] if it is detected that there is a vehicle machine or an in-vehicle device with an offline or online device state, and the vehicle side has not completed the deployment of the target model, returning to the step of obtaining the target model information and the resource occupation information sent by the vehicle side;
[0032] and / or, if it is detected that there is a vehicle machine or an in-vehicle device with an offline device state, and the vehicle side has completed the deployment of the target model, obtaining an offline third target sub-model, and deploying the offline third target sub-model on the current edge access point; the offline third target sub-model is a second target sub-model deployed on the vehicle machine or the in-vehicle device with the offline device state;
[0033] and / or, if it is detected that there is a vehicle machine or an in-vehicle device with an online device state, and the vehicle side has completed the deployment of the target model, after receiving a new model deployment request corresponding to the online vehicle machine or the in-vehicle device sent by the vehicle side, establishing a connection with the online vehicle machine or the in-vehicle device, and performing the step of obtaining the target model information and the resource occupation information sent by the vehicle side.
[0034] Thus, by monitoring the online or offline state of the vehicle machine and the in-vehicle device, the deployment and execution process of the target model are dynamically adjusted, and the dynamic change of the device is better adapted.
[0035] In an optional implementation, before obtaining the target model information and the resource occupation information sent by the vehicle side, the method further includes:
[0036] after receiving the model deployment request sent by the vehicle side, performing legality verification on the vehicle side;
[0037] after detecting that the legality verification is passed, establishing a connection with the vehicle side, and performing the step of obtaining the target model information and the resource occupation information sent by the vehicle side.
[0038] Thus, by performing legality verification on the vehicle side requesting to access the edge access side, the security and legality of the access device are ensured, malicious access is avoided, and the safe and effective operation of the system is ensured.
[0039] In an optional implementation, based on the target model information, the target model requested to be deployed by the vehicle side is searched and obtained, including:
[0040] detecting whether there is a target model corresponding to the target model information stored locally;
[0041] If the target model is not stored locally, a model acquisition request is sent to the cloud, and after receiving the response feedback sent by the cloud based on the model acquisition request, the target model is downloaded from the cloud;
[0042] If the target model is stored locally, the target model is obtained locally.
[0043] Thus, in the case where the target model is stored locally, the target model requested by the vehicle end to deploy is conveniently and quickly obtained by local searching, and in the case where the target model is not stored locally, the target model is obtained by sending a request to the cloud, thereby meeting the diversified needs of the vehicle end.
[0044] In a second aspect, the present application provides a model deployment method applied to a vehicle end, which comprises:
[0045] sending target model information and resource occupation information to an edge access end;
[0046] receiving a policy file sent by the edge access end and a second target sub-model for the vehicle end, and deploying the second target sub-model according to the policy file; the second target sub-model is obtained by the edge access end based on resource occupation information by splitting a target model corresponding to the target model information, and the policy file is used to represent the cooperative execution information between the second target sub-model and a first target sub-model for the edge access end.
[0047] Thus, by splitting the target model corresponding to the target model information by the edge access end, the second target sub-model for the vehicle end and the policy file used to represent the cooperative execution information between the first target sub-model for the edge access end and the second target sub-model are obtained, so that the deployment of the second target sub-model is completed according to the policy file, the distributed deployment of large models is realized, the resources of each device are fully utilized, the resource utilization rate is improved, and the hardware resource requirements for a single device are reduced.
[0048] In an optional implementation, the vehicle end comprises a plurality of vehicles, each vehicle corresponding to one vehicle machine and at least one in-vehicle device; deploying the second target sub-model according to the policy file comprises:
[0049] parsing the cooperative execution information in the policy file to obtain a target vehicle machine or a target in-vehicle device corresponding to the second target sub-model;
[0050] deploying the second target sub-model and the policy file on the target vehicle machine or the target in-vehicle device.
[0051] Thus, by deploying the second target sub-model on the corresponding vehicle machine or in-vehicle device according to the policy file, the cooperative deployment of the edge access end, the vehicle machine and the in-vehicle device is realized.
[0052] In an optional implementation, the method further comprises:
[0053] The target vehicle machine or the target in-vehicle device parses the policy file to obtain data flow conversion rules and node connection rules applied by the second target sub-model when cooperatively executed;
[0054] According to the data flow conversion rules, the node connection rules, and the second target sub-model, the first target sub-model deployed at the edge access end is cooperatively deployed for model training and / or model inference.
[0055] Thus, by using the target model at the edge access end based on the data flow conversion rules and the node connection rules in the policy file, model training and / or model inference are performed, which greatly utilizes the resource advantages of different devices, improves resource utilization, and accelerates the efficiency of model training and / or model inference.
[0056] In an optional implementation, before sending the target model information and the resource occupation information to the edge access end, the method further comprises:
[0057] Sending a model deployment request to the edge access end;
[0058] After detecting that a connection is established with the edge access end, the step of sending the target model information and the resource occupation information to the edge access end is performed.
[0059] Thus, before sending information to the edge access end, a model deployment request is sent to the edge access end, and whether a connection is established is detected, thereby ensuring safe and effective operation of the system.
[0060] In an optional implementation, the edge access end comprises a plurality of edge access points, and the method further comprises:
[0061] Monitoring a packet loss rate of a currently connected current edge access point;
[0062] When it is detected that the packet loss rate is greater than a first packet loss rate threshold value within a preset duration, a candidate edge access point adjacent to the current edge access point is found;
[0063] When it is detected that the packet loss rate is greater than a second packet loss rate threshold value within the preset duration, a node switching request is sent to the current edge access point to disconnect the connection with the current edge access point and establish a connection with the candidate edge access point;
[0064] The second packet loss rate threshold value is greater than the first packet loss rate threshold value.
[0065] Thus, by continuously monitoring the packet loss rate of the current edge access point currently connected by the vehicle end, when the packet loss rate cannot meet the communication demand, a node switching request is sent to the edge access end, the switching of the edge access point is realized, the mobility requirement of the vehicle in the traffic scene is better adapted, and the reliability of the system is ensured.
[0066] In a third aspect, the present application provides a model deployment system, which comprises an edge access end and a vehicle end.
[0067] The edge access end obtains target model information and resource occupation information sent by the vehicle end; based on the target model information, a target model requested to be deployed by the vehicle end is found and obtained; based on the resource occupation information, the target model is split to obtain a first target sub-model deployed by the edge access end and a second target sub-model deployed by the vehicle end, and a strategy file used to represent cooperative execution information between the first target sub-model and the second target sub-model is generated; the first target sub-model is deployed, and the strategy file and the second target sub-model are sent to the vehicle end.
[0068] The vehicle end sends target model information and resource occupation information to the edge access end; receives the strategy file and the second target sub-model deployed by the vehicle end sent by the edge access end, and deploys the second target sub-model according to the strategy file.
[0069] By splitting the target model corresponding to the target model information based on the resource occupation information by the edge access end, a first target sub-model deployed by the edge access end, a second target sub-model deployed by the vehicle end, and a strategy file used to represent cooperative execution information between the first target sub-model and the second target sub-model are obtained, the edge access end deploys the first target sub-model on itself, and sends the strategy file and the second target sub-model to the vehicle end, so that the vehicle end completes the deployment of the second target sub-model, thereby realizing the distributed deployment of the large model, fully utilizing the resources of each device, improving the resource utilization rate, and reducing the hardware resource requirement of a single device.
[0070] In an optional implementation, the edge access end comprises a plurality of edge access points, and the vehicle end comprises a plurality of vehicles.
[0071] Each edge access point corresponds to at least one vehicle, each vehicle corresponds to one vehicle machine and at least one in-vehicle device, and the vehicle machine and the in-vehicle device are loaded in the vehicle.
[0072] In an optional implementation, the system further comprises a cloud end.
[0073] The cloud end receives a model acquisition request sent by the edge access end; based on the model acquisition request, a response feedback is sent to the edge access end, so that the edge access end downloads the target model from the cloud end.
[0074] Thus, by allowing the edge access end to download the target model, the diversified needs of the vehicle end are met.
[0075] In a fourth aspect, the present application provides an edge access device, comprising: a first memory and a first processor, which are in communication connection with each other, and the first memory stores computer instructions; the first processor executes the computer instructions to perform the model deployment method of the first aspect or any of the corresponding embodiments thereof.
[0076] In a fifth aspect, the present application provides a vehicle, comprising: an in-vehicle device and a car machine; the car machine or the in-vehicle device comprises a second memory and a second processor, which are in communication connection with each other, and the second memory stores computer instructions; the second processor executes the computer instructions to perform the model deployment method of the second aspect or any of the corresponding embodiments thereof.
[0077] In a sixth aspect, the present application provides a computer readable storage medium, which stores computer instructions for making a computer execute the model deployment method of the first aspect or any of the corresponding embodiments thereof, or the model deployment method of the second aspect or any of the corresponding embodiments thereof.
[0078] The present application has the following beneficial effects:
[0079] By using the edge access end to receive the target model information and the resource occupation information sent by the vehicle end, and based on the resource occupation information, the target model corresponding to the target model information is split to obtain a first target sub-model with the edge access end as the deployment end, a second target sub-model with the vehicle end as the deployment end, and a strategy file for representing the cooperative execution information between the first target sub-model and the second target sub-model, the edge access end deploys the first target sub-model in itself, and sends the strategy file and the second target sub-model to the vehicle end, so that the vehicle end completes the deployment of the second target sub-model, thereby realizing the distributed dynamic deployment of the large model, fully utilizing the resources of each device, improving the resource utilization, and reducing the hardware resource requirement of a single device. BRIEF DESCRIPTION OF DRAWINGS
[0080] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0081] Figure 1is a structural schematic diagram of a model deployment system according to an embodiment of the present application;
[0082] Figure 2 is a scenario schematic diagram of a model deployment system according to an embodiment of the present application;
[0083] Figure 3A is an architectural schematic diagram of a model deployment system according to an embodiment of the present application;
[0084] Figure 3B is a module schematic diagram of a model deployment system according to an embodiment of the present application;
[0085] Figure 4 is an interaction process schematic diagram of a model deployment system according to an embodiment of the present application;
[0086] Figure 5 is an interaction process schematic diagram of another model deployment system according to an embodiment of the present application;
[0087] Figure 6 is an interaction process schematic diagram of still another model deployment system according to an embodiment of the present application;
[0088] Figure 7 is a flow schematic diagram of a model deployment method according to an embodiment of the present application;
[0089] Figure 8 is a schematic diagram of a target model splitting result according to an embodiment of the present application;
[0090] Figure 9 is a flow schematic diagram of a model online training or inference according to an embodiment of the present application;
[0091] Figure 10 is a flow schematic diagram of a device online or offline processing method according to an embodiment of the present application;
[0092] Figure 11 is a flow schematic diagram of an edge access point switching method according to an embodiment of the present application;
[0093] Figure 12 is a hardware structure schematic diagram of an edge access device according to an embodiment of the present application;
[0094] Figure 13 is a structural schematic diagram of a vehicle according to an embodiment of the present application. DETAILED DESCRIPTION
[0095] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0096] In recent years, the intelligent level of automobiles is continuously improved, and intelligent cockpits and automatic driving have become the main selling points of new automobiles. Artificial intelligence large models have strong feature extraction capability and generalization capability, can explore deeper semantic information of things, and have aroused widespread interest in the academic and industrial circles, and also provide a solution for intelligent cockpits and automatic driving. However, the parameter quantity of hundreds of millions of large models and the complex network structure increase the requirements for storage and computing resources of intelligent and networked automobiles, and the competition in the automobile industry also prompts automobile enterprises to demand lower automobile manufacturing costs, and the contradiction between the two poses new challenges to the large-scale commercial use of vehicle-mounted large models.
[0097] Edge computing is a technology that sinks computing tasks in the cloud to edge servers closer to the end side, provides on-site end services with integrated network, computing, storage and application core capabilities, and has characteristics of low latency, scalability, distribution and privacy protection. Although the service capability of the edge server is not as powerful as that of the cloud, the low-latency characteristic still provides a solution for the deployment of large models on intelligent and networked automobiles. In addition, the vehicle machine and the in-vehicle mobile terminal device (mobile phone, tablet, notebook computer, etc.) have realized interconnection and intercommunication. At present, there are mainly two methods for deploying and executing large models based on edge computing: one is to deploy and execute models on edge access points through edge computing and Internet of Things technology to reduce data transmission delay and bandwidth consumption, but this method deploys or executes large models on a single node, and feeds back the model execution result through data return, which has high requirements for the hardware of the single node. The other is to deploy each network layer of the large model on multiple computing nodes, and obtain the data flow relationship between the nodes through the configuration file, so as to realize the rapid deployment and application of the large model on the device side, but this method only considers the static distributed deployment of the large model, and is not suitable for the complex dynamic environment that changes all the time in the intelligent transportation scenario.
[0098] In addition, these methods do not process the parameters of the large model itself and the intermediate layer, resulting in great bandwidth consumption of data transmission between layers, reducing the efficiency of model reasoning, and increasing the system cost.
[0099] Therefore, the model deployment method provided by the embodiment of the present application can split the target model requested by the vehicle end to deploy into multiple sub-models by using the edge access end and generate a strategy file, so that the edge access end and the vehicle end perform distributed deployment of the sub-models based on the strategy file, the resources of each device are fully utilized, the resource utilization rate is improved, and the hardware resource requirement of a single device is reduced.
[0100] According to the embodiment of the present application, a model deployment system is provided, as shown in the figure, which comprises an edge access end 101 and a vehicle end 102. Figure 1
[0101] The edge access end 101 obtains the target model information and resource occupation information sent by the vehicle end 102, finds and obtains the target model requested by the vehicle end 102 to deploy based on the target model information, splits the target model based on the resource occupation information to obtain a first target sub-model with the edge access end as the deployment end and a second target sub-model with the vehicle end as the deployment end, and generates a strategy file for representing the cooperative execution information between the first target sub-model and the second target sub-model; the first target sub-model is deployed, and the strategy file and the second target sub-model are sent to the vehicle end 102.
[0102] The vehicle end 102 sends the target model information and resource occupation information to the edge access end 101, receives the strategy file and the second target sub-model with the vehicle end as the deployment end sent by the edge access end 101, and deploys the second target sub-model according to the strategy file.
[0103] It should be noted that the target model in the embodiment of the present application mainly includes a large model with large-scale parameters and complex computing structure.
[0104] The model deployment system provided by the embodiment of the present application can split the target model corresponding to the target model information based on the resource occupation information by using the edge access end 101, obtain the first target sub-model with the edge access end as the deployment end, the second target sub-model with the vehicle end as the deployment end, and the strategy file for representing the cooperative execution information between the first target sub-model and the second target sub-model, deploy the first target sub-model in the edge access end 101, and send the strategy file and the second target sub-model to the vehicle end 102, so that the vehicle end 102 completes the deployment of the second target sub-model, thereby realizing distributed deployment of the large model, fully utilizing the resources of each device, improving the resource utilization rate, and reducing the hardware resource requirement of a single device.
[0105] For specific working principles and working processes of the edge access end 101 and the vehicle end 102, refer to the related descriptions of the method embodiments below, which will not be described here.
[0106] In some optional embodiments, such as Figure 2 As shown, the above-mentioned model deployment system may further include a cloud 103, an edge access terminal 101 may include multiple edge access points, and a vehicle terminal 102 may include multiple vehicles, wherein each edge access point corresponds to at least one vehicle, and each vehicle corresponds to a vehicle computer and at least one in-vehicle device, and the vehicle computer and the in-vehicle device are installed inside the vehicle. For example, the vehicle is a smart car with Internet access; the vehicle computer may be the center console inside the vehicle. Since there is a one-to-one relationship between the vehicle and the vehicle computer, the model deployment method applicable to the vehicle terminal in the embodiment of the present invention is also applicable to the vehicle computer; the in-vehicle device may be a mobile phone, tablet computer, laptop computer, or other in-vehicle mobile terminal device installed inside the vehicle. Such devices may be carried by the owner or passengers of the vehicle. In the embodiment of the present invention, unless otherwise specified, such devices are collectively referred to as in-vehicle devices.
[0107] Figure 3A The schematic diagram of the architecture of the model deployment system provided by the embodiment of the present invention is as follows: Figure 3A As shown in Figure 1, the architecture can be broken down into the cloud service layer, edge access layer, terminal application layer, and physical device layer. In the cloud service layer, the cloud platform plays a crucial role, primarily providing cloud services. The edge access layer, primarily comprised of edge access points, is the core of the entire system. Its responsibilities include managing the vehicle and its in-vehicle devices, playing a key role in large-scale model deployment. The terminal application layer is responsible for initiating large-scale model applications and sensing and acquiring external information to form input data for the target model. The physical device layer comprises the system's physical infrastructure, primarily including basic hardware components that provide computing, storage, and network resources.
[0108] See again Figure 2 and Figure 3A , the cloud 103 can be a cloud platform with cloud service functions. The cloud platform (i.e., cloud service layer) is equipped with numerous servers, has excellent computing and storage performance, and can provide huge Internet data resources, thereby providing comprehensive storage, computing, network and model services for various facilities in the system. When necessary, the perception data, model intermediate data, model parameters, and authentication data from edge access points, vehicle computers, and in-vehicle devices can be stored in the cloud platform through storage services, thereby reducing the storage pressure of edge access points, vehicle computers, and in-vehicle devices. For non-real-time computing tasks, such as image processing, visual analysis, semantic recognition, and video content segmentation, edge access points, vehicle computers, and in-vehicle devices can use the computing services of the cloud platform for processing to reduce their own computing load.
[0109] The network service provided by the cloud platform can retrieve and obtain data and resources on the Internet for edge access points, vehicle infotainment systems, and in-vehicle devices, such as video streaming, real-time hotspot information, weather updates, and traffic conditions. In addition, the cloud platform provides a rich storage of large models and supports the local deployment or remote invocation of models by edge access points, vehicle infotainment systems, and in-vehicle devices through various invocation interfaces and tool chains. After the model is executed, the cloud platform returns the inference or training results for further processing and analysis.
[0110] Referring again to Figure 2 , the edge access point 101 includes a plurality of edge access points, which can be deployed at the roadside and mainly include edge servers, edge gateways, edge controllers, switches, and other physical facilities, and are key components of the system. The edge access point is responsible for providing edge computing services for the vehicle end, coordinating its own and the intelligent vehicle and in-vehicle intelligent mobile terminal device resources, and deploying large models in a distributed manner for vehicle adaptation. Referring again to Figure 3A , the edge access point includes, but is not limited to, the following functional modules: a device authentication module, a device management module, a resource monitoring module, a resource scheduling module, a model pruning module, a model quantization module, a model splitting module, a parameter quantization module, and an automobile identity ID synchronization module, and the like.
[0111] Specifically, referring again to Figure 3A , the device authentication module completes the security authentication of the vehicle and its in-vehicle devices requesting to access the edge access point, ensures the legality of the access device, avoids malicious access, and ensures the safe and effective operation of the system. The device management module is used for the edge access point to manage the accessed devices, and the online and offline states of the devices can be determined by this module. The resource monitoring module is responsible for real-time monitoring of the computing, storage, and network resources of the edge access point, providing prior information for the resource scheduling module to dynamically schedule and allocate various resources for each automobile large model working group, and improving the system efficiency. The model pruning module and the model quantization module are mainly responsible for reducing the storage and computing cost of the model to save the storage and computing resources of the edge access point and the intelligent vehicle and in-vehicle devices. The model splitting module splits a large model into several sub-models, which can be inter-layer splitting, intra-layer splitting, and mixed splitting, and can be selected as needed according to the actual situation of the model. The parameter quantization module mainly serves the model execution process, which quantizes the intermediate data in the model training or inference, reduces the data transmission amount, saves the communication bandwidth, improves the interaction efficiency between sub-models, and improves the overall performance of the system. The automobile identity ID synchronization module is used to synchronize the access vehicle end information between edge access points, and when the edge access point connected by the vehicle end is switched, the computing task can be seamlessly continued, and the dynamic changing traffic scene is adapted.
[0112] It should be noted that one edge access point in the above system can cover multiple vehicles, and each vehicle includes at least one in-vehicle device in addition to its corresponding vehicle machine, wherein the vehicle machine and the in-vehicle device are connected to each other through Bluetooth, star flash, WiFi or other communication modes before model deployment.
[0113] Specifically, the in-vehicle device can be a smart mobile terminal device, referring again to Figure 3A These terminal devices include but are not limited to the following functional modules: resource monitoring module, resource virtualization module, information perception module and parameter quantization module, etc. The resource monitoring module and the parameter quantization module have the same function as the corresponding modules in the edge access point, and will not be described again here. The resource virtualization module is mainly used by the vehicle machine and the in-vehicle device. In addition to performing large model calculation tasks, various other applications are also running synchronously. In order to avoid the competition for resources between the large model application and other applications during running, causing the model running efficiency to be reduced, when the vehicle machine initiates a large model application request, the vehicle machine and the in-vehicle device abstract the computing, storage and network resources according to their own resource occupation, forming a virtualization resource pool to ensure the reliable operation of the model. The information perception module completes the perception and collection of vehicle information, road information, environmental information, pedestrian information, in-vehicle atmosphere information, in-vehicle voice information and other information, providing data input for the large model, involving hardware including but not limited to vehicle-mounted radar, camera, sensor, etc.
[0114] The embodiment of the present application provides a module schematic diagram of a model deployment system, as Figure 3B shown, the edge access point includes a data management unit, a device management unit, a resource management unit, a model preprocessing unit, a model deployment unit and a communication unit. The data management unit is mainly responsible for managing the device identity and IP data connected to the edge access point, and the temporary data of model calculation, etc. The device management unit includes a device management and device authentication module. The resource management unit includes a resource monitoring and resource scheduling module, and transmits existing device resource occupation information to the model preprocessing unit. The model preprocessing unit includes a strategy generation, model pruning, model quantization and model splitting module, wherein the strategy generation module generates model pruning, quantization, splitting, data flow rules, etc. using the resource occupation information. The model deployment unit includes a model deployment, model calculation and parameter quantization module, responsible for the deployment and execution of the large model. The communication unit communicates with the vehicle machine and the in-vehicle device in a wireless connection manner. The vehicle machine and the in-vehicle device include a data management unit, a resource management unit, a model deployment unit, an information perception unit and a communication unit, wherein the resource management unit includes a resource monitoring and resource virtualization module, and the other units are the same as the edge access point and will not be described again.
[0115] The model deployment system provided by the embodiment of the present invention takes the edge access end as the core, decomposes the target model into several sub-models, and distributes the sub-models to the edge access point, vehicle computer and in-vehicle equipment, making full use of the computing power, storage and network resources of each device, improving resource utilization, reducing the hardware resource requirements for a single device, and reducing production costs.
[0116] According to an embodiment of the present invention, an embodiment of a model deployment method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0117] In this embodiment, a model deployment method is provided, which can be used to Figure 1 The edge access terminal 101 and the vehicle terminal 102 are shown. Figure 4 1 is a schematic diagram of the interaction process of the model deployment system according to an embodiment of the present invention, wherein the edge access terminal 101 is used to execute steps S101 to S104, and the vehicle terminal 102 is used to execute steps S201 to S202. The specific interaction process between the edge access terminal 101 and the vehicle terminal 102 is as follows:
[0118] Step S201: Send target model information and resource occupancy information to an edge access terminal.
[0119] Specifically, the target model information mainly includes the identification information of the target model that the vehicle side wants to deploy, and the resource occupancy information mainly includes the storage, computing and network resource occupancy information of the vehicle side and the in-vehicle equipment. This information can be obtained through Figure 3A The resource monitoring module shown is used to obtain it.
[0120] In some optional implementations, before executing step S201, the vehicle first sends a model deployment request to the edge access terminal. After detecting that a connection has been established with the edge access terminal, step S201 is executed again. Sending a model deployment request to the edge access terminal and detecting whether a connection has been established before sending information to the edge access terminal ensures the safe and efficient operation of the system.
[0121] Step S101: Acquire target model information and resource occupancy information sent by the vehicle.
[0122] In some optional implementations, after receiving the model deployment request sent by the vehicle end, the edge access end can Figure 3AThe device authentication module shown here verifies the legitimacy of the vehicle. Once the verification is successful, it establishes a connection with the vehicle's head unit and in-vehicle devices. By securely authenticating vehicles and their in-vehicle devices requesting access to the edge access point, the legitimacy of the accessed devices is ensured, malicious access is prevented, and the system's safe and efficient operation is guaranteed.
[0123] In particular, the present invention refers to the vehicle computer and in-vehicle devices in the same vehicle and the current edge access point to which the vehicle is currently connected as a large vehicle model execution workgroup, and the execution and deployment of the model are both performed in the devices of the same workgroup.
[0124] Step S102: Based on the target model information, search and obtain the target model requested to be deployed by the vehicle side.
[0125] Specifically, after receiving the target model information, the edge access terminal first checks whether the local model library stores the target model corresponding to the identification information in the target model information. If the target model is not stored locally, it sends a model acquisition request to the cloud. After receiving a response feedback from the cloud based on the model acquisition request, it can download the target model from the cloud through the base station. If the target model is stored locally, the target model is obtained locally. The cloud receives the model acquisition request sent by the edge access terminal and sends a response feedback to the edge access terminal based on the model acquisition request, so that the edge access terminal can download the target model from the cloud.
[0126] Therefore, when the target model is stored locally, the target model requested to be deployed by the vehicle side can be obtained quickly and conveniently through local search. When the target model is not stored locally, the target model can be obtained by sending a request to the cloud, meeting the diverse needs of the vehicle side.
[0127] In step S103, the target model is split based on the resource occupancy information to obtain a first target sub-model whose deployment end is the edge access end and a second target sub-model whose deployment end is the vehicle end, and a policy file is generated for characterizing the collaborative execution information between the first target sub-model and the second target sub-model.
[0128] In some optional implementations, the above step S103 includes:
[0129] Step a1: perform model pruning and model quantization on the target model to obtain an optimized target model.
[0130] Specifically, it can be achieved through Figure 3A The model pruning module shown prunes the target model, such as the sparse vector method, etc. Model pruning significantly reduces the size of the target model itself by eliminating redundant parameters or channels in the target model that are not important for training or inference results. Figure 3AThe model quantification module quantifies the pruned target model. Model quantification reduces the storage resource occupation space of the target model by quantizing the model parameters of the floating-point type.
[0131] In step a2, the optimized target model is split into multiple sub-models based on the storage resource occupation information, the computing resource occupation information, and the network resource occupation information.
[0132] Specifically, the storage space occupied by the model parameters, weights, and structures is analyzed according to the storage resource occupation information, to determine which parts of the target model occupy larger storage resources. The computing complexity of the target model in the operation process, including convolution operation, matrix multiplication, etc., is evaluated according to the computing resource occupation information. The network resource occupation of the model in the data transmission and synchronization process is analyzed according to the network resource occupation information, which is crucial for determining which parts need to be split into independent sub-models to reduce the network transmission burden.
[0133] Further, after analyzing the storage, computing, and network resource occupation information, the target model can be split into multiple sub-models by Figure 3A The model splitting module splits the optimized target model into multiple sub-models. The specific splitting methods mainly include inter-layer splitting, intra-layer splitting, and mixed splitting, which can be selected as needed according to the actual situation of the target model. It should be noted that inter-layer splitting is to distribute different layers of the target model to different computing devices (such as edge access terminals, vehicle terminals, or in-vehicle devices). Intra-layer splitting is to split the internal operations of a layer of the target model to multiple devices for execution, which is suitable for a layer of the target model that is particularly large or has particularly high computing complexity. Mixed splitting is a combination of inter-layer splitting and intra-layer splitting, that is, both inter-layer and intra-layer splitting strategies are used to split the target model.
[0134] In step a3, the deployment ends corresponding to the multiple sub-models are determined. The sub-models deployed on the edge access terminal are taken as the first target sub-model, and the sub-models deployed on the vehicle terminal are taken as the second target sub-model.
[0135] Thus, by performing model pruning and model quantification on the target model, the size of the target model itself is greatly reduced, and the storage occupation space is reduced. According to the resource occupation information of the vehicle terminal, the target model after model pruning and model quantification is split into multiple sub-models, and the deployment end of each sub-model is determined, so as to facilitate the deployment of each sub-model to a suitable device.
[0136] Specifically, the policy file mainly includes model splitting rules (connection relationship between layers or neurons when the target model is split into sub-models, network structure connection rules, etc.), data flow rules (data transmission rules between different device layers in the training or inference process, data synchronization rules, etc.), information sources, parameter data types, parameter matrix dimensions, required storage and computing resources, and all necessary information required for the deployment of each sub-model when the model is cooperatively executed. It should be noted that the number of the first target sub-model and the second target sub-model is multiple, and the policy file further includes cooperative execution information between the first target sub-model and the second target sub-model.
[0137] In particular, in the process of model splitting and policy file generation, the edge access point gives priority to the vehicle machine and the in-vehicle device as the main nodes for model deployment, and the storage or computing resources required by the split sub-model do not exceed 80% of the idle resources of the current vehicle machine and in-vehicle device, thereby ensuring the reliability of model deployment and execution.
[0138] In step S104, the first target sub-model is deployed, and the policy file and the second target sub-model are sent to the vehicle end.
[0139] Specifically, the edge access end deploys the first target sub-model according to the policy file.
[0140] In step S202, the policy file and the second target sub-model deployed by the vehicle end are received, and the second target sub-model is deployed according to the policy file.
[0141] Specifically, the vehicle end parses the cooperative execution information in the policy file to obtain the target vehicle machine or the target in-vehicle device corresponding to the second target sub-model, and then deploys the second target sub-model and the policy file on the target vehicle machine or the target in-vehicle device. By deploying the second target sub-model on the corresponding vehicle machine or in-vehicle device according to the policy file, the edge access end, the vehicle machine, and the in-vehicle device are cooperatively deployed.
[0142] In some optional embodiments, after the edge access end and the vehicle end complete the cooperative deployment of the target model, the edge access end cooperatively deploys the second target sub-model deployed on the vehicle end based on the policy file and the first target sub-model to perform model training; and / or, cooperatively deploys the second target sub-model deployed on the vehicle end based on the policy file and the first target sub-model to perform model inference, and transmits the obtained model inference result to the vehicle end.
[0143] Thus, the resources of multiple devices are utilized for model training and / or model inference of the target model, thereby improving the resource utilization rate and the efficiency of model execution.
[0144] In some optional embodiments, after the collaborative deployment of the target model is completed at the edge access end and the vehicle end, the vehicle end first controls the target vehicle machine or the target in-vehicle device to parse the policy file to obtain data flow transfer rules and node connection rules applied by the second target sub-model in collaborative execution, and then performs model training and / or model inference on the first target sub-model deployed at the edge access end according to the data flow transfer rules, the node connection rules and the second target sub-model. By using the data flow transfer rules and the node connection rules in the policy file, the edge access end collaboratively uses the target model to perform model training and / or model inference, which greatly utilizes the resource advantages of different devices, improves the resource utilization, and accelerates the efficiency of model training and / or model inference.
[0145] It should be noted that, in the process of model training and / or model inference, the intermediate parameters transmitted by each node in the model execution process can be quantized by the parameter quantization module shown in Figure 3A to reduce the data amount of the intermediate parameters in the model, save bandwidth resources, and improve the data flow transfer rate between sub-models, thereby accelerating the efficiency of model training and / or model inference.
[0146] The model deployment method provided in this embodiment receives the target model information and the resource occupation information sent by the vehicle end at the edge access end, splits the target model corresponding to the target model information based on the resource occupation information, obtains the first target sub-model deployed at the edge access end, the second target sub-model deployed at the vehicle end, and the policy file used to represent the collaborative execution information between the first target sub-model and the second target sub-model, deploys the first target sub-model at the edge access end, and sends the policy file and the second target sub-model to the vehicle end to enable the vehicle end to complete the deployment of the second target sub-model, thereby realizing the distributed dynamic deployment of the large model, fully utilizing the resources of each device, improving the resource utilization, and reducing the hardware resource requirements of a single device.
[0147] In this embodiment, a model deployment method is provided, which can be used in an edge access end 101 and a vehicle end 102 as shown in Figure 1 , and the specific interaction process between the edge access end 101 and the vehicle end 102 is as follows: Figure 5 is a schematic diagram of the interaction process of the model deployment system according to an embodiment of the present application, wherein the edge access end 101 is used to perform steps S301 to S309, the vehicle end 102 is used to perform steps S401 to S405, and the specific interaction process between the edge access end 101 and the vehicle end 102 is as follows:
[0148] Step S401, sending target model information and resource occupation information to the edge access end. For details, refer to the specific description of step S201 as shown in Figure 4 , which will not be described here.
[0149] Step S301, obtaining the target model information and resource occupation information sent by the vehicle end. For details, please refer to the detailed description of step S101 as shown in the Figure 4 specific description of step S101, which will not be repeated here.
[0150] Step S302, based on the target model information, finding and obtaining the target model requested by the vehicle end for deployment. For details, please refer to the detailed description of step S102 as shown in the Figure 4 specific description of step S102, which will not be repeated here.
[0151] Step S303, based on the resource occupation information, splitting the target model to obtain a first target sub-model for the edge access end and a second target sub-model for the vehicle end, and generating a policy file for representing the cooperative execution information between the first target sub-model and the second target sub-model. For details, please refer to the detailed description of step S103 as shown in the Figure 4 specific description of step S103, which will not be repeated here.
[0152] Step S304, deploying the first target sub-model, and sending the policy file and the second target sub-model to the vehicle end. For details, please refer to the detailed description of step S104 as shown in the Figure 4 specific description of step S104, which will not be repeated here.
[0153] Step S402, receiving the policy file and the second target sub-model for the vehicle end sent by the edge access end, and deploying the second target sub-model according to the policy file. For details, please refer to the detailed description of step S202 as shown in the Figure 4 specific description of step S202, which will not be repeated here.
[0154] Step S403, monitoring the packet loss rate of the current edge access point currently connected.
[0155] For example, the infotainment system and in-vehicle devices of the vehicle end periodically monitor the packet loss rate of the current edge access point, which is the edge access point currently connected with the vehicle end.
[0156] Step S404, when it is detected that the packet loss rate is greater than the first packet loss rate threshold within a preset duration, finding an alternative edge access point adjacent to the current edge access point.
[0157] For example, if the packet loss rate of the current edge access point is greater than the first packet loss rate threshold A within a preset duration T, the infotainment system and in-vehicle devices start to listen to the surrounding signal quality and search for adjacent edge access points.
[0158] Step S405, when it is detected that the packet loss rate is greater than the second packet loss rate threshold within a preset duration, sending a node switching request to the current edge access point.
[0159] Exemplarily, when the packet loss rate is greater than the second packet loss rate threshold B within the preset duration T, if a neighboring candidate edge access point is detected, the vehicle end where the car machine and the in-vehicle device are located sends a node switching request to the edge access end to disconnect the connection with the current edge access point and establish a connection with the candidate edge access point. It should be noted that the second packet loss rate threshold is greater than the first packet loss rate threshold.
[0160] Thus, by continuously monitoring the packet loss rate of the current edge access point to which the vehicle end is currently connected, when the packet loss rate cannot meet the communication requirements, a node switching request is sent to the edge access end to switch the edge access point, so as to better adapt to the mobility requirements of the vehicle in the traffic scene and ensure the reliability of the system.
[0161] Step S305, receiving the node switching request sent by the vehicle end, and obtaining the candidate edge access point to be connected next by the vehicle end based on the node switching request.
[0162] Specifically, the node switching request contains the candidate edge access point to be connected next by the vehicle end.
[0163] Step S306, obtaining intermediate data generated when the current edge access point cooperates with the vehicle end to perform model training and / or model inference.
[0164] Specifically, the intermediate data mainly includes intermediate results stored in the current edge access point and model calculation tasks when the model training and / or model inference is cooperatively performed.
[0165] Step S307, controlling the current edge access point to disconnect with the vehicle end, and controlling the candidate edge access point to establish a connection with the vehicle end.
[0166] Specifically, when controlling the candidate edge access point to establish a connection with the vehicle end, the legality of the vehicle also needs to be verified, and the automobile ID identity is synchronized to ensure the security of the device.
[0167] Step S308, sending the policy file, the first target sub-model, and the intermediate data to the candidate edge access point.
[0168] Thus, the candidate edge access point synchronously stores the intermediate data, and deploys the first target sub-model according to the policy file, so as to switch the edge access point, so that the vehicle end continues to perform the calculation task of model training and / or model inference through the candidate edge access point, and ensures the reliability and effectiveness of model execution.
[0169] Step S309, sending the target model, the policy file, and the intermediate data to the cloud end.
[0170] Specifically, the cloud has excellent computing and storage performance and can provide huge Internet data resources, thereby providing comprehensive storage, computing, network and model services for various facilities in the system. In necessary cases, the edge access end can send the target model, the policy file and the intermediate data to the cloud, so that the cloud deploys the target model according to the policy file, and performs model training and / or model inference based on the intermediate data and the target model, thereby ensuring that the vehicle end can perform model training and / or model inference through the cloud, meeting the use requirements of the vehicle end and ensuring the reliability of the system.
[0171] The model deployment method provided in this embodiment splits the target model through the edge access end to obtain a first target sub-model with the deployment end being the edge access end, a second target sub-model with the deployment end being the vehicle end, and a policy file for representing cooperative execution information between the first target sub-model and the second target sub-model. The edge access end deploys the first target sub-model in itself and sends the policy file and the second target sub-model to the vehicle end, so that the vehicle end completes the deployment of the second target sub-model, thereby realizing distributed dynamic deployment of a large model, fully utilizing the resources of each device, improving the resource utilization rate, and reducing the hardware resource requirements for a single device.
[0172] Moreover, considering the moving requirements of the vehicle in the traffic scene, when the vehicle faces edge access point switching, the model calculation task being executed and the intermediate data are synchronized to the adjacent candidate edge access point or uploaded to the cloud for execution, thereby ensuring the continuous execution of the calculation task and guaranteeing the reliability and effectiveness of the model execution.
[0173] In this embodiment, a model deployment method is provided, which can be used in an edge access end 101 and a vehicle end 102 as shown in Figure 1 The interaction process diagram of the model deployment system according to the embodiment of the present application is shown in Figure 6 The specific interaction process between the edge access end 101 and the vehicle end 102 is as follows:
[0174] In step S601, the target model information and the resource occupation information are sent to the edge access end. For details, refer to the specific description of step S401 as shown in Figure 5 The specific description of step S401 is not repeated here.
[0175] In step S501, the target model information and the resource occupation information sent by the vehicle end are obtained. For details, refer to the specific description of step S301 as shown in Figure 5 The specific description of step S301 is not repeated here.
[0176] Step S502, based on the target model information, find and obtain the target model requested by the vehicle end to deploy. For details, please refer to the specific description of step S302 as shown in Figure 5 The detailed description of step S302 is not repeated here.
[0177] Step S503, based on the resource occupation information, split the target model to obtain the first target sub-model deployed in the edge access end and the second target sub-model deployed in the vehicle end, and generate a policy file for representing the cooperative execution information between the first target sub-model and the second target sub-model. For details, please refer to the specific description of step S303 as shown in Figure 5 The detailed description of step S303 is not repeated here.
[0178] Step S504, deploy the first target sub-model, and send the policy file and the second target sub-model to the vehicle end. For details, please refer to the specific description of step S304 as shown in Figure 5 The detailed description of step S304 is not repeated here.
[0179] Step S602, receive the policy file and the second target sub-model deployed in the vehicle end sent by the edge access end, and deploy the second target sub-model according to the policy file. For details, please refer to the specific description of step S402 as shown in Figure 5 The detailed description of step S402 is not repeated here.
[0180] Step S505, monitor the device state of each vehicle machine or in-vehicle device in the vehicle end, and determine whether the vehicle end has completed the deployment of the target model.
[0181] Specifically, the vehicle end includes a plurality of vehicles, each vehicle corresponding to a vehicle machine and at least one in-vehicle device, and the edge access end can monitor the online or offline state of the vehicle machine and the in-vehicle device through its own device management module.
[0182] Step S506, if it is detected that there is a vehicle machine or in-vehicle device with an offline or online device state, and the vehicle end has not completed the deployment of the target model, return to step S501.
[0183] Specifically, when the device is offline or online and in the model deployment stage, the edge access end re-performs model splitting and policy file generation according to the existing device resource occupation in the working group, and completes the distributed deployment of the target model.
[0184] Step S507, if it is detected that there is a vehicle machine or in-vehicle device with an offline device state, and the vehicle end has completed the deployment of the target model, obtain the offline third target sub-model, and deploy the offline third target sub-model in the current edge access point.
[0185] The offline third target sub-model is a second target sub-model deployed on an in-vehicle device or a car machine in an offline state.
[0186] Specifically, when the device is offline and in the model execution stage, the edge access terminal continues the related computing task of the offline device and performs model synchronization and interaction, and continues to execute model training and / or model inference.
[0187] Step S508, if it is detected that there is a car machine or in-vehicle device in an online state, and the vehicle end has completed the deployment of the target model, after receiving the new model deployment request sent by the vehicle end corresponding to the online car machine or in-vehicle device, a connection is established with the online car machine or in-vehicle device, and step S501 is executed.
[0188] Specifically, when the new device is online and in the model execution stage, the new device does not participate in the current model execution process and waits for the initiation of the next model deployment request.
[0189] Thus, by monitoring the online or offline state of the car machine and the in-vehicle device, the deployment and execution process of the target model are dynamically adjusted to better adapt to the dynamic changes of the device.
[0190] The model deployment method of the present application will be described in detail below in conjunction with a specific application example, as shown in the following table: Figure 7 The specific application example includes the following steps:
[0191] Step S701, the vehicle sends a model deployment request to the edge access point.
[0192] Step S702, after receiving the model deployment request, the edge access point verifies the legality of the vehicle through the device authentication module. If the vehicle is legal, it jumps to step S703, otherwise the process is aborted.
[0193] Step S703, the car machine, the in-vehicle device and the edge access point establish network connection with each other. In particular, the car machine and the in-vehicle device in the same vehicle and the currently connected current edge access point are referred to as an automobile model execution workgroup, and the execution and deployment of the model are performed in the same workgroup device.
[0194] Step S704, the vehicle sends the target model information and the storage, computing and network resource occupation information of the car machine and the in-vehicle device to the edge access point.
[0195] Step S705, after receiving the target model information, the edge access point searches its own model library. If it is already pre-stored, it directly enters step S707, otherwise it enters step S706.
[0196] Step S706, the edge access point downloads the model through the base station in the cloud platform.
[0197] Step S707, the edge access point splits the model into multiple sub-models according to the resource occupation of itself and the vehicle machine and in-vehicle devices and generates a strategy file.
[0198] Step S708, the edge access point sends the strategy file to the vehicle machine and in-vehicle devices.
[0199] Step S709, the edge access point, the vehicle machine and the in-vehicle devices respectively deploy the sub-models to themselves according to the strategy file.
[0200] Step S710, the edge access point, the vehicle machine and the in-vehicle devices perform online training or inference of the model according to the strategy file.
[0201] Step S711, after the model training or inference is completed, the connection between the vehicle machine, the in-vehicle devices and the edge access point is disconnected.
[0202] Exemplarily, as shown in Figure 8 , the deployment end of the sub-model 1 obtained after splitting is the edge access point, and the deployment ends of the sub-models 2, 3, 4, …, N are the vehicle machine or the in-vehicle devices. As shown in Figure 9 , step S710 includes the following steps:
[0203] Step S7101: the edge access point parses the strategy file and deploys the sub-model 1.
[0204] Step S7102: the edge access point respectively transmits the strategy file and the sub-models 2, 3, 4, …, N to the vehicle machine and the in-vehicle devices.
[0205] Step S7103: the vehicle machine and the in-vehicle devices parse the strategy file to obtain necessary data flow and node connection rules, load and deploy the sub-models 2, 3, 4, …, N.
[0206] Step S7104: the vehicle machine and the in-vehicle devices complete the deployment of the sub-models and feed back to the edge access point, and start to perform model training and inference.
[0207] Step S7105: the roadside perception device, the vehicle-mounted perception device and the in-vehicle perception device respectively input the perception information (in-vehicle user behavior information, facial expression information, external radar signal, air quality, temperature and humidity information, etc.) into the corresponding sub-models 1, 2, 3, …, N.
[0208] Step S7106: the edge access point schedules the computing tasks according to the resource occupation to improve the parallel execution efficiency of the model.
[0209] Step S7107: The edge access point, the vehicle machine and the in-vehicle device perform nonlinear quantization on the intermediate parameters in the model execution process, reducing the transmission bandwidth. The nonlinear quantization method here is the same as the quantization method of the model parameter quantization module, and only the difference in the quantization object exists, so it will not be described again. It should be noted that the operation of nonlinear quantization needs to add a nonlinear quantization function as an intermediate layer to the network structure.
[0210] Step S7108: The edge access point, the vehicle machine and the in-vehicle device perform data flow conversion and interaction according to the policy file.
[0211] Step S7109: For the inference task, the edge access point or the in-vehicle device transmits the model inference result to the vehicle machine according to the policy file; for the training task, the edge access point, the vehicle machine and the in-vehicle device update the model online.
[0212] The model deployment method provided by the application further reduces the storage resource demand on a single device node through model pruning and model quantization operation, wherein the model pruning greatly reduces the size of the model itself by eliminating redundant parameters or channels in the model that are not important to the training or inference result, and the model quantization replaces the floating-point type model parameters with bit quantization, reducing the storage resource occupation space of the model. Moreover, in view of the problem of additional bandwidth consumption caused by the transmission of intermediate parameters in the model execution of the distributed deployment of the large model by multiple devices, the embodiments of the application perform nonlinear quantization on the intermediate parameters in the model execution, replacing the continuous and floating-point type intermediate parameters with multi-bit data, thereby reducing the data amount of the intermediate parameters in the model transmission, saving the bandwidth resources while improving the data flow rate between the sub-models, and accelerating the model training or inference efficiency.
[0213] Further, due to the uncertainty of the in-vehicle device, the specific application example of the application provides a processing strategy corresponding to the online and offline state of the in-vehicle device, as shown in Figure 10 The device management module of the edge access point monitors the state of the connected devices in real time. When the device is offline and in the model deployment stage, the edge access point re-performs model splitting and generation of the policy file according to the existing device resource occupation in the group and completes the distributed deployment of the model; when the device is offline and in the model execution stage, the edge access point continues the related computing tasks of the offline device and performs model synchronization and interaction, and the model continues to execute; when a new device is online and in the model deployment stage, the edge access point re-performs model splitting and generation of the policy file according to the existing device resource occupation in the group and completes the distributed deployment of the model; when a new device is online and in the model execution stage, the new device does not participate in the current model execution process and waits for the initiation of the next model deployment request.
[0214] Further, in order to adapt to the mobility of the vehicle in the traffic scene and ensure the reliability of the model deployment system, the specific application example of the present application further provides a coping strategy when the edge access point switches. Figure 11 As shown in the figure, the vehicle end refers to the vehicle machine and its in-vehicle equipment, and the scene layout refers to Figure 2 The coping strategy when the edge access point switches specifically includes the following steps:
[0215] Step S111, the vehicle machine and the in-vehicle equipment periodically monitor the packet loss rate of the edge access point 1, and when the packet loss rate of a certain connection is greater than a first packet loss rate threshold A within a preset duration T, step S112 is entered.
[0216] Step S112, the vehicle machine starts to listen to the surrounding signal quality to search for a nearby edge access point. When the packet loss rate is greater than a second packet loss rate threshold B within a preset duration T, if a nearby access point 2 is detected, step S113 is entered, otherwise, step S116 is jumped to.
[0217] Step S113, the vehicle machine and the in-vehicle equipment disconnect the connection with the edge access point 1 and establish a network connection with the edge access point 2, and at the same time, the vehicle ID identity is synchronized to ensure the safety of the equipment.
[0218] Step S114, the edge access point 1 sends a strategy file to the edge access point 2 and synchronizes the intermediate data of the current model training or inference.
[0219] Step S115, the edge access point 2 locally deploys a response sub-model according to the strategy file and continues to perform the task of the edge access point 1, and the switching work is completed.
[0220] Step S116, the edge access point 1 uploads the strategy file to the cloud platform through the base station and synchronizes the intermediate data of the current model training or inference.
[0221] Step S117, the vehicle machine and the in-vehicle equipment disconnect the connection with the edge access point 1 and unload the current model calculation task.
[0222] Step S118, the cloud deploys the entire large model according to the strategy file and continues to perform the current training or inference task, and the switching work is completed.
[0223] As can be seen, the online and offline coping strategies of the in-vehicle equipment and the processing method of the vehicle when facing the switching of the edge access point provided by the present application synchronize the model calculation task being executed to the adjacent edge access point or upload it to the cloud for execution, ensuring the continuous execution of the calculation task and guaranteeing the reliability and effectiveness of the model execution, adapting to the mobility of the vehicle in the traffic scene and the uncertainty of the in-vehicle equipment.
[0224] Please refer to Figure 12 ,Figure 12 Schematic diagram of the structure of an edge access device provided by an optional embodiment of the present invention, such as Figure 12 As shown, the edge access device includes: one or more first processors 10, a first memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the edge access device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple edge access devices can be connected, each device providing part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 12 A first processor 10 is taken as an example.
[0225] The first processor 10 may be a central processing unit, a network processor, or a combination thereof. The first processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0226] The first memory 20 stores instructions that can be executed by at least one first processor 10, so that the at least one first processor 10 executes the method shown in the above embodiment.
[0227] The first memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the edge access device, etc. In addition, the first memory 20 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some optional embodiments, the first memory 20 may optionally include a memory remotely located relative to the first processor 10, and these remote memories may be connected to the edge access device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0228] The first memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the first memory 20 may also include a combination of the above types of memory.
[0229] The edge access device further comprises a communication interface 30 for the edge access device to communicate with other devices or communication networks.
[0230] The embodiments of the present application also provide a vehicle, as shown in the drawings, comprising a vehicle machine and at least one in-vehicle device, wherein the vehicle machine and the in-vehicle device comprise one or more second processors, a second memory, and an interface for connecting components, including a high-speed interface and a low-speed interface, and the specific communication process of the second processor and the second memory can refer to the description of the first processor and the first memory above, which will not be repeated here. Figure 13
[0231] The embodiments of the present application also provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium and downloaded to a local storage medium through network, so that the method described herein can be processed by such software on a storage medium using a general-purpose computer, a special-purpose processor or programmable or special-purpose hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk or a solid-state disk, etc. Further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that the computer, processor, microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code, which, when accessed and executed by the computer, processor or hardware, implements the method shown in the above embodiments.
[0232] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be invoked or provided. Those skilled in the art should understand that the form of computer program instructions in a computer readable medium includes but is not limited to source files, executable files, installation package files, etc. Correspondingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.
[0233] While embodiments of the application have been described in connection with the preferred embodiments of the various figures, those of ordinary skill in the art will appreciate that various modifications and changes can be made without departing from the spirit and scope of the application, and that such modifications and changes fall within the scope of the appended claims.
Claims
1. A model deployment method, applied to an edge access end, and having the steps of, The method comprises: obtaining target model information and resource occupation information sent by a vehicle end; based on the target model information, finding and obtaining a target model requested to be deployed by the vehicle end; based on the resource occupation information, splitting the target model to obtain a first target sub-model with an edge access end as a deployment end and a second target sub-model with a vehicle end as a deployment end, and generating a policy file for representing collaborative execution information between the first target sub-model and the second target sub-model; deploying the first target sub-model, and sending the policy file and the second target sub-model to the vehicle end, so that the vehicle end deploys the second target sub-model based on the policy file.
2. The method of claim 1, wherein, The splitting of the target model based on the resource occupation information to obtain the first target sub-model with the edge access end as the deployment end and the second target sub-model with the vehicle end as the deployment end comprises: model pruning and model quantization are performed on the target model to obtain an optimized target model; based on storage resource occupation information, computing resource occupation information and network resource occupation information in the resource occupation information, the optimized target model is split into multiple sub-models; determining the deployment end corresponding to each of the multiple sub-models, taking the sub-model with the edge access end as the deployment end as the first target sub-model, and taking the sub-model with the vehicle end as the deployment end as the second target sub-model.
3. The method according to claim 2, characterized in that The method further comprises: based on the policy file and the first target sub-model, model training is performed on the second target sub-model deployed in the vehicle end in collaboration; and / or, based on the policy file and the first target sub-model, model inference is performed on the second target sub-model deployed in the vehicle end in collaboration, and the obtained model inference result is transmitted to the vehicle end.
4. The method of claim 1, wherein, The edge access end comprises multiple edge access points, and the method further comprises: receiving a node switching request sent by the vehicle end, and obtaining a candidate edge access point to be connected next by the vehicle end based on the node switching request; obtaining intermediate data generated when the current edge access point collaborates with the vehicle end to perform model training and / or model inference, the current edge access point being an edge access point currently connected with the vehicle end; controlling the current edge access point to disconnect with the vehicle end, and controlling the candidate edge access point to establish a connection with the vehicle end; sending the policy file, the first target sub-model and the intermediate data to the candidate edge access point, so that the candidate edge access point synchronously stores the intermediate data, and deploys the first target sub-model according to the policy file.
5. The method of claim 4, wherein, After controlling the current edge access point to disconnect with the vehicle end, the method further comprises: sending the target model, the policy file and the intermediate data to a cloud end, so that the cloud end deploys the target model according to the policy file, and performs model training and / or model inference based on the intermediate data and the target model.
6. The method of claim 4, wherein, The vehicle end comprises multiple vehicles, each of which corresponds to an in-vehicle device and at least one vehicle machine; the method further comprises: Monitoring a device state of each in-vehicle device or in-vehicle equipment in a vehicle end, and determining whether the vehicle end completes deployment of the target model; If it is detected that there is an in-vehicle device or in-vehicle equipment with an offline or online device state, and the vehicle end does not complete deployment of the target model, returning to the step of obtaining target model information and resource occupation information sent by the vehicle end; And / or, if it is detected that there is an in-vehicle device or in-vehicle equipment with an offline device state, and the vehicle end has completed deployment of the target model, obtaining an offline third target sub-model, and deploying the offline third target sub-model on the current edge access point; the offline third target sub-model is a second target sub-model deployed on an in-vehicle device or in-vehicle equipment with an offline device state; And / or, if it is detected that there is an in-vehicle device or in-vehicle equipment with an online device state, and the vehicle end has completed deployment of the target model, after receiving a new model deployment request corresponding to the online in-vehicle device or in-vehicle equipment sent by the vehicle end, establishing a connection with the online in-vehicle device or in-vehicle equipment, and performing the step of obtaining target model information and resource occupation information sent by the vehicle end.
7. The method of claim 1, wherein, Before obtaining the target model information and the resource occupation information sent by the vehicle end, the method further comprises: After receiving the model deployment request sent by the vehicle end, performing legality verification on the vehicle end; After detecting that the legality verification is passed, establishing a connection with the vehicle end, and performing the step of obtaining the target model information and the resource occupation information sent by the vehicle end.
8. The method of claim 1, wherein, The step of searching for and obtaining the target model requested to be deployed by the vehicle end based on the target model information comprises: Detecting whether there is a target model corresponding to the target model information stored locally; If the target model is not stored locally, sending a model obtaining request to the cloud, and after receiving a response feedback sent by the cloud based on the model obtaining request, downloading and obtaining the target model from the cloud; If the target model is stored locally, obtaining the target model locally. 9.A model deployment method, applied to a vehicle side, the method comprising: The method comprises: Sending target model information and resource occupation information to an edge access end; Receiving a policy file and a second target sub-model for the vehicle end sent by the edge access end, and deploying the second target sub-model according to the policy file; the second target sub-model is obtained by splitting a target model corresponding to the target model information based on the resource occupation information by the edge access end, and the policy file is used to represent cooperative execution information between the second target sub-model and a first target sub-model for the edge access end.
10. The method of claim 9, wherein, The vehicle end comprises a plurality of vehicles, each of which corresponds to an in-vehicle device and at least one in-vehicle equipment; the step of deploying the second target sub-model according to the policy file comprises: Analyzing the cooperative execution information in the policy file to obtain a target in-vehicle device or a target in-vehicle equipment corresponding to the second target sub-model; Deploying the second target sub-model and the policy file on the target in-vehicle device or the target in-vehicle equipment.
11. The method of claim 10, wherein, The method further comprises: The target vehicle machine or the target in-vehicle device parses the policy file to obtain data flow conversion rules and node connection rules applied by the second target sub-model when cooperatively executed; According to the data flow conversion rules, the node connection rules, and the second target sub-model, a first target sub-model deployed at an edge access end is cooperatively deployed for model training and / or model inference.
12. The method of claim 9, wherein, Before sending the target model information and the resource occupation information to the edge access end, the method further comprises: sending a model deployment request to the edge access end; After detecting that a connection with the edge access end is established, the step of sending the target model information and the resource occupation information to the edge access end is performed.
13. The method of claim 12, wherein, The edge access end comprises a plurality of edge access points, and the method further comprises: monitoring a packet loss rate of a currently connected current edge access point; when detecting that the packet loss rate is greater than a first packet loss rate threshold within a preset duration, finding an alternative edge access point adjacent to the current edge access point; when detecting that the packet loss rate is greater than a second packet loss rate threshold within a preset duration, sending a node switching request to the current edge access point to disconnect the connection with the current edge access point and establish a connection with the alternative edge access point; wherein the second packet loss rate threshold is greater than the first packet loss rate threshold.
14. A model deployment system, comprising: The system comprises an edge access end and a vehicle end; The edge access end acquires target model information and resource occupation information sent by the vehicle end; based on the target model information, finding and obtaining a target model requested by the vehicle end to be deployed; based on the resource occupation information, splitting the target model to obtain a first target sub-model deployed at the edge access end and a second target sub-model deployed at the vehicle end, and generating a policy file for representing cooperative execution information between the first target sub-model and the second target sub-model; deploying the first target sub-model, and sending the policy file and the second target sub-model to the vehicle end; The vehicle end sends target model information and resource occupation information to the edge access end; receiving the policy file and the second target sub-model deployed at the vehicle end sent by the edge access end, and deploying the second target sub-model according to the policy file.
15. The system of claim 14, wherein, The edge access end comprises a plurality of edge access points, and the vehicle end comprises a plurality of vehicles; wherein each edge access point corresponds to at least one vehicle, each vehicle corresponds to a vehicle machine and at least one in-vehicle device, and the vehicle machine and the in-vehicle device are loaded in the vehicle.
16. The system of claim 14, wherein, The system further comprises a cloud end; The cloud end receives a model acquisition request sent by the edge access end; based on the model acquisition request, sends a response feedback to the edge access end, so that the edge access end downloads a target model from the cloud end.
17. An edge access device, characterized by comprise: a first memory and a first processor, which are communicatively connected with each other, the first memory stores computer instructions, and the first processor executes the computer instructions to perform the model deployment method in any one of claims 1 to 8.
18. A vehicle characterized by comprising: comprise: a vehicle machine and at least one in-vehicle device; The car machine or the in-vehicle device includes a second memory and a second processor, which are in communication connection with each other, the second memory stores computer instructions, and the second processor executes the computer instructions to perform the model deployment method in any one of claims 9 to 13.
19. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a computer to perform the model deployment method in any one of claims 1 to 8, or the model deployment method in any one of claims 9 to 13.