Model deployment method and system

By concurrently processing model deployment requests between cloud devices and end-side devices, dynamically matching and batch transmission of model data packets, the problems of low file transfer efficiency and insufficient resource reuse capabilities during model deployment are solved, and efficient model deployment and improved system performance are achieved.

CN120012046AInactive Publication Date: 2025-05-16ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510459416.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as low file transfer efficiency and insufficient resource reuse capabilities during the model deployment process, which has affected service response speed and system stability.

Method used

By implementing a method of concurrently processing multiple model deployment requests between cloud devices and end-side devices, the target model data packets are dynamically matched, and a batch transfer mechanism is used to reduce repetitive file pull operations.

Benefits of technology

It significantly improves the distribution efficiency and resource scheduling capabilities of model files, reduces overall deployment delay, improves system throughput in high-frequency, multi-node deployment scenarios, and reduces storage resource occupation and network load pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012046A_ABST
    Figure CN120012046A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model deployment method and system, the model deployment method is applied to a cloud device, the cloud device comprises at least one piece of model information, the model information comprises at least one model data packet, and the method comprises the following steps: receiving at least two model deployment requests sent by at least one end side device; determining at least one target model data packet corresponding to each model deployment request; and sending each target model data packet to an end side device corresponding to each target model data packet, so that each end side device completes model deployment. Intelligent matching of the target model data packets is implemented for the concurrent requests of the multi-end-side equipment, so that the deployment efficiency is improved, the storage resources are intensively utilized, and the service response stability in a multi-concurrent scene is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a model deployment method. Background Art

[0002] With the rapid development of large model technology and the continuous expansion of its application scale, file transfer efficiency and resource reuse capabilities in the model deployment process have gradually become core factors affecting service response speed and system stability. The exponential growth of large model file size has significantly increased the transmission time, and the high-frequency model deployment requirements have further magnified the efficiency bottleneck's restrictive effect on the overall task link.

[0003] At present, the mainstream model deployment solution meets the basic deployment requirements by integrating file transfer and task execution modules. However, in actual application scenarios, the repeated deployment scenarios of the same node lack an effective file reuse mechanism, resulting in waste of storage resources and superposition of performance loss. Therefore, in order to solve the above problems, technicians in this field are in urgent need of a model deployment method. Summary of the invention

[0004] In view of this, the embodiments of this specification provide a model deployment method applied to a cloud device, a model deployment method applied to a terminal device, and a model deployment system. This specification also relates to a model deployment device applied to a cloud device, a model deployment device applied to a terminal device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a model deployment method is provided, which is applied to a cloud device, wherein the cloud device includes at least one model information, and the model information includes at least one model data package, including: Receive at least two model deployment requests sent by at least one end-side device; determine at least one target model data packet corresponding to each model deployment request; send each target model data packet to the end-side device corresponding to each target model data packet, so that each end-side device completes the model deployment.

[0006] According to a second aspect of an embodiment of this specification, a model deployment method is provided, which is applied to a terminal device, including: Generate at least one model deployment request, and send each model deployment request to a cloud device; determine a target model deployment request, and receive at least one target model data packet corresponding to the target model deployment request, wherein the target model deployment request is any one of the model deployment requests, and each target model data packet is sent by the cloud device; generate a model transmission result based on each target model data packet; when the model transmission result is that the transmission is completed, deploy a target data processing model based on each target model data packet.

[0007] According to a third aspect of the embodiments of this specification, a model deployment system is provided, the system comprising a cloud device and at least one end-side device, wherein the cloud device comprises at least one model information, the model information comprises at least one model data packet; a target end-side device is configured to send at least one model deployment request to the cloud device; the cloud device is configured to receive each model deployment request; determine at least one target model data packet corresponding to each model deployment request; and send each target model data packet to the end-side device corresponding to the model data packet; The target end-side device is also configured to receive at least one target model data packet corresponding to a target model deployment request; generate a model transmission result based on each target model data packet, wherein the target model deployment request is any one of the model deployment requests; and deploy a target data processing model based on each target model data packet when the model transmission result is that the transmission is completed.

[0008] According to a fourth aspect of an embodiment of this specification, a computing device is provided, including: A memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned model deployment method.

[0009] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned model deployment method are implemented.

[0010] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned model deployment method when executed by a processor.

[0011] By applying the solution of the embodiments of this specification, the distribution efficiency and resource scheduling capabilities of model files are effectively improved by concurrently processing model deployment requests from multiple end-side devices, and based on the model data packet set stored on the cloud device, the target model data packet is dynamically matched for the deployment requests of different end-side devices, and repetitive file pulling operations are reduced through a batch transmission mechanism, thereby reducing the overall deployment delay. Furthermore, by optimizing the request response mechanism, the system throughput in high-frequency, multi-node deployment scenarios is significantly improved while ensuring the accurate delivery of model files, while reducing the storage resource usage and network load pressure caused by redundant transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flow chart of a model deployment method applied to a cloud device provided by an embodiment of this specification; Figure 2 is a flow chart of a model deployment method applied to a terminal device provided by an embodiment of this specification; Figure 3 This is a schematic diagram of a processing process of model deployment in a terminal device provided by an embodiment of this specification; Figure 4 It is an architecture diagram of a model deployment system provided by an embodiment of this specification; Figure 5 It is a process flow chart of a method for local training of a cloud model provided by an embodiment of this specification; Figure 6 It is a structural diagram of a model deployment device applied to a cloud device provided by an embodiment of this specification; Figure 7 It is a structural diagram of a model deployment device applied to a terminal device provided by an embodiment of this specification; Figure 8 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0013] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0014] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0017] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.

[0018] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.

[0019] First, the terms involved in one or more embodiments of this specification are explained.

[0020] Simple Storage Service (S3) is an object-based distributed storage service that provides high availability and scalability and supports data access operations through standard interfaces. Its core functions include data sharding storage, access control, and version management. It is suitable for scenarios such as large-scale file storage, static resource hosting, and cross-system data sharing. In the model deployment process, S3 is often used to store encrypted model files or training data sets, and implements secure and controllable temporary access rights allocation through pre-signed URLs.

[0021] Graphics Processing Unit (GPU): A hardware device designed for parallel computing with large-scale concurrent processing capabilities. Compared with traditional processors, it accelerates matrix operations through thousands of computing cores, which can significantly improve computing efficiency in deep learning model training and reasoning scenarios. Under the containerized deployment framework, GPU resources are isolated and allocated among multiple tasks through device plug-in mechanisms and virtualization technology, ensuring that model reasoning tasks receive stable computing power support.

[0022] Container technology: It is a lightweight virtualization solution that uses the process isolation mechanism of the operating system kernel to encapsulate and standardize the application runtime environment. Its core components include image management tools and resource scheduling engines, which can package applications and their dependencies into portable units. In the model deployment scenario, container technology provides independent runtime environments for models of different frameworks to avoid dependency conflicts, while ensuring the stability of multi-model parallel services through resource quota restrictions.

[0023] MD5: is a widely used cryptographic hash function that can generate a 128-bit hash value to verify data integrity. It generates a highly unique fingerprint value by performing block processing and nonlinear transformation on the input data. In file transfer scenarios, MD5 is often used to verify the hash consistency before and after the model file is transferred. For example, when sending an encrypted model package from the cloud, the end-side device compares the locally calculated MD5 value with the verification code provided by the server to confirm that the file has not been damaged or tampered with.

[0024] Hook task: It is a preset sequence of operations that the system executes before and after a specific event is triggered, which is used to implement process expansion or status monitoring. Its implementation methods include registering callback functions or configuring automation scripts, such as loading environment variables when the container starts, or triggering log archiving at the end of the task. In the model deployment framework, the Hook task can be associated with the file decryption completion event, automatically perform model loading verification, or trigger the cloud synchronization process after the model is updated to ensure the integrity and traceability of the deployment process.

[0025] In this specification, a model deployment method applied to a cloud device, a model deployment method applied to a terminal device, and a model deployment system are provided, which are described in detail one by one in the following embodiments.

[0026] In the overall link of large model tasks, each model deployment needs to complete the pulling of model files from the database, decryption and decompression of model files, and running model services (including reasoning or training). Among them, the transmission of model files usually takes a long time, leading to several problems. On the one hand, due to the low download efficiency and the long pre-preparation time for each task, it is necessary to decouple the file transfer tool from the model task function. The file transfer tool is responsible for unified download and pre-processing, while the model task focuses on providing services to the outside world; on the other hand, when the same model is repeatedly deployed on the same node, the deployment instances in different containers will cause the model file to be pulled multiple times, thus affecting the overall performance and user experience. From the perspective of model tasks, projects often require certain functional extensions to model tasks, such as web page downloads, encryption and decryption of model algorithm files, and recording of environmental information.

[0027] See also Figure 1 , Figure 1 A flowchart of a model deployment method applied to a cloud device provided according to an embodiment of the present specification is shown. The cloud device includes at least one model information, and the model information includes at least one model data package. The above method specifically includes the following steps.

[0028] Step 102: Receive at least two model deployment requests sent by at least one end-side device.

[0029] In actual applications, cloud devices are service clusters that store model information and execute model deployment logic; model information is a collection of model data packages and their associated parameters; model data packages are structured encapsulated file units that include model architecture, weight parameters, and operating dependencies; end-side devices are edge computing nodes that initiate model deployment requests and receive data packages; model deployment requests are service instructions containing configuration parameters issued by end-side devices to the cloud.

[0030] Cloud devices can be understood as core service nodes deployed in a distributed architecture, which are responsible for the storage management and scheduling and distribution of model resources. They process requests from multiple terminals through a multi-threaded concurrent mechanism. For example, when receiving an image recognition model deployment request, they determine at least one target model data packet corresponding to the image recognition model deployment request. In specific implementations, cloud devices use container technology to perform file transfer initialization, multi-shard downloading, and environmental monitoring functions, such as dynamically adjusting the number of concurrent connections during the transmission process and monitoring memory usage to ensure service stability.

[0031] Model information can be understood as a collection of model resources managed by task scenarios in cloud devices. Each model information corresponds to the complete deployment elements of a specific functional version. For example, in a natural language processing scenario, a model information may contain the network structure file, vocabulary data package, and quantization configuration file of the natural language processing model. This information is stored in the shared volume at a directory level. When multiple terminals are detected requesting the same model version, the file lock status is used to determine whether to trigger the batch distribution process to avoid repeated pulling of basic files.

[0032] A model data packet can be understood as an independent transmission unit corresponding to a certain model information, which contains the atomic components required for model operation. For example, in a computer vision task, a model data packet may include a serialized file of the network structure, pre-trained weight parameters and binary files of their corresponding dependent libraries. When transmitting, the cloud device executes a differentiated decompression strategy based on the parameters in the deployment request (such as whether to overwrite files with the same name), and pulls multiple data packet fragments from the object storage service in parallel through a multi-threaded pool to improve the efficiency of large file transmission.

[0033] The end-side device can be understood as a physical unit that carries the model operation and executes specific task logic. After receiving the model data packet distributed from the cloud, this type of device can parse and generate an executable data processing model based on the local environment, such as a smartphone, industrial robot, or edge computing server. Taking the industrial scenario as an example, when the intelligent quality inspection terminal on the production line needs to deploy a defect detection model, it initiates a model deployment request to the cloud as an end-side device, and then receives a model data packet adapted to its hardware framework to complete the construction of the local reasoning service. For example, after the end-side device obtains the model file through a shared volume or S3 (Simple Storage Service), the auxiliary container completes the decryption, decompression, and environment adaptation, and finally triggers the service operation of the main application container.

[0034] The model deployment request can be understood as the model resource acquisition request proposed by the end-side device to the cloud device, which usually includes information such as the target model version identifier, deployment parameters, and environmental constraints. For example, when a medical imaging terminal needs to upgrade the lesion segmentation model, its request may include parameters such as the model version number, GPU (Graphics Processing Unit) memory limit, and encrypted transmission requirements. In an implementation logic provided in this specification, the request will trigger the cloud device to initialize the transmission task, including setting the number of concurrent threads, file overwrite strategy, etc., and dynamically matching the target model data packet through the read-write mapping table. For high-concurrency scenarios, cloud devices use file lock mechanisms to coordinate access to shared volumes by multiple terminals to avoid transmission congestion caused by resource contention. For example, when multiple terminals request the same model at the same time, the system uses lock logic to ensure the atomicity and consistency of file distribution.

[0035] It should be noted that before receiving at least two model deployment requests sent by at least one end-side device, one or more model deployment requests can be generated for each end-side device. The way in which the end-side device generates the model deployment request can be to determine the corresponding configuration information according to the actual requirements of the task, and then generate a model deployment request based on the various configuration information.

[0036] Through the concurrent processing capabilities of cloud devices for multiple terminal requests, combined with the status marking of model data packets and the shared volume distribution mechanism, redundant data transmission is reduced while ensuring the integrity of file transfers. The dynamic resource allocation strategy enables the system to automatically switch transmission paths according to network conditions, and the modular data processing process (such as automatic decryption and environment adaptation) effectively shortens the overall delay from request issuance to service readiness. This design significantly improves the model version iteration efficiency and resource utilization in cross-regional, multi-node deployment scenarios.

[0037] Step 104: Determine at least one target model data package corresponding to each model deployment request.

[0038] In actual applications, the target model data packet is a model file that is dynamically screened and transmitted by the cloud device based on the deployment request of the end-side device. The target model data packet can be understood as a model resource that the cloud device matches according to the specific needs of the end-side device, including deployment elements such as model parameters, configuration files, and dependent components. For example, in the intelligent customer service system upgrade scenario, the target model data packet may contain the core parameter file of the speech recognition model, the dialogue strategy configuration file, and encryption verification information. In the implementation logic provided by an embodiment provided in the present application, the data packet will be sliced ​​according to the task parameters before transmission, such as splitting a large-size model file into multiple slices, and improving efficiency through multi-threaded concurrent transmission. For deployment scenarios that support shared volumes, the target model data packet can quickly locate the storage location through the read-write mapping table, and ensure data consistency during concurrent access by multiple terminals based on the file lock mechanism. For example, when multiple edge computing nodes request the same image classification model, the cloud device distributes the existing target data packet through the shared volume to avoid repeated pulling from remote storage.

[0039] By dynamically matching the terminal deployment requirements with the model resources stored in the cloud, model deployment requests corresponding to multiple models can be processed concurrently, thereby improving the efficiency of model deployment. After the terminal device receives the target model data packet, an end-to-end automated deployment link can be formed.

[0040] Furthermore, at least one target model data packet corresponding to each model deployment request is determined, including: determining at least one initial model data packet corresponding to each model deployment request; obtaining current status information of each initial model data packet; and determining the target model data packet corresponding to each model deployment request based on the current status information corresponding to each initial model data packet.

[0041] In actual applications, the initial model data package is a set of candidate model files preliminarily screened by the cloud device based on the model deployment request; the current status information is a set of dynamic attributes that characterize the real-time availability and integrity of the model data package in a shared storage environment.

[0042] The initial model data packet can be understood as a collection of basic model resources that are filtered out through version matching or task rules when the cloud device responds to the terminal deployment request. For example, in the intelligent customer service scenario, when the terminal requests to deploy the speech recognition model, the initial model data packet may contain model file directories, corresponding configuration files, and dependent libraries of multiple versions from V2.1 to V2.3. According to the read-write mapping table in the shared volume, the system retrieves the model version metadata (such as MD5 value) recorded in the mapping relationship. If it is detected that the shared volume has cached the same version file, the file under the path is marked as the initial candidate set. If the cache is not hit, the file that meets the version requirements is pulled from the remote storage service as the initial data packet, such as downloading a compressed package of the specified version number and its associated verification file.

[0043] The current status information can be understood as real-time feedback data on the accessible status and consistency identification of the model data package in the shared storage system. For example, in an industrial quality inspection scenario, when multiple terminals simultaneously request to deploy a defect detection model, the file lock status in the shared volume will record whether the model data package is being written or distributed. In the specific implementation, the status information includes the number of reentrants of the file lock (allowing concurrent access by multiple threads in the same processing unit), the lock holding timeout (avoiding deadlock), and the file integrity identifier (such as judging whether the file is complete through the MD5 checksum value). When the auxiliary container attempts to obtain the initial model data package, it will determine whether to trigger the waiting logic or switch the transmission path based on the status information. For example, when it detects that the target file is in a locked state and the remaining timeout is greater than the threshold, it starts a polling mechanism instead of directly interrupting the request.

[0044] By dynamically screening the initial model data package and making secondary decisions based on its real-time status information, the system can effectively identify available resources that are ready in the shared volume and reduce deployment delays caused by file conflicts or transmission interruptions. The precise matching mechanism based on version metadata avoids the misdistribution of wrong version files, and at the same time, through the state-aware lock management strategy, balances the resource contention problem in multi-terminal concurrent access scenarios. This dual screening mechanism not only ensures the accuracy of model deployment, but also improves the throughput efficiency of resource distribution in high-concurrency scenarios.

[0045] Further, based on the current state information corresponding to each initial model data package, the target model data package corresponding to each model deployment request is determined, including: Determine a reference model deployment request in each model deployment request, and determine a first initial model data packet in each initial model data packet corresponding to the reference model deployment request, wherein the reference model deployment request is any one of the model deployment requests, and the first initial model data packet is any one of the initial model data packets corresponding to the reference model deployment request; When the current state information corresponding to the first initial model data packet is in an idle state, determining that the first initial model data packet is a target model data packet corresponding to the reference model deployment request, and setting the current state information corresponding to the first initial model data packet to a read state; When the current state information corresponding to the first initial model data packet is a read state, determining that the first initial model data packet is a target model data packet corresponding to the reference model deployment request; When the current status information corresponding to the first initial model data packet is a write status, a target waiting time is determined, and based on the target waiting time, it is determined whether the first initial model data packet is a target model data packet corresponding to the reference model deployment request.

[0046] In actual applications, the idle state is the file lock state in which the model data package in the shared volume is not occupied and access is allowed; the read state is the file lock state in which the model data package is being called by the terminal device and concurrent operations by multiple threads of the same source are allowed; the write state is the file lock state in which the model data package is being modified or updated and access by other terminals is prohibited; the target waiting time is the maximum continuous waiting time threshold set before the retry mechanism is triggered in the write state.

[0047] The idle state can be understood as the model data package in the shared storage system is in a ready state for immediate distribution. For example, in an industrial quality inspection scenario, when the defect detection model file is not occupied by any terminal and the integrity check has been completed, its file lock is marked as idle. At this time, the auxiliary container can directly locate the physical storage path of the file through the read-write mapping table and start the multi-process shard transfer process. In the specific implementation, the determination of the idle state must meet two conditions: the reentrant count of the file lock is zero (indicating no active operation) and the MD5 checksum value is consistent with the record in the mapping table.

[0048] The read state can be understood as an intermediate state in which the model data packet is being read concurrently by one or more edge devices but write operations are prohibited. For example, when multiple edge computing nodes simultaneously request to deploy the same image classification model, the auxiliary container is marked as read by increasing the number of reentrants of the file lock, allowing multiple coroutines in the same processing unit to transmit shard data in parallel without waiting for the preceding thread to release resources. In this state, the mapping table of the shared volume will update the file access count and the last operation timestamp in real time for state recovery after abnormal interruption.

[0049] The write state can be understood as an exclusive operation state in which the model data package is in the process of content update or version upgrade, and other terminals are prohibited from accessing it. For example, when the cloud device pulls a new version of the speech recognition model file from the remote storage service, the auxiliary container marks the target file lock as a write state and starts a timeout countdown mechanism. At this time, if other terminals initiate a deployment request, the system will trigger a polling detection based on the target waiting time instead of directly returning an error. The release of the write state requires two conditions to be met: the reentrant count of the file lock is reset to zero (indicating that the write is complete) and the newly generated MD5 value is successfully synchronized to the mapping table.

[0050] The target waiting time can be understood as the maximum tolerable waiting time window set by the system for the model data package in the write state. For example, in the intelligent customer service system upgrade scenario, if the dialogue model file requested by a terminal is being updated by other nodes, the auxiliary container will start the countdown (such as setting 30 seconds) and detect the file lock status change at fixed intervals. If it is still in the write state after the timeout, it will automatically switch to the backup transmission path (such as pulling the original file from the remote storage service), and update the version association information in the mapping table to avoid subsequent request conflicts.

[0051] By hierarchically managing the state marking and waiting strategies of model data packets, the system can accurately coordinate the order of resource access in high-concurrency scenarios and reduce deployment interruptions caused by state conflicts. Based on the reentrant number mechanism and timeout release rules, the contradiction between multi-terminal resource contention and transmission efficiency is balanced. At the same time, through the state-driven path switching logic, it ensures that the terminal side devices can always obtain available model files during the version update process. This state-aware decision-making mechanism effectively improves the robustness and response speed of model deployment in resource sharing scenarios.

[0052] Considering that after the target model data packet is sent, there is no other client operating the target model data packet. Therefore, after sending each target model data packet to the end-side device corresponding to each target model data packet, the method also includes: determining that the current status information corresponding to each target model data packet is an idle state.

[0053] After the model data packet is transmitted, its status information needs to be updated synchronously to maintain the accuracy of system resource scheduling. The core of this step is to ensure that the model data packet that has been successfully delivered to the end-side device can release the occupied shared resources in time to avoid misjudgment of other deployment requests due to delayed status information.

[0054] In the specific implementation, when the end-side device confirms that all data shards have been received and returns the integrity check result (such as hash value match), the cloud processing unit will start the status reset process: first check the reentrant count of the current file lock. If the value returns to zero, it indicates that no other active threads depend on the resource. At this time, the status mark of the corresponding model data packet in the shared volume mapping table is updated from "read status" to "idle status", and the latest operation timestamp is recorded as a reference for subsequent resource scheduling.

[0055] In one embodiment provided in this specification, the target task is speech recognition, and the corresponding model deployment request is an intelligent speech model deployment request. When an edge node completes receiving and verifying a file, the system automatically removes the occupancy mark of the file in the shared storage, so that other nodes can directly reuse the ready resources for subsequent requests, reducing repeated transmission overhead. For abnormal situations that occur during the transmission process (such as some fragments not receiving confirmation signals), the system will trigger the retransmission logic instead of directly updating the status, ensuring that the resource release operation is only performed under the premise that the data is fully delivered.

[0056] Through this mechanism, the system ensures efficient resource reuse while maintaining the consistency of status information and physical resources in multi-terminal concurrent scenarios, effectively reducing the risk of deployment delays or conflicts caused by state asynchrony.

[0057] Step 106: Send each target model data packet to the end-side device corresponding to each target model data packet, so that each end-side device completes the model deployment.

[0058] The process of sending the target model data packet to the corresponding end-side device is a key step to ensure that the model file is accurately and efficiently transmitted from the cloud to the terminal. This step is necessary because model deployment needs to solve the problem of reliable transmission of large-scale files in a complex network environment, and at the same time, it needs to adapt to the differences in the operating environments of different end-side devices. Through the fragmented transmission strategy and dynamic resource scheduling mechanism, the distribution efficiency in multi-terminal concurrent scenarios can be improved while ensuring the integrity of the transmission.

[0059] In the specific implementation, the cloud device first pre-processes the target model data packet according to the parameters set in the deployment request (such as the number of shards and transmission priority). For example, for large-size model files, the system automatically splits them into multiple shard units, each shard independently generates verification information, and transmits it to the end-side device through multiple threads concurrently. During the transmission process, the network bandwidth fluctuation is monitored in real time, and the shard size and transmission order are dynamically adjusted: when it is detected that the network bandwidth is sufficient, the core model weight shard is transmitted first to accelerate the partial startup of the end-side inference service; when the network quality decreases, the shard size is reduced and the number of concurrent threads is reduced to reduce the risk of packet loss. For encrypted and compressed data packets, the end-side device automatically triggers the decryption process after receiving it, and verifies the integrity of each shard through hash verification. If a shard that fails the verification is found, an incremental retransmission request for the specified shard is initiated to the cloud to avoid the waste of resources for retransmitting the entire file.

[0060] The technical effects of this step are mainly reflected in three aspects: First, the fragmented transmission mechanism significantly shortens the overall transmission delay in a high-latency network environment by splitting large files into independent units that can be processed in parallel; second, the dynamic adjustment strategy can optimize resource allocation according to the real-time network status, while ensuring the transmission success rate and improving bandwidth utilization; third, the terminal-side automated decryption and verification process reduces the deployment failure rate caused by file damage or version inconsistency, ensuring the rapid readiness of model services. These technical features jointly support the needs of large-scale model deployment across regions and heterogeneous terminal-side device groups, forming a complete closed loop from cloud resource scheduling to terminal-side service activation.

[0061] Taking into account that the tasks of the application model corresponding to the terminal device may modify the files corresponding to the model, and the modified model needs to be uploaded to the cloud for updating, after sending each target model data packet to the terminal device corresponding to each target model data packet, the method also includes: receiving at least two model update requests sent by at least one terminal device, wherein the model update request includes at least one model update data packet; determining the reference model data packet corresponding to each model update data packet, and updating each reference model data packet based on each model update data packet.

[0062] In actual applications, the model update request is an operation instruction initiated by the end-side device to the cloud to synchronously update the model file after modification; the model update data package is the model file modified locally on the end-side device and its associated version identification data set; the reference model data package is the original model file to be updated in the cloud storage environment and its metadata set.

[0063] A model update request can be understood as a structured request for file synchronization update initiated to the cloud by the terminal device after completing local model optimization or parameter adjustment. For example, in an industrial quality inspection scenario, an edge computing device adjusts the threshold parameters of a defect detection model and generates request data containing the new model file path, version number, and modification timestamp. The request uses the weak lock mechanism of the shared volume for conflict detection. When the cloud detects that the same reference model data packet is not occupied by other terminals, it triggers the update process. In the specific implementation, the request data includes the MD5 checksum value of the target file, the changed field index, and the operation type identifier, which are used by the cloud to match the corresponding reference model data packet.

[0064] The model update data package can be understood as a packaged collection of the model files and their integrity verification information that have been modified locally on the end device. For example, in the intelligent driving scenario, the vehicle terminal generates a data package containing the new model weight file, configuration file, and corresponding verification code after lightweight compression of the image recognition model. When the data package is uploaded to the cloud shared volume through the coroutine pool concurrently, the system compares the version information of the reference model data package (such as the original MD5 value). If a version difference is detected, the file overwrite write process is started. During the upload process, the auxiliary container ensures that multi-threaded update operations on the same terminal do not cause internal resource competition by maintaining the reentrant count of the file lock.

[0065] The reference model data package can be understood as an associated collection of original model files to be updated and their version metadata in the cloud storage system. For example, in the speech recognition service, when a terminal uploads an optimized acoustic model, the cloud locates the storage path of the original model file and its version identifier (such as the initial MD5 value) by querying the read-write mapping table of the shared volume. The metadata of the reference model data package includes the file lock status, the last access time, and the associated terminal device list, which are used to determine the legitimacy of the update operation. During the update process, if it is detected that the reference model data package is in a write state (such as being modified by other terminals), the system will start polling detection or trigger version branch management based on the target waiting time.

[0066] By establishing a precise matching mechanism between terminal update requests and cloud reference models, the system can achieve dynamic two-way synchronization of model files and avoid version conflicts caused by concurrent modifications on multiple terminals. Based on the weak lock strategy and version verification mechanism of shared volumes, it supports high-frequency model iteration updates while ensuring data consistency. This end-cloud collaborative update mode not only retains the local optimization capabilities of terminal devices, but also maintains the global validity of model resources through centralized version management, thereby improving the flexibility and reliability of distributed model deployment systems.

[0067] Further, determining the reference model data packet corresponding to each model update data packet includes: determining that the current status information corresponding to each reference model data packet is a write status; correspondingly, after each model update data packet updates each reference model data packet, the method also includes: determining that the current status information corresponding to each reference model data packet is an idle status.

[0068] Marking and releasing the reference model data package during the model update process is the core mechanism to ensure data consistency when multiple terminals are collaboratively updating. This step is necessary to prevent file content conflicts or version confusion caused by multiple terminals initiating updates to the same model data package at the same time. Through closed-loop management of status marking and release, it is ensured that only one terminal in the cloud storage environment writes to the reference model at the same time, avoiding data overwrite or verification failure caused by concurrent modifications. For example, when an edge device is updating a defect detection model, update requests from other devices to the model will be temporarily suspended or redirected to an alternative path until the current operation is completed and the resources are released.

[0069] In the specific implementation process, the current status of the target reference model data package is first queried through the read-write mapping table of the shared volume. If the status is detected to be idle, it is marked as a write state and the reentrant count of the file lock is incremented (allowing concurrent operations of multiple coroutines in the same terminal), and then the file content update process is started. For example, in the intelligent customer service scenario, when the terminal uploads the optimized semantic understanding model, the cloud locks the original model file as a write state, and merges the difference data package uploaded by the terminal with the original file through the coroutine pool. After the update is completed, the system performs an MD5 check on the merged file. If the check passes, the version information in the mapping table is updated and the status is reset to idle. If the status is detected to be write, polling detection is triggered or a conflict prompt is returned according to the preset target waiting time. For example, in an industrial quality inspection system, if a model file has been locked for update, the new request will be delayed for 5 seconds before re-detecting the status.

[0070] The technical effects of this step are mainly reflected in three aspects: First, through the exclusive marking mechanism of the write status, the atomicity of the model update operation is ensured to prevent file damage caused by cross-writing of multiple terminals; second, based on the state-driven update queue management, the execution order of multi-terminal update requests is effectively coordinated to reduce the response delay caused by resource contention; third, the linkage mechanism of state reset and version verification updates the global version identifier while releasing resources, so that subsequent requests can accurately identify the latest available model. For example, in the autonomous driving model iteration scenario, when a vehicle-mounted terminal completes the perception model optimization and releases the write lock, other terminals can immediately obtain the latest model based on the updated version number without repeating the version comparison operation. This state-aware update process improves the utilization of cloud resources while maintaining the orderliness of model version evolution in a distributed environment.

[0071] By applying the solution of the embodiments of this specification, accurate resource scheduling in high-concurrency scenarios is achieved by dynamically matching the model data packet status with the real-time needs of deployment requests. Based on the status information marking mechanism of the model data packet, the idle, read and write states are intelligently distinguished when multiple deployment requests are processed concurrently, and the concurrent access conflicts of the data packet are avoided through the status judgment logic, ensuring the orderly execution of file transfer and model update. For high-frequency deployment scenarios, this method reduces the number of repeated pulls of the same model data packet through a batch transmission strategy, and combines the dynamic update mechanism of status information to achieve rapid reuse of the transmission channel, effectively reducing the average waiting time in multi-terminal deployment scenarios. In addition, through the status collaborative management mechanism of model update requests and deployment requests, while ensuring the integrity of the data packet, the seamless connection between model iteration updates and online services is achieved, avoiding the risk of service interruption due to data version conflicts, thereby improving the stability and resource utilization of the system in multi-task parallel scenarios.

[0072] Corresponding to the above method embodiment, this specification also provides a model deployment method embodiment applied to a terminal device, see Figure 2 , Figure 2 A flowchart of a model deployment method applied to a terminal device provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0073] Step 202: Generate at least one model deployment request, and send each model deployment request to a cloud device.

[0074] In actual applications, the terminal device will execute at least one task that requires model deployment, and each character corresponds to at least one model deployment request. When deploying the model, each model deployment request corresponding to each character needs to be sent to the cloud device so that each task can be executed smoothly.

[0075] Step 204: determine a target model deployment request, and receive at least one target model data packet corresponding to the target model deployment request, wherein the target model deployment request is any one of the model deployment requests, and each target model data packet is sent by the cloud device.

[0076] In practical applications, the target model deployment request is the deployment request corresponding to the target task. The target model deployment request can be understood as a request corresponding to the deployment of the data processing model used to achieve the target task. For example, if the target task is a model training task, its corresponding model deployment request is a request corresponding to the deployment of the data processing model that needs to be trained; for example, if the target task is a text processing task, its corresponding model deployment request is a request corresponding to the deployment of the text processing model that needs to perform text processing.

[0077] By obtaining multiple target model data packets corresponding to the target model deployment request, the target task in the terminal device can be smoothly implemented. In addition, since the cloud device can process multiple model deployment requests at the same time, when the terminal device executes multiple tasks at the same time, the deployment of data processing models corresponding to multiple tasks can be performed concurrently, thereby improving the efficiency of model deployment.

[0078] Step 206: Generate a model transmission result based on each target model data packet.

[0079] In actual applications, the model transmission result is the transmission status judgment conclusion generated after the end device receives the cloud model data packet, which includes transmission completion and transmission error; transmission completion is the operation result that the model data packet is completely received and passes the integrity check; transmission error is an abnormal state in which data loss, verification failure or process interruption occurs during the reception of the model data packet.

[0080] The model transmission result can be understood as the status feedback formed after the end-side device performs full-link monitoring of the model data packet reception process sent from the cloud. For example, when deploying a speech recognition model in an intelligent customer service system, the end-side device performs segment verification on the received file based on pre-set environment variables (such as whether to enable MD5 verification). If the hash values ​​of all segments are consistent with the cloud records, a transmission completion mark is generated. This result is linked to the Hook task configured in the auxiliary container to trigger subsequent decryption and decompression operations and send a ready signal. If a segment verification failure or a network connection interruption is detected, a transmission error result containing an error code and interruption location information is generated.

[0081] The completion of the transmission can be understood as the state in which the model data packet is fully transmitted from the cloud to the end device and passes the integrity verification. For example, in an industrial quality inspection scenario, when the edge device receives all the compressed fragments of the defect detection model, the auxiliary container automatically calls the decompression tool to restore the original file, and determines whether to overwrite the local file with the same name based on the preset parameters. If the file is decrypted successfully and the version information is consistent with the mapping table, the task of obtaining the task file is triggered to end, and the ready state is synchronized to the terminal device, so that it starts the model inference service. In the specific implementation, the judgment of the completion of the transmission must simultaneously meet the three conditions that the number of fragments is complete, the MD5 check passes, and there are no abnormalities in the decryption process.

[0082] Transmission errors can be understood as abnormal termination states caused by network fluctuations, file corruption, or configuration conflicts during the transmission of model data packets. For example, in the scenario of updating the autonomous driving model, if the sensor fusion model file downloaded by the end device is missing due to a transmission timeout, the auxiliary container will record the current environment variables (such as the download task ID, failed shard index) and execute the task error task of obtaining the task file, and generate an analysis report containing the error type (such as lost shards, decryption failure) and repair suggestions. For errors triggered by inconsistent MD5 checks, the system will automatically trigger a retransmission mechanism instead of directly discarding the received data, thereby reducing the overhead of repeated transmissions.

[0083] Through the hierarchical processing mechanism of model transmission results by the end-side device, the system can dynamically adapt to different network environments and file characteristics to improve the success rate of model deployment. The automatic decryption and ready callback after the transmission is completed reduce the manual intervention link, and the refined error location and retransmission strategy during transmission errors effectively reduce the risk of overall deployment failure due to occasional exceptions. This state-driven processing logic, combined with the collaborative operation of Hook tasks and tool chains, realizes seamless connection from file transfer to model service, ensuring the reliability and efficiency of the end-side model deployment process.

[0084] Step 208: When the model transmission result is transmission completion, deploy the target data processing model based on each target model data packet.

[0085] In actual applications, the target data processing model is an executable model instance formed after the end-side device successfully receives and processes the cloud model data packet; the target data processing model can be understood as the model data packet sent by the end-side device based on the cloud, and the available model entity deployed after the environment is configured. For example, in the industrial quality inspection scenario, after the edge device receives and verifies the encrypted compressed package of the defect detection model, the auxiliary container automatically calls the decryption tool to remove the encryption layer of the file with the ".suffix1" suffix, and then restores it to a loadable model weight file and configuration file through the decompression tool. Subsequently, the model inference service is started according to the environment variable configuration to form a quality inspection model instance that can process image data in real time. The generation of this model depends on the automated processing chain after the transmission is completed, including the execution of the preprocessing script triggered by the Hook task, file format conversion, and resource ready status synchronization.

[0086] Through the hierarchical processing of the model transmission status and automated task chain by the end-side device, the system realizes closed-loop management of the entire process from data packet reception, integrity verification to model service readiness. The dynamic generation mechanism of the target data processing model, combined with the collaboration of Hook tasks and multi-tool chains, significantly reduces the reliance of model deployment on manual operations. The refined diagnosis and recovery suggestions for transmission errors effectively shorten the troubleshooting cycle and ensure service continuity in distributed model deployment scenarios. This deployment mode based on state-driven and automated scripts improves the model iteration efficiency and system robustness of end-side devices in complex network environments.

[0087] Further, a target data processing model is generated based on each target model data packet, including: Obtain data type information corresponding to each target model data packet; obtain at least one target model deployment file based on the data type information corresponding to each target model data packet; and generate a target data processing model according to each target model deployment file.

[0088] In actual applications, the data type information is a set of file attributes and processing requirement identifiers of the model data package; the target model deployment file is a standardized file that can be directly used for model deployment after decryption, decompression and format conversion.

[0089] The data type information can be understood as a set of processing rules that are determined based on the file suffix, metadata, and configuration parameters of the model data packet after the transmission is completed. For example, when the model data packet received by the end-side device contains a file with the ".suffix1" suffix, the system identifies the file as an encrypted type based on predefined environment variables (such as whether automatic decryption is enabled), and then triggers the decryption tool to perform key matching and content restoration. In the specific implementation, the data type information also includes a compression identifier (such as the ".suffix2" suffix), a version compatibility identifier, and a list of associated dependencies to guide the auxiliary container to select the corresponding decompression tool or dependent library loading strategy. In the intelligent voice processing scenario, if it is detected that the data packet contains a combination of configuration files and weight files, the input format requirements of the model loader are automatically matched.

[0090] The target model deployment file can be understood as a standardized file that can be directly loaded into the inference engine after the original model data packet is converted and preprocessed. For example, in the industrial visual inspection scenario, after the encrypted compressed package is decrypted into a model file in the ".suffix2" format, the auxiliary container calls the runtime environment according to the framework identifier in the data type information to convert the file into a memory-mappable binary stream. During the generation of this file, the system decides whether to overwrite the existing local version based on preset parameters, and completes operations such as verification code comparison and dependency injection through the file processing tool chain. The final generated deployment file sends a ready signal to the main application container through the Hook task, triggering the model service startup process.

[0091] By dynamically parsing the type attributes of model data packets and performing targeted processing, the system achieves automatic adaptation and deployment of heterogeneous model files. The precise identification mechanism of data type information enables special format files such as encryption and compression to complete format conversion without human intervention, and the standardized generation process of target model deployment files ensures seamless loading of different framework models on end-side devices. This type-driven file processing logic significantly reduces the risk of deployment failure due to format mismatch, and at the same time improves the processing efficiency and reliability in distributed model update scenarios through the linkage control of environment variables and Hook tasks.

[0092] Considering that in actual applications, data packet transmission may be caused by various reasons, after generating the model transmission result based on each target model data packet, the method further includes: In the case where the model transmission result is a transmission error, current environment information and data transmission log information are acquired, and an error analysis result is generated according to the current environment information and the data transmission log information.

[0093] In actual applications, the error analysis result is a set of fault diagnosis and recovery strategies generated by the end-side device when the model data packet transmission fails; the error analysis result can be understood as a comprehensive fault report generated by recording the environment snapshot and executing the exception handling script when the end-side device detects an anomaly during the model data packet transmission process. For example, when an edge device downloads an image recognition model and the shard is lost due to network interruption, the auxiliary container will capture the number of concurrent threads of the current task, the MD5 checksum status of the received shards, and the file lock occupancy, and trigger the task of obtaining the model file error to generate an analysis report containing the error type (such as network interruption), failed shard index, and recommended retransmission times. The result is fed back to the operation and maintenance platform, and the decryption tool log path is associated with the transmission task timestamp, so that technicians can quickly locate specific problems such as failure to decrypt encrypted files or damage to compressed packages.

[0094] The system significantly improves the fault tolerance of the model deployment process through the automatic diagnosis and recovery strategy generation mechanism of transmission errors by the end-side devices. The environment snapshots and script execution logs integrated in the error analysis results enable operation and maintenance personnel to accurately identify the root causes of failures such as network fluctuations, file corruption or configuration conflicts without having to manually reproduce abnormal scenarios.

[0095] In actual applications, the current environment information is a snapshot set of system operating parameters and resource configuration status during the model transmission process; the data transmission log information is a set of operation records and status change trajectories generated during the model data packet transmission process.

[0096] The current environment information can be understood as the real-time operating parameters and resource configuration snapshots captured by the end-side device during the execution of the model transmission task. For example, when an industrial quality inspection system is interrupted in transmission when deploying a defect detection model, the auxiliary container will record the number of concurrent threads, memory usage, file lock holding status, and environment variables of the current task (such as the upper limit of concurrent tasks defined by preset parameters). In the specific implementation, this information covers key indicators such as the process ID of the transmission task, the validity period of the decryption key, and the remaining space of the temporary storage path. In the autonomous driving model update scenario, if the model loading fails due to insufficient GPU video memory, the environmental information will include a video memory usage distribution map and a container resource quota limit value, providing a diagnostic basis for error analysis at the hardware resource configuration level.

[0097] Data transmission log information can be understood as the status record and abnormal event sequence of each operation node during the transmission of the model data packet. For example, when the intelligent customer service system downloads the semantic understanding model, the log will record the shard reception progress, MD5 verification results, network reconnection times, and decryption tool call status by timestamp. When it is detected that the hash value of a shard is inconsistent with the cloud record, the shard index, locally calculated hash value, and failed verification time point will be marked in the log. In the medical image analysis scenario, if the encrypted model file fails to be decrypted due to an expired key version, the log will associate the key identifier, decryption tool version number, and stack trace information when the error is triggered to form a complete abnormal link traceability chain.

[0098] It should be noted that generating error analysis results based on current environmental information and data transmission log information can be understood as locating the root cause of the exception and forming an executable repair strategy by associating the system operation status and operation records of the transmission process. The specific method can be to associate the key validity period in the environment snapshot with the decryption failure record in the log to determine the configuration error caused by the unsynchronized key rotation (for example, when the encrypted file decryption fails, the system records the key version number in the environmental information and the decryption tool call sequence in the log to generate an error report containing key update instructions); it can also combine the fragment transmission interruption timestamp in the log and the network bandwidth fluctuation data in the environmental information to analyze the impact of network stability on the transmission success rate; it can also identify processing exceptions caused by resource overload by matching the memory usage peak record with the decompression operation failure node in the log, and this manual does not impose any restrictions on this.

[0099] By integrating the correlation analysis of real-time environment snapshots and transmission process logs, the system can accurately locate the root cause of model transmission anomalies. The current environment information provides the system context when the failure occurs, such as resource contention status or configuration parameter abnormalities, while the data transmission log restores the complete timing and state changes of the transmission operation. The error analysis results generated by the combination of the two can clearly distinguish different types of failures such as network fluctuations, file corruption or configuration conflicts. For example, in the industrial Internet of Things scenario, if the log shows that the shard is received completely but the decryption fails, combined with the key validity period data in the environment information, it can be quickly determined that it is a configuration error caused by the unsynchronized key rotation. This diagnostic mechanism based on multi-dimensional data linkage significantly improves the fault recovery efficiency and system robustness in distributed model deployment scenarios.

[0100] Taking into account that the task of applying the model may modify the file corresponding to the model, and the modified model needs to be uploaded to the cloud for updating, after generating the target data processing model based on each target model data packet, the method also includes: based on the target data processing model, obtaining at least one model update data packet; obtaining a model update request according to each model update data packet, and sending the model update request to the cloud device.

[0101] In the model deployment scenario, when the end-side device runs the target data processing model, the model file may be dynamically adjusted due to task requirements (such as parameter optimization or structural fine-tuning during the inference process). If there is no effective update feedback mechanism, the cloud model version will be out of sync with the actual application version on the end-side, which will affect the accuracy of subsequent model iterations and multi-device collaboration. To this end, after completing the model deployment, the system needs to establish a two-way synchronization channel between the end and the cloud to ensure that model updates can be fed back to the cloud in a timely manner.

[0102] In the specific implementation, during the model operation, the end-side device continuously tracks the modification behavior of the model file (such as weight parameter update, configuration file adjustment, etc.) through the monitoring module, and generates an incremental model update data packet based on the preset difference comparison rules. For example, in the industrial quality inspection scenario, when the edge device optimizes the defect detection threshold according to the real-time data of the production line, the system calls the file processing tool to extract the difference fragments between the modified configuration file and the historical version, and combines the compression tool to generate an update package containing the version identifier. Subsequently, the upload process is triggered by the predefined startup acquisition model file task: first, it is determined whether encryption is enabled based on the environment variable (such as calling the encryption tool with key matching for the ".suffix1" suffix file), and then the update package is uploaded to the cloud specified storage path through the file transfer pool according to the number of concurrent tasks limited by the preset parameters. After the upload is completed, the main application container generates a model update request containing the update package hash value, the end-side device number and the timestamp, and sends it to the cloud device via a secure channel, triggering the synchronization update operation of the cloud model version library.

[0103] refer to Figure 3 , Figure 3A schematic diagram of the processing process of model deployment in a terminal device provided for an embodiment of the present specification, wherein the key steps in the target model deployment process and the organizational relationship of the tool modules are shown, and the efficient deployment and management of the model are realized through the collaboration of multiple tasks and tools. First, the target model deployment request triggers the start of the task of obtaining the model file, which calls the file transfer pool to extract the model file from the database service, and securely processes the file through the encryption and decryption tool according to the requirements. If the file is in a compressed format, the decompression tool is called to restore it to a suitable file structure. Next, the task of obtaining the model file further integrates and optimizes the file with the help of shared volumes and other tools, and confirms whether the file loading environment meets the model operation requirements in combination with the environment detection tool. In the end of the task of obtaining the model file, all the generated processing results are integrated into the basic files required for the target model deployment, providing support for the gradual completion of the target model task. If an error occurs in the process of obtaining the model file, the task of obtaining the model file error report is used to generate an error analysis result for the subsequent timely processing of the model deployment task. Finally, when the target model is deployed, the system returns an end or exit signal to ensure the smooth completion of the task process, while providing conditional support for possible subsequent updates and the use of script tools to achieve task requirements. Each tool module and task in the whole process is closely linked to each other, forming a complete model deployment workflow.

[0104] This step achieves reverse synchronization of the optimization results on the end to the cloud through closed-loop model update management, providing a data foundation for scenarios such as multi-device collaborative reasoning and model federated learning. The differentiated update package generation mechanism reduces the occupancy of network bandwidth, while the upload trigger logic and encrypted transmission strategy based on the Hook task improves the update efficiency while ensuring the security of data transmission. In addition, the introduction of version identification and hash verification mechanism effectively avoids version conflicts caused by network retransmission, ensuring the accuracy and traceability of cloud model iteration.

[0105] The above is a schematic scheme of a model deployment method applied to a terminal device of this embodiment. It should be noted that the technical scheme of the model deployment method applied to the terminal device and the technical scheme of the model deployment method applied to the cloud device belong to the same concept. For details not described in detail in the technical scheme of the model deployment method applied to the terminal device, please refer to the description of the technical scheme of the model deployment method applied to the cloud device.

[0106] By applying the solution of the embodiments of this specification, the reliability and iteration efficiency of the model deployment process are significantly improved through localized transmission status management and intelligent feedback mechanism. After receiving the target model data packet distributed by the cloud, by generating the model transmission results in real time and performing integrity verification, it is possible to quickly identify transmission abnormal scenarios, combine environmental information with intelligent analysis of data transmission logs, accurately locate the root causes of faults such as network fluctuations or data packet damage, and shorten the troubleshooting cycle. In response to the operation and maintenance needs after model deployment, by automatically generating model update data packets and reversely synchronizing them to cloud devices, a closed-loop update link is formed, effectively avoiding service interruptions caused by model version differences. In addition, the technical feature of dynamically generating target model deployment files based on data type information enables the terminal-side device to adaptively parse multi-format model data, reducing the risk of secondary deployment caused by insufficient data adaptability, thereby achieving efficient implementation and stable operation of model services in a multi-terminal heterogeneous environment.

[0107] See also Figure 4 , Figure 4 An architecture diagram of a model deployment system provided by an embodiment of the present specification is shown. The model deployment system may include a client 100 and a server 200. The server 200 includes at least one model information, and the model information includes at least one model data package. The client 100 is used to send at least one model deployment request to the server 200; The server 200 is used to receive each model deployment request; determine at least one target model data packet corresponding to each model deployment request; and send each target model data packet to the client 100; The client 100 is also used to receive at least one target model data packet corresponding to the target model deployment request sent by the server 200; generate a model transmission result based on each target model data packet, wherein the target model deployment request is any one of the model deployment requests; and when the model transmission result is that the transmission is completed, deploy the target data processing model based on each target model data packet.

[0108] The solution of the embodiment of this specification is applied, through the collaborative processing mechanism between the client and the server, to strengthen the end-to-end transmission quality monitoring and exception handling capabilities, and improve the success rate of model deployment in complex network environments. The client generates transmission results in real time based on the received model data packet and triggers the integrity verification process. When a transmission anomaly is detected, it automatically collects environmental information and transmission logs for multi-dimensional fault analysis, quickly locates the causes of typical problems such as network interruption or data packet damage, and reduces the cost of manual intervention. In response to the continuous iteration requirements after model deployment, the system generates model update data packets locally on the client and synchronizes them to the server in reverse, forming a two-way data flow channel to ensure dynamic matching of model versions and service requests. In addition, the client adaptively constructs the data processing model based on the model data packet type, effectively solving the compatibility issues of multi-source heterogeneous data formats, and reducing the probability of repeated deployment due to differences in environmental adaptation, thereby achieving efficient delivery and stable operation and maintenance of model services in cross-platform, multi-terminal application scenarios.

[0109] The model deployment system may include multiple clients 100 and a server 200, wherein the client 100 may be referred to as a client-side device and the server 200 may be referred to as a cloud-side device. Multiple clients 100 may establish a communication connection through the server 200. In the model deployment scenario, the server 200 is used to provide model deployment services between multiple clients 100. Multiple clients 100 may serve as a sender or a receiver, respectively, and achieve communication through the server 200.

[0110] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the model deployment scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 generates a target model data packet based on the data stream and pushes the target model data packet to the server 200. Among them, the client 100 and the server 200 are connected through the network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links or optical fiber cables, etc. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, etc. before being published to the server 200.

[0111] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language Version 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 100 can be based on the software development kit (SDK, Software Development Kit) of the corresponding service provided by the server 200, such as based on the real-time communication (RTC, Real Time Communication) SDK development and acquisition. The client 100 can be deployed in an electronic device and needs to rely on the device to run or some APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0112] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers for background training that support models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0113] The following combination Figure 5 , taking the application of the model deployment method provided in this specification in the local training of the cloud model as an example, the model deployment method is further explained. Figure 5 A processing flow chart of a cloud model local training method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0114] Step 502: The terminal device determines the configuration information corresponding to the model deployment request based on the model training task.

[0115] Step 504: Generate a model deployment request according to the configuration information corresponding to the model deployment request, and send the model deployment request to the cloud-side device.

[0116] Step 506: The cloud-side device receives the above-mentioned model deployment request and determines the corresponding multiple initial model data packets.

[0117] Step 508: When the current status information corresponding to each initial model data packet is an idle state or a reading state, the initial model data packet is sent to the above-mentioned terminal side device as a target model data packet.

[0118] Step 510: The terminal device processes each target model data packet according to the type information corresponding to each received target model data packet, and obtains the target model deployment file.

[0119] Step 512: Generate a target data processing model based on each target model deployment file.

[0120] Step 514: Train the target data processing model, obtain an updated data processing model, and obtain multiple model update data packets based on the updated data processing model.

[0121] Step 516: Pack each model update data packet into a model update request, and send the model update request to the cloud-side device.

[0122] Step 518: The cloud-side device receives the model update request, and determines the corresponding reference model data packet according to each model update data packet in the model update request.

[0123] Step 520: When the current state information corresponding to the reference model data packet is in an idle state, the reference model data packet is updated according to each model update data packet, and the current state information of the reference model data packet is set to a write state.

[0124] Step 522: After each reference model data package is updated, the current state information of each reference model data package is set to an idle state.

[0125] By applying the scheme of the embodiments of this specification, the dynamic configuration and state synchronization strategy in the end-cloud collaborative training process is used to achieve a dual improvement in model iteration efficiency and data integrity. The end-side device generates dynamic configuration information based on the training task requirements and triggers a model deployment request, so that the cloud side can accurately match the basic model data packet that adapts to the current training stage, avoiding the invalid transmission of redundant model versions. In the model update stage, the end-side device reversely synchronizes the updated data packet after training to the cloud side, and performs version coverage in an orderly manner according to the state marking mechanism of the reference model data packet to form a closed-loop iteration link, effectively preventing the risk of model degradation caused by the confusion of training data versions. In view of the type characteristics of the model data packet, the end-side device enhances the data compatibility under the heterogeneous training framework through differentiated parsing and deployment file generation logic, and reduces the computing resource loss caused by format conversion. At the same time, the cloud-side device dynamically controls the read and write permissions of the model data packet based on the state information, while ensuring data consistency in multi-end concurrent training scenarios, shortening the response delay of model version switching, thereby improving the overall resource utilization and iteration stability of distributed training tasks.

[0126] Corresponding to the above method embodiment, this specification also provides a model deployment device embodiment applied to a cloud device, Figure 6 FIG. 1 is a schematic diagram showing a structure of a model deployment device applied to a cloud device provided by an embodiment of the present specification. Figure 6 As shown, the device stores at least one model information, and the model information includes at least one model data package, including: The receiving module 602 is configured to receive at least two model deployment requests sent by at least one end-side device; A determination module 604 is configured to determine at least one target model data package corresponding to each model deployment request; The sending module 606 is configured to send each target model data packet to the end-side device corresponding to each target model data packet, so that each end-side device completes the model deployment.

[0127] Optionally, the determination module 604 is further configured to: determine at least one initial model data packet corresponding to each model deployment request; obtain current status information of each initial model data packet; and determine the target model data packet corresponding to each model deployment request based on the current status information corresponding to each initial model data packet.

[0128] Optionally, the determining module 604 is further configured to: Determine a reference model deployment request in each model deployment request, and determine a first initial model data packet in each initial model data packet corresponding to the reference model deployment request, wherein the reference model deployment request is any one of the model deployment requests, and the first initial model data packet is any one of the initial model data packets corresponding to the reference model deployment request; When the current state information corresponding to the first initial model data packet is in an idle state, determining that the first initial model data packet is a target model data packet corresponding to the reference model deployment request, and setting the current state information corresponding to the first initial model data packet to a read state; When the current state information corresponding to the first initial model data packet is a read state, determining that the first initial model data packet is a target model data packet corresponding to the reference model deployment request; When the current status information corresponding to the first initial model data packet is a write status, a target waiting time is determined, and based on the target waiting time, it is determined whether the first initial model data packet is a target model data packet corresponding to the reference model deployment request.

[0129] Optionally, the model deployment device applied to the cloud device also includes a release module, which is configured to: determine that the current state information corresponding to each target model data packet is an idle state.

[0130] Optionally, the model deployment device applied to the cloud device also includes an update module, which is configured to: receive at least two model update requests sent by at least one end-side device, wherein the model update request includes at least one model update data packet; determine the reference model data packet corresponding to each model update data packet, and update each reference model data packet based on each model update data packet.

[0131] Optionally, the update module is further configured to: determine that the current status information corresponding to each reference model data packet is a write status; the model deployment device applied to the cloud device also includes a write module, which is configured to: determine that the current status information corresponding to each reference model data packet is an idle status.

[0132] The above is a schematic scheme of a model deployment device applied to a cloud device in this embodiment. It should be noted that the technical scheme of the model deployment device applied to the cloud device and the technical scheme of the model deployment method applied to the cloud device belong to the same concept. For details not described in detail in the technical scheme of the model deployment device applied to the cloud device, please refer to the description of the technical scheme of the model deployment method applied to the cloud device.

[0133] The scheme of the embodiment of this specification is applied to realize efficient management and precise scheduling of model resources in high-concurrency scenarios through modular architecture and state collaborative control mechanism. The parallel design of the receiving module and the sending module supports instant response and data packet distribution of multi-channel deployment requests. Combined with the dynamic screening logic based on the model data packet status information in the determination module, it can intelligently avoid read-write conflicts when multiple terminals request concurrent access, and ensure the orderly execution of file transfer and model update. Through the linkage mechanism of the state marking and release module, the state information is automatically reset after the data packet transmission is completed, reducing the idle loss caused by resource locking and improving the recycling rate of cloud storage resources. In addition, the update module strictly guarantees the integrity and consistency of the data version coverage process when receiving the model update request fed back by the terminal through the write state control and reference model data packet matching mechanism, avoiding the risk of data pollution that may occur during multi-terminal collaborative iteration, thereby maintaining the high availability of the model service and the stability of update iteration in a complex deployment environment.

[0134] Corresponding to the above method embodiment, this specification also provides a model deployment device embodiment applied to a terminal device. Figure 7 FIG. 1 shows a schematic diagram of a model deployment device applied to a terminal device provided by an embodiment of the present specification. Figure 7 As shown, the device comprises: A request sending module 702 is configured to generate at least one model deployment request and send each model deployment request to a cloud device; The receiving module 704 is configured to determine a target model deployment request and receive at least one target model data packet corresponding to the target model deployment request, wherein the target model deployment request is any one of the model deployment requests, and each target model data packet is sent by the cloud device; A generating module 706 is configured to generate a model transmission result based on each target model data packet; The deployment module 708 is configured to deploy the target data processing model based on each target model data packet when the model transmission result is transmission completion.

[0135] Optionally, the deployment module 708 is further configured to: obtain data type information corresponding to each target model data packet; obtain at least one target model deployment file based on the data type information corresponding to each target model data packet; and generate a target data processing model according to each target model deployment file.

[0136] Optionally, the model deployment device applied to the terminal side device also includes an error reporting module, which is configured to: when the model transmission result is a transmission error, obtain current environment information and data transmission log information, and generate an error analysis result based on the current environment information and the data transmission log information.

[0137] Optionally, the model deployment device applied to the terminal device also includes an update module, which is configured to: obtain at least one model update data packet based on the target data processing model; obtain a model update request according to each model update data packet, and send the model update request to the cloud device.

[0138] The above is a schematic scheme of a model deployment device applied to a terminal device of this embodiment. It should be noted that the technical scheme of the model deployment device applied to the terminal device and the technical scheme of the model deployment method applied to the terminal device belong to the same concept. For details not described in detail in the technical scheme of the model deployment device applied to the terminal device, please refer to the description of the technical scheme of the model deployment method applied to the terminal device.

[0139] The scheme of the embodiment of this specification is applied, and the fault tolerance and localized execution efficiency of the model deployment process are significantly enhanced through modular division of labor and adaptive processing mechanism. After the receiving module captures the target model data packet distributed by the cloud in real time, the generating module triggers the differentiated processing flow through the status judgment of the transmission result. When the transmission is completed, the deployment module intelligently parses and generates a model deployment file adapted to the current operating environment based on the data type information, effectively solving the compatibility problem between multi-format model data and the end-side framework, and reducing the secondary deployment cost caused by data format conflicts. The error reporting module accurately identifies the root causes of abnormal scenarios such as network jitter and missing data packets by integrating the multi-dimensional analysis capabilities of environmental information and transmission logs, shortening the fault diagnosis cycle and improving the pertinence of error repair. In addition, the update module automatically generates an update request after local model training and feeds it back to the cloud, forming a closed-loop feedback link of end-cloud bidirectional collaboration, avoiding the problem of inconsistent model services caused by version iteration delays, thereby achieving stable connection between model deployment and dynamic updates under complex network conditions, and improving the autonomous operation and maintenance capabilities and resource utilization of end-side devices.

[0140] Figure 8 The structure block diagram of a computing device 800 provided according to an embodiment of the present specification is shown. The components of the computing device 800 include but are not limited to a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and the database 850 is used to store data.

[0141] The computing device 800 also includes an access device 840 that enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0142] In one embodiment of the present specification, the above components of the computing device 800 and Figure 8 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 8 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0143] The computing device 800 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 800 may also be a mobile or stationary server.

[0144] Among them, the processor 820 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned model deployment method.

[0145] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned model deployment method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above-mentioned model deployment method.

[0146] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned model deployment method.

[0147] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned model deployment method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned model deployment method.

[0148] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned model deployment method when executed by a processor.

[0149] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the above-mentioned model deployment method belong to the same concept, and the details not described in detail in the technical scheme of the computer program can be referred to the description of the technical scheme of the above-mentioned model deployment method.

[0150] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0151] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0152] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0153] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0154] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A model deployment method, applied to a cloud device, wherein the cloud device includes at least one model information, the model information includes at least one model data package, and the method includes: Receiving at least two model deployment requests sent by at least one end-side device; Determine at least one target model data package corresponding to each model deployment request; Each target model data packet is sent to the end-side device corresponding to each target model data packet, so that each end-side device completes the model deployment.

2. The method of claim 1, determining at least one target model data package corresponding to each model deployment request, comprising: Determine at least one initial model data packet corresponding to each model deployment request; Obtain the current status information of each initial model data package; Based on the current status information corresponding to each initial model data package, the target model data package corresponding to each model deployment request is determined.

3. The method of claim 2, determining the target model data package corresponding to each model deployment request based on the current state information corresponding to each initial model data package, comprising: Determine a reference model deployment request in each model deployment request, and determine a first initial model data packet in each initial model data packet corresponding to the reference model deployment request, wherein the reference model deployment request is any one of the model deployment requests, and the first initial model data packet is any one of the initial model data packets corresponding to the reference model deployment request; When the current state information corresponding to the first initial model data packet is in an idle state, determining that the first initial model data packet is a target model data packet corresponding to the reference model deployment request, and setting the current state information corresponding to the first initial model data packet to a read state; When the current state information corresponding to the first initial model data packet is a read state, determining that the first initial model data packet is a target model data packet corresponding to the reference model deployment request; When the current status information corresponding to the first initial model data packet is a write status, a target waiting time is determined, and based on the target waiting time, it is judged whether the first initial model data packet is the target model data packet corresponding to the reference model deployment request, wherein the target waiting time is a preset waiting time threshold for obtaining the model data packet.

4. The method according to claim 2, after sending each target model data packet to the terminal side device corresponding to each target model data packet, the method further comprises: Determine that the current state information corresponding to each target model data packet is an idle state.

5. The method according to any one of claims 1 to 4, after sending each target model data packet to the terminal side device corresponding to each target model data packet, the method further comprises: Receiving at least two model update requests sent by at least one end-side device, wherein the model update request includes at least one model update data packet; A reference model data packet corresponding to each model update data packet is determined in each model data packet, and each reference model data packet is updated based on each model update data packet.

6. The method according to claim 5, after determining the reference model data packet corresponding to each model update data packet, the method further comprises: Determine that the current state information corresponding to each reference model data packet is a write state; Correspondingly, after each model update data packet updates each reference model data packet, the method further includes: It is determined that the current state information corresponding to each reference model data packet is an idle state.

7. A model deployment method, applied to a terminal device, comprising: generating at least one model deployment request, and sending each model deployment request to a cloud device; Determine a target model deployment request, and receive at least one target model data packet corresponding to the target model deployment request, wherein the target model deployment request is any one of the model deployment requests, and each target model data packet is sent by the cloud device; Generate model transmission results based on each target model data packet; When the model transmission result is transmission completion, the target data processing model is deployed based on each target model data packet.

8. The method of claim 7, generating a target data processing model based on each target model data packet, comprising: Obtain the data type information corresponding to each target model data packet; Based on the data type information corresponding to each target model data packet, obtaining at least one target model deployment file; Deploy the target data processing model according to each target model deployment file.

9. The method according to claim 7, after generating the model transmission result based on each target model data packet, the method further comprises: In the case where the model transmission result is a transmission error, current environment information and data transmission log information are acquired, and an error analysis result is generated according to the current environment information and the data transmission log information.

10. The method according to claim 7, after generating the target data processing model based on each target model data packet, the method further comprises: Based on the target data processing model, obtaining at least one model update data packet; A model update request is obtained according to each model update data packet, and the model update request is sent to the cloud device.

11. A model deployment system, the system comprising a cloud device and at least one end-side device, wherein: The cloud device includes at least one model information, and the model information includes at least one model data packet; A target-side device, configured to send at least one model deployment request to the cloud device; The cloud device is configured to receive each model deployment request; determine at least one target model data packet corresponding to each model deployment request; and send each target model data packet to the end-side device corresponding to each target model data packet; The target end-side device is further configured to receive at least one target model data packet corresponding to a target model deployment request; generate a model transmission result based on each target model data packet, wherein the target model deployment request is any one of the model deployment requests; When the model transmission result is transmission completion, the target data processing model is deployed based on each target model data packet.

12. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Dispatching method and system of distributed system

    CN101753608A

  • Model deployment method and device, storage medium and electronic equipment

    CN111240698A

  • Data model construction system based on three-dimensional structure

    CN115830220A

  • Data matching device and data matching method

    CN118277460A

  • Space-based microtask application system and method giving consideration to monitoring of various ecological environments

    CN119477182A

Cited By

  • Cloud platform tenant modeling test method and device, equipment and storage medium

    CN116599881A

  • Method, device and equipment for cloud platform tenant modeling test and storage medium

    CN116599881B

  • Cross-domain large file transmission method based on distributed soft bus

    CN120281765A

  • AI reasoning optimization method and system of dynamic model switching framework for edge device

    CN120354954A