Model dynamic publishing method based on TensorFlow
By deploying TensorFlow Serving, model configuration is generated and model is published, and the model deployment is monitored based on the duration and data volume, and the relevant parameters are adjusted, the problem of model deployment exceptions in the existing technology is solved, and the model generation efficiency and stability are improved.
Patent Information
- Application Number
- CN202510748243.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing technology does not consider monitoring the deployment of the model, and the generation parameters of the model are not adjusted when the model deployment exception is detected, which affects the generation efficiency of the model.
By deploying TensorFlow Serving, the model configuration is generated and the model is published, and the model is published is determined based on the time and data amount of publishing the model, and the deployment parameters such as Docker container memory, TCP window or compatibility detection cycle are adjusted when determining deployment exceptions.
It improves the generation efficiency of the model, provides a stable operating environment, reasonably evaluates the deployment of the model, promptly discovers and solves exceptions, and optimizes data transmission and compatibility issues.
Smart Images

Figure CN120256028A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model publishing, and in particular, to a method for dynamically publishing a model based on TensorFlow. Background Art
[0002] At present, machine learning technology is becoming more and more mature. As an open-source machine learning platform, TensorFlow provides functions such as data preprocessing and model training.
[0003] In actual machine learning and deep learning applications, it is often necessary to dynamically publish the trained TensorFlow model for calling and using in different environments.
[0004] How to uniformly deploy the trained model has become an indispensable part.
[0005] Chinese Patent Application Publication No.: CN112799710A discloses a model publishing system and a model publishing method, including a model generation device, a model service forwarding device, and a model publishing and displaying device. The model generation device generates a model to be published according to a design strategy corresponding to a preset operating environment selected by an engineer under the preset operating environment, and sends the generated model to be published to the model service forwarding device; the model service forwarding device determines the publishing status corresponding to the received model to be published, and sends the determined publishing status corresponding to the model to be published to the model publishing and displaying device. When the model publishing and displaying device receives the corresponding publishing status, it displays the model to be published according to the corresponding publishing status. It can be seen that the above technical solution has the following problems: it does not consider monitoring the deployment situation of the model, and does not consider adjusting the generation parameters of the model when detecting abnormal model deployment, which affects the generation efficiency of the model. Summary of the Invention
[0006] Therefore, the present invention provides a method for dynamically publishing a model based on TensorFlow to overcome the problems in the prior art that do not consider monitoring the deployment situation of the model, do not consider adjusting the generation parameters of the model when detecting abnormal model deployment, and affect the generation efficiency of the model.
[0007] To achieve the above object, the present invention provides a method for dynamically publishing a model based on TensorFlow, including: Deploy TensorFlow Serving; Deploy the model; Generate a model configuration and publish the model; Determine whether the model publishing is qualified based on the duration of generating the published model, including: The deployment of the determination model is abnormal, and the deployment parameters of the model are adjusted based on the duration deviation amount, where adjusting the deployment parameters of the model includes adjusting the allocated memory of the Docker container to the corresponding value, adjusting the TCP window to the corresponding value, or adjusting the compatibility detection period for compatibility check to the corresponding value; Or, it is determined that the deployment of the model is qualified, and an API address and a parameter template are generated for the deployed model.
[0008] Further, the process of determining whether the deployment of the model is qualified based on the monitored release duration of the model includes, when the release duration is greater than the second preset release duration, determining that the deployment of the model is abnormal, and adjusting the deployment parameters of the model based on the duration deviation amount.
[0009] Further, the process of determining whether the deployment of the model is qualified based on the monitored release duration of the model includes, when the release duration is less than or equal to the second preset release duration and greater than the first preset release duration, determining whether the deployment of the model is qualified based on the data volume of the model file, and when the data volume is greater than the preset data volume, adjusting the first preset release duration to the corresponding value based on the data volume, where the increase amplitude of the first preset release duration is proportional to the data volume.
[0010] Further, the process of determining whether the deployment of the model is qualified based on the data volume of the model file further includes, when the data volume is less than or equal to the preset data volume, determining that the deployment of the model is abnormal, and adjusting the deployment parameters of the model based on the duration deviation amount.
[0011] Further, the process of adjusting the deployment parameters of the model based on the duration deviation amount includes: Determining the difference between the calculated release duration and the second preset release duration as the duration deviation amount; When the duration deviation amount is less than or equal to the first preset duration deviation amount, adjusting the allocated memory of the Docker container to the corresponding value based on the model generation frequency of the server; When the duration deviation amount is less than or equal to the second preset duration deviation amount and greater than the first preset duration deviation amount, adjusting the TCP window to the corresponding value based on the current data transmission rate; When the duration deviation amount is greater than the second preset duration deviation amount, adjusting the compatibility detection period for compatibility check to the corresponding value based on the abnormal frequency.
[0012] Further, adjusting the allocated memory of the Docker container to the corresponding value based on the model generation frequency of the server, where The increase amplitude of the allocated memory is proportional to the model generation frequency.
[0013] Further, adjusting the TCP window to the corresponding value based on the current data transmission rate, where The increase amplitude of the TCP window is inversely proportional to the data transmission rate.
[0014] Further, after completing the adjustment of the TCP window, the compression ratio of the model file is adjusted to the corresponding value based on the difference between the adjusted TCP window and the initial TCP window, where the difference between the adjusted TCP window and the initial TCP window is denoted as the window difference; the increase amplitude of the compression ratio is proportional to the window difference.
[0015] Further, based on the abnormal frequency, the compatibility detection period for compatibility check is adjusted to the corresponding value, where the decrease amplitude of the compatibility detection period is proportional to the abnormal frequency.
[0016] Further, when completing the adjustment of the compatibility detection period, the release duration is re-obtained, and when the re-obtained release duration is greater than the first preset release duration, the configuration file of TensorFlow Serving is updated.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: deploy TensorFlow Serving; deploy the model; generate a model configuration and publish the model; determine whether the release of the model is qualified based on the duration of generating the model. When it is determined that the deployment of the model is abnormal, the deployment parameters of the model are adjusted based on the duration deviation amount. The deployment parameters include adjusting the allocated memory of the Docker container to the corresponding value, adjusting the TCP window to the corresponding value, or adjusting the compatibility detection period for compatibility check to the corresponding value; or, determine that the deployment of the model is qualified, and generate an API address and a parameter template for the deployed model; monitor the deployment situation of the model, and when it is detected that the model deployment is abnormal, adjust the generation parameters of the model, improving the generation efficiency of the model.
[0018] Furthermore, deploy the TensorFlow Serving model to provide a stable running environment. Place the trained model in the environment managed by TensorFlow Serving, generate a model configuration file, and inform TensorFlow Serving of information such as the location, name, and platform of the model so that the model can be correctly loaded and managed. Determine whether the model release is qualified based on the duration of generating the model; when the release duration is less than or equal to the first preset release duration, since operations such as model loading and initialization can be quickly completed within the current time range, in this case, resources are sufficient, and the compatibility between the model itself and the environment is good, it is determined that the model deployment is qualified. When the release duration is less than or equal to the second preset release duration and greater than the first preset release duration, further judge whether the deployment is qualified based on the data volume of the model file. The data volume size will affect the model loading and initialization time, so a more detailed evaluation needs to be combined with the data volume. When the data volume is less than or equal to the preset data volume, if the release duration is still long, in this case, there are other abnormal factors, and it is determined that the deployment is abnormal. When the data volume is greater than the preset data volume, since a large data volume will cause the deployment time to extend, in this case, adjust the first preset release duration based on the data volume. Consider the impact of the data volume on the deployment duration and reasonably evaluate the model deployment situation. While improving the accuracy of model evaluation, the model generation efficiency is further improved.
[0019] Furthermore, adjust the deployment parameters based on the duration deviation amount, which reflects the degree to which the model deployment exceeds the normal time. According to the magnitude of the duration deviation amount, judge the severity of the problem and adopt different adjustment strategies. When the duration deviation amount is less than or equal to the first preset duration deviation amount, in this case, the model deployment time is too long due to insufficient server resources. Adjust the allocated memory of the Docker container based on the model generation frequency of the server; the model generation frequency of the server characterizes the busyness of the server. When the model generation frequency is high, the server requires more resources to process the generation and deployment of new models. By increasing the allocated memory of the Docker container, improve the processing capacity of the server to handle high-frequency model generation tasks. When the duration deviation amount is less than or equal to the second preset duration deviation amount and greater than the first preset duration deviation amount, in this case, the model deployment time is too long due to network transmission problems, and the data transmission rate is low due to network congestion or unreasonable TCP window settings. In this case, adjust the TCP window based on the current data transmission rate, and the current data transmission rate characterizes the network transmission capacity. By adjusting the TCP window size, optimize the data transmission efficiency and improve the transmission speed of the model file. When the duration deviation amount is greater than the second preset duration deviation amount, in this case, the model deployment is abnormal due to compatibility problems, resulting in an excessive model release duration. In this case, by reducing the compatibility detection period, perform compatibility checks more frequently, promptly discover and solve compatibility problems, and further improve the model generation efficiency. Description of the Drawings
[0020] Figure 1 It is a flowchart of the steps of the method for dynamically releasing a model based on TensorFlow according to an embodiment of the present invention; Figure 2 It is a logical decision diagram for determining whether the deployment of a model based on the release duration of the monitored model is qualified according to an embodiment of the present invention; Figure 3 It is a logical decision diagram for determining whether the deployment of a model is qualified based on the data volume of the model file according to an embodiment of the present invention; Figure 4 It is a logical decision diagram for adjusting the deployment parameters of a model based on the duration deviation amount according to an embodiment of the present invention. Detailed Embodiments
[0021] In order to make the purpose and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0022] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0023] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention.
[0024] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0025] Please refer to Figure 1 , Figure 2 , Figure 3 and Figure 4 as shown, which are respectively the step flow chart of the model dynamic publishing method based on TensorFlow in the embodiments of the present invention, the logical decision chart for determining whether the deployment of the model is qualified based on the publishing duration of the monitored model, the logical decision chart for determining whether the deployment of the model is qualified based on the data volume of the model file, and the logical decision chart for adjusting the deployment parameters of the model based on the duration deviation amount; An embodiment of the present invention provides a model dynamic publishing method based on TensorFlow, including: S1, deploy TensorFlow Serving; S2, deploy the model; S3, generate a model configuration and publish the model; S4, determine whether the publishing of the model is qualified based on the duration of generating the model by publishing, including: Determine that the deployment of the model is abnormal, and adjust the deployment parameters of the model based on the duration deviation amount. The deployment parameters include adjusting the allocated memory of the Docker container to the corresponding value, adjusting the TCP window to the corresponding value, or adjusting the compatibility detection period for compatibility check to the corresponding value; Or, determine that the deployment of the model is qualified, and generate an API address and a parameter template for the deployed model.
[0026] Specifically, during the process of releasing a model, the model to be released can be selected on the page, and its deployment configuration file can be dynamically generated based on Java.
[0027] Specifically, the code is as follows: model_config_list { config { name: "mymodels" / / Specify the name of the model. This name is used to uniquely identify the model in the TensorFlow Serving service. When the client sends an inference request to TensorFlow Serving, it needs to use this name to specify the model to be called.
[0028] base_path: " / root / python_soft / models / mymodel" / / Specify the base storage path of the model files. TensorFlow Serving will look for model files in this path. The model files will be organized by version number. There can be multiple subdirectories under the base path, and each subdirectory corresponds to a model version.
[0029] model_platform: "tensorflow" / / Specify the platform type used by the model. Here it is set to "tensorflow", indicating that the model is trained using the TensorFlow framework. TensorFlow Serving will correctly load and process the model based on this information.
[0030] model_version_policy{ specific { versions: 1 versions: 2} / / Specify the loading policy for model versions to determine the model versions that TensorFlow Serving should load. Use the specific policy to specify the specific version numbers.
[0031] / / Under the specific version policy, clearly specify that the model version numbers to be loaded are 1 and 2. When the client requests inference, it can choose to use either of these two versions.
[0032] } } } Specifically, in the step S1, for the deployment of TensorFlow Serving, local deployment based on Docker can be used, which includes installing the Docker engine, then pulling the official image of TensorFlow Serving from Docker Hub, creating a Docker container, mounting the local model file inside the container, and specifying the port mapping and environment variables when the container starts; optionally, in this embodiment, the memory allocated to the Docker container is 3GB.
[0033] Specifically, TensorFlow Serving can encapsulate the model trained using the TensorFlow framework into a form that can provide services externally and deploy the model to the production environment.
[0034] Specifically, Docker containers achieve isolation, and the applications running inside the Docker containers are isolated from the host machine and do not interfere with each other.
[0035] Specifically, in the step S1, the compatibility between TensorFlow Serving and the running environment is periodically detected to ensure the stable generation and release of the model.
[0036] Specifically, for the deployment of the model in the step S2, it can include saving the trained TensorFlow model in the SavedModel format and placing it in a directory accessible by TensorFlow Serving. Ensure that the input and output formats of the model meet the requirements of TensorFlow Serving.
[0037] Specifically, in the process of generating the model configuration and publishing the model in the step S3, it can include creating a model configuration file, specifying the name, version, and path of the model. Pass the model configuration file to the TensorFlow Serving service, start the service, and load the model.
[0038] Specifically, deploy TensorFlow Serving; deploy the model; generate a model configuration and publish the model; determine whether the model publication is qualified based on the duration of generating the model. When it is determined that the model deployment is abnormal, adjust the model deployment parameters based on the duration deviation amount. The deployment parameters include adjusting the allocated memory of the Docker container to the corresponding value, adjusting the TCP window to the corresponding value, or adjusting the compatibility detection period for compatibility check to the corresponding value; or, determine that the model deployment is qualified and generate an API address and parameter template for the deployed model; monitor the model deployment situation, and when it is detected that the model deployment is abnormal, adjust the model generation parameters, improving the model generation efficiency.
[0039] Specifically, in step S4, the process of determining whether the model deployment is qualified based on the monitored model publication duration includes: If the publication duration is less than or equal to the first preset publication duration, it is determined that the model deployment is qualified, and an API address and parameter template are generated for the deployed model; If the publication duration is less than or equal to the second preset publication duration and greater than the first preset publication duration, determine whether the model deployment is qualified based on the data volume of the model file; If the publication duration is greater than the second preset publication duration, it is determined that the model deployment is abnormal, and the model deployment parameters are adjusted based on the duration deviation amount.
[0040] Specifically, the first preset publication duration T1 is selected within the interval [55s, 65s], and the second preset publication duration T2 is selected within the interval [80s, 86s].
[0041] It should be noted that the data in this embodiment are obtained through comprehensive analysis and evaluation of the historical detection data and corresponding historical detection results in the four months before this detection. The present invention determines the numerical values of the preset parameter standards for this detection based on 2,857 actual models cumulatively deployed in the previous four months, the deployment time corresponding to each actual model, the actual abnormal conditions of each actual model, the data transmission rate, and the file size of the model before this detection. Those skilled in the art can understand that the determination method for each single parameter of the present invention can be to select the value with the highest proportion according to the data distribution as the preset standard parameter, use weighted summation to obtain the value as the preset standard parameter, or other selection methods, as long as it satisfies that the present solution can clearly define different specific situations in the single-item determination process through the obtained values.
[0042] Specifically, through statistics and analysis, it is determined that the deployment time of most models with simple to medium complexity in an ideal environment is concentrated around 55 - 65 seconds. Considering various abnormal situations encountered in actual deployment and through multiple tests and experience summaries, the interval of 80 - 86 seconds is determined to be able to better define these situations. It is understandable that those skilled in the art can set each preset value according to the actual implementation situation.
[0043] Specifically, the duration from the start of the recorded TensorFlow Serving service to the model release is recorded as the model release duration.
[0044] Specifically, the process of determining whether the deployment of a model is qualified based on the data volume of the model file includes: If the data volume is less than or equal to the preset data volume, it is determined that the model deployment is abnormal, and the model deployment parameters are adjusted based on the duration deviation. If the data volume is greater than the preset data volume, the first preset release duration is adjusted to the corresponding value based on the data volume.
[0045] Specifically, the preset data volume S0 is selected within the interval [100MB, 140MB].
[0046] Specifically, the data volume of the model file can be viewed by using a file management tool or command to check the file size, or by finding the folder where the model file is located in the resource manager to determine the size of the model file. This is prior art and will not be elaborated further.
[0047] Specifically, according to the statistical machine learning model scale, the file size of small to medium models is between 100 - 140MB. Therefore, in this embodiment, the interval of S0 is selected as the threshold for judging the influence of data volume. Specifically, deploying the TensorFlow Serving model provides a stable running environment. Place the trained model into the environment managed by TensorFlow Serving, generate a model configuration file, and inform TensorFlow Serving of information such as the location, name, and platform of the model, so that the model can be correctly loaded and managed. Determine whether the model release is qualified based on the duration of generating the model; when the release duration is less than or equal to the first preset release duration, since operations such as model loading and initialization can be quickly completed within the current time range, in this case, resources are sufficient, and the compatibility between the model itself and the environment is good, it is determined that the model deployment is qualified. When the release duration is less than or equal to the second preset release duration and greater than the first preset release duration, further determine whether the deployment is qualified based on the data volume of the model file. The data volume affects the model loading and initialization time, so a more detailed evaluation needs to be combined with the data volume. When the data volume is less than or equal to the preset data volume, if the release duration is still long, there are other abnormal factors in this case, and it is determined that the deployment is abnormal. When the data volume is greater than the preset data volume, since the large data volume will cause the deployment time to extend, in this case, adjust the first preset release duration based on the data volume. Considering the impact of the data volume on the deployment duration, reasonably evaluate the model deployment situation. While improving the accuracy of model evaluation, further improve the model generation efficiency.
[0048] Specifically, the data volume corresponds to the corresponding value of the first preset release duration adjustment value, where the growth rate of the first preset release duration is proportional to the data volume.
[0049] In this embodiment, optionally, Compare the data volume with the first preset data volume comparison threshold and the second preset data volume comparison threshold; If the data volume is less than or equal to the first preset data volume comparison threshold, adjust the first preset release duration to 1.1 times the initial first preset release duration; If the data volume is less than or equal to the second preset data volume comparison threshold and greater than the first preset data volume comparison threshold, adjust the first preset release duration to 1.2 times the initial first preset release duration; If the data volume is greater than the second preset data volume comparison threshold, adjust the first preset release duration to 1.28 times the initial first preset release duration; The first preset data volume comparison threshold is taken as 1.6S0, and the second preset data volume comparison threshold is taken as 2.7S0.
[0050] Specifically, after completing the adjustment for the first preset release duration, determine again whether the deployment of the model is qualified based on the release duration. If the release duration is less than or equal to the adjusted first preset release duration, it is determined that the deployment of the model is qualified, and an API address and parameter template are generated for the deployed model; if the release duration is greater than the adjusted first preset release duration, it is determined that the deployment of the model is abnormal, and the deployment parameters of the model are adjusted based on the duration deviation.
[0051] Specifically, the process of adjusting the deployment parameters of the model based on the duration deviation includes: Record the difference between the calculated release duration and the second preset release duration as the duration deviation. If the duration deviation is less than or equal to the first preset duration deviation, adjust the allocated memory of the Docker container to the corresponding value based on the model generation frequency of the server. If the duration deviation is less than or equal to the second preset duration deviation and greater than the first preset duration deviation, adjust the TCP window to the corresponding value based on the current data transmission rate. If the duration deviation is greater than the second preset duration deviation, adjust the compatibility detection period for compatibility check to the corresponding value based on the abnormal frequency.
[0052] Specifically, the model generation frequency of the server is the ratio of the number of new models generated within a preset monitoring period to the preset monitoring period. Specifically, the abnormal frequency is the ratio of the number of deployments of the model determined to be abnormal within a preset evaluation period to the preset evaluation period.
[0053] Specifically, the current data transmission rate can obtain the traffic information of the network interface using the psutil library to calculate the data transmission rate.
[0054] Specifically, the first preset duration deviation is taken as 1.3T2, and the second preset duration deviation is taken as 2.7T2.
[0055] Specifically, through the study of the relationship between server resource usage and model deployment duration, it is statistically found that when the deviation is within 1.3T2, adjusting the memory allocation can effectively improve the deployment situation. After a large number of network performance tests and model deployment data analyses, it is found that when the deviation is in the range of 1.3T2 - 2.7T2, network transmission problems are the main cause of deployment delays, and adjusting the TCP window can effectively improve the situation.
[0056] Specifically, adjust the allocated memory of the Docker container to the corresponding value based on the model generation frequency of the server, where The increase amplitude of the allocated memory is proportional to the model generation frequency.
[0057] Optionally, in this embodiment, Compare the model generation frequency with the first preset generation frequency and the second preset generation frequency; If the model generation frequency is less than or equal to the first preset generation frequency, adjust the allocated memory of the Docker container to 1.11 times the initial allocated memory; If the model generation frequency is less than or equal to the second preset generation frequency and greater than the first preset generation frequency, adjust the allocated memory of the Docker container to 1.23 times the initial allocated memory; If the model generation frequency is greater than the second preset generation frequency, adjust the allocated memory of the Docker container to 1.31 times the initial allocated memory; The preset monitoring period is 3600s, the first preset generation frequency is 0.01, and the second preset generation frequency is 0.05.
[0058] Specifically, adjust the TCP window to the corresponding value based on the current data transmission rate, where The increase amplitude of the TCP window is inversely proportional to the data transmission rate.
[0059] Optionally, in this embodiment, Compare the data transmission rate with the first preset transmission rate and the second preset transmission rate; If the data transmission rate is less than or equal to the first preset transmission rate, adjust the TCP window to 1.28 times the initial TCP window; If the data transmission rate is less than or equal to the second preset transmission rate and greater than the first preset transmission rate, adjust the TCP window to 1.21 times the initial TCP window; If the data transmission rate is greater than the second preset transmission rate, adjust the TCP window to 1.13 times the initial TCP window; The first preset transmission rate is 1MB / s, and the second preset transmission rate is 0.015MB / s.
[0060] Optionally, in this embodiment, the initial TCP window is 1024.
[0061] Specifically, in the scenario where the model is deployed on TensorFlow Serving, the adjusted TCP window is the TCP window of the network socket of the server where TensorFlow Serving is located. When the client sends a request to the TensorFlow Serving server and the server returns a response to the client, the data transmission is based on the TCP protocol. Each TCP connection has a sending window and a receiving window, which respectively control the amount of data that the sender can send and the amount of data that the receiver can receive. The TCP window adjusted in this solution is the window size of the corresponding TCP connection when the server processes the client request.
[0062] After adjusting the TCP window, the compression ratio of the model file is adjusted to the corresponding value based on the difference between the adjusted TCP window and the initial TCP window, where The difference between the adjusted TCP window and the initial TCP window is denoted as the window difference; The increase in the compression ratio is proportional to the window difference.
[0063] In this embodiment, optionally, The window difference is compared with the first preset window difference and the second preset window difference; If the window difference is less than or equal to the first preset window difference, the compression ratio of the model file is adjusted to 1.13 times the initial compression ratio; If the window difference is less than or equal to the second preset window difference and greater than the first preset window difference, the compression ratio of the model file is adjusted to 1.21 times the initial compression ratio; If the window difference is greater than the second preset window difference, the compression ratio of the model file is adjusted to 1.29 times the initial compression ratio; The first preset window difference is 200B, and the second preset window difference is 300.
[0064] Specifically, the compression ratio is the ratio of the size of the original model file to the size of the compressed model file.
[0065] Specifically, the compatibility detection period for compatibility check is adjusted to the corresponding value based on the exception frequency, where The decrease in the compatibility detection period is proportional to the exception frequency.
[0066] In this embodiment, optionally, The exception frequency is compared with the first preset exception frequency and the second preset exception frequency. If the exception frequency is less than or equal to the first preset exception frequency, the compatibility detection period is adjusted to 0.92 times the initial compatibility detection period; If the abnormal frequency is less than or equal to the second preset abnormal frequency and greater than the first preset abnormal frequency, adjust the compatibility detection period to 0.84 times the initial compatibility detection period; If the abnormal frequency is greater than the second preset abnormal frequency, adjust the compatibility detection period to 0.72 times the initial compatibility detection period; The preset evaluation period is 86400s, the first preset abnormal frequency is 0.007, and the second preset abnormal frequency is 0.012.
[0067] Specifically, when the adjustment of the compatibility detection period is completed, optimize the compatibility, including using a compatibility layer to perform conversion for the model.
[0068] Specifically, when the adjustment of the compatibility detection period is completed, re-obtain the release duration, and when the re-obtained release duration is greater than the first preset release duration, update the configuration file of TensorFlow Serving.
[0069] Specifically, updating the configuration file of TensorFlow Serving includes modifying the configuration path, and after modifying the configuration file, restart the TensorFlow Serving service to make the new configuration take effect.
[0070] Specifically, the deployment parameters of the duration deviation adjustment model are adjusted. The duration deviation reflects the degree to which the model deployment exceeds the normal time. According to the magnitude of the duration deviation, the severity of the problem is judged, and different adjustment strategies are adopted. When the duration deviation is less than or equal to the first preset duration deviation, in this case, the model deployment time is long due to insufficient server resources. The allocated memory of the Docker container is adjusted based on the model generation frequency of the server; the model generation frequency of the server characterizes the busyness of the server. When the model generation frequency is high, the server requires more resources to process the generation and deployment of new models. By increasing the allocated memory of the Docker container, the processing capacity of the server is improved to cope with high-frequency model generation tasks. When the duration deviation is less than or equal to the second preset duration deviation and greater than the first preset duration deviation, in this case, the model deployment time is too long due to network transmission problems. The data transmission rate is low due to network congestion or unreasonable TCP window settings. In this case, the TCP window is adjusted based on the current data transmission rate, and the current data transmission rate characterizes the transmission capacity of the network. By adjusting the size of the TCP window, the data transmission efficiency is optimized, and the transmission speed of the model file is increased. When the duration deviation is greater than the second preset duration deviation, in this case, the model deployment exception due to compatibility problems results in an excessive model release duration. In this case, by reducing the compatibility detection period, compatibility checks are performed more frequently, and compatibility problems are discovered and solved in a timely manner, further improving the model generation efficiency.
[0071] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0072] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for dynamically publishing a model based on TensorFlow, characterized in that, Including: Deploy TensorFlow Serving; Deploy the model; Generate a model configuration and publish the model; Determine whether the model publication is qualified based on the duration of generating the model and the data volume of the model file, including: When it is determined that the model deployment is abnormal, adjust the model deployment parameters based on the duration deviation amount, where adjusting the model deployment parameters includes adjusting the allocated memory of the Docker container to the corresponding value, adjusting the TCP window to the corresponding value, or adjusting the compatibility detection period for compatibility check to the corresponding value; Or, when it is determined that the model deployment is qualified, generate an API address and a parameter template for the deployed model.
2. The method for dynamically publishing a model based on TensorFlow according to claim 1, wherein The process of determining whether the model deployment is qualified based on the monitored publication duration of the model includes, when the publication duration is greater than the second preset publication duration, determining that the model deployment is abnormal and adjusting the model deployment parameters based on the duration deviation amount.
3. The method for dynamically publishing a model based on TensorFlow according to claim 2, wherein The process of determining whether the model deployment is qualified based on the monitored publication duration of the model includes, when the publication duration is less than or equal to the second preset publication duration and greater than the first preset publication duration, determining whether the model deployment is qualified based on the data volume of the model file, and when the data volume is greater than the preset data volume, adjusting the first preset publication duration to the corresponding value based on the data volume, where the increase amplitude of the first preset publication duration is proportional to the data volume.
4. The method for dynamically publishing a model based on TensorFlow according to claim 3, wherein The process of determining whether the model deployment is qualified based on the data volume of the model file further includes, when the data volume is less than or equal to the preset data volume, determining that the model deployment is abnormal and adjusting the model deployment parameters based on the duration deviation amount.
5. The method for dynamically publishing a model based on TensorFlow according to claim 4, wherein The process of adjusting the model deployment parameters based on the duration deviation amount includes: Determine the difference between the calculated publication duration and the second preset publication duration as the duration deviation amount; When the duration deviation amount is less than or equal to the first preset duration deviation amount, adjust the allocated memory of the Docker container to the corresponding value based on the model generation frequency of the server; When the duration deviation amount is less than or equal to the second preset duration deviation amount and greater than the first preset duration deviation amount, adjust the TCP window to the corresponding value based on the current data transmission rate; When the duration deviation amount is greater than the second preset duration deviation amount, adjust the compatibility detection period for compatibility check to the corresponding value based on the exception frequency.
6. The method for dynamically publishing a model based on TensorFlow according to claim 5, wherein Adjust the allocated memory of the Docker container to the corresponding value based on the model generation frequency of the server, where The increase amplitude of the allocated memory is proportional to the model generation frequency.
7. The method for dynamically publishing a model based on TensorFlow according to claim 6, wherein Adjust the TCP window to the corresponding value based on the current data transmission rate, where The increase amplitude of the TCP window is inversely proportional to the data transmission rate.
8. The method for dynamically publishing a model based on TensorFlow according to claim 7, wherein After completing the adjustment of the TCP window, adjust the compression ratio of the model file to the corresponding value based on the difference between the adjusted TCP window and the initial TCP window, where Record the difference between the adjusted TCP window and the initial TCP window as the window difference; The increase amplitude of the compression ratio is proportional to the window difference.
9. The method for dynamically publishing a model based on TensorFlow according to claim 8, wherein Adjust the compatibility detection period for compatibility check to the corresponding value based on the exception frequency, where The decrease amplitude of the compatibility detection period is proportional to the exception frequency.
10. The method for dynamically publishing a model based on TensorFlow according to claim 9, wherein When the adjustment for the compatibility detection period is completed, the release duration is retrieved again, and when the retrieved release duration is greater than the first preset release duration, the configuration file of TensorFlow Serving is updated.
Citation Information
Patent Citations
Docker-based software large-scale-testing method
CN108121654A
Model publishing method and device, model deployment method, equipment and storage medium
CN111580926A
Model online deployment method and device
CN112015519A
Containerized resource management method based on dynamic graph network
CN118550652A
Update anomaly detection method
CN119603015A