Service-based data processing method, apparatus, device, and medium

CN117194027BActive Publication Date: 2026-09-29KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311167322.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-11
Publication Date
2026-09-29
Estimated Expiration
2043-09-11

AI Technical Summary

Benefits of technology

[0011]根据本公开的一个或多个实施例,能够通过处理模块之间的数据依赖关系,从而获取多个数据处理路径;通过测试并统计当前配置下每个路径需要的处理时长,进而确定其中制约服务整体执行效率的关键路径,并针对该关键路径中的性能最低的处理模块进行优化(例如增加其线程数量),从而能够更加有针对性地对服务的整体执行效率进行调整和优化,提升了优化效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117194027B_ABST
    Figure CN117194027B_ABST
Patent Text Reader

Abstract

The present disclosure provides a service-based data processing method and device, equipment and medium, relates to the technical field of artificial intelligence, and in particular to the field of deep learning technology and computer vision. The implementation scheme is: obtaining configuration information and performance parameters of each processing module in a plurality of processing modules under the current configuration; determining a first processing time length of each data processing path in a plurality of data processing paths; determining a first path with the longest first processing time length in the plurality of data processing paths; determining a target processing module on the first path based on the performance parameters of each processing module on the first path; and in response to the device resource occupation condition of the first device meeting a first preset condition, performing a first configuration update operation on the target processing module to perform data processing based on the updated service configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to deep learning technology and computer vision, and specifically to a service-based data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] With the continuous advancement of machine learning technology and the increasing computing power and storage of AI acceleration chips, multiple machine learning models are increasingly being combined in practical applications to solve real-world problems. For example, AI video understanding services require models such as video frame segmentation, image recognition, face recognition, scene recognition, action recognition, optical character recognition (OCR), object recognition, speech recognition, and text recognition based on natural language processing (NLP).

[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0005] This disclosure provides a service-based data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product.

[0006] According to one aspect of this disclosure, a service-based data processing method is provided. The service includes multiple processing modules. The method includes: obtaining configuration information and performance parameters for each of the multiple processing modules under a current configuration; the configuration information includes the number of threads used to execute the corresponding processing module; the performance parameters indicate the data processing efficiency of the corresponding processing module under the current configuration; the multiple processing modules constitute multiple data processing paths connecting the data input and data output ends of the service based on data dependencies between the multiple processing modules; determining a first processing duration for each of the multiple data processing paths; the first processing duration instructs the corresponding data processing path to perform data processing and... The total data transmission time; among multiple data processing paths, the first path with the longest processing time is determined; based on the performance parameters of each processing module on the first path, the target processing module is determined on the first path; and in response to the device resource occupancy of the first device meeting the first preset condition, a first configuration update operation is performed on the target processing module to perform data processing based on the service with the updated configuration, wherein the first device is at least used to run the target processing module, the first preset condition includes that the CPU memory, GPU memory, and device utilization of the first device do not exceed their respective upper limits, and the first configuration update operation includes increasing the number of threads of the target processing module according to the first preset number.

[0007] According to another aspect of this disclosure, a service-based data processing apparatus is provided, wherein the service includes multiple processing modules. The apparatus includes: a first acquisition unit configured to acquire configuration information and performance parameters of each of the multiple processing modules under a current configuration; the configuration information includes the number of threads used to execute the corresponding processing module, and the performance parameters indicate the data processing efficiency of the corresponding processing module under the current configuration; the multiple processing modules constitute multiple data processing paths based on data dependencies between the multiple processing modules; a first determination unit configured to determine a first processing duration for each of the multiple data processing paths; the first processing duration indicates the total time consumed for data processing and data transmission in the corresponding data processing path; and a second determination unit. The system comprises: a first processing unit configured to determine the first path with the longest processing time among multiple data processing paths; a third determining unit configured to determine a target processing module on the first path based on the performance parameters of each processing module on the first path; and a first updating unit configured to perform a first configuration update operation on the target processing module in response to the device resource occupancy of the first device meeting a first preset condition, so as to perform data processing based on the service after configuration update, wherein the first device is at least used to run the target processing module, the first preset condition includes that the CPU memory, GPU memory, and device utilization of the first device do not exceed their respective upper limits, and the first configuration update operation includes increasing the number of threads of the target processing module according to a first preset number.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned service-based data processing method.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described service-based data processing method.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described service-based data processing method when executed by a processor.

[0011] According to one or more embodiments of this disclosure, multiple data processing paths can be obtained by processing the data dependencies between modules; by testing and statistically analyzing the processing time required for each path under the current configuration, the critical path that restricts the overall execution efficiency of the service can be determined, and the processing module with the lowest performance in the critical path can be optimized (e.g., by increasing its number of threads), thereby enabling more targeted adjustment and optimization of the overall execution efficiency of the service and improving optimization efficiency.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;

[0015] Figure 2 A flowchart of a service-based data processing method according to an embodiment of the present disclosure is shown;

[0016] Figure 3 A schematic diagram of a directed acyclic graph of an AI video understanding service according to an exemplary embodiment of the present disclosure is shown;

[0017] Figure 4 A flowchart of a service-based data processing method according to an embodiment of the present disclosure is shown;

[0018] Figure 5 A flowchart of a service performance optimization method according to an exemplary embodiment of the present disclosure is shown;

[0019] Figure 6 A structural block diagram of a service-based data processing apparatus according to an embodiment of the present disclosure is shown;

[0020] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0022] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0023] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0024] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0025] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0026] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the service-based data processing methods described above.

[0027] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.

[0028] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0029] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to request services. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0030] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0031] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0032] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0033] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0034] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.

[0035] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0036] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0037] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0038] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0039] According to some embodiments, such as Figure 2 As shown, a service-based data processing method is provided. The service includes multiple processing modules. The method includes: Step S201: Obtaining the configuration information and performance parameters of each processing module in the current configuration. The configuration information includes the number of threads used to execute the corresponding processing module, and the performance parameters indicate the data processing efficiency of the corresponding processing module in the current configuration. The multiple processing modules constitute multiple data processing paths connecting the data input and data output ends of the service based on the data dependencies between the multiple processing modules; Step S202: Determining the first processing time of each data processing path in the multiple data processing paths. The first processing time indicates the total time consumed by the corresponding data processing path for data processing and data transmission. In step S203, among multiple data processing paths, determine the first path with the longest first processing time; in step S204, based on the performance parameters of each processing module on the first path, determine the target processing module on the first path; and in step S205, in response to the device resource occupancy of the first device meeting the first preset condition, perform a first configuration update operation on the target processing module to perform data processing based on the service after configuration update, wherein the first device is at least used to run the target processing module, the first preset condition includes that the CPU memory, GPU memory, and device utilization of the first device do not exceed their respective upper limits, and the first configuration update operation includes increasing the number of threads of the target processing module according to the first preset number.

[0040] Therefore, by processing the data dependencies between modules, multiple data processing paths can be obtained; by testing and statistically analyzing the processing time required for each path under the current configuration, the critical path that restricts the overall execution efficiency of the service can be identified, and the lowest-performing processing module in the critical path can be optimized (e.g., by increasing its number of threads). This allows for more targeted adjustments and optimizations to the overall execution efficiency of the service, thereby improving optimization efficiency.

[0041] In some embodiments, a service is software that executes on a server to provide a specific function. Some complex services can provide multiple functions, each implemented through a code module. After the developers have written the code modules, they can deploy them to the server, that is, deploy the service on the server. The server can then provide the service to the user. Hereinafter, the code module in the service that provides a specific function will be referred to as a "processing module".

[0042] For example, video surveillance services can be deployed for intelligent transportation scenarios. These services may include, for instance, three processing modules: a vehicle model recognition module, a human detection module, and a human posture recognition module.

[0043] In some embodiments, each processing module may also be an artificial intelligence model with a specific function. For example, an AI video understanding service may consist of models such as video frame segmentation, image recognition, face recognition, scene recognition, action recognition, OCR recognition, object recognition, speech recognition, and NLP text recognition, based on certain data dependencies.

[0044] In some embodiments, multiple processing modules can form multiple data processing paths for connecting the data input and data output ends of a service based on the data dependencies between the multiple processing modules.

[0045] In this context, data dependencies between multiple processing modules are used to represent the execution order and data flow of these modules. According to some embodiments, these data dependencies can be represented using a Directed Acyclic Graph (DAG). A DAG is a directed graph without cycles. DAGs are an effective tool for describing workflows. Using a DAG to represent the data dependencies between multiple processing modules facilitates the configuration and analysis of these dependencies. For example, DAGs can be used to analyze whether the service composed of each processing module can be executed smoothly, estimate the overall response time of the service, and so on.

[0046] In some embodiments, multiple data processing paths in the service can be constructed based on the data dependencies between the aforementioned processing modules.

[0047] Figure 3 A schematic diagram of a directed acyclic graph of an AI video understanding service according to an exemplary embodiment of the present disclosure is shown.

[0048] In some exemplary embodiments, such as Figure 3 As shown, service 300 includes multiple processing modules such as a pre-processing module 301, a video frame extraction module 302, an audio recognition module 303, a face recognition module 304, a scene recognition module 305, an action recognition module 306, an object recognition module 307, an NLP text recognition module 308, an OCR text recognition module 309, a sentiment classification module 310, and a scoring module 311. Furthermore, the data dependencies between these processing modules are represented by edges in a directed acyclic graph, where each edge connects the data sender and receiver, and its arrow indicates the direction of data transmission.

[0049] like Figure 3As shown, upon receiving a processing request, the data to be processed (e.g., a video clip) is first sent to the pre-processing module 301. After pre-processing, the pre-processing module 301 distributes the processed data to the video frame extraction module 302 and the audio recognition module 303. The video frame extraction module 302 further sends its processed data to the face recognition module 304, scene recognition module 305, action recognition module 306, and object recognition module 307 for different processing. Similarly, the audio recognition module 303 further sends its processed data to the NLP text recognition module 308, OCR text recognition module 309, and sentiment classification module 310 for different processing. Finally, the scoring module 311 receives the processing results from each upstream module and performs a comprehensive score after receiving all the processing results to output the final result.

[0050] In some embodiments, a node path connecting the data input and data output ends of a connection service can be used as a data processing path.

[0051] like Figure 3 As shown, service 300 includes 7 processing paths, namely path 31, path 32, path 33, path 34, path 35, path 36 and path 37.

[0052] In some embodiments, for services with more complex processing logic, such as one or more intermediate nodes in its directed acyclic graph, each of which needs to process data based on the processing results of multiple upstream modules, the intermediate node can also be regarded as a data output end, and its corresponding data processing path can be determined based on the data output end for the intermediate node to perform subsequent configuration update operations.

[0053] In some embodiments, for such a more complex service, the following configuration update operations can be performed first on multiple data processing paths of each intermediate node; after the update is completed, the following configuration update operations can continue to be performed on multiple data processing paths of the overall service. This allows for more targeted adjustments and optimizations to critical paths and nodes that affect the overall performance of the service, thereby improving the efficiency of the optimization operation while ensuring the optimization effect.

[0054] In some embodiments, the multiple processing modules in the above service may be deployed on the same device. In some embodiments, the multiple processing modules in the above service may be deployed on multiple device nodes on a distributed cluster.

[0055] In some embodiments, the configuration information for each processing module may include the number of threads used to execute the corresponding processing module. The number of threads may be the number of threads included in the thread pool used to execute the corresponding processing module.

[0056] In some embodiments, a service can be initialized with its configuration information, for example, by setting the initial number of threads for each processing module to 1. In some embodiments, a pre-configured service can be obtained, its existing configuration information can be acquired, and subsequent adjustments and optimizations can be made based on this.

[0057] In some embodiments, performance parameters may include parameters that indicate the data processing efficiency of the corresponding processing module under the current configuration.

[0058] In some embodiments, the performance parameter can be the average response time of a single request when the corresponding processing module is executed by a single thread. In some embodiments, the performance parameter can be the number of requests processed per unit time when the corresponding processing module is executed by a single thread, i.e., the absolute throughput as described below.

[0059] The performance parameters of each processing module will differ depending on the configuration. In some embodiments, the performance parameters of each processing module under different configuration parameters can be obtained through prior experiments or through real-time testing, and no limitation is imposed here.

[0060] In some embodiments, the first processing time can be the total time consumed by data processing and data transmission on each of the above data processing paths.

[0061] In some embodiments, the first processing time of each data processing path under the current configuration can be obtained, and the data processing path with the longest first processing time can be determined as the first path, which is the critical path that affects the overall performance of the service under the current configuration.

[0062] In some embodiments, after determining the critical path, the target processing module on the path, i.e. the key node that affects the processing efficiency of the path under the current configuration, can be further determined based on the performance parameters of each processing module on the path.

[0063] In some exemplary embodiments, the processing module with the longest average response time in the first path can be selected as the target processing module.

[0064] In some exemplary embodiments, the processing module with the lowest absolute throughput in the first path can be selected as the target processing module.

[0065] Understandably, relevant technical personnel can also evaluate the processing efficiency and determine the target processing module based on other performance parameters, without any restrictions.

[0066] In some embodiments, after determining the target processing module, it is first necessary to determine whether the device resource usage of the first device used to execute the target processing module meets the first preset condition, that is, whether its central processing unit memory, graphics processing unit memory and device utilization rate do not exceed their respective upper limits.

[0067] In some embodiments, the aforementioned upper limit may be, for example, 100%. It is understood that those skilled in the art can also set the upper limits for the CPU memory, GPU memory, and device utilization rate according to actual needs, and no restrictions are imposed here.

[0068] In some embodiments, device utilization can be at least one of central processing unit utilization and graphics processing unit utilization.

[0069] In some embodiments, when the first device meets the first preset condition, a first configuration update operation can be performed on the target processing module, that is, the number of threads of the target processing module is increased according to the first preset number.

[0070] In some embodiments, the first preset quantity can be 1. It is understood that those skilled in the art can also set the first preset quantity according to actual needs, and no limitation is imposed here.

[0071] Therefore, the updated service, through the above methods, can achieve a more balanced processing efficiency across its various data processing paths, with improved processing efficiency on critical paths. This enables the optimization of the overall service processing performance and enhances the data processing efficiency based on the service.

[0072] In some embodiments, such as Figure 4As shown, the above-mentioned service-based data processing method may further include: step S401, repeatedly executing the following first operation until, before performing the first configuration update operation on the current target processing module, it is detected that the CPU memory and GPU memory of the first device corresponding to the current target processing module both exceed their respective upper limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate upper limit. The first operation includes: step S4011, updating the performance parameters of the target processing module and the first processing time of each data processing path in the multiple data processing paths for the updated service; step S4012, redetermining the first path with the first processing time based on the updated first processing time of each data processing path; step S4013, redetermining the target processing module on the first path based on the updated performance parameters of each processing module on the redetermined first path; and step S4014, in response to the device resource occupancy of the first device corresponding to the current target processing module meeting the first preset condition, performing the first configuration update operation on the current target processing module; and step S405, performing data processing based on the service with the updated configuration.

[0073] Therefore, by repeatedly determining the first path and the target processing module, and optimizing the performance of the target processing module by increasing the number of threads, the overall performance of the service can be further optimized through continuous and targeted iterative optimization.

[0074] In some embodiments, after the above configuration update, the performance parameters of each processing module can be updated accordingly, and a new round of determination of the first path and its target processing unit, checking of the device resource usage of the first device and updating of the target processing unit configuration can be performed. The above operations are repeated until the CPU memory and GPU memory of the first device exceed their respective upper limits, or the device utilization rate exceeds the device utilization rate upper limit.

[0075] In some embodiments, when the entire service is deployed on the same device, the above configuration update operation can be stopped and the current configuration information can be recorded when it is detected that the CPU memory and GPU memory of the device both exceed their respective limits, or the device utilization rate exceeds the device utilization rate limit.

[0076] In some embodiments, multiple processing modules in the service are deployed on multiple device nodes in a distributed cluster. The service-based data processing method described above may further include: responding to the fact that the CPU memory and GPU memory of the first device corresponding to the current target processing module both exceed their respective upper limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate upper limit, and the current first path also includes other processing modules, traversing the other processing modules one by one from low to high according to the performance parameters of each processing module, until a second processing module capable of performing a first configuration update operation or a second configuration update operation is obtained, or it is determined that each processing module on the current first path cannot perform a configuration update operation; and responding to the acquisition of the second processing module, performing a corresponding configuration update operation on the second processing module to perform data processing based on the service with the updated configuration.

[0077] In some embodiments, when multiple processing modules in the service are deployed on different device nodes in a distributed cluster, the processing module with the lowest data processing efficiency among the other processing modules on the current first path can be further selected as the target processing module. The above-mentioned device resource usage check and target processing unit configuration update are performed until the configuration information of a certain processing module on the first path is updated, or it is detected that each processing module on the first path cannot be updated. Then the above-mentioned configuration update operation can be stopped and the current configuration information is recorded.

[0078] Therefore, by deploying different processing modules in the service on different device nodes in a distributed cluster, it is possible to further improve the data processing capabilities of the service, and when the current target processing module on the first path cannot be adjusted by threading, the performance of that path can be optimized by adjusting other processing modules on that path, thereby optimizing the overall performance of the service.

[0079] In practical applications, some processing modules can perform batch data processing. For example, some artificial intelligence models use batch input of training data during the training phase, and then process and output the results in batches.

[0080] For the aforementioned models capable of batch data processing, merging multiple data sets into a single batch and inputting them into the processing module for batch processing can typically improve the data processing performance of that module to some extent.

[0081] In some embodiments, the improvement in data processing performance can specifically manifest as an increase in the absolute throughput of the processing module, or a reduction in its average response time per data item.

[0082] In some embodiments, based on the findings of the above experiments, the data batch size can be used as another configuration information, and the processing efficiency of the processing module can be optimized by further updating this configuration information. The data batch size is used to instruct the corresponding processing module to merge input data according to the data batch size for batch processing of the merged input data.

[0083] In some embodiments, the first operation may further include: in response to the device resource occupancy of the first device corresponding to the current target processing module meeting a second preset condition, performing a second configuration update operation on the current target processing module, wherein the second configuration update operation includes updating the data batch number of the current target processing module according to a preset strategy, and the second preset condition includes that either the central processing unit memory or the graphics processing unit memory of the first device corresponding to the current target processing module does not exceed the upper limit, the device utilization rate does not exceed the upper limit, and the current target processing module supports batch processing of data.

[0084] Therefore, when it is not possible to optimize the performance of the processing module by increasing the number of threads, the overall performance of the service can be further improved by adjusting the number of data batches processed in batches, thereby jointly adjusting the number of threads and the number of data batches.

[0085] In some embodiments, when the number of threads in the target processing module cannot be adjusted, it can be further detected whether the device resource usage of the first device meets a second preset condition. If the second preset condition is met, the data batch size of the current target processing module can be updated according to a preset strategy.

[0086] In some embodiments, a preset value (e.g., 1) can be added to the current data batch size to update the data batch size.

[0087] In some embodiments, the updated batch number can be determined based on the current batch number. For example, the data batch number can be set to a power of 2. When the current batch number is 1, the batch number can be updated to 2; when the current batch number is 4, the batch number can be updated to 8, and so on.

[0088] In some embodiments, due to the different data processing logic of the processing modules, for individual processing modules, increasing the batch size may produce a reverse speedup ratio, that is, when the batch size is increased, the average response time per data item will actually increase.

[0089] In some embodiments, experiments can be conducted beforehand to determine whether each processing module generates a positive speedup when performing different batch number updates. Based on this, when it is determined that a target processing module needs to perform a data batch number update, it can first be determined whether the target processing module will generate a positive speedup after this update. If it can, then the target processing module will perform the data batch number update; otherwise, it will not be updated, and the detection and update operations of other processing modules in the first path will be performed. This further ensures the effectiveness of service performance optimization.

[0090] In some embodiments, updating the current data batch number of the target processing module according to a preset strategy may include: determining a first data batch number based on the current data batch number of the target processing module; determining whether the overall processing time of the updated service exceeds a preset processing time threshold after the current data batch number of the target processing module is updated to the first data batch number; and updating the current data batch number of the target processing module to the first data batch number in response to the overall processing time not exceeding the preset processing time threshold.

[0091] In practical applications, when the processing module performs batch data processing, because the data is input and output in batches, this often leads to a situation where, although the average response time per data point decreases, the total waiting time for batch data from input to output increases compared to the processing time for a single data point. In this case, while it can improve the overall performance of the service, it also increases the waiting time for some users.

[0092] Therefore, to ensure a good user experience, a maximum waiting time for the entire service (i.e., the preset processing time threshold below) can be set as another constraint. When it is determined that a target processing module needs to perform a batch data update, it can first be determined whether the total waiting time of the entire service after the target processing module performs this update will exceed the preset processing time threshold. If it does not exceed the threshold, then the target processing module will perform the batch data update; otherwise, it will not be updated, and the detection and update operations of other processing modules in the first path will be performed.

[0093] Therefore, by setting a preset processing time threshold, it is possible to avoid excessive waiting time for individual users due to overall performance optimization, thereby improving overall performance while taking into account the processing time of individual requests and the experience of each user.

[0094] In some embodiments, updating the data batch number of the current target processing module according to a preset strategy may further include: in response to the overall processing time exceeding a preset processing time threshold, and other processing modules being included on the current first path, traversing the other processing modules one by one from low to high according to the performance parameters of each processing module, until a first processing module capable of updating the data batch number is obtained, or it is determined that each processing module on the current first path cannot update the data batch number; and in response to obtaining the first processing module, updating the data batch number of the first processing module to a second data batch number, wherein the second data batch number is determined based on the data batch number of the first processing module before the update.

[0095] Therefore, if the current target processing unit cannot adjust the batch size, the performance of the path can be further optimized by adjusting other processing modules on the first path, thereby optimizing the overall performance of the service.

[0096] In some embodiments, when multiple processing modules of a service are deployed on the same device, it can be directly determined whether other processing modules on the first path can further adjust the batch size.

[0097] In some embodiments, when multiple processing modules of a service are deployed on different device nodes in a distributed cluster, after determining that the current target processing module cannot adjust the batch number, it can first determine whether the module with the lowest processing efficiency among the other processing modules on the first path can adjust the number of threads. If not, it can then be further determined whether it can adjust the data batch number.

[0098] In some embodiments, when it is determined that each processing module on the current first path is unable to perform any configuration update operation, the above configuration update operation can be stopped and the current configuration information can be recorded.

[0099] In some embodiments, the service can be deployed and launched based on the configuration information recorded above, thereby achieving optimal performance of the service based on its optimal configuration.

[0100] Figure 5 A flowchart of a service performance optimization method according to an exemplary embodiment of the present disclosure is shown.

[0101] In some exemplary embodiments, such as Figure 5As shown, the service performance optimization method may include: Step S501, determining the target processing module; Step S502, detecting whether the CPU memory of the first device corresponding to the target processing module has not exceeded its upper limit; Step S503, detecting whether the GPU memory of the first device has not exceeded its upper limit; Step S504, detecting whether the device utilization rate of the first device has not exceeded its upper limit; Step S505, in response to the fact that the CPU memory, GPU memory, and device utilization rate of the first device have not exceeded their respective upper limits, updating the number of threads of the target processing module; Step S506, in response to the fact that either the CPU memory or GPU memory of the first device has not exceeded its upper limit, and the device utilization rate has not exceeded its upper limit, updating the data batch number of the target processing module; Step S507, in response to the configuration information of the target processing module being updated, updating the performance parameter ranking of each processing module and the device resource usage of each device node; Step S508, repeating the above operations until the CPU memory and GPU memory of the first device both exceed their upper limits, or the device utilization rate exceeds its upper limit, ending the above operations, and recording the current configuration information of each processing module.

[0102] In some embodiments, such as Figure 6 As shown, a service-based data processing apparatus 600 is provided. The service includes multiple processing modules, and the apparatus 600 includes:

[0103] The first acquisition unit 610 is configured to acquire the configuration information and performance parameters of each of the multiple processing modules under the current configuration. The configuration information includes the number of threads used to execute the corresponding processing module, and the performance parameters are used to indicate the data processing efficiency of the corresponding processing module under the current configuration. The multiple processing modules constitute multiple data processing paths based on the data dependencies between the multiple processing modules.

[0104] The first determining unit 620 is configured to determine a first processing duration for each of a plurality of data processing paths, wherein the first processing duration indicates the total time spent on data processing and data transmission for the corresponding data processing path.

[0105] The second determining unit 630 is configured to determine the first path with the longest first processing time among multiple data processing paths;

[0106] The third determining unit 640 is configured to determine a target processing module on the first path based on the performance parameters of each processing module on the first path; and

[0107] The first update unit 650 is configured to perform a first configuration update operation on the target processing module in response to the device resource occupancy of the first device meeting a first preset condition, so as to perform data processing based on the service after the configuration update. The first device is at least used to run the target processing module. The first preset condition includes that the CPU memory, GPU memory, and device utilization of the first device do not exceed their respective upper limits. The first configuration update operation includes increasing the number of threads of the target processing module according to a first preset number.

[0108] The operations performed by units 610 to 650 in the service-based data processing apparatus 600 are similar to the operations of steps S201 to S205 of the service-based data processing method, and will not be described in detail here.

[0109] In some embodiments, the data processing apparatus for the above-mentioned service may further include: a first execution unit configured to repeatedly execute the first operation of the following subunits until, before performing a first configuration update operation on the current target processing module, it is detected that the central processing unit memory and graphics processing unit memory of the first device corresponding to the current target processing module both exceed their respective upper limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate upper limit. The first execution unit includes: a first update subunit configured to update the performance parameters of the target processing module and the first processing time of each of the multiple data processing paths for the updated service; a first determination subunit configured to redetermine the first path with the first processing time based on the updated first processing time of each data processing path; a second determination subunit configured to redetermine the target processing module on the first path based on the updated performance parameters of each processing module on the redetermined first path; and a second update subunit configured to perform a first configuration update operation on the current target processing module in response to the device resource occupancy of the first device corresponding to the current target processing module meeting a first preset condition; and a data processing unit configured to perform data processing based on the service with the updated configuration.

[0110] In some embodiments, the configuration information further includes a data batch number, which is used to enable the corresponding processing module to merge input data according to the data batch number to perform batch processing on the merged input data. The first execution unit may further include a third update subunit configured to perform a second configuration update operation on the current target processing module in response to the device resource occupancy of the first device corresponding to the current target processing module meeting a second preset condition. The second configuration update operation includes updating the data batch number of the current target processing module according to a preset strategy. The second preset condition includes that either the CPU memory or GPU memory of the first device corresponding to the current target processing module does not exceed an upper limit, the device utilization rate does not exceed an upper limit, and the current target processing module supports batch data processing.

[0111] In some embodiments, updating the current data batch number of the target processing module according to a preset strategy may include: determining a first data batch number based on the current data batch number of the target processing module; determining whether the overall processing time of the updated service exceeds a preset processing time threshold after the current data batch number of the target processing module is updated to the first data batch number; and updating the current data batch number of the target processing module to the first data batch number in response to the overall processing time not exceeding the preset processing time threshold.

[0112] In some embodiments, updating the data batch number of the current target processing module according to a preset strategy may further include: in response to the overall processing time exceeding a preset processing time threshold, and other processing modules being included on the current first path, traversing the other processing modules one by one from low to high according to the performance parameters of each processing module, until a first processing module capable of updating the data batch number is obtained, or it is determined that each processing module on the current first path cannot update the data batch number; and in response to obtaining the first processing module, updating the data batch number of the first processing module to a second data batch number, wherein the second data batch number is determined based on the data batch number of the first processing module before the update.

[0113] In some embodiments, multiple processing modules in a service can be deployed on multiple device nodes in a distributed cluster. The service-based data processing apparatus may further include: a second acquisition unit configured to, in response to the following conditions: the CPU memory and GPU memory of the first device corresponding to the current target processing module both exceed their respective upper limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate upper limit, and the current first path also includes other processing modules, traverse the other processing modules one by one from low to high according to the performance parameters of each processing module until a second processing module capable of performing a first configuration update operation or a second configuration update operation is acquired, or it is determined that each processing module on the current first path cannot perform a configuration update operation; and a second update unit configured to, in response to the acquisition of the second processing module, perform a corresponding configuration update operation on the second processing module to perform data processing based on the service with the updated configuration.

[0114] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0115] refer to Figure 7 The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0116] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0117] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 707 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0118] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the service-based data processing methods described above. For example, in some embodiments, the service-based data processing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the service-based data processing method can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the service-based data processing method by any other suitable means (e.g., by means of firmware).

[0119] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0120] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0123] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0124] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0125] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0126] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A service-based data processing method, wherein the service includes multiple processing modules, the method comprising: The configuration information and performance parameters of each of the multiple processing modules under the current configuration are obtained. The configuration information includes the number of threads used to execute the corresponding processing module. The performance parameters are used to indicate the data processing efficiency of the corresponding processing module under the current configuration. The multiple processing modules form multiple data processing paths that connect the data input and data output ends of the service based on the data dependencies between the multiple processing modules. Determine a first processing duration for each of the plurality of data processing paths, wherein the first processing duration indicates the total time spent on data processing and data transmission for the corresponding data processing path; Among the multiple data processing paths, the first path with the longest first processing time is determined; Based on the performance parameters of each processing module on the first path, the target processing module is determined on the first path; as well as In response to the first device's resource usage meeting a first preset condition, a first configuration update operation is performed on the target processing module to perform data processing based on the updated service. The first device is at least used to run the target processing module. The first preset condition includes that the CPU memory, GPU memory, and device utilization of the first device do not exceed their respective upper limits. The first configuration update operation includes increasing the number of threads in the target processing module by a first preset amount.

2. The method according to claim 1, further comprising: Repeat the following first operation until, before performing the first configuration update operation on the current target processing module, it is detected that the CPU memory and GPU memory of the first device corresponding to the current target processing module both exceed their respective limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate limit. The first operation includes: For the service after the configuration update, update the performance parameters of the target processing module and the first processing time of each data processing path in the plurality of data processing paths; Based on the updated first processing time of each data processing path, the first path with the longest first processing time is re-determined. Based on the updated performance parameters of each processing module on the redefined first path, the target processing module on that first path is redefined; and In response to the fact that the device resource usage of the first device corresponding to the current target processing module meets the first preset condition, the first configuration update operation is performed on the current target processing module; and Data processing is performed based on the service with the updated configuration.

3. The method according to claim 2, wherein, The configuration information also includes a data batch size, which is used to enable the corresponding processing module to merge input data according to the data batch size in order to perform batch processing on the merged input data. Furthermore, the first operation further includes: In response to the fact that the device resource usage of the first device corresponding to the current target processing module meets the second preset condition, a second configuration update operation is performed on the current target processing module. The second configuration update operation includes updating the data batch number of the current target processing module according to a preset strategy. The second preset condition includes that either the central processing unit memory or the graphics processing unit memory of the first device corresponding to the current target processing module does not exceed the upper limit, the device utilization rate does not exceed the upper limit, and the current target processing module supports batch processing of data.

4. The method according to claim 3, wherein, The step of updating the current data batch size of the target processing module according to the preset strategy includes: Determine the first data batch size based on the current data batch size of the target processing module; After determining whether the overall processing time of the updated service exceeds a preset processing time threshold, the data batch size of the current target processing module is updated to the first data batch size; and In response to the overall processing time not exceeding the preset processing time threshold, the current data batch number of the target processing module is updated to the first data batch number.

5. The method according to claim 4, wherein, The step of updating the current data batch size of the target processing module according to the preset strategy also includes: In response to the overall processing time exceeding the preset processing time threshold, and the current first path also includes other processing modules, the other processing modules are traversed one by one from low to high according to the performance parameters of each processing module until a first processing module capable of performing data batch number updates is obtained, or it is determined that each processing module on the current first path is unable to perform data batch number updates; and In response to the acquisition of the first processing module, the data batch number of the first processing module is updated to a second data batch number, wherein the second data batch number is determined based on the data batch number of the first processing module before the update.

6. The method according to any one of claims 3 to 5, wherein, The multiple processing modules in the service are deployed on multiple device nodes in a distributed cluster, and the method further includes: In response to the following: the CPU memory and GPU memory of the first device corresponding to the current target processing module both exceed their respective limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate limit, and the current first path also includes other processing modules, the other processing modules are traversed one by one from low to high according to the performance parameters of each processing module until a second processing module capable of performing the first configuration update operation or the second configuration update operation is obtained, or it is determined that each processing module on the current first path cannot perform a configuration update operation; and In response to obtaining the second processing module, a corresponding configuration update operation is performed on the second processing module to perform data processing based on the service after the configuration update.

7. A service-based data processing apparatus, wherein the service includes multiple processing modules, the apparatus comprising: The first acquisition unit is configured to acquire configuration information and performance parameters of each of the plurality of processing modules under the current configuration. The configuration information includes the number of threads used to execute the corresponding processing module, and the performance parameters are used to indicate the data processing efficiency of the corresponding processing module under the current configuration. The plurality of processing modules constitute multiple data processing paths based on the data dependencies between the plurality of processing modules. The first determining unit is configured to determine a first processing duration for each of the plurality of data processing paths, wherein the first processing duration indicates the total time spent on data processing and data transmission for the corresponding data processing path. The second determining unit is configured to determine the first path with the longest first processing time among the plurality of data processing paths; The third determining unit is configured to determine the target processing module on the first path based on the performance parameters of each processing module on the first path. as well as The first update unit is configured to perform a first configuration update operation on the target processing module in response to the device resource usage of the first device meeting a first preset condition, so as to perform data processing based on the service after the configuration update. The first device is at least used to run the target processing module. The first preset condition includes that the CPU memory, GPU memory, and device utilization of the first device do not exceed their respective upper limits. The first configuration update operation includes increasing the number of threads of the target processing module by a first preset number.

8. The apparatus according to claim 7, further comprising: A first execution unit is configured to repeatedly execute the first operation of the following subunit until, before performing the first configuration update operation on the current target processing module, it is detected that the CPU memory and GPU memory of the first device corresponding to the current target processing module both exceed their respective upper limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate upper limit. The first execution unit includes: The first update subunit is configured to update the performance parameters of the target processing module and the first processing time of each of the plurality of data processing paths for the service after the update configuration. The first determining subunit is configured to redetermine the first path with the longest first processing time based on the first processing time of each updated data processing path. The second determining subunit is configured to redetermine the target processing module on the first path based on the updated performance parameters of each processing module on the redetermined first path; and The second update subunit is configured to perform the first configuration update operation on the current target processing module in response to the device resource occupancy status of the first device corresponding to the current target processing module meeting the first preset condition; and The data processing unit is configured to perform data processing based on the service after the configuration update.

9. The apparatus according to claim 8, wherein, The configuration information also includes a data batch number, which is used to enable the corresponding processing module to merge input data according to the data batch number in order to perform batch processing on the merged input data. Furthermore, the first execution unit further includes: The third update subunit is configured to perform a second configuration update operation on the current target processing module in response to the device resource occupancy of the first device corresponding to the current target processing module meeting the second preset condition. The second configuration update operation includes updating the data batch number of the current target processing module according to a preset strategy. The second preset condition includes that either the central processing unit memory or the graphics processing unit memory of the first device corresponding to the current target processing module does not exceed the upper limit, the device utilization rate does not exceed the upper limit, and the current target processing module supports batch processing of data.

10. The apparatus according to claim 9, wherein, The step of updating the current data batch size of the target processing module according to the preset strategy includes: Determine the first data batch size based on the current data batch size of the target processing module; After determining whether the overall processing time of the updated service exceeds a preset processing time threshold, the data batch size of the current target processing module is updated to the first data batch size; and In response to the overall processing time not exceeding the preset processing time threshold, the current data batch number of the target processing module is updated to the first data batch number.

11. The apparatus according to claim 10, wherein, The step of updating the current data batch size of the target processing module according to the preset strategy also includes: In response to the overall processing time exceeding the preset processing time threshold, and the current first path also includes other processing modules, the other processing modules are traversed one by one from low to high according to the performance parameters of each processing module until a first processing module capable of performing data batch number updates is obtained, or it is determined that each processing module on the current first path is unable to perform data batch number updates; and In response to the acquisition of the first processing module, the data batch number of the first processing module is updated to a second data batch number, wherein the second data batch number is determined based on the data batch number of the first processing module before the update.

12. The apparatus according to any one of claims 9 to 11, wherein, The multiple processing modules in the service are deployed on multiple device nodes in a distributed cluster, and the device further includes: The second acquisition unit is configured to respond to situations where the CPU memory and GPU memory of the first device corresponding to the current target processing module both exceed their respective limits, or the device utilization rate of the first device corresponding to the current target processing module exceeds the device utilization rate limit, and the current first path also includes other processing modules, by iterating through the other processing modules one by one from low to high according to the performance parameters of each processing module, until a second processing module capable of performing the first configuration update operation or the second configuration update operation is acquired, or it is determined that each processing module on the current first path cannot perform a configuration update operation; and The second update unit is configured to perform a corresponding configuration update operation on the second processing module in response to obtaining the second processing module, so as to perform data processing based on the service after the configuration update.

13. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Service deployment method and device, electronic equipment and storage medium

    CN113885956A

  • Allocating Resources to New Programs in a Cloud Computing Environment

    US20210124613A1