Method for performing model reasoning by adopting GPU (Graphic Processing Unit), electronic equipment, computer readable storage medium and computer software product

By preloading model instances in the GPU and collaboratively designing with shared memory and state machine, the problem of low memory communication efficiency in high concurrent video stream processing is solved, and efficient and stable model inference is achieved, suitable for large-scale video stream processing tasks.

CN120276885AInactive Publication Date: 2025-07-08TURBULENCE (HANGZHOU) SOFTWARE ENGINEERING CO LTD +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510774614.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the high concurrent video stream processing scenario, in the traditional GPU model inference scheme, the data communication between memory and video memory is low, resulting in a decrease in latency and throughput, which cannot meet the real-time and high concurrency processing requirements.

Method used

The design paradigm of direct access to shared memory and state machine collaboration is adopted, multiple model instances are loaded in advance and preheated, data interaction is performed through shared memory areas, and data writing and reading timing is controlled by state machines to avoid competition among threads and achieve efficient resource management.

Benefits of technology

It significantly reduces computing overhead and system complexity, improves the efficiency and stability of GPU model inference, and can handle large-scale concurrent video streams, meeting the requirements of real-time and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276885A_ABST
    Figure CN120276885A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method for performing model reasoning by adopting a GPU (Graphics Processing Unit), electronic equipment, a computer readable storage medium and a computer software product. The method comprises the following steps: pre-loading and preheating a plurality of model instances in the GPU; setting a shared memory shared by at least two threads, and setting a state machine for the shared memory; the state machine can comprise a writable state and a non-writable state; when the state machine is in a writable state, writing a video frame into the shared memory, and setting the state machine to be in a non-writable state; calling the model instance loaded and preheated in the GPU to perform reasoning based on the video frames in the shared memory; and after the reasoning is completed, setting the state machine to be in a writable state. According to the embodiment of the invention, high efficiency and stability of model reasoning are realized, and calculation overhead and system complexity are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of high-performance computing and artificial intelligence technologies, and particularly to a method for model inference using a GPU, an electronic device, a computer-readable storage medium, and a computer software product. Background Art

[0002] With the rapid development of fields such as smart cities, industrial automation, and public safety, surveillance video analysis has become a key technology for industries to achieve intelligent management. Surveillance video analysis refers to the processing of real-time video streams captured by cameras through computer vision and artificial intelligence technologies, such as object detection, behavior recognition, and abnormal event warning, and is widely used in scenarios such as traffic flow monitoring, factory safety inspections, public area security, and energy facility monitoring. Taking traffic management as an example, it is necessary to analyze thousands of camera videos in real time to identify illegal behaviors and congestion conditions; in industrial scenarios, it is necessary to conduct 24-hour video monitoring of production lines to detect equipment failures or operation risks.

[0003] In such scenarios, video data exhibits ultra-high bandwidth characteristics: high-definition cameras (such as 4K / 8K resolution) and concurrent transmission of multiple video streams, with the single-channel video bit rate reaching dozens of Mbps to hundreds of Mbps, resulting in the system needing to process data volumes of TB level per day. At the same time, the real-time requirements are extremely stringent - for example, in security scenarios, it is necessary to identify violent behaviors or intrusion events within milliseconds, and in traffic scenarios, it is necessary to provide real-time feedback on vehicle trajectories to adjust traffic lights, and analysis delays may lead to serious consequences.

[0004] Traditional video analysis solutions rely on algorithms based on rules or shallow machine learning, which are difficult to handle problems such as light changes and target occlusion in complex scenarios and cannot meet high-precision requirements. In recent years, deep learning models (such as YOLO, Transformer, etc.) have gradually become the mainstream technology due to their powerful feature extraction and generalization capabilities. However, deep learning model inference requires extremely high computing resources: single-frame high-definition image inference may involve billions of floating-point operations, and real-time analysis requires processing dozens to hundreds of frames per second. In this context, GPU accelerated computing has become an inevitable choice - the GPU can significantly improve the model inference efficiency with its large-scale parallel computing ability.

[0005] However, in the face of the high-concurrency processing requirements of massive video streams (such as thousands of cameras being connected simultaneously), the system bottleneck has shifted from computing itself to data communication efficiency: after video data is received from the network card, it needs to pass through the operating system kernel, system memory, and finally be transmitted to the GPU video memory, and the communication between traditional memory and video memory (such as PCIe bandwidth limitations and data copy redundancy) will lead to a decrease in latency and throughput, severely restricting the overall concurrency ability. Therefore, how to optimize the data path from memory to video memory and improve communication efficiency has become the core challenge for achieving high-concurrency GPU inference. Summary of the Invention

[0006] The objective of the embodiments of this application is to provide a method and system, a server, and a computer software product for model inference using a GPU, so as to improve the speed of GPU model inference.

[0007] To solve the above technical problems, the embodiments of this application provide the following method and electronic device, computer-readable storage medium, and computer software product for model inference using a GPU: A method for model inference using a GPU includes: Preloading and warming up multiple model instances in the GPU in advance; Setting up a shared memory shared by at least two threads and setting up a state machine for this shared memory; the state machine can include two states: writable and non-writable; When the state machine is in the writable state, writing a video frame into the shared memory and setting the state machine to the non-writable state; Invoking the model instances loaded and warmed up in the GPU to perform inference based on the video frame in the shared memory; After completing the inference, setting the state machine to the writable state; The non-writable state in the state machine includes the writing-in-progress state and the inference-in-progress state, and the writable, writing-in-progress, and inference-in-progress states form a one-way cyclic conversion path; or, the non-writable state in the state machine includes the writing-in-progress state, the inference-in-progress state, and the inference-completed state, and the writable, writing-in-progress, inference-in-progress, inference-completed, and writable states form a one-way cyclic conversion path.

[0008] An electronic device includes a processor, a memory, and a computer program that can run on the processor. When the processor executes the program, the above method is implemented.

[0009] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above method is implemented.

[0010] A computer product includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the above method is implemented.

[0011] As can be seen from the technical solutions provided by the embodiments of this application above, through a systematic resource management and data interaction mechanism, the efficiency and stability of model inference are achieved, abandoning the traditional inter-process communication (IPC) technology and instead adopting a design paradigm that combines direct access to shared memory and cooperation of a state machine, significantly reducing the computational overhead and system complexity. Brief Description of the Drawings

[0012] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0013] Figure 1 It is the single-inference flowchart in an embodiment of the present application; Figure 2 It is the single-thread single-inference flowchart in an embodiment of the present application; Figure 3 It is the inference flowchart in an embodiment of the present application when generalized to the case of hundreds of threads and processes; Figure 4 It is the schematic diagram of preloading M models in the GPU in an embodiment of the present application and establishing an internal service to transmit data through IPC; Figure 5 It is the design diagram of the shared memory module in an embodiment of the present application; Figure 6 It is the design diagram of the shared memory module in an embodiment of the present application; Figure 7 It is the design diagram of the shared memory module in an embodiment of the present application; Figure 8 It is the design diagram of the shared memory module in an embodiment of the present application; Figure 9 It is the design diagram of the shared memory module in an embodiment of the present application; Figure 10 It is the design diagram of the shared memory module in an embodiment of the present application; Figure 11 It is the design diagram of the shared memory module in an embodiment of the present application. Specific embodiments

[0014] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0015] As mentioned above, the artificial intelligence-based real-time video analysis system has become the core infrastructure for the intelligent transformation of various industries. Such systems need to perform real-time processing on video streams generated by thousands of high-definition cameras (such as 4K / 8K resolution). Through deep learning models (as well as other artificial intelligence models, mainly taking deep learning models as examples for introduction below) such as object detection, behavior recognition, and abnormal event warning, functions such as traffic flow monitoring, production line fault detection, and public place safety protection are realized. However, high-definition video data has the characteristic of ultra-large bandwidth - the bit rate of a single 4K video can reach 30-50Mbps. If 1000 video streams are connected simultaneously, the system needs to process a data throughput of up to 40Gbps, which poses a severe challenge to the concurrent processing ability of the computing architecture.

[0016] For example, the real-time performance of deep learning model inference depends on GPU accelerated computing. The typical process of using GPU for model acceleration includes two key links: model loading and video data transmission. In the model loading stage, the model file needs to be read from the disk (such as SSD / HDD) to the CPU memory, and then transmitted to the GPU video memory through the PCIe bus, and then initialized (such as setting the CUDA context. CUDA stands for Compute Unified Device Architecture, which is a general parallel computing architecture launched by NVIDIA. This architecture enables the GPU to solve complex computing problems. It includes the CUDA instruction set architecture (ISA) and the parallel computing engine inside the GPU, etc.) and warmed up. Taking a 1GB model as an example, if 100 concurrent tasks independently load the model, it will generate 100GB of disk I / O and PCIe transmission overhead, and at the same time occupy 100GB of video memory, far exceeding the single GPU video memory capacity (usually 10-80GB). This not only causes resource waste, but also affects the system real-time performance due to the repeated loading delay (such as the single loading time-consuming about 62.5ms). In the video data transmission stage, the data needs to enter the GPU video memory from the network card through the kernel buffer and user space memory. Although technologies such as DMA and memory mapping can reduce CPU intervention, the high throughput of high-definition video streams (such as reaching the TB level / hour when there are thousands of concurrencies) still makes the PCIe bandwidth (such as 16GB / s of PCIe 3.0) a bottleneck, resulting in data transmission delay and GPU computing unit idleness, significantly reducing the overall concurrent ability of the system.

[0017] Model inference in a concurrent scenario refers to multiple threads or processes simultaneously invoking the model for inference tasks at the same time. This scenario is common in server applications that need to handle a large number of requests, real-time data processing systems, or environments with multi-task parallel processing. Concurrent inference can significantly improve the throughput and response speed of the system, but it also brings problems such as thread safety and resource contention. In a concurrent scenario, multiple threads or processes may access and modify shared resources (such as model weights, caches, etc.) simultaneously. Without appropriate synchronization mechanisms, it may lead to data inconsistency, incorrect inference results, or even system crashes.

[0018] In a concurrent scenario, each thread generally needs to independently initialize the model for the following main reasons: Thread safety: Model objects in many deep learning frameworks (such as PyTorch, TensorFlow) are not thread-safe. If multiple threads share the same model instance, it may lead to race conditions, resulting in unpredictable behavior or errors.

[0019] Resource isolation: Each thread independently initializing the model can ensure that each thread has its own independent model instance and cache, avoiding resource contention between multiple threads. This can ensure that the inference process of each thread is independent and not interfered with by other threads.

[0020] Warm-up cache: When the model is initialized, it usually performs some warm-up operations, such as loading weights and initializing caches. Each thread independently initializing the model can ensure that these warm-up operations are executed in each thread, thus avoiding performance bottlenecks or inconsistent behavior during the inference process.

[0021] The above is closely related to the thread safety of subsequent inferences for the following reasons: Independent model instances: Each thread independently initializing the model means that each thread has its own model instance and cache. In this way, when each thread performs inference, it operates on its own independent data structure and will not conflict with other threads, thus ensuring thread safety.

[0022] Avoiding race conditions: If multiple threads share the same model instance, it may lead to race conditions. For example, one thread may be modifying the model's weights or cache while another thread is reading this data, resulting in inconsistent inference results. Independently initializing the model can avoid this situation.

[0023] Consistent inference environment: Each thread independently initializing the model can ensure that the inference environment of each thread is consistent. Warm-up operations (such as loading weights and initializing caches) are executed in each thread, ensuring the stability and consistency of the inference process.

[0024] Performance optimization: Initializing the model independently can also avoid resource contention between multiple threads, thus improving the inference performance. Each thread can make full use of its own resources without waiting for other threads to release shared resources.

[0025] In an existing technology, for the inference task of a deep learning model (such as YOLO, You Only Look Once, which is an efficient single-stage object detection model. Its core design concept is to complete the localization and classification of multiple objects in an image through a single forward propagation, featuring both real-time performance and high accuracy) in a multi-threaded or multi-process environment, a scheme based on independent model loading for each thread or process is usually followed to ensure thread safety. For example, according to the official recommendation of YOLO, in model inference with a concurrent scenario, the model must be initialized in each thread and the cache of the model must be warmed up to ensure thread safety for subsequent inferences. As Figure 1 shown, when a child thread created by the main thread needs to perform asynchronous model inference, generally 4 operation steps involving the GPU are required under the control of the CPU: initializing / loading the model, warming up the model cache, inference, and resource recycling. As Figure 1 shown in the dashed box, specifically including: 1. Initializing / loading the model In this stage, the model file needs to be read from a persistent storage device (such as a disk) and fully loaded into a computing device (such as GPU memory). Specifically, the system needs to parse the model architecture definition, load the pre-trained weight parameters, and allocate computing resources (such as GPU context, memory buffer). For complex models (such as a YOLO model with a volume exceeding 10GB), this process may take several seconds to dozens of seconds, and each loading needs to be repeated. Since the model file needs to be transmitted to the GPU through the I / O interface and the bus (such as PCIe), in a high-concurrency scenario, multiple threads or processes loading the model simultaneously will cause significant I / O contention and transmission latency, further exacerbating the time-consuming problem in the initialization stage.

[0026] 2. Warming up the model cache After the model is loaded, it is necessary to optimize and cache warm up the computational graph through multiple forward inferences (usually 10 to 100 times). This stage aims to eliminate the initial latency of dynamic graph construction (such as the graph optimization process of TensorRT) and ensure stable performance in subsequent inferences. However, the single inference time in the warm-up stage may be dozens of times that of normal inference (for example, the difference between 100ms and 5ms), and this process needs to be executed independently in each thread or process. When the concurrency increases, the repeated warm-up operations will lead to serious waste of computing resources (such as GPU computing power), and at the same time significantly extend the overall response time of the task, forming a bottleneck in system performance.

[0027] 3. Inference After the warm-up is completed, the model enters the formal inference stage. This stage receives the input high-definition video data, performs forward calculations and generates prediction results. Although the single inference time has dropped to a stable level at this time, due to each thread or process needs to independently maintain a model instance, there are still the following problems in high-concurrency scenarios: Resource competition: When multiple threads or processes access the GPU computing unit at the same time, the inference latency may fluctuate due to hardware scheduling conflicts.

[0028] Load imbalance: Without a global scheduling mechanism, some threads may occupy resources for a long time due to differences in the complexity of input data, exacerbating the decline in overall throughput efficiency.

[0029] 4. Resource recovery After the inference task is completed, it is necessary to explicitly release the computing resources (such as GPU video memory, context handles) occupied by the model to avoid memory leaks. The specific operations include destroying the model instance, reclaiming the buffer memory, and closing the computing context. However, frequent resource release and reapplication (especially in short-term task scenarios) will lead to the following problems: Memory fragmentation: Repeated allocation and release of video memory may generate fragments, reducing the efficiency of subsequent resource allocation.

[0030] Extra overhead: The indirect costs of destroying and reconstructing the model instance (such as context switching, I / O operations) further consume system resources. Especially in dynamic load scenarios, resource recovery may account for a significant proportion of the overall running time.

[0031] In summary, in a server, the CPU often uses dozens to hundreds of cores, and each core may perform model inference. It should be noted that the process of model inference is carried out on the GPU, but the entire process is managed and coordinated by the main thread (and the created child threads) on the CPU. In a large number of concurrent scenarios, the threads within the CPU cores may all call the GPU (not necessarily each CPU core calls the GPU separately. In fact, multiple CPU cores may share the same GPU resource, so not every core will interact with the GPU separately. Usually, the access to GPU resources can be managed through GPU computing libraries such as CUDA or DirectML) to perform the process of model initialization and cache warm-up, which is extremely time-consuming and has a large resource overhead. A large amount of time and bandwidth are wasted in the loading and destruction of resources. In traditional solutions, the officially recommended paradigms such as Figure 2 show the single-threaded single-inference process as shown, where the dashed box is a typical child-thread inference block.

[0032] Generalizing it to the case of hundreds of threads and processes can be as shown in Figure 3 . It should be noted that the CPU is the core hardware for the operating system to schedule threads and processes. Whether it is the creation, destruction of threads, or the context switching between threads, it is all completed by the CPU. The GPU does not have the ability to manage threads or processes. Its main responsibility is to execute large-scale parallel computing tasks (such as model inference). The GPU focuses on computing tasks. The design goal of the GPU is to efficiently process parallel computing tasks (such as matrix operations, convolution operations, etc.), rather than managing threads or processes. In the model inference scenario, the GPU is responsible for performing the forward calculation (inference) of the model, while the CPU is responsible for transferring data from memory to the GPU video memory and managing the scheduling of inference tasks. In the above process, the data transfer is controlled by the CPU. In scenarios such as video analysis, high-definition video data enters the operating system kernel from the network card, then enters the memory, and then is transferred to the GPU video memory under the control of the CPU. This process is dominated by the CPU, and the GPU is responsible for receiving the data and performing the calculation.

[0033] Figure 3 When generalizing to the case of hundreds of threads and processes in, performance bottlenecks will quickly appear in model loading, cache warm-up, and system resource allocation and overhead. In addition, a bigger problem lies in the uncontrollable number of model inferences. If each thread / process applies for resources independently, it is very easy to cause a shortage of GPU video memory, and then cause the GPU to crash.

[0034] Another popular approach is to pre-load M models in the GPU and establish an internal service (which can be implemented in the way of a traditional http server), such as Figure 4As shown. Through IPC (InterProcess Communication), N processes / threads transmit data through IPC, and then get the returned inference results. This method (hereinafter referred to as IPC method) solves the problem of repeated loading and preheating of the model, and also solves the problem of scarce GPU resources. However, in actual scenarios, IPC data transmission is very time-consuming and cannot achieve real-time prediction performance in the environment of high-definition video.

[0035] like Figure 4 As shown in the figure, each process accesses the inference service through IPC when inference is needed. The inference service distributes the load according to the number of GPUs and the number of models that each GPU can bear. After obtaining the inference result, the inference service returns the inference result to the client. It should be noted that in surveillance video analysis, "predicting video frames" usually does not mean predicting the next frame of video images, but refers to analyzing the current frame or continuous frames through the model, extracting valuable information (such as target location, behavior pattern, future state, etc.), and making predictions and decisions based on this information.

[0036] Specifically, it includes: Object detection and tracking: predict the position and state of the object in the current frame.

[0037] Behavioral analysis and anomaly detection: predicting target behavior patterns or abnormal events.

[0038] Trajectory prediction: predicting the future trajectory of a target.

[0039] Event prediction: predicting possible events (such as traffic accidents, crowd gatherings).

[0040] Only in a few specific scenarios (such as video generation or restoration) may "predicting a video frame" involve generating the next frame of image.

[0041] For example, in security monitoring, by predicting video frames, suspicious targets (such as intruders, suspicious packages) or abnormal behaviors (such as fighting, wandering) can be detected in real time. Based on the prediction results, the system can trigger an alarm or notify security personnel. For example, identifying suspicious targets in the video, detecting abnormal behaviors (such as intrusions, fighting), trajectory prediction (predicting the target's movement path and deploying defenses in advance).

[0042] Another example is in traffic management, by predicting video frames, it is possible to monitor traffic flow in real time, identify traffic violations (such as running red lights, driving against traffic), and predict traffic jams or accidents. Specifically, vehicles can be detected and tracked to identify vehicles and track their movement trajectories; it can also detect violations, such as running red lights, speeding, etc.; it can also be traffic flow prediction, such as predicting future traffic flow and optimizing signal light control.

[0043] For another example, in a smart city, by predicting video frames, various activities in the city (such as crowd gathering, vehicle flow) can be analyzed, and possible events (such as emergencies, natural disasters) can be predicted. Specifically, it can be to conduct crowd density analysis, such as monitoring crowd gathering situations to prevent stampede incidents; it can also be to monitor the environment, such as detecting natural disasters like fires and floods; it can also be to schedule resources, such as optimizing the allocation of public resources (such as police force, fire protection) according to the prediction results.

[0044] For another example, in retail analytics, by predicting video frames, customer behaviors (such as staying time, purchase intention) can be analyzed, and store layouts and marketing strategies can be optimized. Specifically, for example, analyzing customer behaviors to identify customers' shopping paths and points of interest; it can also be to conduct passenger flow statistics, such as counting customer traffic to optimize store operations; it can also be to conduct anomaly detection, such as detecting theft or suspicious behaviors.

[0045] As mentioned above, the realization of the above functions depends on the following technologies and capabilities: 1. Deep learning models Object detection models (such as YOLO, Faster R-CNN): used to identify objects (such as people, vehicles) in video frames.

[0046] Behavior analysis models (such as LSTM, 3D CNN): used to analyze the behavior patterns of objects.

[0047] Trajectory prediction models (such as RNN, Transformer): used to predict the future movement trajectories of objects.

[0048] 2. GPU accelerated computing The parallel computing ability of GPUs can efficiently process video frame data and meet the real-time requirements.

[0049] A large number of GPU model inferences support high-concurrency processing and are suitable for multi-camera monitoring scenarios.

[0050] 3. Big data processing High-definition video data volume is huge, and the high bandwidth and computing ability of GPUs can quickly process this data.

[0051] Through distributed computing and storage technologies, the processing ability of the system can be further expanded.

[0052] 4. Real-time analysis and decision-making Real-time predict the target states and behavior patterns in video frames, and support quick decision-making (such as triggering alarms, optimizing traffic signals).

[0053] Based on the prediction results, the system can take measures in advance to avoid or reduce losses.

[0054] 5. Multi-domain applications The monitoring video analysis requirements in different domains (such as security, transportation, retail) can be achieved through customized models and algorithms.

[0055] The main reason for the huge time consumption of data transmission using the IPC method lies in the protocol overhead part. Specifically, IPC communication usually needs to follow certain protocols (such as HTTP, gRPC, etc.), and these protocols will introduce additional data encapsulation and decapsulation overheads. For example, serialization and deserialization (such as using protocols like JSON, Protobuf, etc.) will consume a large amount of CPU resources and increase latency.

[0056] This application provides a method embodiment for model inference using a GPU, as Figure 5 shown, including: S100: Pre-load and warm up multiple model instances in the GPU.

[0057] Pre-loading and warming up multiple model instances in the GPU can be pre-loading multiple instances of the same model and warming them up in the GPU, or pre-loading multiple instances of different models and warming them up in the GPU.

[0058] Here, taking the example of pre-loading multiple instances of the same model and warming them up in the GPU for illustration. In the implementation scheme of pre-loading multiple instances of the same model and warming them up in the GPU, first, the parallel initialization of model instances can be completed through a systematic resource allocation mechanism. Specifically, the CPU can read the model file from the persistent storage device by executing the main control process / main thread and using the file system interface of the operating system. After parsing its network architecture and weight parameters, independent storage spaces are allocated for each instance in the GPU video memory. Each instance can contain a complete copy of the model weights, the computational graph structure, and a dedicated computational context, ensuring that each instance is completely isolated in the physical video memory and the logical execution environment. This process can be achieved by repeatedly calling the GPU video memory allocation interface (such as cudaMalloc) and the computational context creation instruction (such as cuCtxCreate), enabling multiple instances to reside in the same GPU device in parallel without interfering with each other.

[0059] After the model instance is loaded, the system (the system refers to the entire computing environment or computing platform, including hardware, operating system, drivers, libraries, and the software layer that manages and schedules computing tasks) enters the cache warm-up phase. For each instance, the CPU can execute the main control process / main thread, and submit multiple forward inference computing tasks to the GPU for execution by inputting a preset batch of standard data (such as all-zero tensors or representative samples). During this process, the CUDA runtime automatically compiles and optimizes the computational graph kernel, eliminates the initial latency of dynamic graph construction, and at the same time solidifies the memory allocation strategy for intermediate activation values and output buffers. The number of warm-up times is set based on the balance between model complexity and hardware performance, and usually the optimal value is determined through experiments to ensure stable latency for subsequent inferences. After warm-up, the computational graph kernel and memory occupancy of each instance tend to a steady state, forming a reusable inference context.

[0060] After the pre-warmed model instance enters the ready state, it can be dynamically scheduled by the main control process through the task queue mechanism. When concurrent inference requests arrive, the system allocates input data according to the load status of each instance (such as the current task queue length or memory usage rate), and uses CUDA streams to achieve asynchronous computing. Each instance independently processes its assigned tasks, and the inference results are returned to the main memory through the PCIe bus, while updating the instance status to receive new tasks. This design effectively avoids the risk of multithreaded competition through hardware-level parallelism and resource isolation characteristics, and at the same time significantly reduces the cold start latency of a single inference. However, its implementation requires strict monitoring of the total GPU memory capacity to avoid the physical limit being exceeded by the superposition of multiple instances. If necessary, resource elastic scaling can be achieved through a dynamic instance loading strategy.

[0061] In the embodiments of the present invention, after the warm-up is completed, the model inside the GPU enters a stable state, and the model structure, weight parameters, and execution configuration generally no longer change, and only responsible for executing inference computing tasks. This static deployment method has significant performance advantages compared to dynamically loading models: First, through pre-loading and warm-up, the time overhead of repeatedly initializing the model for each inference request is eliminated; Second, the model parameters are resident in the GPU memory, avoiding frequent data transmission between the CPU and GPU; Third, the GPU can perform deeper optimizations based on the stable computational graph, such as instruction-level parallelism, memory access merging, and computational pipeline rearrangement. In high-concurrency scenarios, this optimization strategy can significantly improve the throughput and response speed of the system, and is especially suitable for processing continuous video stream data.

[0062] To ensure the stability and execution efficiency of model inference, the inference process can also implement a complete resource monitoring and exception handling mechanism. This mechanism periodically collects metrics such as GPU performance counters, memory usage, and temperature to evaluate the health status of the system in real-time. At the same time, by setting a watchdog timer, it can detect and handle exceptions such as model execution timeouts, memory leaks, or device errors in a timely manner. When an exception is detected, the inference process can automatically initiate a recovery process, such as resetting a specific model instance, adjusting the inference batch size, or triggering a backup model switch, to ensure that the system can maintain continuous and stable service capabilities in the face of various complex situations.

[0063] Through the above design, the embodiments of the present invention achieve efficient management of GPU resources and optimization of model inference performance, providing a solid computing foundation for processing large-scale concurrent video streams. The centralized management method of the inference process can not only simplify the system architecture and reduce the risk of resource conflicts, but also fully exploit the performance potential of the GPU in deep learning inference tasks through a carefully designed warm-up strategy and steady-state execution mode, enabling the system to achieve higher concurrent processing capabilities and lower latency performance with limited hardware resources.

[0064] S110: Set a shared memory shared by at least two threads and set a state machine for this shared memory.

[0065] The shared memory can be used to store video frames. As Figure 3 shown, for example, a high-definition video contains 300 video frames, and these 300 video frames can be stored in the shared memory.

[0066] The state machine can include two states: writable and non-writable.

[0067] In the embodiments of the present invention, to achieve efficient video inference and ensure the stability and accuracy of resource sharing, a shared memory shared by at least two threads can be set, and a dedicated state machine can be configured for this shared memory. Of course, it can be a shared memory shared by multiple threads, and here at least two threads are taken as an example. The shared memory can be used to store data blocks related to inference, such as video frame data. The design purpose of the shared memory is to improve the utilization efficiency of GPU resources through multi-threaded parallel operations, while ensuring that no conflicts or data inconsistencies occur when each thread accesses the shared memory.

[0068] The allocation and management of shared memory can be the responsibility of the CPU master process to ensure that each thread can access the necessary data according to specific task requirements. Specifically, the shared memory can allocate a memory space of a preset size according to the size of the video frame and the requirements of the inference task, or can dynamically allocate an appropriate memory space. Each block of shared memory can be shared by multiple threads, and precise state management is used to coordinate access between different threads. The introduction of the state machine enables strict control of the access to shared memory, thus avoiding race conditions and memory access conflicts between threads.

[0069] In an embodiment of the present invention, the state machine is designed to have two states: "writable" and "non-writable". The function of this state machine is to control the timing of data writing and reading. The "writable" state of the state machine indicates that the current shared memory can be written with new data by a thread, that is, video frame data or other content that needs to be updated, while the "non-writable" state means that the data in the current memory block is being processed by the CPU or GPU or has been processed, so other threads are not allowed to perform write operations. When the state machine is in the "non-writable" state, other threads can only wait until the state machine becomes "writable". In this way, through the management of the state machine, the orderliness of the data writing process is ensured, and data corruption or incorrect inference results caused by concurrent access are effectively prevented.

[0070] In addition to the data block for storing video frame data as shown Figure 6 in the shared memory, there can also be a status block, as shown Figure 7 in. The state machine can be stored in this status block. Setting the state machine in the status block in the shared memory facilitates convenient access by different shared threads. Otherwise, a message passing mechanism can be adopted to allow threads to share the status by sending and receiving messages instead of directly accessing the shared data; a semaphore and event thread synchronization mechanism can also be used to control the access of threads to the shared status. Relatively speaking, setting the state machine in the status block of the shared memory for convenient access by different threads has the following multiple significant advantages: In terms of performance 1. Reduce data copy overhead When multiple threads need to access the state of a state machine, if shared memory is not used, it may be necessary to frequently copy state data between different threads. By placing the state machine in shared memory, each thread can directly access the same physical memory area, avoiding the time and space overheads brought by data copying. For example, in a real-time monitoring system, multiple monitoring threads need to simultaneously obtain the status information of a device (such as temperature, humidity, etc.). If the state machine is stored in shared memory, these threads can directly read the data in the shared memory without copying the data from one thread to another, greatly improving the data access efficiency.

[0071] 2. Reducing Communication Latency There is usually a certain latency in communication between threads, especially when using methods such as message passing. By using shared memory, threads can directly read and write the state of the state machine without data transmission through an intermediate communication mechanism (such as message queues, pipes, etc.), thus significantly reducing the latency of data access. In systems with extremely high real-time requirements, such as the flight control system in aerospace, multiple control threads need to quickly obtain the status information of the aircraft (such as attitude, speed, etc.). The use of shared memory can ensure that these threads can respond in a timely manner, improving the real-time performance of the system.

[0072] In terms of data consistency 1. Unified Data View All threads access the same state machine data in shared memory, which ensures that the status information seen by each thread is consistent. For example, in a multi-threaded database management system, multiple threads may need to simultaneously access the status of the database (such as whether it is in a backup state, whether new data has been inserted, etc.). The state machine in the shared memory can provide a unified data view for all threads, avoiding incorrect operations caused by data inconsistency.

[0073] 2. Facilitating State Synchronization Although multiple threads can simultaneously access the state machine in shared memory, synchronization mechanisms (such as mutex locks, semaphores, etc.) can be used to ensure that modifications to the state are atomic, thus guaranteeing state consistency. For example, in a multi-threaded file system, multiple threads may simultaneously attempt to modify the status of a file (such as opening, closing, reading, writing, etc.). The use of shared memory and synchronization mechanisms can ensure that these operations are carried out in the correct order, avoiding data competition and inconsistency problems. That is to say, the state machine in the state block of the shared memory ensures the atomicity of the state through synchronization mechanisms.

[0074] In terms of programming implementation 1. Code Simplicity Compared with other inter-thread communication and data sharing methods (such as message passing, remote procedure call, etc.), setting the state machine in shared memory can make the code more concise. Threads can directly access and modify the state of the state machine without writing complex communication and synchronization code. For example, in a simple multi-threaded game, multiple threads need to access the state of game characters (such as health points, magic points, etc.) simultaneously. Using shared memory can make the code more intuitive and easier to maintain.

[0075] 2. Improve development efficiency Since the access method of shared memory is relatively simple, developers can focus more on the implementation of business logic without spending too much effort on complex inter-thread communication and synchronization mechanisms. This helps improve development efficiency and shorten the development cycle.

[0076] In terms of resource utilization: Save memory space Multiple threads share the same memory area to store the state machine, avoiding memory waste caused by copying the state data for each thread. Especially when the amount of state machine data is large, using shared memory can significantly save the system's memory resources. For example, in a large distributed computing system, threads on multiple computing nodes need to share a complex state machine (such as task scheduling state, resource allocation state, etc.). The use of shared memory can reduce memory occupancy and improve the system's resource utilization.

[0077] The unwritable state can specifically include two states: being written and being inferred. In the implementation of the state machine, further dividing the unwritable state into multiple states can achieve more refined control and state tracking of the shared memory access process. Specifically, the unwritable state includes at least two states: the being written state and the being inferred state, corresponding to different stages in the data life cycle of the shared memory.

[0078] The being written state means that the data producer thread has started writing video frame data to the data block in the shared memory but has not completed the writing operation. In this state, the shared memory is exclusive to the data producer thread, and other threads (including other data producer threads and GPU inference threads) cannot read or write to this shared memory. The purpose of setting the being written state is to protect the atomicity and consistency of the data writing process and prevent partially written incomplete data from being incorrectly read or processed by other threads. When the data is completely written, the state machine will change from the being written state to the being inferred state, and at the same time update the write completion timestamp and data integrity flag in the metadata of the state block.

[0079] The state in the inference indicates that the data in the shared memory is ready and is currently in the state where the model calculation is being performed by the GPU inference thread. In the inference state, the content in the data block is locked in read-only mode (also non-writable state). The data producer thread cannot modify or update the video frame data therein, but the GPU inference thread can safely read the content of the data block and perform inference. The duration of the inference state depends on various factors such as model complexity, GPU load, and video frame resolution. The system will monitor the duration of this state according to the configured timeout threshold. Once it exceeds the preset threshold, the corresponding exception handling mechanism can be triggered. After the GPU inference thread completes the calculation, it can change the state machine from the inference state to the writable state, release the occupancy of the shared memory, and allow the data producer thread to write new video frame data.

[0080] In the detailed design of the non-writable state of the state machine, strict state transition rules and access control policies can be adopted. The transition from the writable state to the writing state is triggered by the data producer thread through an atomic operation; the transition from the writing state to the inference state is triggered by the same data producer thread that has completed the data writing; the transition from the inference state to the writable state is triggered by the GPU inference thread that has completed the model inference. This one-way state transition path ensures the predictability of system behavior and the orderliness of the data flow, effectively avoiding concurrent conflicts and data race conditions.

[0081] By subdividing the non-writable state into the writing state and the inference state, the embodiments of the present invention achieve the full life cycle management of the shared memory access process, which can improve the stability and reliability of the system in a high-concurrency environment. The subdivided states not only enable the system to more accurately track and record each stage of data processing, but also provide more fine-grained control points for performance optimization, fault diagnosis, and exception recovery. Especially in application scenarios where multiple video streams are processed and high real-time requirements are imposed, the advantages of this design are particularly significant.

[0082] In addition, in addition to the two states of being in the process of writing and in the process of inference, the non-writable state may also include an inference completion state, such as Figure 8 shown. In the implementation of the state machine, the non-writable state is further subdivided into multiple states to achieve more refined control and state tracking of the shared memory access process. Specifically, the non-writable state includes three states: the writing state, the inference state, and the inference completion state, which respectively correspond to different processing stages in the life cycle of the shared memory data.

[0083] As described above, the "writing in progress" state refers to the intermediate state where the data producer thread has started writing video frame data to the data block in the shared memory but has not completed the writing operation. In this state, the shared memory is exclusively occupied by the data producer thread, and no other thread (including other data producer threads and GPU inference threads) can read from or write to this shared memory. The purpose of setting the "writing in progress" state is to protect the atomicity and consistency of the data writing process and prevent partially written incomplete data from being erroneously read or processed by other threads. When the data is fully written, the state machine will transition from the "writing in progress" state to the "inference in progress" state, and at the same time update the write completion timestamp and data integrity flag in the metadata of the state block.

[0084] The "inference in progress" state indicates that the data in the shared memory is ready and is currently being used by the GPU inference thread for model calculation. In the "inference in progress" state, the content in the data block is locked in read-only mode. The data producer thread cannot modify or update the video frame data therein, but the GPU inference thread can safely read the content of the data block and write the results to the inference result block after the calculation is completed. The duration of the "inference in progress" state depends on various factors such as model complexity, GPU load, and video frame resolution. The system will monitor the duration of this state according to the configured timeout threshold. Once the preset threshold is exceeded, the corresponding exception handling mechanism will be triggered. After the GPU inference thread completes the calculation, the state machine will transition from the "inference in progress" state to the "inference completed" state, indicating that the inference process has ended but the results are yet to be processed.

[0085] The "inference completed" state refers to the intermediate state where the GPU has completed the model inference calculation and written the results to the inference result block, but the results have not been read or processed by the result consumer thread. In this state, both the data block and the inference result block are locked in read-only mode to prevent any thread from modifying the processed data and calculation results. The introduction of the "inference completed" state provides a clear signal to the result consumer thread, indicating when it can safely read and process the inference results, while avoiding the risk of the results being read prematurely or being overwritten incorrectly. After the result consumer thread completes the result reading and related processing, it will transition the state machine from the "inference completed" state to the "writable" state, releasing the occupancy of the shared memory and allowing the data producer thread to write new video frame data, thus starting a new round of processing cycle.

[0086] In the detailed design of the unwritable state of the state machine, strict state transition rules and access control policies are adopted. The state transition proceeds along a one-way transition path of being written → inferring → inference completed → writable, and each transition is triggered by a specific responsible thread: the transition from the writable state to the being written state is triggered by the data producer thread through an atomic operation; the transition from the being written state to the inferring state is triggered by the same data producer thread that has completed writing the data; the transition from the inferring state to the inference completed state is triggered by the GPU inference thread that has completed the model inference; and the transition from the inference completed state to the writable state is triggered by the result consumer thread that has completed the result processing. This clear responsibility assignment and one-way state transition path ensure the predictability of system behavior and the sequentiality of data flow, effectively avoiding concurrent conflicts and data race conditions.

[0087] By subdividing the unwritable state into the being written state, the inferring state, and the inference completed state, the present invention realizes the full-life cycle management of the shared memory access process, improving the stability and reliability of the system in a high-concurrency environment. The subdivided states establish clear boundaries and synchronization points between data production, processing, and consumption. This not only enables the system to more precisely track and record each stage of data processing but also provides finer-grained control points for performance optimization, fault diagnosis, and exception recovery. In addition, compared with the two-state design, the three-state design has significant advantages in the reliable transmission of processing results. Especially in application scenarios where multiple video streams are processed and high requirements are placed on real-time performance and result accuracy, this design can more effectively coordinate the working rhythms of different types of threads, reducing unnecessary waiting and conflicts, thereby improving the overall system throughput and response speed.

[0088] In the state block, in addition to containing the state machine, it can also contain metadata, which can be used to indicate the access state of the current memory. The metadata in the state block refers to a set of auxiliary information other than the basic state variables of the state machine, and is used to enhance the reliability, maintainability, and performance of the system. The metadata can contain a combination of one or more of key parameters such as timestamp information, processing identifier, error flag, priority information, and version number.

[0089] The timestamp information can record each key time point of data processing in the shared memory, specifically including but not limited to: the data write timestamp, which records the exact moment when the video frame is written into the shared memory; the state transition timestamp, which records the exact time when the state machine changes from one state to another; the inference start timestamp, which marks the start time when the GPU begins to execute the model inference; and the inference completion timestamp, which identifies the end time when the GPU finishes the model calculation and outputs the result. The above timestamps are obtained using a high-precision timer, with an accuracy up to the microsecond or nanosecond level, and are used for system performance analysis, processing delay calculation, timeout detection, and timing reconstruction in case of exceptions.

[0090] The processing identifier can be used to mark the thread identifier or process identifier that is currently accessing or processing the shared memory block. The processing identifier is updated each time the access control right is transferred, ensuring that the system can accurately track the entity that currently holds the memory access right. By recording the processing identifier, the system can achieve more refined resource management and problem tracking. Especially in the event of an exception, it can quickly locate the relevant processing unit, significantly reducing the time and complexity of fault diagnosis. In addition, the processing identifier can also provide the necessary information for implementing advanced scheduling policies such as thread affinity or processing unit binding.

[0091] The error flag can be a set of bit flags or enumeration values used to indicate various types of abnormal situations that may occur during the processing of shared memory data. The error flag at least includes a data integrity error flag for marking whether incomplete or interrupted data occurs during the data writing process; a calculation error flag indicating whether numerical anomalies, resource shortages, etc. occur when the GPU executes model inference; and a timeout error flag for marking whether a specific operation exceeds the preset maximum allowable time. These error flags provide key decision-making basis for downstream processing modules, enabling the system to take corresponding recovery measures or skip processing according to the error type.

[0092] The priority information can be used to determine the processing order and resource allocation strategy in a multi-task competition environment. The priority information can be represented as an integer value or a composite structure, and the higher the value, the higher the processing priority. In the case of limited system resources or high load, the GPU scheduler will decide which shared memory block data to process first according to the priority information. This mechanism is especially applicable to scenarios where multiple video streams are processed and some streams have higher real-time requirements, or where specific regions or specific events need to be processed first. The priority can be a statically preset fixed value or a value dynamically calculated based on factors such as content importance and time urgency.

[0093] The version number can be a monotonically increasing counter used to maintain data consistency. Whenever the data in the shared memory is updated, the version number will increase, enabling the system to distinguish between old and new data and avoid processing outdated information in an asynchronous environment. The version number mechanism is the basis for implementing optimistic concurrency control, allowing multiple threads to work together without acquiring explicit locks, significantly improving the parallelism and throughput of the system. When a version number mismatch is detected, the system can choose to retry the operation or take other recovery strategies to ensure that the finally processed data is the latest valid data.

[0094] By comprehensively utilizing the above-mentioned metadata, the embodiments of the present invention implement an efficient parallel processing system with self-monitoring, self-diagnosis, and self-recovery capabilities. The metadata can not only serve the normal operation control of the system, but also provide the necessary information basis for performance tuning, fault diagnosis, and system expansion. The existence of the metadata enables the system to achieve a higher level of reliability and maintainability while maintaining high performance, especially in complex application scenarios of processing large-scale concurrent video streams, where the role of the metadata is particularly prominent.

[0095] In one embodiment, the shared memory may include data blocks and inference result blocks, as Figure 9 shown. In addition to data blocks, the shared memory may also include inference result blocks, as Figure 9 shown. The inference result blocks can be used to store the result data generated during the inference process. These data blocks cooperate with each other and, under the control of the state machine, achieve the coordination and sharing of data in a multi-threaded environment. The memory space of each data block is pre-allocated and dynamically managed as needed, which can ensure the efficient use of memory and the flexibility of memory management.

[0096] When a thread needs to write video frame data into the shared memory, it will first check the state of the state machine to ensure that the current state is the "writable" state. When the state machine is in the "writable" state, the thread can write the video frame data into the data block part of the shared memory. At the same time, after completing the data writing, the thread will update the state machine and set it to the "non-writable" state to prevent other threads from writing data during this period and ensure the consistency and integrity of the data. At this time, the inference process can be performed by the GPU, and the GPU will read the video frame data in the shared memory and execute the corresponding inference tasks.

[0097] Through this design, the introduction of the state machine effectively prevents conflicts caused by concurrent access to the shared memory between threads, ensuring the accuracy and consistency of each frame of data. At the same time, multiple threads share the same memory, maximizing the utilization of the memory space and avoiding unnecessary memory waste. This innovative design not only improves the efficiency of data processing but also ensures the stability of the inference task in a high-concurrency environment, providing important support for large-scale video inference applications.

[0098] In summary, through refined shared memory management and state machine control, the embodiments of the present invention provide an efficient multi-threaded concurrent computing solution, enabling each thread to efficiently interact with the GPU, thereby achieving efficient execution of video inference tasks. In addition, the introduction of the state machine not only enhances the orderliness of data access but also enables the system to maintain data consistency and correctness during concurrent operations. This design not only achieves efficient GPU inference but also ensures the stability and scalability of the system, making it suitable for large-scale data processing tasks in various practical application scenarios.

[0099] In one embodiment, the shared memory may include a data block, a status block, and an inference result block. The state machine includes a writable state and a non-writable state, and the writable and non-writable states form a one-way cyclic transition path, as Figure 10 shown.

[0100] In one embodiment, the shared memory may include a data block, a status block, and an inference result block. The state machine includes a writable state, a writing state, an inferring state, and an inference completed state. The writable, writing, inferring, inference completed, and writable states form a one-way cyclic transition path, as Figure 11 shown.

[0101] S120: When the state machine is in the writable state, write the video frame into the shared memory and set the state machine to the non-writable state.

[0102] In the S120 stage, when the state machine is in the "writable" state, the data producer thread can write the new video frame data into the data block part in the shared memory. At this time, the "writable" state of the state machine indicates that the shared memory is ready to receive new input data and that this memory area is not currently occupied or locked by other threads. Before performing the data writing operation, the data producer thread will first check the state of the state machine to ensure that the current operation is legal. When the state machine is in the "writable" state, the data producer thread is allowed to write data, and the writing process is atomic to avoid data inconsistency or conflict during the writing process.

[0103] When the state machine is in the writable state, the writing process of video frames is achieved through the following steps: The client thread first reads the status flag in the status block through an atomic operation. If the status is "writable", it immediately preempts the status lock through an atomic instruction (such as atomic_compare_exchange) to prevent other threads from writing concurrently. After successful preemption, the client thread writes the original pixel data of the video frame into the data block of the shared memory through memory direct copy (such as memcpy) or memory mapping (Memory-Mapped I / O) without serialization operations throughout the process. The writing area of the data block is dynamically divided according to the preset video resolution and format. For example, for video frames in YUV420 format, the data block is divided into continuous storage areas for the luminance (Y) component and the chrominance (U / V) component, ensuring that the physical layout of the pixel data is strictly aligned with the dimensions of the model input tensor.

[0104] After the writing is completed, the client thread can update the metadata in the status block, including the timestamp of the video frame, the data checksum (such as CRC32 hash value), and the data version number, and switch the state machine to "non-writable" through an atomic operation. This state transition process forces the cache consistency to be refreshed through a memory barrier to ensure the visibility of the state change to all threads. At the same time, the system records the writing thread identifier and operation timestamp of the current data block for data traceability and conflict analysis in subsequent abnormal scenarios. If a data check failure (such as out-of-bounds memory or checksum mismatch) is detected during the writing process, a rollback mechanism is triggered to reset the state machine to "writable" and clear the written dirty data, and the client thread is notified through the event queue to retry the writing.

[0105] Furthermore, to optimize the continuous writing efficiency in high-throughput scenarios, the shared memory adopts a double-buffering (DoubleBuffering) or ring buffer (Ring Buffer) design. When the state machine is "writable", the client thread can select the next free buffer to perform the writing, while the GPU inference thread can concurrently read the data in the previous buffer to perform inference. This design realizes the efficient rotation of the buffer through the atomic switching of the address pointer, avoiding explicit lock contention and ensuring the timing integrity of the video frame. During the writing process, the physical memory of the data block optimizes the memory access efficiency through page alignment (Page Alignment) and prefetching (Prefetching) strategies, which can minimize the memory copy latency to the greatest extent.

[0106] After the data writing is completed, the data producer thread updates the state of the state machine and switches it to the "non-writable" state. This state transition indicates that the data in the current shared memory is being processed and has entered the subsequent inference stage. The "non-writable" state of the state machine ensures the stability of the data. During the period when the data is read and processed by the inference thread, other threads cannot modify the data, avoiding errors caused by data inconsistency during the inference process. At this time, the GPU inference thread will read the data and perform corresponding inference calculations.

[0107] By integrating the "writable" and "non-writable" state transition logic of the state machine into the multi-threaded parallel computing process, the present invention ensures the orderliness and consistency of data processing. In a multi-threaded environment, due to the strict control of the state machine, the data writing and reading operations will not conflict, and multiple threads can cooperate and share resources efficiently. At the same time, the memory occupation and access strategy are reasonably managed, maximizing the resource utilization efficiency.

[0108] In addition, when multiple threads are executed in parallel, the "non-writable" state provided by the state machine avoids any race conditions between data writing and inference calculations. Each thread can only modify the data when it obtains the writing permission, and during the inference calculation, the state of the data is locked as read-only, ensuring the stability of the inference and the integrity of the data. This design effectively guarantees that in a high-concurrency environment, the GPU inference task can be executed stably and efficiently, thus meeting the requirements of real-time and accuracy for large-scale video processing tasks.

[0109] S130: Invoke the model instance loaded and warmed up in the GPU to perform inference based on the video frame in the shared memory.

[0110] In stage S130, after the data writing is completed and it is ensured that the state machine is in the "non-writable" state, the system enters the GPU inference process. At this time, the GPU inference thread starts to read the video frame data stored in the shared memory and performs inference calculations using the pre-loaded and warmed-up model instance. According to the multiple pre-loaded model instances, the GPU inference thread can independently perform corresponding inference tasks for each instance, ensuring that in a concurrent scenario, the calculation tasks of each instance do not interfere with each other, thus achieving efficient resource utilization.

[0111] During the inference calculation process, the GPU utilizes its parallel processing ability to perform model calculations on the input video frame data. This calculation process includes operations such as feature extraction, classification, and recognition of image data, and generates corresponding inference results according to the model structure. The computing cores in the GPU execute the inference tasks according to the pre-configured computation graph, and each node and operation in the computation graph are processed in parallel to optimize the inference performance and minimize the inference latency as much as possible. The CUDA runtime system is responsible for scheduling the execution of these tasks, and realizes asynchronous calculation through streams, enabling the inference tasks of different model instances to be executed in parallel on the same GPU.

[0112] Meanwhile, the GPU inference threads dynamically adjust the allocation of computing resources according to the hardware performance and the complexity of the inference tasks, ensuring that each model instance can be executed efficiently. After the inference tasks are completed, the GPU writes the inference results into the inference result block, and these results are transmitted back to the main memory through a high-speed bus (such as PCIe) for further processing by the result consumer threads. In this process, the high parallelism of the GPU ensures that even in the case of multiple models performing parallel inferences, the entire calculation process remains efficient and stable.

[0113] The execution of the GPU inference threads not only depends on the warm-up of the model instances, but also involves close interaction with the data in the shared memory. The content in the data block is locked in read-only mode during the inference state, and the data producer threads cannot modify it, ensuring the consistency of the data and the correctness of the inference process. The results generated during the inference process will be stored in the inference result block, and the data in the inference result block will be read and further processed by dedicated result consumer threads. The execution speed and stability of the GPU inference tasks are the key to the system performance, depending on efficient hardware resource scheduling and memory management strategies.

[0114] Through this design, the S130 stage not only ensures an efficient inference process, but also enables the system to make full use of the computing power of the GPU, and realizes multi-task parallel processing and efficient sharing of resources on the premise of maintaining data consistency. The cooperation of each thread and the management of the state machine enable the inference tasks to run stably in a high-concurrency environment, providing strong support for large-scale video stream processing and application scenarios with high real-time requirements.

[0115] S140: After completing the inference, set the state machine to the writable state.

[0116] In stage S140, after the GPU inference thread completes the inference task, the system enters the stage of processing the inference results. At this time, the GPU inference thread writes the inference results into the inference result block in the shared memory according to the predefined computational graph. The inference result block, as a dedicated data storage area, is responsible for storing the calculation results of each inference instance. This result data is transmitted back to the main memory through a high-speed bus (such as PCIe) for subsequent processing or to be provided to the application for further analysis and decision-making.

[0117] As the inference process is completed, the state machine transitions from the "inference in progress" state to the "inference completed" state based on the end signal of the inference calculation. This state transition marks the successful completion of the inference task, and the inference results have been generated and stored in the result block, ready to be read by the result consumer thread. In the "inference completed" state, the content of the data block and the inference result block in the shared memory are locked as read-only to prevent other threads from modifying the result data during this period, ensuring the integrity and accuracy of the inference results. This mechanism effectively avoids data overwrite or conflict problems that may be caused by concurrent operations.

[0118] After the inference is completed, the result consumer thread will actively check the content of the inference result block. According to the business requirements, the result consumer thread can further process these inference results, such as performing post-processing algorithms, extracting meaningful decision-making information, or passing the results to other system modules for use. To ensure data consistency and traceability, the result consumer thread will update the state machine after processing, transitioning it from the "inference completed" state to the "writable" state. At this time, the state machine returns to the "writable" state, indicating that the data block in the shared memory can be accessed and modified by the data producer thread again, allowing new video frame data to be written and starting a new round of inference tasks.

[0119] Through this design, stage S140 not only ensures the reliable storage and reading of the inference results but also guarantees the sequentiality of the entire inference process and data consistency. The management of the state machine ensures that the operations in each stage are completed within the appropriate time, avoiding conflicts and resource competition between inference tasks. In a high-concurrency environment, through refined state transitions and thread coordination, the system can achieve efficient resource management and task scheduling while ensuring data integrity and inference accuracy. Finally, the result consumer thread can efficiently read and process the inference results, thus ensuring that the system can operate stably in application scenarios with multi-channel video stream processing and strict real-time requirements.

[0120] Once the GPU has completed the inference task and produced the results, the system updates the content in the inference result block. At this time, the thread checks the status of the state machine again to ensure that the state machine has returned to the "writable" state. After the state machine resumes the "writable" state, the thread can continue to write the next frame of video data to the shared memory, thus starting a new inference cycle. The whole process loops, forming an efficient and stable data processing and inference flow.

[0121] Once the GPU has completed the inference task and generated the inference results, the system first updates the content in the inference result block to ensure that the inference results are accurately recorded and ready for subsequent processing. The data stored in the inference result block includes the output of the model, the intermediate data generated during the inference process, and the decision-making information of the inference. These data are transferred back to the main memory through a high-speed data transfer mechanism (such as PCIe) and are further used or processed by the result consumer thread. During this process, the write operation of the inference results is controlled by the GPU inference thread to ensure the integrity and accuracy of the inference results.

[0122] After the update of the inference result block is completed, the thread checks the status of the state machine again to ensure that the shared memory is in the "writable" state. At this time, the transition of the state machine is triggered by the GPU inference thread, indicating that the inference calculation has been successfully completed and the inference results have been successfully written to the shared memory. Only when the state machine is in the "writable" state can the data producer thread continue to perform the data write operation, thus ensuring the orderly progress of the data stream and avoiding data inconsistency or conflicts caused by concurrent operations. This state transition ensures a strict order in the data access process, prevents the inference results from being modified or overwritten, and provides a stable environment for the input of a new round of data.

[0123] After the state machine resumes the "writable" state, the data producer thread can continue to write the new video frame data to the shared memory, starting a new inference cycle. After the data writing is completed, the state machine turns into the "non-writable" state again to ensure that the data at this time will not be modified during the inference process. At this time, the GPU inference thread will read the new video frame data and process it to generate new inference results. The whole process proceeds in a loop. Whenever the GPU completes an inference and updates the results, the data producer thread can write new input data to the shared memory, thus starting the next round of inference tasks.

[0124] The loop execution of this process enables the system to continuously and stably process and infer video data in a concurrent environment. By precisely controlling the state transitions of the state machine, the system can ensure the reliability and consistency of data, avoiding resource competition and data errors that may occur due to multiple threads accessing the same shared memory. In application scenarios with high concurrency and high real-time requirements, this strict state control mechanism can significantly improve the throughput and response speed of the system, ensure the efficient execution of inference tasks, and at the same time ensure the stability and reliability of the system when processing large-scale video streams.

[0125] In the multi-threaded sharing mechanism of the shared memory, the system achieves efficient collaborative access to the same memory area through logical isolation and atomic state control. The shared memory is divided into functionally independent data blocks, state blocks, and inference result blocks. Among them, the state block embeds a state machine and associated metadata, which is used to coordinate the read-write timing and permission allocation of multiple threads. The state machine adopts a binary state identifier ("writable" and "non-writable") supported by atomic operations, and ensures the global visibility of state updates through a memory barrier (MemoryBarrier). When the state machine is in the "writable" state, any associated thread can preempt the state lock through an atomic operation, perform the write operation on the data block, and switch the state to "non-writable" after completion, while recording the data version number and verification information, so as to achieve mutual exclusion access and data consistency guarantee among multiple threads.

[0126] For the same piece of memory shared by n threads, the system optimizes resource utilization through time-sharing multiplexing and event-driven mechanisms. Specifically, each thread listens for state changes of the state machine through polling or interrupt notifications. When the state switches to "inference completed", all associated threads can concurrently read the structured data in the inference result block without competing for access permissions. After the reading is completed, the first thread to complete the operation is responsible for resetting the state machine to "writable" and triggering the memory cleaning and data block initialization process to prepare for the next round of writing tasks. During this process, the metadata (such as timestamps, thread identifiers) in the state block is used to track the data life cycle, preventing dirty reads or data overwrite problems caused by differences in thread execution timing.

[0127] Furthermore, the thread safety of the shared memory depends on the coordinated design of hardware-level atomic instructions and software-level synchronization strategies. For example, the atomic_compare_exchange instruction is used to implement lock-free switching of the state machine, and a mutex or semaphore is used to perform fine-grained control over the batch writing of data blocks. In addition, the write operation of the data block adopts a double buffering or ring buffer structure, so that when one thread executes the reading of the inference result, another thread can fill in new input data in parallel, thereby maximizing the memory bandwidth utilization. This mechanism significantly reduces the synchronization overhead of multi-threaded collaboration while ensuring data integrity through the dynamic flow of the state machine and the optimized design of the memory access mode, and is suitable for high-throughput, low-latency real-time inference scenarios.

[0128] In the technical solution of the above embodiment of the present invention, the efficiency and stability of model reasoning are achieved through systematic resource management and data interaction mechanism, and the traditional inter-process communication (IPC) technology is abandoned. Instead, the design paradigm of shared memory direct access combined with state machine collaboration is adopted, which significantly reduces the computing overhead and system complexity, as follows: First, the system loads the target model into the GPU memory at one time during the initialization phase, and completes the preheating and curing of the model through multiple rounds of forward reasoning. During the preheating process, the CUDA runtime compiles and optimizes the computational graph kernel, and pre-allocates the memory space for model weights, intermediate activation values, and output buffers. After the preheating is completed, the network parameters and computing context of the model instance are locked to form a static reasoning unit. This design ensures that all concurrent reasoning tasks reuse the same model instance, avoiding memory fragmentation and I / O bandwidth competition caused by repeated loading, and preventing multi-threaded access conflicts (such as through hardware-level memory protection mechanisms (such as cudaMemAttachGlobal)), fundamentally solving the risk of memory overflow caused by insufficient GPU resources.

[0129] Secondly, thread-safe inference prediction is achieved through the cooperation of shared memory and an atomic state machine. The shared memory area is divided into a data block, a state block, and a result block. The state block embeds binary state identifiers ("writable" and "non-writable") and metadata (such as data version numbers and check codes). When the state machine is in the "writable" state, the client process directly writes the pixel data of the video frame into the data block without serialization operations, and then switches the state to "non-writable" through atomic instructions, triggering the GPU inference task. The GPU scheduling process monitors the changes in the state machine, calls the pre-loaded model instance to perform inference, and writes the results into the result block. After completion, the state is restored to "writable" through an atomic operation. During this process, the data block and the result block use memory mapping technology to achieve zero-copy transmission between the CPU and the GPU, completely eliminating the computational latency and memory redundancy caused by serialization and deserialization.

[0130] Furthermore, the resource dynamic balancing mechanism is achieved through video memory occupancy quantization monitoring and task queue scheduling. The main control process continuously collects the video memory usage rate and computational load of each model instance, and uses a priority scheduling algorithm (such as minimum video memory occupancy first) to allocate inference requests. When the number of concurrent tasks exceeds the GPU video memory capacity, the system delays the execution of low-priority tasks through task queuing and resource reservation strategies to avoid GPU crashes caused by video memory overload. At the same time, the double-buffer design of the shared memory allows the write thread and the read thread to operate in parallel, ensuring seamless connection of the data stream and the computational stream in high-throughput scenarios.

[0131] In summary, through the four-fold technical cooperation of single model loading, direct storage and retrieval of shared memory, state machine cooperation, and dynamic resource scheduling, the present invention realizes low-latency and high-throughput real-time inference capabilities while ensuring thread safety and system stability. Its core advantage lies in avoiding the performance bottlenecks of traditional IPC technology through the efficient reuse of hardware resources and the extremely simplified design of data interaction, providing an extensible solution for model deployment in high-concurrency scenarios.

[0132] In the implementation case of the present invention, for a single model executing 100,000 video frame inference tasks in a single GPU environment, the efficiency and stability of the technical solution are systematically verified through comparison. The following analysis is carried out from three dimensions: the traditional paradigm, the IPC optimization scheme, and the present invention's scheme: 1. Performance Bottlenecks and Risks of the Traditional Paradigm In traditional implementations, each inference task needs to independently execute model loading, cache warming, and inference calculation. Taking the example that the single model loading takes 2 seconds, the warming takes 2 seconds, and the single-frame inference takes 30 ms, the total latency for completing 100,000 inferences (each containing 300 frames) can be decomposed as follows: Model loading and warm-up overhead: The model needs to be loaded and warmed up repeatedly for each inference task, with a total time consumption of about 1.2 hours. However, this calculation assumes that the model needs to be loaded independently for each batch, which cannot be achieved in the actual scenario due to GPU video memory limitations. Therefore, the system will crash due to exhausted video memory after the first task. In actual tests, the GPU process also crashed due to video memory overflow, and the task could not be completed.

[0133] 2. Limitations of the IPC Optimization Scheme The IPC scheme reduces the number of model loadings through inter-process communication, but still needs to perform data serialization and deserialization operations for each inference. Assume that in the actual IPC scheme, the model is only loaded once, and batch processing optimization reduces the serialization frequency. If N frames are processed in each batch, the total time consumption is: (10 5 * 300 / N) * (10ms * N + 30ms * N) + 4s ≈ 1.1 hours (assuming N = 1000) Even so, the serialization operation still occupies a significant proportion, and multi-process communication introduces scheduling complexity, making it difficult to avoid the risk of GPU video memory competition.

[0134] 3. Technical Advantages of the Solution of the Present Invention The present invention completely eliminates the key bottlenecks of the traditional scheme through direct storage and retrieval of shared memory, single model loading and warm-up, and cooperative scheduling of state machines: Model loading and warm-up overhead: It is only executed once globally, with a time consumption of 4 seconds, which can be ignored.

[0135] Data interaction overhead: Zero-copy transmission is achieved through shared memory, and the serialization / deserialization time consumption is reduced to 0.

[0136] Inference calculation time consumption: 10 5 ×300×30ms = 9×10 5 s (250 hours) of theoretical value. Through GPU multi-instance parallelism, pipelined scheduling, and saturated utilization of computing resources, it is actually compressed to 0.8 hours. The core optimizations include: Staticization of model instances: After warm-up, the model resides in the video memory, and the computational graph is solidified to reduce the single-frame inference time consumption to 25ms.

[0137] Double buffering and asynchronous pipeline: The writing thread and the inference thread operate in parallel to effectively hide the memory access latency.

[0138] Dynamic resource scheduling: The task queue is adjusted in real time according to the GPU utilization rate to avoid video memory overload and idle computing units.

[0139] 4. Summary of Performance Comparison Through techniques such as model instance resident in video memory, zero-copy interaction in shared memory, and lock-free cooperation driven by a state machine, the present invention reduces the time consumption by more than 50% in 100,000 inference tasks, while completely avoiding the risk of system crashes caused by GPU resource competition. Its technical value lies in providing a scalable and highly reliable solution for high-concurrency real-time inference scenarios through the efficient reuse of hardware resources and the minimalist design of software logic.

[0140] An embodiment of this application also provides an electronic device, including a processor, a memory, and a computer program that can run on the processor. When the processor executes the program, the above-described method is implemented.

[0141] An embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-described method is implemented.

[0142] An embodiment of this application also provides a computer product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the above-described method is implemented.

[0143] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0144] For the convenience of description, when describing the above device, various units are described separately according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0145] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the specified functions in one block or multiple blocks.

[0146] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0148] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0149] Those skilled in the art will appreciate that the embodiments of the present application may be provided as a method, system or computer program product. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0150] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0151] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0152] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for model inference using a GPU, comprising: Preloading and warming up multiple model instances in the GPU; Setting up a shared memory shared by at least two threads and setting up a state machine for the shared memory; The state machine may include two states: writable and non - writable; When the state machine is in the writable state, writing a video frame into the shared memory and setting the state machine to the non - writable state; Invoking the model instances loaded and warmed up in the GPU to perform inference based on the video frame in the shared memory; After completing the inference, setting the state machine to the writable state; The non - writable state in the state machine includes the writing - in state and the inference - in state, and the writable, writing - in, and inference - in states form a one - way cyclic conversion path; or, the non - writable state in the state machine includes the writing - in state, the inference - in state, and the inference - completed state, and the writable, writing - in, inference - in, inference - completed, and writable states form a one - way cyclic conversion path.

2. The method according to claim 1, wherein the state machine is set in a state block of the shared memory.

3. The method according to claim 2, wherein the state machine in the state block of the shared memory ensures atomicity of the state through a synchronization mechanism, and the synchronization mechanism includes a mutex lock and a semaphore.

4. The method according to claim 3, wherein the state block in the shared memory further includes metadata for indicating the current access state of the shared memory, and the metadata includes a combination of one or more of timestamp information, a processing identifier, an error flag, priority information, and a version number.

5. The method according to claim 4, wherein when the state machine is in the writable state, writing a video frame into the shared memory includes: When the state machine is in the writable state, reading the state identifier in the state block through an atomic operation. If the state is writable, preempting the state lock through an atomic instruction. After successful preemption, writing the original pixel data of the video frame into the data block of the shared memory by means of direct memory copy, memory mapping, double buffering, or a circular queue; After writing is completed, updating the metadata in the state block.

6. The method according to claim 1, wherein the shared memory further includes an inference result block for storing the result data generated during the inference process.

7. The method according to claim 1, wherein preloading and warming up multiple model instances in the GPU includes: Preloading and warming up multiple instances of the same model in the GPU; Or, Preloading and warming up multiple instances of different models in the GPU.

8. An electronic device, comprising a processor, a memory, and a computer program that can run on the processor, wherein when the processor executes the program, the method according to any one of claims 1 - 7 is implemented.

9. A computer - readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of claims 1 - 7 is implemented.

10. A computer software product, comprising a computer program / instructions, and when the computer program / instructions are executed by a processor, the method according to any one of claims 1 - 7 is implemented.

Citation Information

Cited By

  • Shared memory communication method and device, equipment, storage medium and program product

    CN121255500A

  • Shared memory communication method, apparatus, device, storage medium and program product

    CN121255500B

  • Image data processing method, system and device, storage medium and program product

    CN121280497A

  • Large-scale monocular vision SLAM-GS method and system based on depth prior and subgraph management

    CN122306046A