Flexible and Efficient Communication in a Microservice-Based Stream Analysis Pipeline

DataXc addresses inefficiencies in existing microservice communication by implementing a pull-based method that optimizes frame processing and reduces network bandwidth, enhancing the performance and accuracy of real-time video analysis.

JP2025522320AActive Publication Date: 2025-07-15NEC LABORATORIES AMERICA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024570435
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-23
Filing Date
2023-05-24
Publication Date
2025-07-15
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

Existing communication mechanisms for microservices in real-time streaming video analysis pipelines, such as NATS, gRPC, and Linkerd, face challenges in flexibility, network bandwidth utilization, and uneven frame processing, leading to inefficiencies and reduced accuracy in video analysis applications.

Method used

A pull-based communication method called DataXc, which uses a mesh controller to selectively assign inputs to detector and extractor sidecars, allowing microservices to pull data items for processing, thereby optimizing frame processing and reducing network bandwidth usage.

Benefits of technology

DataXc achieves higher processing speed, lower latency, and reduced network bandwidth consumption while maintaining flexibility, improving the accuracy and efficiency of real-time video analysis applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522320000001_ABST
    Figure 2025522320000001_ABST
Patent Text Reader

Abstract

A pull-based communication method for a micro-service-based real-time streaming video analysis pipeline is provided. This method receives a plurality of frames from a plurality of cameras each including a camera sidecar (1001), arranges a plurality of detectors in layers (1003), such that the first detector layer includes detectors having a detector sidecar and detector business logic, and the second detector layer includes detectors having only a sidecar, arranges a plurality of extractors in layers (1005), such that the first extractor layer includes extractors having an extractor sidecar and extractor business logic, and the second extractor layer includes extractors having only a sidecar, and during registration, a mesh controller selectively assigns inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer to enable pulling data items for processing (1007).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Application Information This application claims the benefit of U.S. Provisional Patent Application No. 63 / 351,803, filed Jun. 13, 2022, and U.S. Patent Application No. 18 / 321,880, filed May 23, 2023, the entire contents of each of which are incorporated herein by reference.

Background Art

[0002] The present invention relates to micro-service-based applications, and more specifically, to flexible and efficient communication in a micro-service-based stream analysis pipeline. Description of Related Art

[0003] A major challenge in changing monolithic applications to high-performance micro-service-based applications is designing an efficient mechanism for the micro-services to communicate with each other. Previous approaches have ranged from custom point-to-point communication between micro-services using protocols such as gRPC, to flexible many-to-many communication networks using service meshes such as Linkerd and broker-based messaging systems such as NATS.

Summary of the Invention

[0004] A pull-based communication method for a microservice-based real-time streaming video analysis pipeline is presented. This method includes receiving a plurality of frames from a plurality of cameras, each corresponding to a camera driver and each including a camera sidecar; arranging a plurality of detectors in layers such that a first detector layer includes detectors with a detector sidecar and detector business logic and a second detector layer includes detectors with only a sidecar; arranging a plurality of extractors in layers such that a first extractor layer includes extractors with an extractor sidecar and extractor business logic and a second extractor layer includes extractors with only a sidecar; enabling a mesh controller to communicate only with the detector sidecar and the extractor sidecar of the first layer, and the mesh controller selectively allocating inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer during registration to pull data items for processing.

[0005] A non-transitory computer-readable storage medium containing a computer-readable program for a microservices-based real-time streaming video analysis pipeline is presented. When the computer-readable program is executed on a computer, it causes the computer to perform steps including receiving a plurality of frames from a plurality of cameras, each corresponding to a camera driver and each including a camera sidecar; arranging a plurality of detectors in layers such that a first detector layer includes detectors having a detector sidecar and detector business logic, and a second detector layer includes detectors having only a sidecar; arranging a plurality of extractors in layers such that a first extractor layer includes extractors having an extractor sidecar and extractor business logic, and a second extractor layer includes extractors having only a sidecar; enabling a mesh controller to communicate only with the detector sidecars and extractor sidecars of the first layer, and during registration, the mesh controller selectively assigns inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer to pull out data items for processing.

[0006] A system is presented for a microservices-based real-time streaming video analysis pipeline. The system includes a processor and a memory storing a computer program that, when executed by the processor, causes the processor to receive a plurality of frames from a plurality of cameras, each corresponding to a camera driver and each including a camera sidecar, and to arrange a plurality of detectors in layers such that a first detector layer includes a detector having a detector sidecar and detector business logic, a second detector layer includes a detector having only a sidecar, arrange a plurality of extractors in layers such that a first extractor layer includes an extractor having an extractor sidecar and extractor business logic, a second extractor layer includes an extractor having only a sidecar, and enable a mesh controller to communicate only with the detector sidecar and the extractor sidecar of the first layer, the mesh controller selectively allocating inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer during registration to draw data items for processing.

[0007] These and other features and advantages will become apparent from the following detailed description of its exemplary embodiments, read in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0008] The present disclosure provides details in the following description of preferred embodiments with reference to the following figures.

[0009]

Figure 1

[0010]

Figure 2

[0011]

Figure 3

[0012]

Figure 4

[0013]

Figure 5

[0014]

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0015] In a monolithic application executed in a single process, software components communicate with each other using language-level method or function calls. A major challenge in converting a monolithic application into a high-performance microservice-based application is designing an efficient mechanism for microservices to communicate with each other. An exemplary approach focuses on a microservice-based real-time streaming video analysis pipeline and presents a new communication architecture for flexible and efficient communication between microservices.

[0016] Thanks to the remarkable progress of computer vision and machine learning technologies, the growth of the Internet of Things (IoT), and more recently, the advancements in edge computing and 5G networks, video analysis systems have become widespread. This significant progress in independent yet related fields has led to the large-scale introduction of cameras worldwide, enabling new camera-based use cases in various market segments such as retail, healthcare, transportation, automotive, entertainment, security, and safety. The global video analysis market is expected to grow to $21 billion by 2027.

[0017] Real-time video analysis applications are pipelines for various video analysis tasks. Until recently, such pipelines were implemented as monolithic applications. A recent trend is to change monolithic applications into microservice-based applications where each analysis task is a microservice. In one example, the analysis pipeline starts with a frame decoding task, followed by preprocessing and video analysis tasks. Subsequently, postprocessing and the generation of analysis outputs are performed, and insights are disseminated. These tasks, i.e., microservices, rely on an efficient communication architecture to communicate with each other. For communication between microservices, three common well-known communication mechanisms called service meshes such as NATS, gRPC, and Linkerd are used.

[0018] NATS uses a broker for communication, and all data exchanges are carried out via the broker. This makes the system very flexible as new services only need to interact with the broker, allowing new microservices to be added on the fly. However, due to the presence of the broker in the middle, additional copies of messages are created, significantly increasing the utilization of network bandwidth.

[0019] In gRPC-based communication, there is no broker, and microservices communicate directly with each other. However, due to this direct communication, microservices need to recognize each other, and it is not possible to add new microservices on the fly (while the system is running), which reduces the flexibility of the system. When a new microservice is added, it is necessary to notify other microservices and change to start communicating with the new microservice. Therefore, this is a more rigid system, but because direct communication is carried out, it is more efficient in terms of network bandwidth utilization.

[0020] gRPC or Linkerd tend to exhibit better performance than NATS in terms of other performance metrics such as processing speed, latency, and jitter.

[0021] In an exemplary approach, a new communication architecture DataXc is proposed that is as flexible as NATS while being more efficient than NATS, gRPC, and Linkerd.

[0022] In real-time video analysis applications, several scenarios and challenges arise that are not considered in technologies such as NATS, gRPC, and Linkerd. Video cameras have a configuration parameter called the frame rate, which varies in the range of 1 to 60 frames per second (FPS). Once set, these cameras continue to send frames at the specified frame rate throughout the deployment's active period. Cameras generate a large number of frames per second, but not all frames are processed throughout the pipeline. For example, consider a face recognition application pipeline. The camera driver can decode and process all the frames generated by the camera, but subsequent microservices in the pipeline (such as face detection and feature extraction) take approximately 200 to 250 milliseconds per frame to process and thus cannot keep up. Therefore, since not all frames are processed throughout the pipeline, processing all frames by the camera driver is a waste of resources. In NATS, gRPC, or Linkerd-based communication, it is a "push"-based communication that blindly "pushes" data items to the next microservice in the chain for processing regardless of whether the data items can be further processed, and thus this insight cannot be utilized. In contrast, the exemplary technology DataXc is a "pull"-based technology, where the consumer of data items fetches only the data items it can process, avoiding waste of communication resources.

[0023] Since not all frames can be processed, another scenario and challenge in real-time video analysis is which frames among the processable frames to process and which to drop. Using NATS or gRPC, it can be seen that frames are processed unevenly. That is, there may be cases where 4 to 5 consecutive frames are processed, cases where 2 to 3 frames are dropped, cases where 10 to 12 frames are dropped, and even cases where 20 to 23 frames are dropped. Since this frame drop is variable, it is not guaranteed that a particular camera will process frames every 6 frames or every 10 frames depending on the processing speed. Since object tracking depends greatly on the order of the processed frames, random frames are processed, directly affecting the accuracy of the analysis. Such tracking is necessary in video analysis applications such as people counting. When random frames are processed, person tracking is lost and the same person may be counted multiple times, resulting in a loss of application accuracy. This unevenness in frame processing observed with NATS and gRPC is intentionally avoided in DataXc, which focuses on processing frames at equal intervals.

[0024] In summary, the contributions are as follows.

[0025] A new communication method called DataXc is introduced. It has the same flexibility as NATS, high utilization efficiency of network bandwidth, and exhibits better performance than NATS, gRPC, and Linkerd in terms of other metrics such as processing speed, latency, and jitter. Various issues occurring in communication between microservices, especially those occurring in real-time video analysis applications, are introduced, and it is shown that the design of DataXc is suitable for overcoming these issues (compared to conventional proposals such as NATS, gRPC, and Linkerd). Two actual video analysis applications are implemented, and it is shown that DataXc achieves a processing speed up to 80% higher, a latency ~3 times lower, a jitter ~7.5 times lower, and a network bandwidth consumption ~4.5 times lower than NATS while maintaining the same flexibility as NATS. Compared to gRPC and Linkerd, DataXc is very flexible and achieves a processing speed up to 2 times higher (excellent), low latency, and low jitter, but also increases the consumption of network bandwidth.

[0026] Regarding the issues of microservice communication and the increase in waiting time for processing data items, when the number of consumers and producers is variable and there is direct access between producers and consumers, one consumer may be overloaded with data items from multiple producers. In such a scenario, data items from the first producer are processed immediately, but data items from other producers that arrive a little later need to wait until their turn for processing. This is because consumers can only process data items in the order of arrival first-come, first-served. Therefore, when data items are not evenly distributed and one consumer receives more data items than other consumers, the waiting time for processing data items from producers becomes longer. Furthermore, in the case of real-time applications where it is important to process the latest data items, specific data items may be dropped without being processed because the waiting time is too long to be processed at a sufficient speed.

[0027] Regarding the issues of microservice communication and non-uniform processing among producers, multiple replicas of producer microservices and consumer microservices are running at a specific point in time. When a one-to-one mapping is maintained between a producer and a consumer, the architecture becomes a very simple one where one producer and one consumer are directly connected. However, in such an architecture, artificial restrictions are imposed by the architecture, and the parallel processing that can be achieved by scaling the consumer is not utilized. That is, when multiple consumers are provided, the data items generated by the producer have the potential to be processed in parallel much faster. However, it is unrealistic and in some cases impossible to scale the consumer to the extent that all the data items generated by the producer can be consumed. Therefore, situations are provided where various numbers of producers and various numbers of consumers are used. All of these consumers can process data items from any producer, output the generated insights, and then proceed to process the next data item again. When multiple replicas of consumers are running in parallel, there is a possibility that data items of some producers are processed more than other data items, and the variation in the number of data items processed per producer can become very large.

[0028] Regarding the issues of microservice communication and the difference in consumption speed, consumer microservices typically process data items, generate insights, and output these insights to generate data items for the next microservice in the pipeline. The above procedures (processing, output, etc.) are repeated many times for each new data item. Since the processing logic is the same across all replicas of a particular consumer microservice, the time taken for each replica to process a data item is approximately the same (assuming they are running on equivalent hardware). However, if data items from a single producer microservice are consumed by two different consumer microservices with different processing logics, the speed at which the data items are processed can vary significantly. This difference in processing speed results in a difference in consumption speed across the consumer microservices. Moreover, the processing time may be determined by the content within the frame, and even for the same microservice, the consumption speed may vary depending on the video content.

[0029] As an example, consider a microservice that connects to a camera, captures and decodes frames, and outputs the frames for further use in a video analysis application pipeline. Assume that two video analysis applications, such as face recognition and motion detection, are running on the same camera feed. In the case of the face recognition application, the first step takes approximately 200 - 250 milliseconds. On the other hand, the first step of the motion detection application is background subtraction, which takes only about 20 milliseconds. There is a tenfold difference in the time taken to process a particular data item between these two different video analysis applications running on the same camera feed. This difference in processing time results in a difference in consumption speed.

[0030] According to the number of consumers consuming from a specific producer and the consumption rate of each consumer, the processing speed on the producer side can be easily adjusted. That is, if a consumer can only process 5 data items per second, the producer does not need to generate more than 5 data items, and those data will be wasted and discarded. Similarly, when there are multiple consumers, the generation speed is sufficient to match the maximum consumption speed. Otherwise, the producer will consume wasted computational cycles to generate data items that will never be used. DataXc alleviates this problem.

[0031] Regarding NATS, in an implementation using NATS, there is a NATS message queue where all communications occur. The camera driver acquires frames from the camera, decodes them, and writes ( "pushes") them to the queue. There are multiple detector replicas that acquire frames from the queue and write back the detected faces ( "push"). Next, multiple extractor replicas acquire the detected faces, extract the features of these faces, and write them back to the queue ( "push"). There are tags corresponding to the camera and the frame, which are used to identify frames from a specific camera during processing. In this implementation, the camera driver, detector replicas, and extractor replicas do not need to recognize each other. All that needs to be recognized is the message format consumed from the NATS message queue. This provides the flexibility to add new data streams on the fly without the need to modify the existing system. Also, data is continuously written to the message queue, and there is always work that can be executed by any microservice, so the processing speed, that is, the number of messages processed per second, is also very high. However, this implementation has a drawback. When communication occurs via the message queue, an additional copy of each message will be created. This increases the overall network bandwidth usage of the application.

[0032] Regarding gRPC, unlike NATS in gRPC-based communication, each component needs to recognize subsequent components in the pipeline, and code changes in the microservices are required to change the number of replicas or add new streams. In this implementation, the detector stream and the extractor stream are set as services, and there can be m replicas of the detector service and n replicas of the extractor service. The camera driver queries the Domain Name System (DNS) service discovery to detect the detector service, and each driver receives a response regarding the available detector services. The camera driver receives all available detectors, but it is important to note that the lists of these services are received in different orders. Here, each of these camera drivers obtains a list of detector services and starts to "push" data to the detector services in a round-robin manner according to the same sequence received in the DNS query. Therefore, all camera drivers start to "push" data, and the data is input into the detector services in a round-robin manner, starting from different detector services. The same applies when the detector service "pushes" data to the extractor service. In this implementation, there is no extra copy by direct communication, but the flexibility provided by NATS is lost, and the entire application pipeline becomes very rigid.

[0033] Regarding Linkerd using gRPC, NATS provides deployment flexibility at the cost of high network bandwidth usage, while gRPC using client - based load balancing requires less network bandwidth usage but is very robust. Another implementation using gRPC is to deploy Linkerd and let it handle all gRPC - based communications between microservices. In such a configuration, there is a "Linkerd controller" to assist with load balancing, and an additional "proxy" for actual communication between microservices is added to the microservices. These "proxies" are implemented as "sidecars" in a Kubernetes - based cluster deployment. These "proxy" replicas of detectors receive frames and either process them locally or send them to other replica proxies for processing. All proxies are interconnected, and based on load and associated policies, one of the proxies performs the processing and "pushes" the output to subsequent microservices in the pipeline.

[0034] It is noted that NATS has the drawback of excessive network bandwidth consumption despite its flexibility. In implementations using gRPC with client - based load balancing, the network bandwidth consumption is reduced, but the application pipeline becomes very robust and flexibility is lost. Linkerd using gRPC alleviates rigidity and has slightly more flexibility, but is not as excellent as NATS in terms of processing speed and fine - grained control. Therefore, none of the existing well - known communication mechanisms work well in the scenario of a real - time microservice - based video analysis application pipeline. To mitigate these problems, an exemplary approach proposes a new communication mechanism called DataXc that uses "pull" - based communication (Figure 1). Figure 1 shows push - based communication NATS and gPRC10 and pull - based communication in DataXc100.

[0035] To support "pull"-based communication, DataXc provides a Software Development Kit (SDK) with specific Application Programming Interfaces (APIs) that need to be used in microservices. The SDK is available in various programming languages such as Go, Python, C++, Java. These APIs are implemented using the idioms of the programming languages and are very simple and lightweight. The three main APIs provided by the DataXc SDK are as follows.

[0036] get-confguration(): Gets the configuration provided at the time of registration of a specific sensor or stream.

[0037] next(): Receives the first available data item from the input stream. This function returns the data item and the name of the stream that generated the data item. If there are multiple input streams, the name of the stream can be used to identify the source of the input data item.

[0038] emit(data item): Publishes a data item to the output stream. All data items from a specific driver or AU are sent to the same output stream. These three APIs are sufficient for DataXc to manage communication between microservices.

[0039] Regarding the DataXc architecture 300 (Figure 3), similar to the Linkerd "proxy", there is a sidecar connected to the microservice pods within DataXc 300, and this sidecar processes all communication within DataXc 300. One of the major differences in the DataXc architecture 300 is that microservices do not "push" data to a message queue (NATS) or other microservices (gRPC), but actually "pull" (fetch) data from other microservices for processing.

[0040] For example, as shown in Figure 2, there is a camera driver C having frames available as data item 210 z There are two microservice replicas that consume these frames. That is, the "converter" replica 220 and the "background subtractor" replica 230. In this case, there are two slots 205 in the sidecar of the camera driver, and the latest frame is copied into the slot each time. Different replicas of a specific microservice move to the same slot and "pull" the latest frame available for processing.

[0041] Figure 3 shows the architecture implemented within DataXc300 for the logical view created by the user. Each of cameras A through Z (305) has a corresponding camera driver 310, detector stream 320, and extractor stream 330. The detector stream 320 and the extractor stream 330 are "sidecar only" pods, and there are multiple replicas of detectors (D1 through D m ) and extractors (E1 through E n ). All of these are packaged as Kubemetes pods together with the business logic and the sidecar container. The sidecar implements the DataXc API (part of the SDK) used by the microservices and manages all communications between various microservices by communicating with the DataXc mesh controller 350. Each time a pod is launched, the sidecar container is registered with the mesh controller 350. As part of the registration process, the stream sidecar provides the name of the stream and the name of the input stream. In the case of replicas, the sidecar provides the name of the replica and the pod name at the time of registration. In response to the registration, the mesh controller 350 assigns inputs to the sidecars that should connect and "pull" the data items.

[0042] Each sidecar can obtain a different set of inputs. For example, the sidecar of D1 is assigned C A and C B , while the sidecar of D m is assigned C B and C Zare assigned. These sidecars move to their assigned inputs, pull data items, process them with business logic, and then push them to the corresponding data stream sidecars (the "sidecar-only" pods). This prepares the output in the slot for subsequent consumers such as extractors in the pipeline to use. To E1, D A and D Z are assigned, and to E n D A and D B are assigned. Therefore, the detector stream D A has two slots, one for E1 and one for E n , and D B and D Z each have one slot, one for E n and one for E1. All outputs from the streams associated with the cameras are properly routed via the sidecars to the streams of their respective cameras. Thus, new microservices can directly utilize the intermediate stream outputs and build new applications by reusing the outputs from all existing computations without having to reconstruct everything from scratch.

[0043] The sidecars to which inputs are assigned via the mesh controller 350 periodically connect to the mesh controller 350 to obtain any new assignments. If there is a new assignment, they begin to "pull" data items from the newly assigned input. Otherwise, they continue to "pull" from the previously assigned input. Thus, by design, the application does not go down even if the mesh controller 350 restarts. For example, if the node on which the mesh controller 350 is running goes down, Kubemates starts it on another node, and the application continues to run using the previous inputs assigned to the sidecars by the mesh controller 350.

[0044] Since DataXc300 adopts a "pull"-based design, it can control the production speed at the producer and avoid unnecessary generation of data items that will not be consumed. This design choice helps to address the above issues. Also, by equipping with the mesh controller 350, DataXc300 can periodically adjust the input assignment, so that data items from different producers are processed uniformly, the processing waiting time of data items is minimized, and the above issues can be addressed.

[0045] Another advantage of the architecture and design of DataXc300 is that with the "pull"-based design, data items that are not processed do not pass through the network. In this case, if the data items are not processed fast enough for real-time applications, they are dropped locally on the producer side. In contrast, in a "push"-based design, data items reach the consumer through the network regardless of whether they are processed by the consumer. This "pull"-based design where data items pass through the network only when needed significantly saves wasted network bandwidth.

[0046] In addition to the above advantages, using the DataXc SDK makes the actual mechanism of communication between microservices transparent to developers. Developers do not need to worry about or learn about various communication methods, nor do they need to judge which one is optimal for themselves. Instead, DataXc300 automatically processes communication in the background through the sidecar. Furthermore, even if the underlying communication mechanism used by DataXc changes in the future, developers do not need to modify the code they have written.

[0047] In conclusion, in an exemplary approach, a new communication mechanism, DataXc, is proposed that is more efficient than previous proposals with respect to message latency, jitter, message processing speed, and network resource usage. DataXc is the first communication design that combines the desirable flexibility of a broker-based messaging system such as NATS and the high performance of a robust custom point-to-point communication scheme such as gRPC. DataXc proposes a new "pull"-based communication method (e.g., a consumer retrieves messages from a producer). This is different from conventional proposals such as NATS, gRPC, Linkerd, etc. These are all "push"-based (e.g., a producer sends messages to a consumer). With such a communication method, it becomes difficult to utilize the different processing speeds of consumers such as video analysis tasks. In contrast, DataXc300 proposes a "pull"-based design that avoids unnecessary communication of messages that will ultimately be discarded by the consumer. Also, unlike previous proposals, DataXc300 appropriately addresses several important issues in a streaming video analysis pipeline that all negatively impact the quality of insights from streaming video analysis, such as non-uniform processing of frames from multiple cameras and large variations in the latency of frames processed by the consumer.

[0048] Figure 4 is a block / flow diagram of an exemplary action recognition pipeline according to an embodiment of the present invention.

[0049] The real-world action recognition application 400 is shown in Figure 4. In a practical example, 16 camera drivers 420 corresponding to 16 cameras 410, 2 replicas of object detection 430, 4 replicas of feature extraction 450, and 16 replicas of object tracking 440 and action recognition 460 (each associated with a camera) are used to obtain actions 470. The object detection 430, feature extraction 450, and action recognition 460 microservices utilize a graphics processing unit (GPU) for speedup. It can be seen that DataXc300 (Figure 3) achieves the highest processing speed, the lowest average latency, the lowest jitter, and a slightly higher but much lower network bandwidth utilization rate compared to NATS. In the case of the action recognition application 400, while maintaining the same flexibility as NATS, the processing speed is improved by ~80%, the average latency is reduced by ~2 times, the jitter is reduced by ~7.5 times, and the consumption of network bandwidth is reduced by ~4 times.

[0050] The processing speed is defined as the number of messages processed per second. Generally, the higher the processing speed, the higher the accuracy of the analysis.

[0051] Latency is the time it takes to process one frame end-to-end across the entire pipeline. To obtain analysis information as early as possible, it is recommended to reduce latency.

[0052] Jitter defines how accurately the system processes equally spaced frames. A low jitter means that the system is approaching the processing of equally spaced frames. Since this directly affects the accuracy of the analysis, the lower the jitter, the higher the accuracy of the analysis.

[0053] The utilization rate of network bandwidth is the total utilization rate of network bandwidth within the system. It is desirable to reduce the network bandwidth so that the system does not consume network bandwidth unnecessarily, causing network congestion or increasing network latency.

[0054] FIG. 5 shows an exemplary processing system for a microservice-based real-time streaming video analysis pipeline according to an embodiment of the present invention.

[0055] The processing system includes at least one processor (CPU) 904 operably coupled to other components via a system bus 902. A GPU 905, a cache 906, a read-only memory (ROM) 908, a random access memory (RAM) 910, an input / output (I / O) adapter 920, a network adapter 930, a user interface adapter 940, and a display adapter 950 are operably coupled to the system bus 902. Further, DataXc300 is connected to the bus 902.

[0056] The storage device 922 is operably coupled to the system bus 902 by an I / O adapter 920. The storage device 922 can be any of a disk storage device (e.g., a magnetic or optical disk storage device), a solid-state magnetic device, etc.

[0057] The transceiver 932 is operably coupled to the system bus 902 by a network adapter 930.

[0058] The user input device 942 is operably coupled to the system bus 902 by a user interface adapter 940. The user input device 942 can be any of a keyboard, a mouse, a keypad, an image capture device, a motion sensing device, a microphone, a device incorporating at least two functions of the aforementioned devices, etc. Of course, other types of input devices can also be used while maintaining the spirit of the present invention. The user input device 942 can be the same type of user input device or different types of user input devices. The user input device 942 is used to input and output information to and from the processing system.

[0059] The display device 952 is operably coupled to the system bus 902 by a display adapter 950.

[0060] Of course, the processing system can also include other elements (not shown) and can omit certain elements, as can be easily contemplated by those skilled in the art. For example, as can be easily understood by those skilled in the art, various other input devices and / or output devices can be included in the system according to its specific implementation. For example, various types of wireless and / or wired input and / or output devices can be used. Further, additional processors, controllers, memories, etc. of various configurations can also be utilized as can be easily understood by those skilled in the art. These and other variations of the processing system can be easily contemplated by those skilled in the art given the teachings of the present invention provided herein.

[0061] FIG. 6 is a block / flow diagram of an exemplary pull-based communication method for a microservice-based real-time streaming video analysis pipeline according to an embodiment of the present invention.

[0062] In block 1001, a plurality of frames are received from a plurality of cameras. Each camera corresponds to its respective camera driver, and each camera includes a camera-sidecar.

[0063] In block 1003, a plurality of detectors are arranged in layers such that the first detector layer includes a detector having a detector-sidecar and detector business logic, and the second detector layer includes a detector having only a sidecar.

[0064] In block 1005, a plurality of extractors are arranged in layers such that the first extractor layer includes an extractor having an extractor-sidecar and extractor business logic, and the second extractor layer includes an extractor having only a sidecar.

[0065] In block 1007, the mesh controller is configured such that it can communicate only with the detector sidecar and the extractor sidecar of the first layer. During registration, the mesh controller selectively assigns inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer, and extracts data items for processing.

[0066] The latest applications are described as a collection of microservices that interact with each other. The performance of communication between microservices has a significant impact on the overall performance of the application. An exemplary method focuses particularly on real-time video analysis applications, proposes a new communication method called DataXc, and compares it with well-known gRPC and NATS-based communication methods.

[0067] As used herein, the terms "data," "content," "information," and similar terms can be used interchangeably to refer to data that can be captured, transmitted, received, displayed, and / or stored according to various exemplary embodiments. Accordingly, the use of such terms should not be construed as limiting the spirit and scope of the present disclosure. Further, when an arithmetic device is described herein as receiving data from another arithmetic device, the data can be received directly from the other arithmetic device or indirectly via one or more intermediate arithmetic devices such as, for example, one or more servers, relays, routers, network access points, base stations, and / or the like. Similarly, when an arithmetic device is described herein as transmitting data to another arithmetic device, the data can be transmitted directly to the other arithmetic device or indirectly via one or more intermediate arithmetic devices such as one or more servers, relays, routers, network access points, base stations, etc.

[0068] As will be understood by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may generally be referred to herein as a "circuit," "module," "computer," "device," or "system." Further, aspects of the present invention may take the form of a computer program product embodied on one or more computer-readable media having computer-readable program code embodied thereon.

[0069] Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical data storage device, a magnetic data storage device, or any suitable combination of the foregoing. As used herein, a computer-readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0070] A computer-readable signal medium can include a propagated data signal in which computer-readable program code is embodied, for example, as part of a baseband or a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electromagnetic waves, optical, or suitable combinations thereof. A computer-readable signal medium may be any computer-readable medium that can communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device, rather than a computer-readable storage medium.

[0071] The program code embodied on the computer-readable medium can be transmitted using any appropriate medium, including, but not limited to, wireless, wired, fiber optic cable, RF, or any suitable combination of the foregoing.

[0072] The computer program code for performing operations for aspects of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0073] Aspects of the present invention will be described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0074] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.

[0075] Alternatively, the computer program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0076] It should be understood that the term "processor" as used herein is intended to include any processing device, such as, for example, a CPU (Central Processing Unit) and / or other processing circuits. Also, it should be understood that the term "processor" may refer to multiple processing devices, and various elements associated with the processing device may be shared by other processing devices.

[0077] The term "memory" as used herein is intended to include memory associated with a processor or CPU, such as, for example, RAM, ROM, fixed memory devices (e.g., hard drives), removable memory devices (e.g., floppy disks), flash memory, etc. Such memory is considered to be a computer-readable storage medium.

[0078] Furthermore, the phrase "input / output device" or "I / O device" as used herein is intended to include, for example, one or more input devices (e.g., keyboard, mouse, scanner, etc.) for inputting data to the processing unit, and / or one or more output devices (e.g., speaker, display, printer, etc.) for presenting results associated with the processing unit.

[0079] The above is illustrative and exemplary in all respects and not limiting. It is understood that the scope of the invention disclosed herein is determined from the claims construed in accordance with the full breadth permitted by patent law, rather than from the detailed description. The embodiments shown and described herein are merely illustrative of the principles of the invention, and it should be understood that those skilled in the art can make various modifications without departing from the scope and spirit of the invention. Those skilled in the art can implement various other combinations of features without departing from the scope and spirit of the invention. Thus, while the aspects of the invention have been described with the detail and particularity required by patent law, what is desired to be claimed and protected by patent is as set forth in the appended claims.

Claims

1. A computer-implemented method for a microservice-based real-time streaming video analysis pipeline, comprising: Receiving a plurality of frames from a plurality of cameras, each corresponding to a camera driver and each including a camera sidecar (1001); Arranging a plurality of detectors in layers such that a first detector layer includes detectors with detector sidecars and detector business logic, and a second detector layer includes detectors with sidecars only (1003); Arranging a plurality of extractors in layers such that a first extractor layer includes extractors with extractor sidecars and extractor business logic, and a second extractor layer includes extractors with sidecars only (1005); Enabling a mesh controller to communicate only with the detector sidecars and extractor sidecars of the first layer, and during registration, the mesh controller selectively assigns inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer to pull data items for processing (1007).

2. The computer-implemented method according to claim 1, wherein: The plurality of detectors and the plurality of extractors are packaged as Kubernetes pods.

3. The computer-implemented method according to claim 1, wherein: The camera sidecars of the plurality of cameras communicate directly only with the detector sidecars of the first detector layer.

4. The computer-implemented method according to claim 1, wherein: The sidecars of the second detector layer communicate directly only with the extractor sidecars of the first extractor layer.

5. The computer-implemented method according to claim 1, wherein: The first detector layer includes detector replicas, the first extractor layer includes extractor replicas, and the detector replicas are prevented from direct communication with the extractor replicas.

6. The computer-implemented method according to claim 1, wherein: During registration, the second detector layer provides the name of the detector stream and the name of the input stream to the mesh controller.

7. The computer-implemented method according to claim 1, wherein: A method for the first detector layer to provide the replica name and pod name to the mesh controller during registration.

8. A computer program product for a microservice-based real-time streaming video analysis pipeline, comprising a non-transitory computer-readable storage medium having program instructions incorporated therein, the program instructions being executable by a computer, and causing the computer to receive a plurality of frames from a plurality of cameras, each corresponding to a camera driver and each including a camera sidecar (1001); layer a plurality of detectors such that the first detector layer includes detectors with a detector sidecar and detector business logic, and the second detector layer includes detectors with only a sidecar (1003); layer a plurality of extractors such that the first extractor layer includes extractors with an extractor sidecar and extractor business logic, and the second extractor layer includes extractors with only a sidecar (1005); enable the mesh controller to communicate only with the detector sidecar and the extractor sidecar of the first layer, and during registration, the mesh controller selectively assigns inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer to pull data items for processing (1007) for a computer program product for executing a method including.

9. The computer program product according to claim 8, wherein the plurality of detectors and the plurality of extractors are packaged as Kubernetes pods.

10. The computer program product according to claim 8, wherein the camera sidecars of the plurality of cameras communicate directly only with the detector sidecars of the first detector layer.

11. The computer program product according to claim 8, wherein the sidecars of the second detector layer communicate directly only with the extractor sidecars of the first extractor layer.

12. The computer program product according to claim 8, The first detector layer includes a detector replica, and the first extractor layer includes an extractor replica. The detector replica is a computer program product that prevents direct communication with the extractor replica.

13. The computer program product according to claim 8, During registration, a computer program product in which the second detector layer provides the name of the detector stream and the name of the input stream to the mesh controller.

14. The computer program product according to claim 8, During registration, a computer program product in which the first detector layer provides the name of the replica and the pod name to the mesh controller.

15. A computer processing system for a microservice-based real-time streaming video analysis pipeline, A storage device for storing program code, A processor device operatively coupled to the storage device, the processor device Receives a plurality of frames from a plurality of cameras (1001), each corresponding to a camera driver and each including a camera sidecar, Arrange a plurality of detectors in layers (1003) such that the first detector layer includes detectors with a detector sidecar and detector business logic, and the second detector layer includes detectors with only a sidecar, Arrange a plurality of extractors in layers (1005) such that the first extractor layer includes extractors with an extractor sidecar and extractor business logic, and the second extractor layer includes extractors with only a sidecar, A computer processing system that enables the mesh controller to communicate only with the detector sidecar and the extractor sidecar of the first layer, and the mesh controller selectively assigns inputs to one or more detector sidecars of the first detector layer and one or more extractor sidecars of the first extractor layer during registration to pull out data items for processing (1007) to execute the program code.

16. The computer processing system according to claim 15, The computer processing system in which the plurality of detectors and the plurality of extractors are packaged as Kubernetes pods.

17. The computer processing system according to claim 15, The camera sidecars of the plurality of cameras are computer processing systems that communicate directly only with the detector sidecar of the first detector layer. **Claim 18** The computer processing system according to claim 15, wherein the sidecar of the second detector layer is a computer processing system that communicates directly only with the extractor sidecar of the first extractor layer. **Claim 19** The computer processing system according to claim 15, wherein the first detector layer includes a detector replica, the first extractor layer includes an extractor replica, and the detector replica is prevented from directly communicating with the extractor replica. **Claim 20** The computer processing system according to claim 15, wherein during registration, the second detector layer provides the name of the detector stream and the name of the input stream to the mesh controller, and during registration, the first detector layer provides the name of the replica and the pod name to the mesh controller.

Citation Information

Patent Citations

  • Video monitoring system, control method of video monitoring system, and video monitoring device

    JP2018093324A

  • Information processing apparatus, information processing system, and program

    JP2020107268A

  • Differentiated smart sidecars in a service mesh

    US20200329114A1