Performance optimization method and device of edge box and electronic equipment

By dividing the video stream data into multiple video frames for parallel processing within the edge box and preloading configuration information, the problem of limited image frame rate in edge boxes is solved, achieving higher frame rates and lower algorithm switching time.

CN120915974APending Publication Date: 2025-11-07BEIJING TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511131204.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Edge boxes are limited in processing image frame rates. In existing technologies, configuration information needs to be reloaded when switching algorithm streams, which increases processing time and reduces frame rate.

Method used

The video stream data is divided into multiple video frame data and processed in parallel by multiple processing units. The configuration information of all computing nodes is preloaded at once to ensure that the frame rate meets the preset threshold before outputting the target video stream.

Benefits of technology

It improves the frame rate of image processing, reduces the time spent switching algorithm streams, and ensures the stability and efficiency of edge boxes when switching between different algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915974A_ABST
    Figure CN120915974A_ABST
Patent Text Reader

Abstract

The invention provides a performance optimization method and device of an edge box and electronic equipment, and the method comprises the steps: responding to video stream data received by the edge box, dividing the video stream data into a plurality of pieces of video frame data according to a time sequence, and processing the plurality of pieces of video frame data through a plurality of processing units, the method comprises the steps of obtaining a plurality of pieces of processed video frame data, determining whether the frame rate of each piece of processed video frame data meets a preset frame rate threshold, and obtaining a target video stream based on processing results of a plurality of processing units after determining the frame rate of each piece of processed video frame data. Parallel processing of multiple pieces of video frame data obtained through division through different processing units is achieved, configuration information needed by the processing units is loaded in advance at a time, when different algorithms of the processing units are switched, the configuration information does not need to be loaded again, time consumed for algorithm processing when the processing units are switched is shortened, and the switching efficiency of the processing units is improved. Therefore, the frame rate of image processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, and particularly relates to a performance optimization method and device of an edge box and an electronic device. BACKGROUND

[0002] With the rise of edge computing, as a hardware device of edge computing, more and more edge boxes begin to be built-in chips or support inference computing, which can intelligently analyze and process data. However, due to the resource limitation of the edge box, the image frame rate after processing by the edge box is not high, and therefore, it is crucial to improve the image frame rate after processing by the edge box.

[0003] In the related art, the edge box runs the multi-thread technology through an algorithm stream, which includes three links of loading a model, runtime environment initialization, and algorithm inference. The algorithm stream sequentially executes the three links in a serial manner. Since different algorithm streams correspond to independent execution environments, when switching the algorithm stream, each link and the corresponding parameters in the algorithm stream need to be reloaded. This loading process increases the algorithm time consumption, and further causes the reduction of the image frame rate. SUMMARY

[0004] The present disclosure provides a performance optimization method and device of an edge box and an electronic device to solve the problems in the related art, which realizes the parallel processing of the divided multiple video frame data through different processing units, and the configuration information required by the processing units is preloaded at one time. When switching the algorithms of the processing units, the configuration information does not need to be reloaded, which reduces the algorithm processing time consumption when switching the processing units, and thus improves the image processing frame rate.

[0005] According to a first aspect embodiment of the present disclosure, a performance optimization method of an edge box is provided, which comprises:

[0006] In response to the video stream data received by the edge box, the video stream data is divided into multiple video frame data in time sequence, and the multiple video frame data are processed through multiple processing units respectively to obtain multiple processed video frame data. Each processing unit includes a computing node for processing video frame data, and the computing node preloads all configuration information corresponding to the computing node at one time before processing the video frame data.

[0007] It is determined whether the frame rate of each processed video frame data meets a preset frame rate threshold.

[0008] If it is determined that there is a case that does not meet the preset frame rate threshold, the multiple processed video frame data are reprocessed through the multiple processing units respectively until the frame rate of all processed video frame data meets the preset frame rate threshold.

[0009] Based on the processing results of the plurality of processing units, a target video stream is obtained.

[0010] In some embodiments of the present disclosure, the processing of the plurality of video frame data by the plurality of processing units respectively to obtain a plurality of processed video frame data comprises:

[0011] Based on the loading model, the configuration information corresponding to each computing node is loaded;

[0012] The plurality of video frame data are transmitted to the plurality of processing units in sequence, and the plurality of processing units process target video frame data in parallel;

[0013] In each processing unit, the target video frame data are processed according to the execution order of each computing node.

[0014] In some embodiments of the present disclosure, the computing nodes comprise a pre-processing node, an inference node and a post-processing node;

[0015] The processing of the target video frame data in each processing unit according to the execution order of each computing node comprises:

[0016] The target video frame data are processed by the pre-processing node to obtain first data, and the first data are stored in a first cache queue;

[0017] In a case where the inference node is determined to be in an idle state, the first data are read from the first cache queue to the inference node, and the first data are processed by the inference node to obtain second data, and the second data are stored in a second cache queue;

[0018] In a case where the post-processing node is determined to be in an idle state, the second data are read from the second cache queue to the post-processing node, and the second data are processed by the post-processing node to obtain third data.

[0019] In some embodiments of the present disclosure, the determination of whether the frame rate of each processed video frame data conforms to a preset frame rate threshold comprises:

[0020] According to the second data obtained by the inference node of each processing unit, the response time of algorithm inference is determined respectively;

[0021] Based on the response time of algorithm inference, it is determined whether the frame rate of the second data processed by each inference node conforms to the preset frame rate threshold;

[0022] if the preset frame rate threshold is met, the second data is transmitted to the post-processing node of each processing unit respectively;

[0023] if the preset frame rate threshold is not met, the second data is frame-dropped according to a preset frame-dropping strategy to obtain processed second data, wherein the preset frame-dropping strategy is associated with the total frame rate of the inference node of each processing unit.

[0024] In some embodiments of the present disclosure, the re-processing of the plurality of processed video frame data by the plurality of processing units comprises:

[0025] The processed second data is input into each pre-processing node, and each pre-processing node processes the processed second data until the frame rate of the processed second data meets the preset frame rate threshold.

[0026] In some embodiments of the present disclosure, the frame-dropping processing of the second data according to the preset frame-dropping strategy comprises:

[0027] If the first total frame rate of all video frame data is greater than the second total frame rate of the inference node of each processing unit, the redundant video frame data whose first total frame rate is greater than the second total frame rate is determined.

[0028] Frame-dropping operation is performed on the redundant video frame data.

[0029] In some embodiments of the present disclosure, before the video stream data is divided into a plurality of video frame data in sequence, the method further comprises:

[0030] The estimated time consumption of each processing unit is compared with the processing time consumption of the video stream data determined according to the frame rate of the video stream data;

[0031] If the estimated time consumption is greater than the processing time consumption, the video stream data is frame-skipped according to a preset frame-skipping strategy to obtain frame-skipped video stream data;

[0032] The video stream data is divided into a plurality of video frame data in sequence, comprising:

[0033] The frame-skipped video stream data is divided into the plurality of video frame data in sequence according to the video stream data.

[0034] In some embodiments of the present disclosure, the target video stream is obtained based on the processing results of the plurality of processing units, comprising:

[0035] The processing results output by the post-processing node in the plurality of processing units are summarized in sequence to obtain the target video stream.

[0036] In some embodiments of the present disclosure, the loading of the configuration information corresponding to all computing nodes respectively based on the loading model comprises:

[0037] obtaining the weights and configuration items corresponding to the inference nodes in each processing unit and the required maximum memory of each computing node from the loading model of the edge box;

[0038] configuring the inference nodes based on the obtained weights and configuration items corresponding to the inference nodes in each processing unit and the required maximum memory of each computing node.

[0039] According to a second aspect of the present disclosure, a performance optimization device of an edge box is provided, and the device comprises:

[0040] a division unit configured to divide video stream data received by the edge box into a plurality of video frame data in time sequence;

[0041] a first processing unit configured to process the plurality of video frame data through a plurality of processing units respectively to obtain a plurality of processed video frame data; each processing unit comprises a computing node configured to process video frame data, and the computing node loads configuration information corresponding to all computing nodes respectively in advance before processing the video frame data;

[0042] a first determination unit configured to determine whether the frame rate of each processed video frame data meets a preset frame rate threshold;

[0043] a second processing unit configured to re-process the plurality of processed video frame data through the plurality of processing units respectively until the frame rate of all processed video frame data meets the preset frame rate threshold when it is determined that there is a case not meeting the preset frame rate threshold;

[0044] a second determination unit configured to obtain target video stream based on the processing results of the plurality of processing units.

[0045] In some embodiments of the present disclosure, the first processing unit comprises:

[0046] a loading module configured to load the configuration information corresponding to all computing nodes respectively based on a loading model;

[0047] a first transmission module configured to transmit the plurality of video frame data to the plurality of processing units in time sequence respectively, and the plurality of processing units process target video frame data in parallel;

[0048] The processing module is configured to process the target video frame data in each processing unit according to an execution sequence of each computing node.

[0049] In some embodiments of the present disclosure, the computing nodes include a pre-processing node, an inference node, and a post-processing node.

[0050] The processing module includes:

[0051] The first processing submodule is configured to process the target video frame data by the pre-processing node to obtain first data.

[0052] The first storage submodule is configured to store the first data to a first cache queue.

[0053] The first reading submodule is configured to read the first data from the first cache queue to the inference node when the inference node is determined to be in an idle state.

[0054] The second processing submodule is configured to process the first data by the inference node to obtain second data.

[0055] The second storage submodule is configured to store the second data to a second cache queue.

[0056] The second reading submodule is configured to read the second data from the second cache queue to the post-processing node when the post-processing node is determined to be in an idle state.

[0057] The third processing submodule is configured to process the second data by the post-processing node to obtain third data.

[0058] In some embodiments of the present disclosure, the first determination unit includes:

[0059] The first determination module is configured to determine a response time of algorithm inference according to the second data obtained by the inference node of each processing unit.

[0060] The second determination module is configured to determine whether a frame rate of the second data processed by each inference node meets a preset frame rate threshold based on the response time of algorithm inference.

[0061] The second transmission module is configured to transmit the second data to the post-processing node of each processing unit when the preset frame rate threshold is met.

[0062] The frame loss module is configured to perform frame loss processing on the second data according to a preset frame loss strategy to obtain processed second data when the preset frame rate threshold is not met, wherein the preset frame loss strategy is associated with a total frame rate of the inference node of each processing unit.

[0063] In some embodiments of the present disclosure, the second processing unit is further configured to input the processed second data into each of the pre-processing nodes, and each of the pre-processing nodes processes the processed second data until a frame rate of the processed second data meets the preset frame rate threshold.

[0064] In some embodiments of the present disclosure, the frame dropping module comprises:

[0065] A determination sub-module is configured to determine, when a first total frame rate of all video frame data is greater than a second total frame rate of the inference node of each processing unit, redundant video frame data for which the first total frame rate is greater than the second total frame rate.

[0066] A frame dropping sub-module is configured to perform a frame dropping operation on the redundant video frame data.

[0067] In some embodiments of the present disclosure, the device further comprises:

[0068] A comparison unit is configured to compare an estimated time consumption of each processing unit with a processing time consumption of the video stream data determined according to a frame rate of the video stream data.

[0069] A frame skipping unit is configured to perform a frame skipping process on the video stream data according to a preset frame skipping strategy when the estimated time consumption is greater than the processing time consumption, to obtain frame-skipped video stream data.

[0070] In some embodiments of the present disclosure, the division unit is further configured to divide the frame-skipped video stream data into the plurality of video frame data according to a time sequence of the video stream data.

[0071] In some embodiments of the present disclosure, the second determination unit is further configured to collect the processing results output by the post-processing nodes in the plurality of processing units in a time sequence to obtain the target video stream.

[0072] In some embodiments of the present disclosure, the loading module comprises:

[0073] An acquisition sub-module is configured to acquire, from the loading model of the edge box, weights and configuration items corresponding to the inference nodes in each processing unit, and a maximum required memory of each computing node.

[0074] A configuration sub-module is configured to configure the inference nodes based on the acquired weights and configuration items corresponding to the inference nodes in each processing unit, and the maximum required memory of each computing node.

[0075] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0076] at least one processor; and

[0077] a memory communicatively connected with the at least one processor; wherein

[0078] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the preceding first aspect embodiment.

[0079] According to a fourth aspect embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the preceding first aspect embodiment.

[0080] According to a fifth aspect embodiment of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the preceding first aspect embodiment.

[0081] In summary, according to the performance optimization method, device and electronic equipment of the edge box provided by the present disclosure, the method comprises: in response to the video stream data received by the edge box, dividing the video stream data into a plurality of video frame data according to time sequence, and processing the plurality of video frame data through a plurality of processing units respectively to obtain a plurality of processed video frame data, wherein each processing unit comprises a computing node for processing video frame data, and the computing node loads all configuration information corresponding to the computing node in advance before processing the video frame data, determines whether the frame rate of each processed video frame data meets a preset frame rate threshold, and in the case where there is a situation that does not meet the preset frame rate threshold, re-processes the plurality of processed video frame data through the plurality of processing units respectively until the frame rate of all processed video frame data meets the preset frame rate threshold, and obtains a target video stream based on the processing result of the plurality of processing units. The scheme of the present disclosure realizes parallel processing of the plurality of video frame data divided through different processing units, and the configuration information required by the processing units is loaded in advance, so that when the algorithms of the processing units are switched, the configuration information does not need to be loaded again, the time consumption of algorithm processing when each processing unit is switched is reduced, and the frame rate of image processing is improved.

[0082] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0083] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:

[0084] Figure 1 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0085] Figure 2 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0086] Figure 3 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0087] Figure 4 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0088] Figure 5 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0089] Figure 6 A flowchart of inference node processing of an edge box provided by an embodiment of the present disclosure;

[0090] Figure 7 A three-dimensional schematic diagram of a convolution operation of a convolutional neural network provided by an embodiment of the present disclosure;

[0091] Figure 8 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0092] Figure 9 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0093] Figure 10 A flowchart of an adaptive frame loss strategy provided by an embodiment of the present disclosure;

[0094] Figure 11 A flowchart of a video stream data scheduling processing provided by an embodiment of the present disclosure;

[0095] Figure 12 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0096] Figure 13 A flowchart of a performance optimization method of an edge box provided by an embodiment of the present disclosure;

[0097] Figure 14 A flowchart of switching of algorithms corresponding to each inference node provided by an embodiment of the present disclosure;

[0098] Figure 15A flowchart of another method for optimizing performance of an edge box according to an embodiment of the present disclosure is provided.

[0099] Figure 16 A structural diagram of a device for optimizing performance of an edge box according to an embodiment of the present disclosure is provided.

[0100] Figure 17 A structural diagram of another device for optimizing performance of an edge box according to an embodiment of the present disclosure is provided.

[0101] Figure 18 A schematic block diagram of an example electronic device according to an embodiment of the present disclosure is provided. DETAILED DESCRIPTION

[0102] Embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. Examples of the embodiments are illustrated in the drawings, in which the same or similar components are denoted by the same or similar reference numerals, and therefore repeated description is omitted. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present disclosure, and cannot be understood as a limitation of the present disclosure.

[0103] With the rise of edge computing, edge boxes as a kind of hardware device of edge computing, more and more edge boxes begin to be built-in chips or support inference computing, which can intelligently analyze and process data. However, due to the resource limitation of edge boxes, the image frame rate after processing by edge boxes is not high, and therefore, it is crucial to improve the image frame rate after processing by edge boxes.

[0104] In the related art, the edge box runs the multi-thread technology through an algorithm stream, which includes three links of loading a model, runtime environment initialization, and algorithm inference. The algorithm stream sequentially executes the three links in series. Since different algorithm streams correspond to independent execution environments, when switching the algorithm stream, each link and the corresponding parameters in the algorithm stream need to be reloaded. This loading process increases the algorithm time consumption, and further causes the reduction of the image frame rate.

[0105] Therefore, in order to solve the problems in the related art, the present disclosure proposes a performance optimization method of an edge box, which comprises: in response to video stream data received by the edge box, dividing the video stream data into a plurality of video frame data in time sequence, and processing the plurality of video frame data through a plurality of processing units respectively to obtain a plurality of processed video frame data, wherein each processing unit comprises a computing node for processing video frame data, and all the computing nodes correspondingly pre-loaded configuration information once, and determining whether the frame rate of each processed video frame data meets a preset frame rate threshold, and in the case of determining that there is a case of not meeting the preset frame rate threshold, re-processing the plurality of video frame data through the plurality of processing units respectively until the frame rate of all the processed video frame data meets the preset frame rate threshold, and obtaining a target video stream based on the processing result of the plurality of processing units.

[0106] The scheme of the present disclosure divides the video stream data received by the edge box into a plurality of video frame data in time sequence, and processes the plurality of video frame data through a plurality of processing units respectively to obtain a plurality of processed video frame data, wherein each processing unit comprises a computing node for processing video frame data, and all the computing nodes correspondingly pre-loaded configuration information once, and determining whether the frame rate of each processed video frame data meets a preset frame rate threshold, and in the case of determining that there is a case of not meeting the preset frame rate threshold, re-processing the plurality of video frame data through the plurality of processing units respectively until the frame rate of all the processed video frame data meets the preset frame rate threshold, and obtaining a target video stream based on the processing result of the plurality of processing units, which realizes parallel processing of the plurality of video frame data divided by different processing units, and the configuration information required by the processing units is pre-loaded once, and in the case of switching different algorithms of the processing units, there is no need to load the configuration information again, which reduces the time consumption of algorithm processing when each processing unit is switched, thereby improving the frame rate of image processing.

[0107] The present disclosure is not exhaustive, and only some embodiments are shown, which are not specific limitations on the protection scope of the present disclosure. In the case of no contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, the scheme after removing some steps in an embodiment can also be implemented as an independent embodiment, and the order of the steps in an embodiment can be exchanged arbitrarily, in addition, the optional implementation manners in an embodiment can be combined arbitrarily; in addition, the embodiments can be combined arbitrarily, for example, some or all steps of different embodiments can be combined arbitrarily, and an embodiment can be combined with optional implementation manners of other embodiments.

[0108] In the embodiments of the present disclosure, the terms and / or descriptions among the embodiments are consistent and can be referred to each other if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0109] The terms used in the embodiments of the present disclosure are only for the purpose of describing particular embodiments and are not used as limitations of the present disclosure.

[0110] In the embodiments of the present disclosure, unless otherwise specified, the elements expressed in singular form, such as "one", "a", "the", "above", "said", "preceding", "this" and the like, can represent "one and only one", or "one or more", "at least one" and the like. For example, in the case of using articles such as "a", "an", "the" in English, the noun after the article can be understood as singular expression, or can be understood as plural expression.

[0111] In some embodiments, the terms "in response to", "in response to determining", "in the case of", "when", "when", "if", "if" and the like can be replaced with each other.

[0112] In some embodiments, the terms "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not lower than", "above" and the like can be replaced with each other, and the terms "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", "below" and the like can be replaced with each other.

[0113] The prefix words "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different description objects, and do not constitute limitations on the position, order, priority, quantity or content of the description objects. The description of the description object should be referred to the description in the claims or embodiments, and should not constitute redundant limitations because of the use of the prefix word.

[0114] In the embodiments of the present disclosure, "a plurality of" means two or more.

[0115] In the embodiments of the present disclosure, the terms "import", "input", "read in" and the like can be replaced with each other.

[0116] In some embodiments, an apparatus or the like can be interpreted as an entity, and can also be interpreted as virtual, and the name thereof is not limited to the name described in the embodiments. The terms "apparatus", "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject" and the like can be replaced with each other.

[0117] In some embodiments, the terms "terminal", "terminal device", "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", "subscriber station", "mobile unit", "subscriber unit", "wireless unit", "remote unit", "mobile device", "wireless device", "wireless communication device", "remote device", "mobile subscriber station", "access terminal", "mobile terminal", "wireless terminal", "remote terminal", "handset", "user agent", "mobile client", "client" and the like can be replaced with each other.

[0118] Figure 1 A flowchart of a performance optimization method of an edge box provided by the embodiments of the present disclosure is shown in Figure 1 The performance optimization method of the edge box contains steps 101-104.

[0119] In step 101, in response to the video stream data received by the edge box, the video stream data is divided into a plurality of video frame data in time sequence, and the plurality of video frame data is processed by a plurality of processing units respectively to obtain a plurality of processed video frame data; wherein each processing unit includes a computing node for processing video frame data, and the computing node loads all configuration information corresponding to the computing node in advance before processing the video frame data.

[0120] In the embodiment of the present disclosure, the video stream data received by the edge box is dynamic video data transmitted in continuous time sequence, which is composed of a plurality of single-frame video frame data. The video stream data is divided into a plurality of video frame data by time sequence, and the plurality of video frame data is processed by a plurality of processing units to obtain a plurality of processed single-frame video frame data.

[0121] The continuous video stream data is divided into a plurality of video frame data, and the plurality of video frame data is processed by a plurality of processing units in parallel, which breaks the bottleneck of single-thread serial processing, loads the running parameters of the algorithm model for each computing node in advance, avoids the time-consuming of repeated model loading and environment initialization, and improves the video processing efficiency of the edge box.

[0122] In step 102, it is determined whether the frame rate of each processed video frame data meets a preset frame rate threshold.

[0123] In the embodiment of the present disclosure, each processed video frame data is the video frame data processed by each processing unit. The frame rate threshold of the video frame data is preset as a reference for judging whether the processing of the video frame data meets the requirements. The preset frame rate threshold is usually determined by the application scenario. By comparing the frame rate of the processed video frame data with the preset frame rate threshold, it is determined whether the edge box has sufficient computing power to support the current task.

[0124] Determining whether it meets the preset frame rate threshold can early warn potential performance bottlenecks (including but not limited to hardware failure, algorithm memory leakage, etc.). Through log recording and abnormal alarm, it helps the operation and maintenance personnel to timely troubleshoot problems and ensures the long-term stable operation of the edge computing system.

[0125] In step 103, in a case where it is determined that there is a condition that does not meet the preset frame rate threshold, the plurality of processed video frame data is reprocessed by the plurality of processing units respectively until the frame rate of all processed video frame data meets the preset frame rate threshold.

[0126] The embodiments of the present disclosure reprocess the multiple video frame data through multiple processing units to compensate for the contingency of processing, avoid image frame freezing or data loss caused by single abnormality, and continuous performance degradation caused by short-term interference.

[0127] The multiple processing units compensate for the contingency of processing by reprocessing the multiple video frame data, avoiding image frame freezing or data loss caused by single abnormality, and continuous performance degradation caused by short-term interference.

[0128] In step 104, the target video stream is obtained based on the processing results of the multiple processing units.

[0129] In the embodiments of the present disclosure, the processing results of the multiple processing units are the final results output by the multiple processing units after computing the single frame or multiple frames of video data, and the target video stream is the continuous dynamic video data obtained by recombining the final results output after computing in the original time sequence.

[0130] By integrating the parallel computing results of the multiple processing units, the required continuous video stream is generated, thereby ensuring the integrity and availability of the output.

[0131] As an implementable manner of the embodiments of the present disclosure, as shown in Figure 2 The embodiments of the present disclosure provide a flowchart of a performance optimization method of an edge box. First, initialization operation is performed to complete pre-allocation of hardware resources, building of inference framework and basic environment, etc. After initialization, the loaded model is run. When the loaded model is run, all configuration information (such as weight, configuration cache, etc.) required by all inference nodes (inference node 1, inference node 2, …, inference node n) is loaded at one time, so that all inference nodes do not need to be loaded multiple times when performing inference tasks. The video stream data is divided into multiple video frame data through video decoding, and the multiple video frame data is input to all pre-processing nodes (pre-processing node 1, pre-processing node 2, …, pre-processing node n) after frame rate control. Then, all inference nodes are passed. When the multiple video frame data is input to all inference nodes, the frame rate of the video frame data that does not meet the preset frame rate threshold is input to all pre-processing nodes again through algorithm time-consuming feedback until the frame rate of the video frame data meets the preset frame rate threshold. The video frame data whose frame rate meets the preset frame rate threshold is input to all post-processing nodes (post-processing node 1, post-processing node 2, …, post-processing node n). The processing results of all post-processing nodes are summarized and the algorithm result is output. After the process is completed, the cache is released.

[0132] In summary, the performance optimization method of the edge box provided by the present disclosure comprises, in response to video stream data received by the edge box, dividing the video stream data into a plurality of video frame data in time sequence, and processing the plurality of video frame data through a plurality of processing units respectively to obtain a plurality of processed video frame data, wherein each processing unit comprises a computing node for processing video frame data, and the computing node loads all configuration information corresponding to each computing node in advance before processing the video frame data, determines whether the frame rate of each processed video frame data meets a preset frame rate threshold, and re-processes the plurality of processed video frame data through the plurality of processing units respectively until the frame rate of all processed video frame data meets the preset frame rate threshold, and obtains a target video stream based on the processing result of the plurality of processing units. The scheme of the present disclosure realizes parallel processing of the plurality of video frame data divided by different processing units, and the configuration information required by the processing units is loaded in advance, so that the configuration information does not need to be loaded again when the algorithms of the processing units are switched, the time consumption of algorithm processing when each processing unit is switched is reduced, and the frame rate of image processing is improved.

[0133] Figure 3 Further, a flowchart of a performance optimization method of an edge box is shown. Based on the embodiment shown in Figure 1 The step 101 is further explained, Figure 3 may comprise the following steps:

[0134] Step 201, loading all configuration information corresponding to each computing node based on a loading model.

[0135] In the present embodiment of the present disclosure, all computing nodes load the required configuration information from the loading model, wherein the required configuration information includes but is not limited to cfg / weight resources (also known as weights), configuration cache and the like.

[0136] The loading model is loaded into the hardware resources of the computing nodes in advance to avoid the delay of runtime loading, reduce runtime configuration operations and resource conflicts, and ensure long-term stable operation of the edge box. The loading model is a complete model resource set deployed to all computing node hardware resources in the system initialization stage according to configuration information (such as weights, configuration cache, etc.).

[0137] Step 202, transmitting the plurality of video frame data to the plurality of processing units in time sequence respectively, and the plurality of processing units processing target video frame data in parallel.

[0138] The plurality of processing units simultaneously perform computing tasks in parallel on the plurality of video frame data in time sequence.

[0139] The video frame data is distributed to the plurality of processing units in the original time sequence, ensuring that the generated target video frame data is picture continuous, and the end-to-end delay of the plurality of video frame data is shortened through parallel processing.

[0140] In step 203, the target video frame data is processed in each processing unit according to the execution order of each computing node.

[0141] The logical order of processing tasks of each computing node in each processing unit ensures that the video frame data format and content output by the plurality of processing units during parallel processing are consistent.

[0142] In the embodiments of the present disclosure, the computing nodes include a pre-processing node, an inference node, and a post-processing node.

[0143] Each computing node has an input and an output port, and performs different processing tasks in a pipeline manner. Each computing node is an independent computing node, and each computing node is transmitted through a cache queue.

[0144] The time of the pre-processing node is t1, the time of the inference node is t2, and the time of the post-processing node is t3, all in ms, so that t3

[0145] The cache queue size of 2 and 10 above is only an example, and in actual application, it can be configured according to the experimental results, and the specific configuration is not limited.

[0146] Figure 4 A processing flowchart of a performance optimization method of an edge box according to the embodiments of the present disclosure is shown. The plurality of target video frame data entering the pre-processing node is numbered as c1f1, c1f2,..., cnfn, where c represents a channel, and f represents a frame number. The plurality of target video frame data realizes zero-copy in the entire channel, reduces the copying time of the plurality of target video frame data between transmission, inputs the numbered plurality of target video frame data to the pre-processing node, and transmits the pre-processing node to the inference node through the first cache queue, and then transmits the second cache queue to the post-processing node.

[0147] Figure 5 Further, a flowchart of a performance optimization method of an edge box according to the embodiments of the present disclosure is shown. Based on the embodiments shown in Figure 3 The step 203 is further explained, Figure 5 may include the following steps:

[0148] Step 301, processing the target video frame data by the pre-processing node to obtain first data, and storing the first data to a first cache queue.

[0149] In the embodiment of the present disclosure, the pre-processing node processes the target video frame data by using a multimedia unit, which includes scaling, color space conversion, data padding, and converting the YUV format target video frame data decoded from the video into ARGB format target video frame data to obtain the first data, and storing the first data into the first cache queue, which is a first-in-first-out buffer temporarily storing the first data.

[0150] Step 302, in the case of determining that the inference node is in an idle state, reading the first data from the first cache queue to the inference node, processing the first data by the inference node to obtain second data, and storing the second data to a second cache queue.

[0151] When the inference node is in an idle state, the inference node processes the first data, and the inference node performs algorithm inference on the loaded cfg / weight resource and the first data by using the RNE-RT-lib library in the Reconfigurable Neuron Engine (RNE) to obtain the second data, and stores the second data to the second cache queue, which is a first-in-first-out buffer temporarily storing the second data.

[0152] In the embodiment of the present disclosure Figure 6 The flowchart of the inference node processing of the edge box is shown in FIG. 6. The first data is input to the algorithm inference to obtain the output second data. The algorithm inference is full-load running, which realizes zero-waiting of upstream and downstream resources and maximum performance of computing power.

[0153] Step 303, in the case of determining that the post-processing node is in an idle state, reading the second data from the second cache queue to the post-processing node, and processing the second data by the post-processing node to obtain third data.

[0154] The post-processing node includes dequantization and Non-Maximum Suppression (NMS) processing. When the post-processing node is in an idle state, the post-processing node processes the second data to obtain the third data.

[0155] In the embodiment of the present disclosure Figure 7A three-dimensional diagram of a convolution operation of a convolutional neural network, where h (height) represents the size of the second data in the height dimension, w (width) represents the size of the second data in the width dimension, c (channel) represents the number of channels of the second data, c stride represents the stride (channel step) in the channel dimension, c pad represents the padding in the channel dimension, which can be additional elements or bytes added in the channel direction to meet specific memory alignment requirements or algorithmic needs, h*w*c stride represents the total amount of stride required in the channel dimension after crossing the entire height h and width w plane, and HW (c+c pad) represents the total amount of data of all channels (c+c pad) in the entire height (H, equivalent to h) and width (W, equivalent to w) plane considering the channel padding (c pad).

[0156] The RNE algorithm score in the dequantization process converts fixed-point to floating-point, and converts NHWC four-dimensional data to one-dimensional data through a four-layer loop. In the process of taking data, the c channel needs to be aligned with the stride.

[0157] The blob represents the hwc stacked along the channel axis ( Figure 7 The numbers 1, 2, and 3 represent the direction). The multi-dimensional tensor hwc of the data block blob is usually large and time-consuming to calculate. By setting the NMS filtering threshold, only data greater than the threshold is taken, which can reduce the data extraction and calculation time, eliminate redundant detection boxes or feature points, and make the final detection result clearer and more accurate. The NMS processing makes the final detection result clearer and more accurate, while reducing the amount of calculation and improving the efficiency of the algorithm.

[0158] Figure 8 Further, a flowchart of a performance optimization method of an edge box is shown. Based on the embodiment shown in Figure 5 The step 102 is further explained, Figure 8 The method can include the following steps:

[0159] Step 401, according to the second data obtained by each processing unit of the inference node, respectively determine the response time of the algorithm inference.

[0160] The response time of the algorithm inference is the time difference from the first data being received by the inference node to the second data being output.

[0161] By monitoring the response time of the algorithm inference, the performance of the edge box is dynamically adjusted.

[0162] At step 402, it is determined whether the frame rate of the second data processed by each inference node meets the preset frame rate threshold based on the response time of the algorithm inference.

[0163] By converting the response time of the algorithm inference in the inference node into the actual second data processing frame rate, and comparing the actual second data processing frame rate with the preset frame rate threshold, a closed-loop monitoring and optimization system for the inference performance of the edge box is constructed.

[0164] At step 403, if the preset frame rate threshold is met, the second data is transmitted to the post-processing nodes of each processing unit respectively.

[0165] If the frame rate of the second data processed by each inference node meets the preset frame rate threshold, the second data is transmitted to the post-processing nodes of each processing unit respectively.

[0166] At step 404, if the preset frame rate threshold is not met, the second data is frame-dropped according to a preset frame-dropping strategy to obtain processed second data, wherein the preset frame-dropping strategy is associated with the total frame rate of the inference nodes of each processing unit.

[0167] If the frame rate of the second data processed by each inference node does not meet the preset frame rate threshold, the part of the second data that does not meet the preset frame rate threshold is frame-dropped, and the second data that meets the preset frame rate threshold is transmitted to each processing unit after the frame-dropping.

[0168] After it is determined that the frame rate of the processed second data does not meet the preset frame rate threshold, the processed second data is input to each pre-processing node, and each pre-processing node processes the processed second data until the frame rate of the processed second data meets the preset frame rate threshold. Since the frame rate of the processed second data is reduced after the frame-dropping, the processed second data is input to each pre-processing node again for processing until the frame rate of the processed second data meets the preset frame rate threshold.

[0169] Figure 9 Further, a flowchart of a performance optimization method of an edge box is shown. Figure 8 The step 404 is further explained based on the embodiment shown in the figure, Figure 9 may include the following steps:

[0170] Step 501, if the first total frame rate of all video frame data is greater than the second total frame rate of the inference node of each processing unit, determine the redundant video frame data whose first total frame rate is greater than the second total frame rate.

[0171] The total time consumption of the algorithm in each inference node in the embodiment of the disclosure is T, the first total frame rate of all video frame data is F, and each video frame data is represented as fa1, fa2…fan, wherein T=1 / fa1+2 / fa2+n / fan, when F>1 / T, the frame dropping strategy processing is performed, and the frame dropping strategy is shown in formula 1.

[0172]

[0173] Wherein, y(t) is the output effective frame rate, u(t) is the frame rate of the input video frame data, G is the algorithm weight coefficient, and the weight coefficients of different algorithms are also different, the feedback coefficient β represents the proportion of the algorithm inference time consumption and the time consumption of each frame of video frame data, when the algorithm inference time consumption increases, the feedback coefficient β also increases, and y(t) decreases, when the algorithm inference time consumption decreases, the feedback coefficient β also decreases, and y(t) increases.

[0174] Step 502, performing frame dropping operation on the redundant video frame data.

[0175] The redundant video frame data whose first total frame rate is greater than the second total frame rate is subjected to frame dropping operation, Figure 10 An adaptive frame dropping strategy flowchart is shown in FIG. 1, which takes multiple video frame data decoded from multiple videos as a starting point, dynamically discards part of frames through adaptive frame dropping strategy, and manages the video frame data memory processed by the pre-processing node. Figure 6 The inference node processes the first data in parallel, outputs the second data, and processes the second data through the post-processing node.

[0176] Since the memory size of the pre-processing node is fixed, N blocks of VB memories are applied through a media processing platform (MPP), each block of VB memory is marked to distinguish data of different channels, and the memory block management can be managed through an array A[n], the length of the array is N, and when the data is full, the data is overwritten from the head A[0]; the media processing platform MPP provides memory application, hardware acceleration and other functions, the VB memory is a video buffer area, which can store memory blocks of video frame data, and A[n] is a memory block management array with a length of N.

[0177] Figure 11A video stream data scheduling process flowchart is shown. The multi-channel video stream data is input to a decoder to obtain YUV video frame data. The YUV video frame data is transmitted to a pre-processing node through an adaptive frame dropping strategy. The pre-processing node converts the YUV video frame data into ARGB format and caches the ARGB format video frame data in a memory pool. The pre-processing node processed video frame data is transmitted to an inference node through a cache queue for processing. The inference node processed video frame data is transmitted to a post-processing node through a cache queue for processing.

[0178] Figure 12 Further, a flowchart of a performance optimization method of an edge box is shown. Based on the embodiment shown in Figure 1 Before step 101, Figure 12 may include the following steps:

[0179] Step 601: comparing the estimated time consumption of each processing unit with the processing time consumption of the video stream data determined according to the frame rate size of the video stream data.

[0180] In the embodiment of the present disclosure, the estimated processing time of each processing unit is compared with the actual processing time calculated based on the video stream frame rate to realize fine evaluation and dynamic optimization of video processing.

[0181] Step 602: if the estimated time consumption is greater than the processing time consumption, performing frame skipping processing on the video stream data according to a preset frame skipping strategy to obtain frame skipping processed video stream data.

[0182] When the estimated time consumption is greater than the processing time consumption, since the algorithm cannot process the frame rate of the original video stream data, frame skipping processing is performed at the input end of the video stream data, and the time consumption of each algorithm is calculated as t1, t2,... tn, t1 corresponds to the time consumption of the first algorithm, t2 corresponds to the time consumption of the second algorithm, and tn corresponds to the time consumption of the nth algorithm, with the unit of milliseconds, and 1 / tn is the maximum frame rate of the nth algorithm. If the input frame rate of the video stream data is fi, and the frame rate after frame skipping is fa, then fa=fi / k, where k=fi*tn / 1000.

[0183] As an implementation manner of the embodiment of the present disclosure, when step 101 is executed, the frame skipping processed video stream data can be divided into the plurality of video frame data according to the time sequence of the video stream data.

[0184] According to the time sequence order of the video stream data, the continuous video stream data after frame skipping processing is divided into independent plurality of video frame data.

[0185] As an implementation manner of the embodiment of the present disclosure, based on Figure 5In the embodiment shown, when step 104 is executed, the processing results output by the post-processing nodes in the plurality of processing units can be sequentially aggregated to obtain the target video stream.

[0186] The processing results output by the post-processing nodes will be converted from dispersed video stream data into a target video stream that is time-sequentially aligned, information-fused, and format-unified.

[0187] Figure 13 Further shown is a flowchart of a performance optimization method of an edge box according to an embodiment of the present disclosure. Based on the edge box, the performance optimization method comprises the following steps of: Figure 3 In the embodiment shown, step 201 is further explained as follows, Figure 13 The method can comprise the following steps:

[0188] In step 701, the weights and configuration items corresponding to the inference nodes in each processing unit and the maximum memory required by each computing node are obtained from the loading model of the edge box.

[0189] When the loading model is completed, the weights and configuration items of each algorithm and the maximum memory of each computing node are preloaded to each inference node at one time. The inference nodes corresponding to each algorithm do not need to be preloaded with weights and reconfigured with configuration items during the running process. The memory size required by each algorithm is read, and the maximum memory in each algorithm is obtained as shared memory, wherein the shared memory can meet the memory requirements of each algorithm, such as Figure 14 As shown is a switching flowchart of algorithms corresponding to each inference node. According to the construction storage system, the inference nodes in each processing unit directly obtain the weights and configuration caches corresponding to each algorithm (algorithm 1-weight and configuration cache, algorithm 2-weight and configuration cache, …, algorithm n-weight and configuration cache) by means of shared memory blob, and do not need to apply temporary memory and output memory for each algorithm again.

[0190] For better understanding, the following examples are given: video frame 1 is processed by inference node 1, and video frame 2 is processed by inference node 2. When inference node 1 needs algorithm 1 to process video frame 1, and after algorithm 1 is processed, video frame 1 still needs to be processed by algorithm 2. At this time, algorithm 1 is switched to algorithm 2. If the memory size required by algorithm 1 is greater than the memory size required by algorithm 2, then algorithm 1 is used as shared memory as the maximum memory. When inference node 2 needs algorithm 1 to process video frame 2, and after algorithm 1 is processed, video frame 2 needs to be processed by algorithm 2, and after algorithm 2 is processed, video frame 2 needs to be processed by algorithm 3. At this time, algorithm 1 is switched to algorithm 2, and algorithm 2 is switched to algorithm 3. If the memory size required by algorithm 3 is greater than the memory size required by algorithm 1 or algorithm 2, then algorithm 3 is used as shared memory as the maximum memory. It should be noted that the above examples are only illustrative, and the specific embodiments are not limited.

[0191] At step 702, the inference nodes are configured based on the weights and configuration items corresponding to the inference nodes in each processing unit and the required maximum memory of each computing node.

[0192] The configuration information of the weights and configuration items corresponding to the inference nodes that have been loaded when loading the model and the required maximum memory of each computing node is configured.

[0193] The embodiments of the present disclosure can achieve the following beneficial effects:

[0194] 1. Reduce the total time consumption of RNE algorithm execution pipeline operation, reduce system load, improve frame rate, and run more algorithms.

[0195] 2. Fast switching between algorithms, saving memory resources and realizing parallel running of different algorithms.

[0196] Corresponding to the performance optimization method of the edge box described above, the present application also proposes a performance optimization device of the edge box. Since the device embodiments of the present application correspond to the method embodiments described above, for details not disclosed in the device embodiments, reference can be made to the method embodiments described above, which will not be described in detail in the present application.

[0197] In practical applications, the performance optimization method of the edge box described in the embodiments of the present application can be applied in various industries, for example, in the food safety supervision of the catering industry, as shown in Figure 15 , video stream data is collected by multiple cameras (camera 1, camera 2, camera 3) respectively, the video stream data is transmitted to the edge box, the received video stream data is subjected to image frame rate optimization processing in the edge box, and the target video stream meeting the preset frame rate threshold is obtained, the target video stream is transmitted to the background, and the management end of the monitoring center automatically identifies the scenes such as whether the kitchen staff wears or not, whether the animal intrudes, whether the outside personnel intrude, etc. Through the display end of the background, the information is intuitively displayed and published, and the screen user can see whether the operation of the back kitchen staff is regular, whether the hygiene is qualified, whether there are illegal articles, etc. The specific application scenarios of the embodiments of the present application are not limited.

[0198] Figure 16 The structure diagram of a performance optimization device of an edge box provided by the embodiments of the present application is shown in Figure 16 , which includes a division unit 151, a first processing unit 152, a first determination unit 153, a second processing unit 154, and a second determination unit 155.

[0199] The division unit 151 is configured to divide the video stream data received by the edge box into multiple video frame data in time sequence.

[0200] The first processing unit 152 is used to process the multiple video frame data through multiple processing units to obtain multiple processed video frame data; wherein, each processing unit includes a computing node for processing video frame data, and the computing node preloads the configuration information corresponding to all computing nodes at one time before processing the video frame data;

[0201] The first determining unit 153 is used to determine whether the frame rate of each processed video frame data meets the preset frame rate threshold.

[0202] The second processing unit 154 is used to reprocess the multiple processed video frame data through the multiple processing units when it is determined that there are no frames that do not meet the preset frame rate threshold, until the frame rate of all processed video frame data meets the preset frame rate threshold.

[0203] The second determining unit 155 is used to obtain the target video stream based on the processing results of the plurality of processing units.

[0204] In summary, the edge box performance optimization device provided in this disclosure includes, in response to video stream data received by the edge box, dividing the video stream data into multiple video frame data according to the time sequence, and processing the multiple video frame data separately through multiple processing units to obtain multiple processed video frame data. Each processing unit includes a computing node for processing the video frame data. Before processing the video frame data, the computing node preloads the configuration information corresponding to all computing nodes at once, determines whether the frame rate of each processed video frame data meets a preset frame rate threshold, and if it is determined that there are any that do not meet the preset frame rate threshold, reprocesses the multiple processed video frame data separately through the multiple processing units until the frame rate of all processed video frame data meets the preset frame rate threshold. Based on the processing results of the multiple processing units, a target video stream is obtained. This achieves parallel processing of the divided multiple video frame data through different processing units, and the configuration information required by these processing units is preloaded at once. When switching between different algorithms of processing units, there is no need to reload the configuration information, reducing the processing time of the algorithms when switching between processing units, thereby improving the frame rate of image processing.

[0205] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 17 As shown, the first processing unit 152 includes:

[0206] Loading module 1521 is used to load the configuration information corresponding to all computing nodes based on the loading model;

[0207] The first transmission module 1522 is configured to transmit the plurality of video frame data in time sequence to the plurality of processing units respectively, and the plurality of processing units process the target video frame data in parallel.

[0208] The processing module 1523 is configured to process the target video frame data in each processing unit according to the execution sequence of each computing node.

[0209] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 17 The computing node includes a pre-processing node, an inference node and a post-processing node.

[0210] The processing module 1523 includes:

[0211] The first processing submodule 15231 is configured to process the target video frame data by the pre-processing node to obtain first data.

[0212] The first storage submodule 15232 is configured to store the first data to a first cache queue.

[0213] The first reading submodule 15233 is configured to read the first data from the first cache queue to the inference node in a case where it is determined that the inference node is in an idle state.

[0214] The second processing submodule 15234 is configured to process the first data by the inference node to obtain second data.

[0215] The second storage submodule 15235 is configured to store the second data to a second cache queue.

[0216] The second reading submodule 15236 is configured to read the second data from the second cache queue to the post-processing node in a case where it is determined that the post-processing node is in an idle state.

[0217] The third processing submodule 15237 is configured to process the second data by the post-processing node to obtain third data.

[0218] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 17 The first determination unit 153 includes:

[0219] The first determination module 1531 is configured to determine the response time of algorithm inference according to the second data obtained by the inference node of each processing unit respectively.

[0220] The second determining module 1532 is configured to determine whether the frame rate of the second data processed by each inference node meets the preset frame rate threshold based on the response time determined by the algorithm inference.

[0221] The second transmission module 1533 is configured to transmit the second data to the post-processing nodes of the processing units respectively when the preset frame rate threshold is met.

[0222] The frame dropping module 1534 is configured to perform frame dropping processing on the second data according to a preset frame dropping strategy to obtain processed second data when the preset frame rate threshold is not met, wherein the preset frame dropping strategy is associated with the total frame rate of the inference nodes of the processing units.

[0223] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 17 The second processing unit 154 is further configured to input the processed second data into each of the pre-processing nodes, and each of the pre-processing nodes processes the processed second data until the frame rate of the processed second data meets the preset frame rate threshold.

[0224] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 17 The frame dropping module 1534 includes:

[0225] The determining sub-module 15341 is configured to determine redundant video frame data when the first total frame rate of all video frame data is greater than the second total frame rate of the inference nodes of the processing units.

[0226] The frame dropping sub-module 15342 is configured to perform frame dropping operation on the redundant video frame data.

[0227] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 17 The apparatus further includes:

[0228] The comparison unit 156 is configured to compare the estimated time consumption of each processing unit with the processing time consumption of the video stream data determined according to the frame rate of the video stream data.

[0229] The frame skipping unit 157 is configured to perform frame skipping processing on the video stream data according to a preset frame skipping strategy to obtain frame skipping processed video stream data when the estimated time consumption is greater than the processing time consumption.

[0230] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 17As shown, the dividing unit 151 is further configured to divide the video stream data after the frame skipping processing into the plurality of video frame data according to the time sequence of the video stream data.

[0231] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 17 As shown, the second determining unit 155 is further configured to collect the processing results output by the post-processing nodes in the plurality of processing units according to the time sequence to obtain the target video stream.

[0232] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 18 As shown, the loading module 1521 comprises:

[0233] The obtaining sub-module 15211 is configured to obtain, from the loading model of the edge box, the weights and configuration items corresponding to the inference nodes in each processing unit and the required maximum memory of each calculation node.

[0234] The configuration sub-module 15212 is configured to configure the inference nodes based on the obtained weights and configuration items corresponding to the inference nodes in each processing unit and the required maximum memory of each calculation node.

[0235] It should be noted that the foregoing explanation and description of the method embodiments are also applicable to the apparatus of the present disclosure, and the principles are the same, which are not limited in the present disclosure.

[0236] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0237] Figure 18 A schematic block diagram of an example electronic device 1700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0238] As ​As shown, the electronic device 1700 includes a computing unit 1701 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 1702 or a computer program loaded into a RAM (Random Access Memory) 1703 from a storage unit 1708. Various programs and data required for the operation of the electronic device 1700 can also be stored in the RAM 1703. The computing unit 1701, the ROM 1702, and the RAM 1703 are connected to each other through a bus 1704. An I / O (Input / Output) interface 1705 is also connected to the bus 1704.

[0239] A plurality of components in the electronic device 1700 are connected to the I / O interface 1705, including an input unit 1706 such as a keyboard, a mouse, and the like, an output unit 1707 such as various types of displays, a speaker, and the like, a storage unit 1708 such as a magnetic disk, an optical disk, and the like, and a communication unit 1709 such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 1709 allows the electronic device 1700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0240] The computing unit 1701 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 1701 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphic Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1701 performs various methods and processes described above, such as the performance optimization method of the edge box. For example, in some embodiments, the performance optimization method of the edge box can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1700 via the ROM 1702 and / or the communication unit 1709. When the computer program is loaded onto the RAM 1703 and executed by the computing unit 1701, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 1701 can be configured to perform the aforementioned performance optimization method of the edge box by any other appropriate means, such as by means of firmware.

[0241] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0242] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0243] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage medium can include but are not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical wire, portable computer diskette, hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, fiber optics, CD-ROM (Compact Disc Read-Only Memory), optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0244] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0245] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0246] The computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0247] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of people (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.

[0248] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.

[0249] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.

Claims

1. A method for performance optimization of an edge cassette, characterized by, The method comprises: In response to the video stream data received by the edge box, the video stream data is divided into a plurality of video frame data in time sequence, and the plurality of video frame data is processed by a plurality of processing units respectively to obtain a plurality of processed video frame data; wherein each processing unit comprises a computing node for processing video frame data, and the computing node loads all configuration information corresponding to the computing node in advance before processing the video frame data; Determine whether the frame rate of each processed video frame data meets a preset frame rate threshold; In the case where it is determined that there is no frame rate that meets the preset frame rate threshold, the plurality of processed video frame data is reprocessed by the plurality of processing units respectively until the frame rate of all processed video frame data meets the preset frame rate threshold; Based on the processing result of the plurality of processing units, a target video stream is obtained.

2. The method of claim 1, wherein, The plurality of video frame data is processed by a plurality of processing units to obtain a plurality of processed video frame data, comprising: Load all configuration information corresponding to the computing node based on the loading model; The plurality of video frame data is transmitted to the plurality of processing units in time sequence respectively, and the plurality of processing units process the target video frame data in parallel; In each processing unit, the target video frame data is processed according to the execution order of each computing node.

3. The method of claim 2, wherein, The computing node comprises a pre-processing node, an inference node and a post-processing node; The target video frame data is processed in each processing unit according to the execution order of each computing node, comprising: The first data is obtained by processing the target video frame data by the pre-processing node, and the first data is stored in the first cache queue; In the case where it is determined that the inference node is in an idle state, the first data is read from the first cache queue to the inference node, and the second data is obtained by processing the first data by the inference node, and the second data is stored in the second cache queue; In the case where it is determined that the post-processing node is in an idle state, the second data is read from the second cache queue to the post-processing node, and the third data is obtained by processing the second data by the post-processing node.

4. The method of claim 3, wherein, The determination of whether the frame rate of each processed video frame data meets a preset frame rate threshold comprises: According to the second data obtained by the inference node of each processing unit, the response time of algorithm inference is determined respectively; Based on the response time of algorithm inference, it is determined whether the frame rate of the second data processed by each inference node meets the preset frame rate threshold; If it meets the preset frame rate threshold, the second data is transmitted to the post-processing node of each processing unit respectively; If it does not meet the preset frame rate threshold, the second data is processed according to a preset frame dropping strategy to obtain processed second data, wherein the preset frame dropping strategy is associated with the total frame rate of the inference node of each processing unit.

5. The method of claim 4, wherein, The plurality of processed video frame data is reprocessed by the plurality of processing units respectively, comprising: The processed second data is input into each of the pre-processing nodes, and each of the pre-processing nodes processes the processed second data until a frame rate of the processed second data meets the preset frame rate threshold.

6. The method of claim 4, wherein, The frame dropping processing of the second data according to the preset frame dropping strategy comprises: If the first total frame rate of all the video frame data is greater than the second total frame rate of the inference nodes of each processing unit, it is determined that the first total frame rate is greater than the second total frame rate of the redundant video frame data; Frame dropping operation is performed on the redundant video frame data.

7. The method of claim 1, wherein, Before the video stream data is divided into a plurality of video frame data according to the time sequence, the method further comprises: comparing the estimated time consumption of each processing unit with the processing time consumption of the video stream data determined according to the frame rate of the video stream data; If the estimated time consumption is greater than the processing time consumption, frame skipping processing is performed on the video stream data according to a preset frame skipping strategy to obtain frame-skipped video stream data; The video stream data is divided into a plurality of video frame data according to the time sequence, comprising: The frame-skipped video stream data is divided into the plurality of video frame data according to the time sequence of the video stream data.

8. The method of claim 3, wherein, The target video stream is obtained based on the processing results of the plurality of processing units, comprising: The processing results output by the post-processing nodes in the plurality of processing units are sequentially aggregated to obtain the target video stream.

9. The method of claim 2, wherein, The configuration information corresponding to each of the plurality of processing units is loaded based on the loading model, comprising: The weights and configuration items corresponding to the inference nodes in each processing unit and the required maximum memory of each computing node are obtained from the loading model of the edge box; Each inference node is configured based on the obtained weights and configuration items corresponding to the inference nodes in each processing unit and the required maximum memory of each computing node.

10. A performance optimization device for edge boxes, characterized in that, The device comprises: A division unit is configured to divide video stream data received by the edge box into a plurality of video frame data according to a time sequence; A first processing unit is configured to process the plurality of video frame data through a plurality of processing units to obtain a plurality of processed video frame data; each processing unit includes a computing node for processing video frame data, and all configuration information corresponding to each computing node is loaded in advance; A first determination unit is configured to determine whether the frame rate of each processed video frame data meets a preset frame rate threshold; A second processing unit is configured to re-process the plurality of video frame data through the plurality of processing units until the frame rate of all processed video frame data meets the preset frame rate threshold; A second determination unit is configured to obtain a target video stream based on the processing results of the plurality of processing units.