Data processing method and device and electronic equipment
By integrating AI algorithms at the driver layer, the multimedia data collected by the camera is solved, and the user experience improvement and data security guarantee of sharing AI functions in different applications is achieved.
Patent Information
- Application Number
- CN202510726142.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing video processing solution, the multimedia data collected by the camera is transmitted to the application software through the DMFT framework, resulting in the application volume swelling and the inability to cover the unintegrated artificial intelligence functions, limiting the user experience and functional universality.
AI processing is placed at the driver layer to perform, and AI algorithms are integrated into camera device extension components through the DMFT framework to realize intelligent processing of multimedia data, and interact in the kernel space to ensure data security and low latency, while supporting data output from different applications.
It realizes sharing of AI functions in different applications, improves user experience and functional versatility, reduces data processing delays and ensures data security.
Smart Images

Figure CN120343346A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of data processing technology, and particularly to a data processing method, apparatus, and electronic device. Background Art
[0002] In the current video processing solution, the multimedia data collected by the camera is transmitted to the application software through the DMFT framework, and the artificial intelligence model is called by the application layer to implement video enhancement. However, this method not only causes the application volume to expand, but also cannot cover the application software that does not integrate artificial intelligence functions, thus limiting the user experience and function universality. Summary of the Invention
[0003] In view of the above problems, the present disclosure provides a data processing method, apparatus, and electronic device.
[0004] According to a first aspect of the present disclosure, a data processing method is provided, including: obtaining first multimedia data in response to a target instruction; the target instruction is generated by a target application, and the first multimedia data is collected by a target hardware; processing the first multimedia data based on a target model to obtain corresponding first information; processing the corresponding first multimedia data based on the first information to obtain second multimedia data; and outputting the second multimedia data to the target application.
[0005] According to an embodiment of the present disclosure, the method further includes: for each first data frame in the first multimedia data, performing a first process to obtain a corresponding second data frame, so that the target model processes the second data frame; the first process is used to make the first data frame meet the model input condition; processing the corresponding first multimedia data based on the first information to obtain second multimedia data, including: for each first data frame, superimposing the first information onto the first data frame to obtain a corresponding third data frame.
[0006] According to an embodiment of the present disclosure, the method further includes: for each first data frame in the first multimedia data, during output, in response to not obtaining the corresponding first information, directly outputting the first data frame to the target application or outputting a fourth data frame that has undergone a second process to the target application; the second process is used to make the first data frame meet the application input condition.
[0007] According to an embodiment of the present disclosure, for each first data frame in the first multimedia data, the method further includes: performing a first process to obtain a corresponding second data frame, storing the second data frame in a buffer pool, so that a target model processes based on the data frames in the buffer pool; performing a second process to obtain a corresponding fourth data frame; when outputting, in response to not obtaining a first piece of information of an output frame, directly outputting the fourth data frame to a target application; when outputting, in response to obtaining the first piece of information of the output frame, processing the fourth data frame of the output frame based on the first piece of information, and outputting a fifth data frame to the target application.
[0008] According to an embodiment of the present disclosure, processing the first multimedia data based on a target model includes: determining the number of data frames in the buffer pool that have not been processed by the model; in response to the number of data frames being less than a threshold, controlling the target model to process the data frames that have not been processed by the model based on a first strategy, so that the target model meets a processing accuracy condition; in response to the number of data frames being greater than the threshold, controlling the target model to process the data frames that have not been processed by the model based on a second strategy, so that the target model meets a processing efficiency condition.
[0009] According to an embodiment of the present disclosure, the method further includes: obtaining a status identifier of an output frame based on a communication connection corresponding to a thread of the target model; in response to the status identifier indicating that the first data frame corresponding to the output frame has not been processed by the target model or the processing of the target model is not completed, determining that the first piece of information of the output frame has not been obtained; in response to not obtaining the status identifier, determining that the first piece of information of the output frame has not been obtained.
[0010] According to an embodiment of the present disclosure, the method further includes: for the first process, calling a first target function from an acceleration library to perform intelligent acceleration on the first process; for the second process, calling a second target function from the acceleration library to perform intelligent acceleration on the second process.
[0011] A second aspect of the present disclosure provides a data processing device, including: a first module, configured to obtain first multimedia data in response to a target instruction; the target instruction is generated by triggering of a target application, and the first multimedia data is collected by a target hardware; a second module, configured to process the first multimedia data based on a target model to obtain corresponding first information; a third module, configured to process the corresponding first multimedia data based on the first information to obtain second multimedia data; a fourth module, configured to output the second multimedia data to the target application.
[0012] According to an embodiment of the present disclosure, it further includes: a fifth module, configured to perform a first process on each first data frame in the first multimedia data to obtain a corresponding second data frame, and store the second data frame in a buffer pool, so that the target model processes based on the data frames in the buffer pool; a sixth module, configured to perform a second process on each first data frame in the first multimedia data to obtain a corresponding third data frame.
[0013] A third aspect of the present disclosure provides an electronic device, a memory, at least one processor, and target hardware, where: the target hardware is used to collect first multimedia data; the memory is used to store computer instructions and a target model; the processor is used to load the target model to process the first multimedia data and obtain corresponding first information; the processor is further used to load the computer instructions to implement: processing the corresponding first multimedia data based on the first information to obtain second multimedia data, and outputting the second multimedia data to a target application.
[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0016] Figure 1 Schematically shows an application scenario diagram of a data processing method, device, and electronic device according to an embodiment of the present disclosure;
[0017] Figure 2 Schematically shows a flowchart of a data processing method according to an embodiment of the present disclosure;
[0018] Figure 3 Schematically shows a schematic diagram of preprocessing, AI processing, and post-processing methods for images with AI processing according to an embodiment of the present disclosure;
[0019] Figure 4 Schematically shows a schematic diagram of other processing methods for images without AI processing according to an embodiment of the present disclosure;
[0020] Figure 5 Schematically shows a schematic diagram of preprocessing, inference by an AI thread, and other processing according to an embodiment of the present disclosure;
[0021] Figure 6 Schematically shows a schematic diagram of adjusting a target model processing strategy according to an embodiment of the present disclosure;
[0022] Figure 7 Schematically shows a schematic diagram of determining first information based on a status identifier according to an embodiment of the present disclosure;
[0023] Figure 8 Schematically shows a system architecture diagram of a data processing method according to an embodiment of the present disclosure;
[0024] Figure 9Shows a schematic diagram of performance evaluation after adding AI processing in DMFT according to an embodiment of the present disclosure;
[0025] Figure 10 Shows a schematic diagram of performance evaluation after optimizing by adding AI processing in DMFT according to an embodiment of the present disclosure;
[0026] Figure 11 Schematically shows a structural block diagram of a data processing device according to an embodiment of the present disclosure; and
[0027] Figure 12 Schematically shows a block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure. Detailed implementation manners
[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0031] It should be noted that in the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved all comply with the provisions of relevant laws and regulations, and necessary confidentiality measures are taken, and do not violate public order and good customs. In the technical solution of the present disclosure, before obtaining or collecting user personal information, the authorization or consent of the user is obtained.
[0032] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0033] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means two or more, unless otherwise specifically defined.
[0034] It has been found through research that in some examples, an application (APP) calls an artificial intelligence (AI) model at the application layer for processing to obtain a video that meets the requirements, that is, performing AI operations at the application layer. First, it is necessary to copy the images captured by the camera from the Microsoft multimedia kernel to the user space, that is, user space. This not only poses risks in terms of data security, but also integrates the algorithm with the application, resulting in the application being overly bloated and at the same time being inconvenient for compatibility with different conference application software.
[0035] In view of this, the embodiments of the present disclosure provide a data processing method, apparatus, and electronic device. The data processing method, apparatus, and electronic device will be introduced below with reference to the accompanying drawings.
[0036] Figure 1 FIG. 100 is a schematic diagram showing an application scenario of the data processing method, apparatus, and electronic device according to an embodiment of the present disclosure.
[0037] It should be noted that Figure 1 The shown is only an example of the scenario to which the embodiments of the present disclosure can be applied to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.
[0038] Such as Figure 1As shown, the application scenario 100 according to this embodiment may include a terminal device 101 and a terminal device 102. Data interaction may occur between the terminal device 101 and the terminal device 102 through a network, that is, a medium providing a communication link. The network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The communication link established by the network supports streaming media transmission protocols such as the Real-Time Transport Control Protocol and the Real-Time Messaging Protocol to ensure high-quality and low-latency transmission of multimedia data between devices.
[0039] The terminal device 101 is an intelligent terminal device with multimedia processing capabilities, and can install and run various application programs, including multimedia processing applications such as video conferencing applications, live platform clients, and security monitoring software (only for examples), and is capable of receiving, decoding, and presenting audio and video data streams. The terminal device 101 includes, but is not limited to, electronic devices with display functions such as smartphones, smart TVs, tablets, and in-vehicle infotainment systems.
[0040] The terminal device 102 is a multimedia data collection terminal, and may include image collection devices such as network cameras, intelligent monitoring cameras, and action cameras, and is equipped with a camera device extension component DMFT (Device Media Foundation Transform), which can perform intelligent processing, dynamic coding optimization, and network transmission adaptation on the collected real-time audio and video data, and realize high-quality and low-latency multimedia data transmission to the terminal device 101 through the DMFT framework.
[0041] In some embodiments, the terminal device 102 may further include auxiliary data collection devices such as environmental sensors and infrared camera modules to provide a richer multimedia data source.
[0042] It should be noted that the data processing method provided by the embodiments of the present disclosure can generally be applied to the terminal device 102, specifically to the DMFT. Correspondingly, the devices and electronic devices provided by the embodiments of the present disclosure can be set in the terminal device 102. The data processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the terminal device 102 and capable of communicating with the terminal device 102. Correspondingly, the data processing device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the terminal device 102 and capable of communicating with the terminal device 102. It should be understood that Figure 1 the number of terminal devices in
[0043] is merely illustrative. According to actual needs, there can be any number of terminal devices. Figure 1 Based on the Figures 2 - 8 scenario described below, the data processing method of the embodiments of the present disclosure will be described in detail through
[0044] The execution subject of the data processing method provided by the embodiments of the present disclosure can be DMFT. In the Windows operating system, the camera video capture stack can provide underlying support for the acquisition and preliminary processing of multimedia data. As an extended component of the target hardware in this stack, that is, the camera device extension component, DMFT supports user-mode extensions in the form of DMFT. DMFT can receive the first multimedia data (i.e., the original image frame) from the camera device and can perform post-processing operations on the image frame within DMFT. DMFT can internally include multiple MFT (Media Foundation Transform) image processing plugins, and these different MFT plugins complete different image processing tasks, such as background replacement / blurring, beauty enhancement, face focusing, and effect comparison, etc.
[0045] As Figure 2 shown, the data processing method of this embodiment includes S210 to S240.
[0046] In S210, in response to a target instruction, the first multimedia data is obtained. The target instruction is generated by a target application being triggered, and the first multimedia data is collected by the target hardware.
[0047] The target application can be a software program designed and developed based on specific business requirements, user operation intentions, or system preset rules, and the target instruction can be generated by the user interacting with the target application. When using a specific multimedia application, the user can, based on their own business needs or operation intentions, select a specified AI function, such as face detection, portrait segmentation, etc., on the target application interface, thereby triggering the generation of the target instruction. For example, in a video surveillance analysis application, in order to keep track of the people in the surveillance footage in real time, the user can select the "face detection" function in the application's settings interface and can also set corresponding triggering conditions, such as automatically starting the function when a specific number of people appear in the footage. Once the triggering conditions are met, the application can immediately generate the target instruction.
[0048] The generation process of the target instruction is related to the internal logical judgment and instruction encoding of the target application. Information such as the type of the instruction, parameters (such as the sensitivity of face detection, the accuracy of portrait segmentation, etc.), and the execution timing can be determined according to the user's specific operations or preset rules.
[0049] In different multimedia application scenarios and requirements, different target hardware can be selected or configured to collect the first multimedia data. For example, the target hardware can be an independent video source device such as a camera, or a camera module integrated in a mobile device (such as a smartphone, tablet, etc.), portable device, to start the corresponding acquisition function according to the requirements of the target instruction.
[0050] In some complex application scenarios, the target hardware can also be composed of multiple devices. For example, in the security monitoring scenario, in order to accurately detect the facial information of people in the monitoring video, a high-resolution camera can be used to capture the facial features of the people in the video, ensuring that high-quality first multimedia data can be collected in different environments. When the user specifies the function of portrait segmentation, in addition to the camera, auxiliary devices such as depth sensors can be introduced into the target hardware to obtain the distance information of the objects in the video, providing additional spatial dimension data to help the system more accurately distinguish people from the background.
[0051] In S220, the first multimedia data is processed based on the target model to obtain corresponding first information.
[0052] The target model can be an AI model constructed based on advanced artificial intelligence technologies. Among them, deep learning technology is the mainstream method for constructing such models. The deep learning model can be trained with a large amount of labeled data to learn the complex patterns and features in the data, so as to be able to make accurate predictions and analyses on new input data.
[0053] For example, in the field of face recognition, the target model can be a Convolutional Neural Network (CNN) model trained with a large dataset containing face images, which can identify the faces in the images and extract the feature vectors of the faces for tasks such as face comparison and identity verification. In the field of portrait segmentation, the target model can adopt a semantic segmentation network to classify the image pixel by pixel to distinguish the person and background areas.
[0054] When the first multimedia data (such as video frames or images) is input into the target model, the model can process it according to the algorithms and structures pre-designed inside it, that is, extract and transform the features of the data through a series of neural network layers such as convolutional layers, pooling layers, and fully connected layers. After being processed by the target model, corresponding first information can be obtained, and these information can be organized and stored in a specific data structure, such as the data structures of inferred face detection boxes, portrait segmentation, etc. Among them, the data structure of the face detection box can contain multiple key information to describe the position and size of the detected face in the image. The common components can include the upper left coordinates (x, y) of the face box, the width (width) and height (height) of the face box; the data structure of portrait segmentation can exist in the form of a pixel-level mask (mask).
[0055] In S230, the corresponding first multimedia data is processed based on the first information to obtain second multimedia data.
[0056] Based on the acquired first information, targeted processing can be performed on the first multimedia data, such as operations like enhancement, annotation, or conversion, to obtain second multimedia data that meets the requirements of different application scenarios. Taking a face detection box as an example, it clearly identifies the position and size range of the face in the image. Therefore, the coordinate and size information of the face detection box in the first information can be used to draw a rectangular box at the corresponding position in the original image. At the same time, specific colors, line types, and transparencies can be set for this rectangular box to obtain second multimedia data that distinguishes the face area from the background.
[0057] In addition to overlaying the face detection box, various other processes can be performed on the first multimedia data based on the first information. For example, according to the result of portrait segmentation, beautification processing can be performed on the person area, while blurring processing is performed on the background area to highlight the person as the main subject. Or, according to the position and size of the face detection box, the face area can be locally enlarged to more clearly display the facial expressions of the person.
[0058] In S240, output to the target application is performed based on the second multimedia data.
[0059] After processing the first multimedia data based on the first information and obtaining the second multimedia data, the processed second multimedia data can be output to the target application. Different target applications have different requirements and processing methods for the second multimedia data. Therefore, the output process can be adapted according to the characteristics of the target application. For example, in a video surveillance system, the target application can be the display terminal in the surveillance center, and the second multimedia data with information such as the overlaid face detection box can be displayed on the screen in real time; in a video conferencing application, the target application can be the client devices of the participants, and the processed video data can be smoothly transmitted and displayed.
[0060] The output methods can include local display, network transmission, and file storage, etc. Among them, local display is suitable for directly presenting the processing result on local devices, such as computer monitors and mobile device screens, etc.; network transmission is used to send the processed data to a remote target application, such as transmitting video data to other users' devices through the Internet; file storage is to save the processing result as a file for subsequent viewing, editing, and analysis.
[0061] Exemplarily, in the intelligent security scenario, the target application is a security monitoring software. In the target application, a user can trigger a target instruction to monitor suspicious persons (such as those wearing black hats and black clothes) in a specific area. In response to the target instruction, through the target hardware, i.e., the camera in that area, video, i.e., the first multimedia data, can be collected. Then, after obtaining the first multimedia data from the camera through DMFT, face detection and personnel feature recognition are performed on the video frames of the first multimedia data based on the target model to obtain first information such as the face position and clothing. Then, according to the first information, if a person meeting the characteristics is detected, a red prominent box is drawn outside the corresponding face box and the feature information is marked. After the above processing, the video is the second multimedia data. Finally, the second multimedia data can be transmitted to the monitoring center software in real time, and the marked image is displayed after the software decodes it, so that security personnel can quickly locate suspicious persons.
[0062] It can be understood that when developing streaming media AI processing based on the Microsoft Device Media Foundation Transform (DMFT) framework, the embodiments of the present disclosure propose to place the AI processing, i.e., processing the first multimedia data based on the target model, in the driver layer for execution. Specifically, the AI algorithm can be integrated into the driver service and can be called by the Windows CameraFrame Service. In this way, the first multimedia data can be interacted in the same kernel space, which can not only ensure data security but also reduce data processing latency, enabling different APPs to share the processed second multimedia data and directly output it. At the same time, even when different cameras are connected and different conferencing software is used, users can experience these AI functions, improving generality and practicality.
[0063] Figure 3 Schematically shows a schematic diagram of preprocessing, AI processing, and post-processing methods for images with AI processing according to embodiments of the present disclosure.
[0064] In the embodiments of the present disclosure, as Figure 3 shown, the method further includes: for each first data frame in the first multimedia data, performing a first process to obtain a corresponding second data frame, so that the target model processes the second data frame; the first process is used to make the first data frame meet the model input conditions; processing the corresponding first multimedia data based on the first information to obtain the second multimedia data, including: for each first data frame, superimposing the first information onto the first data frame to obtain a corresponding third data frame.
[0065] During the real-time processing of multimedia data streams, the input stream process can continuously obtain first multimedia data containing multiple first data frames from a data source. For example, in a real-time video surveillance system, a camera captures multiple frames of images per second, and the input stream process can sequentially obtain each first data frame at a preset frame rate. The nth frame can be a specific first data frame that needs to be processed currently.
[0066] Since different target models have different requirements for the format, size, color space, etc. of input data, the first data frame can be subjected to a first process, namely preprocess, to ensure the consistency and compatibility of the obtained second data frame, so that the second data frame meets the input conditions of the target model, thereby improving the accuracy and efficiency of the target model in processing the second data frame.
[0067] Exemplarily, if the target model requires the input image to have a specific size (such as 224×224 pixels), and the original size of the first data frame does not meet the requirement, the first process will use an image scaling algorithm (such as bilinear interpolation or bicubic interpolation, etc.) to adjust the size of the image to obtain a second data frame of 224×224 pixels to meet the input requirements of the model.
[0068] Exemplarily, if the target model requires the input data to be in a specific color space, such as converting RGB to grayscale or YUV color space. The first process can process the first data frame using the corresponding color space conversion formula according to the requirements of the model to obtain a second data frame. For example, convert the first data frame in the RGB color space to grayscale to reduce the data volume and highlight the brightness information of the image.
[0069] The second data frame obtained after the first process can be input into the target model for inference. After the inference of the AI model, corresponding first information can be obtained, including but not limited to the bounding box coordinates, class labels, confidence levels, etc. of the detected objects. Then, the first information can be superimposed on the original first data frame to obtain a third data frame containing the original image data and the result information of AI processing.
[0070] During the superimposing process, for the first information corresponding to the ith frame returned by the target model, the first information can be subjected to real-time superimposing processing with the nth frame image to be processed currently, where: when n = i, that is, when the first information is obtained in real time, spatial position matching and superimposing can be directly performed; when n > i, that is, when there is a processing delay in the first information, a motion estimation compensation algorithm can be used to adjust the superimposing position of the first information, mapping the inference result of the historical frame to the current nth frame; when the first information is missing for multiple consecutive frames, interpolation compensation can be performed based on the inference results of adjacent frames.
[0071] Exemplarily, in a video surveillance system, the first information, i.e., the detected face bounding box and identity information, can be superimposed on each original first data frame to obtain a third data frame that facilitates security personnel to quickly identify and locate the target person.
[0072] After obtaining the third data frame, corresponding processing can also be performed on the third data frame to make the third data frame meet the model output conditions, realize seamless data transfer and efficient utilization, ensure that the finally output data has good usability, compatibility and accuracy, and adapt to different application scenarios and subsequent processing requirements.
[0073] It can be understood that the first processing is performed on each first data frame to meet the model input conditions, so as to avoid inference errors or performance degradation of the model due to inconsistent input data formats, and ensure that the model can process data efficiently and accurately. At the same time, each first data frame is superimposed with the corresponding first information to obtain a third data frame, which can ensure the comprehensive alignment of data and improve the accuracy of data processing.
[0074] Figure 4 A schematic diagram schematically shows other processing methods for images without AI processing according to an embodiment of the present disclosure.
[0075] In an embodiment of the present disclosure, as Figure 4 shown, the method further includes: for each first data frame in the first multimedia data, when outputting, in response to not obtaining the corresponding first information, directly outputting the first data frame to the target application or outputting a fourth data frame that has been second-processed to the target application; the second processing is used to make the first data frame meet the application input conditions.
[0076] During the data frame processing, interaction with the AI Inference Thread may be involved. The system can send a message to notify the AI Inference Thread to perform an inference operation, and then obtain the first information inferred by it. There are two situations when obtaining the inference result: one is that the inference has not been completed due to reasons such as complex AI model calculation and large data volume, resulting in no inference result or processing in progress; the other is that there is a new inference result, that is, the AI Inference Thread has completed the inference of the current data frame and generated the corresponding first information.
[0077] When performing an output operation on a certain first data frame in the first multimedia data, if in response to not obtaining the corresponding first information, the original first data frame can be directly output to the target application. This method is applicable to application scenarios with high requirements for data real-time performance and that do not rely on AI inference results. For example, in some simple real-time monitoring scenarios, simply displaying the original picture can meet the basic requirements. Or, a fourth data frame that has undergone a second process can be output to the target application. The second process can enable the first data frame to meet the application input conditions, such as performing operations like resolution conversion and color space adjustment.
[0078] The Postprocess module can process the current data frame according to the inference result of the AI Inference Thread. If there is no inference result, the Postprocess module can directly output the current frame n, that is, output while maintaining the original state of the data frame or the state after the second process. If there is an inference result, the Postprocess module can superimpose the inferred first information on the current frame n. The data frame processed by Postprocess can be output by the output stream process. That is, when the first information is obtained, the output stream process can obtain the result of the inference process from the AI Inference thread through thread communication and then superimpose it on the video image to be output currently. When the first information is not obtained, the output stream process can also learn through thread communication from the thread that the inference is not completed and directly output the current video image. After completing the output of one frame of data, the system can enter the loop processing stage and continue to perform the same processing flow on the next frame of data, thereby realizing the continuous and automated processing of multimedia data and meeting the requirements of different application scenarios for the real-time processing and analysis of multimedia data.
[0079] Exemplarily, in a video conference, the camera captures the images of the participants in real time to form the first multimedia data, and each frame of the image is the first data frame. Then, the target model can be used to analyze the picture, such as detecting the status of the participants (such as whether someone has left, etc.) to generate the first information. However, the analysis may have a delay, resulting in the inability to output the obtained first information in a timely manner. At this time, the first data frame can be directly output to the terminals of each participant for everyone to watch the real-time picture. If the terminal has requirements for the image format, such as a specific compression ratio, the first data frame can also be second-processed using an image compression algorithm to generate a fourth data frame that meets the conditions and then output.
[0080] It should be noted that after the target model successfully generates the first information, it can be associated or subsequently processed with the already output first data frame as needed.
[0081] It can be understood that directly outputting the original frame without the first information and running the thread processed by the target model independently and in parallel with other processing threads for data frames not processed by the target model can not only greatly improve the system efficiency but also meet diverse application requirements.
[0082] Figure 5 A schematic diagram showing preprocessing, inference by the AI thread, and other processing according to an embodiment of the present disclosure is schematically shown.
[0083] Based on the above embodiments, in this embodiment, as Figure 5 shown, for each first data frame in the first multimedia data, the method further includes: performing a first process to obtain a corresponding second data frame, storing the second data frame in a buffer pool so that the target model processes based on the data frames in the buffer pool; performing a second process to obtain a corresponding fourth data frame; at the time of output, in response to not obtaining the first information of the output frame, directly outputting the fourth data frame to the target application; at the time of output, in response to obtaining the first information of the output frame, processing the fourth data frame of the output frame based on the first information and outputting a fifth data frame to the target application.
[0084] For each first data frame in the first multimedia data, a first process can be performed to meet the input conditions of the target model and obtain a corresponding second data frame. Then, an image copy operation can be performed, that is, copying the second data frame after the first process to a specially constructed buffer pool (video buffer, abbreviated as VB). The buffer pool is associated with the target model processing thread and serves to provide a stable and orderly data source for subsequent target model processing, ensuring that the target model can obtain image frames from the VB frame by frame for processing, and the processing content includes but is not limited to adjusting the brightness, contrast, and granularity of the image. While storing the second data frame in the buffer pool, the threads of other processing normally flow the first multimedia data, ensuring that other threads are not affected by the processing speed of the target model, that is, reducing the impact of AI processing on the DMFT video stream.
[0085] In the DMFT thread of other processing different from the inference thread, for each first data frame in the first multimedia data, a second process can be performed to make the first data frame meet the application input conditions and obtain a corresponding fourth data frame.
[0086] In the data frame output stage, different output strategies can be adopted based on whether the first information of the output frame is obtained. If the first information of the output frame is not obtained, the first data frame can be directly subjected to a second process to obtain a fourth data frame that meets the application input conditions, and output it to the target application. If the first information of the output frame is obtained, the fourth data frame passed by the DMFT thread can be superimposed with the first information after model inference. For example, the first information can be integrated into the fourth data frame in the form of annotation, superimposition, etc. to generate a fifth data frame containing rich information and output it to the target application. In this way, the target application can not only receive image data that meets the format requirements, but also obtain key information related to the image, providing more comprehensive and valuable content for users.
[0087] Exemplarily, in an intelligent video surveillance and analysis system, a camera can continuously collect surveillance images as the first multimedia data, and each frame of the image is the first data frame. Then, the system can perform a first process, such as performing noise reduction, color correction, etc. on the first data frame. For example, the random noise in the image can be removed by the median filtering algorithm, and then the color curve can be adjusted to make the image color more real and natural. After processing, a second data frame is obtained. Subsequently, these second data frames can be stored in the VB, so that the target model can obtain the data frames frame by frame from the VB for analysis, generating first information such as target object recognition results, behavior judgment information, etc. In the output link, if the first information of the output frame is obtained, the system can further process the fourth data frame (obtained by performing a second process such as resolution adaptation on the first data frame) based on this information. For example, if the first information shows that there is a suspicious person in the image, a position box and identity prompt information of the suspicious person can be superimposed on the fourth data frame to generate a fifth data frame containing key information and output it to the target application in the monitoring center for security personnel to view.
[0088] It can be understood that through the first process, the AI thread for inference, and the second process, while ensuring the synchronous operation of the inference thread and other processing threads, it does not affect the normal processing of other threads, realizing efficient and flexible data output.
[0089] It should be noted that for the details not described in this embodiment, please refer to the implementation details of the foregoing embodiments.
[0090] In the embodiments of the present disclosure, the method further includes: for the first process, calling a first target function from an acceleration library to perform intelligent acceleration on the first process; for the second process, calling a second target function from the acceleration library to perform intelligent acceleration on the second process.
[0091] In some exemplary embodiments, an acceleration library based on OpenGL (Open Graphics Library) can be created. As a cross-language and cross-platform professional graphics program interface, OpenGL has powerful computing capabilities and a rich and diverse function library, providing a basis for building an efficient acceleration library. In terms of computing power, OpenGL has excellent parallel computing characteristics. In today's hardware environment where multi-core processors are widely used, it can fully utilize multiple cores of the processor to simultaneously process a large amount of data in parallel. Taking video image processing as an example, when faced with a large amount of pixel data, the OpenGL acceleration library can divide the image into multiple small blocks and allocate them to different processor cores for simultaneous processing. In addition, the OpenGL acceleration library also has a rich and powerful set of functions, covering all aspects of graphics processing, including image rendering, geometric transformation, texture mapping, and lighting processing. When building the acceleration library, developers can flexibly call these functions according to specific application requirements to implement various image processing algorithms. In addition, the OpenGL acceleration library also has good cross-platform compatibility and can run on multiple operating systems, enabling applications developed based on the OpenGL acceleration library to be easily deployed on different hardware platforms without a large amount of code modification and adaptation work.
[0092] In the first processing stage, image optimization tasks such as noise reduction, color correction, and sharpening can be performed on the first data frame to meet the model input conditions. To accelerate the first processing speed, the first target function can be called from the acceleration library to achieve intelligent acceleration of the first processing, enabling the first processing to be completed quickly and with high quality, generating the second data frame, and providing a data basis for subsequent target model inference.
[0093] Exemplarily, during the noise reduction process, the first target function can utilize the parallel computing characteristics of the openGL acceleration library to simultaneously process multiple pixel points in the first data frame image, greatly shortening the processing time compared to the traditional serial processing method. For color correction, the function can quickly adjust parameters such as color balance, contrast, and saturation of the first data frame image to make the image color more realistic and natural.
[0094] In the second processing stage, operations such as image format conversion, resolution adjustment, and encoding compression can be performed on the first data frame to make the first data frame meet the input conditions of the target application. During this process, the second target function can be called from the acceleration library to achieve intelligent acceleration, enabling the second processing to be efficiently completed, generating the fourth data frame, and ensuring that the data can be smoothly output to the target application.
[0095] Taking image format conversion as an example, the second objective function can utilize the graphics processing capabilities of the OpenGL acceleration library to quickly convert an image from one format to another while maintaining the image quality. During the resolution adjustment process, this function can quickly generate a resolution image that meets the requirements of the target application through efficient interpolation algorithms and parallel computing, thus ensuring that in application scenarios with high real-time requirements such as video conferencing, even if the resolutions of the first data frames are different, the real-time requirements can be met without stuttering or frame loss. During encoding and compression, it can utilize the optimized algorithms of the acceleration library to minimize the size of the image data as much as possible while ensuring the image quality, thereby improving the data transmission efficiency.
[0096] It can be understood that since the input and output of a general target model are fixed, in practical applications, preprocessing needs to be performed frame by frame to adapt to the video stream. The toolset for accelerating AI application development only targets the target model and its inference, and additional acceleration processing is required for preprocessing, etc. Otherwise, the processing efficiency of a general CPU is relatively low and time-consuming. Through the intelligent acceleration of the first processing and the second processing provided by the embodiments of the present disclosure, it is ensured that in addition to the acceleration of AI inference, other image processing can also be accelerated in parallel, thereby reducing the latency of image processing.
[0097] Figure 6 Schematically shows a schematic diagram of adjusting the processing strategy of the target model according to an embodiment of the present disclosure.
[0098] Based on the above embodiments, in this embodiment, as Figure 6 shown, processing the first multimedia data based on the target model includes: determining the number of data frames in the buffer pool that have not been processed by the model; in response to the number of data frames being less than a threshold, controlling the target model to process the data frames that have not been processed by the model based on a first strategy so that the target model meets the processing accuracy condition; in response to the number of data frames being greater than the threshold, controlling the target model to process the data frames that have not been processed by the model based on a second strategy so that the target model meets the processing efficiency condition.
[0099] When the target model processes the first multimedia data, it can obtain unprocessed data frames from the buffer pool and process them. During this process, the AI inference thread can wait and obtain messages from the DMFT thread. If there is no new multimedia data message, that is, no new first multimedia data is input into the buffer pool, it indicates that there is no current inference task, and the processing strategy of the target model can be dynamically adjusted according to the number of unprocessed data frames in the buffer pool.
[0100] First, the number of data frames in the buffer pool that have not been processed by the model can be monitored in real time. When the number of data frames is less than the threshold, it indicates that there is resource redundancy in the model processing relative to the current data input, that is, the model can process the existing data frames within the specified time. At this time, the target model can be controlled to process the data frames that have not been processed by the model according to the first strategy, so that the target model meets the processing accuracy condition. The first strategy can be that the AI Inference Thread copies the current data frame (such as the nth frame) and performs normal inference frame by frame to ensure the accuracy of the inference result. When the number of data frames in the buffer pool that have not been processed by the model is greater than the preset threshold, it indicates that the model processing speed is slow. At this time, the target model can be controlled to process the data frames that have not been processed by the model based on the second strategy, so that the target model meets the processing efficiency condition. The second strategy can be to copy the current data frame and perform subsequent inferences based on the inference results of the data frames that have been processed to quickly generate the inference result of the current frame, thereby reducing the computational amount of the model and improving the processing speed. The second strategy can also be to adjust the inference parameters of the target model and perform inferences based on the simplified inference parameters.
[0101] Exemplarily, taking the face detection of the 10th frame currently processed by the target model as an example, if the number of data frames is greater than the threshold at this time, and it is detected that the position differences of the faces in three consecutive frames (the 7th, 8th, and 9th frames) are small. In this case, the model can directly determine the position of the face detection frame in the 10th frame based on the face detection frames inferred from the first 3 frames at the position 10×10 above the face. In addition, the model can also analyze the coordinate information of the face detection frames in the 7th, 8th, and 9th frames obtained, and predict the approximate position of the face detection frame in the 10th frame by analyzing the change trend of these coordinates. Then, the model can only perform local search and verification near the predicted position of the 10th frame instead of performing a full scan of the entire image, so as to quickly determine the accurate position of the face detection frame.
[0102] Exemplarily, still taking the face detection of the 10th frame by the target model as an example, when the number of data frames is greater than the threshold, the system can adjust the inference parameters of the model. For example, the number of convolutional layers used for feature extraction in the model can be reduced from the original 5 layers to 3 layers, or the size of the convolutional kernel can be adjusted from 3×3 to 2×2. In this way, when the model processes the 10th frame, the number of convolutional operations required will be reduced accordingly, thereby reducing the computational complexity.
[0103] During the process of target model inference, when the number of data frames in the buffer pool is greater than the threshold, multi-threading can also be used for processing. That is, according to the characteristics of the inference task and the structure of the data, a complex inference task is decomposed into multiple sub-tasks and assigned to different threads to execute simultaneously, so as to meet the accuracy condition and efficiency condition at the same time. For example, during the process of target model processing, one thread can be responsible for feature extraction to extract the key features in the data frame; another thread can be responsible for target detection and tracking to judge the position and movement trajectory of the target object according to the extracted features.
[0104] It can be understood that by reasonably processing the data frames in the buffer pool based on the first strategy and the second strategy, it is possible to flexibly balance the processing accuracy and efficiency in different scenarios, ensuring the stability and reliability of multimedia data processing.
[0105] Figure 7 Schematically shows a schematic diagram of determining the first information based on the status identifier according to an embodiment of the present disclosure.
[0106] In an embodiment of the present disclosure, the method further includes: as Figure 7 shown, based on the communication connection with the thread corresponding to the target model, obtain the status identifier of the output frame; in response to the status identifier indicating that the first data frame corresponding to the output frame has not been processed by the target model or the processing of the target model is not completed, determine that the first information of the output frame has not been obtained; in response to not obtaining the status identifier, determine that the first information of the output frame has not been obtained.
[0107] In order to obtain the processing progress of the target model for the data frame (the first data frame or the second data frame) and ensure that the system can perform correct subsequent operations according to the processing result of the target model, a communication connection with the thread corresponding to the target model can be established when the target model is performing the current inference task, and a thread message is set, that is, the status identifier of the output frame is obtained, so that the system can obtain the latest status of the target model processing the data frame in real time according to this status identifier. For example, the system can use a flag bit to mark the status of the data frame. When the flag bit is 0, it means that the frame data has not been processed by the target model; when the flag bit is 1, it means that the frame data has been processed. The status identifier formed by this marking method can quickly and accurately judge the processing status of each data frame.
[0108] When the obtained status identifier indicates that the first data frame corresponding to the output frame has not been processed by the target model or the processing of the target model has not been completed, it can be determined that the first information of the output frame has not been obtained. This may be due to reasons such as complex processing tasks and tight computing resources, resulting in a slow processing progress and incomplete processing. In this case, the system can adopt other strategies, such as directly outputting the original frame or the frame after simple processing, to ensure the real-time nature of the video stream.
[0109] If the system fails to obtain the status identifier, it is also possible to determine that the first information of the output frame has not been obtained, which may occur in situations such as a communication connection interruption or a failure of the target model thread. For example, in an unstable network environment, the communication connection between the system and the target model thread may be temporarily interrupted, resulting in the inability to obtain the status identifier. At this time, according to the preset rules, it can be considered that the first information has not been obtained, and corresponding countermeasures can be taken, such as re - establishing the communication connection and performing error handling, to ensure the normal operation of the system.
[0110] Exemplarily, in a video conferencing scenario, the target model can identify and analyze the people in the video frame. Through the communication connection with the target model thread, the system can obtain the status identifier and can timely understand the processing situation of each video frame (i.e., the first data frame) in the target model based on this identifier. If the status identifier indicates that the processing of the target model is not completed, that is, the target model is processing the nth video frame. At this time, it can be determined that the first information of this frame has not been obtained. Then, the first data frame can be directly output to the target application.
[0111] It can be understood that determining whether the first information is obtained through the status identifier of the output frame not only realizes the real - time monitoring and accurate judgment of the processing status of the target model, but also ensures the smooth progress of the data processing task.
[0112] Regarding the data processing method provided in the embodiments of the present disclosure, that is, the design of PDMFT can be executed on a CPU (Central Processing Unit). The CPU can coordinate the work of each thread and undertake tasks such as data flow and control logic, with strong versatility.
[0113] The target model inference thread can have flexibly selectable computing power resources, including a GPU (Graphics Processing Unit), an NPU (Neural Network Processing Unit), and a CPU. When the device (target hardware) is equipped with a GPU or an NPU, since the GPU and NPU are optimized for deep - learning calculations, they can significantly improve the inference speed and reduce the load on the CPU. Therefore, it is possible to preferentially consider deploying the AI inference task on the dedicated computing power resources. For example, in a video conference, using a GPU for real - time face recognition and expression analysis can provide richer interaction functions without affecting the video fluency. For the NPU, its low - power consumption characteristics enable the AI model to run stably for a long time on some portable conference devices, extending the battery life of the device. When the GPU or NPU resources are unavailable, or the AI inference task is relatively simple and using the CPU will not have too much impact on the system performance, the inference task can be reverted to the CPU.
[0114] In the above - mentioned manner, while ensuring the AI inference computing power, the inference tasks can be reasonably allocated according to the real - time load conditions of each computing resource (CPU / GPU / NPU), and the overall computing power can be balanced.
[0115] Figure 8 Schematically shows a system architecture diagram of a data - processing method according to an embodiment of the present disclosure.
[0116] As Figure 8 shown, in a multimedia data - processing architecture based on the Windows system, first, the first multimedia data can be obtained through the target hardware and then transmitted to the user - mode component Devproxy. Devproxy can transmit commands and video frames between the camera driver of the target hardware and the target application (such as video - processing software). In addition, Devproxy also has a powerful multi - output support ability. When the camera driver generates multiple output streams due to different scenario requirements or hardware characteristics, these outputs can be processed and transmitted. In the DMFT thread, first, the input - stream processing can be performed, that is, the input stream (the first multimedia data) can be obtained from the user - mode component Devproxy, and then subsequent pre - processing and post - processing can be carried out. The second data frame obtained after pre - processing can be copied to the buffer pool through thread communication so that the AI inference thread can obtain data from the buffer pool for AI inference. In this process, priority can be given to deploying the AI inference task on the GPU or NPU to improve the inference speed. In the image - processing stages such as pre - processing and post - processing, the target functions in the acceleration library can be called to optimize and accelerate operations such as image copying, size adjustment, and color conversion. The entire process realizes the collaborative optimization of image processing and AI inference through parallelization and hardware - acceleration technologies (such as optimizing image - processing algorithms and reducing latency), and finally, the output - stream processing can be completed.
[0117] Figure 9 and Figure 10 respectively show schematic diagrams of performance evaluation before and after adding AI processing optimization in DMFT according to an embodiment of the present disclosure.
[0118] As Figure 9 and Figure 10 shown, in the AI - processing thread, by shortening the pre - processing time before input to the target model through the first processing and shortening the post - processing time of the output of the target model, thus, for each combination of video - input resolution and frame rate, the time consumption of each stage and the total time consumption added to the AI - processing thread are significantly reduced, the video - output frame rate is increased, and the overall time consumption of DMFT from input to output is also significantly reduced, indicating that the optimization measures effectively improve the performance of adding AI processing in DMFT.
[0119] Figure 11A block diagram of a data processing device according to an embodiment of the present disclosure is schematically shown.
[0120] As Figure 11 shown, the data processing device 1100 includes a first module 1110, a second module 1120, a third module 1130, and a fourth module 1140.
[0121] According to some embodiments of the present disclosure, the data processing device 1100 can be used to implement the data processing method according to the embodiment of the present disclosure described with reference to Figures 2 - 8 description.
[0122] The first module 1110 can execute, for example, operation S210 to obtain first multimedia data in response to a target instruction; the target instruction is generated by a target application, and the first multimedia data is collected by target hardware.
[0123] The second module 1120 can execute, for example, operation S220 to process the first multimedia data based on a target model to obtain corresponding first information.
[0124] The third module 1130 can execute, for example, operation S230 to process the corresponding first multimedia data based on the first information to obtain second multimedia data.
[0125] The fourth module 1140 can execute, for example, operation S240 to output to the target application based on the second multimedia data.
[0126] For example, any multiple of the first module 1110, the second module 1120, the third module 1130, and the fourth module 1140 can be combined and implemented in one module, or any one of the modules can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the first module 1110, the second module 1120, the third module 1130, and the fourth module 1140 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as hardware or firmware for integrating or packaging circuits, or can be implemented in any one of the three implementation manners of software, hardware, and firmware or in any appropriate combination of several of them. Or, at least one of the first module 1110, the second module 1120, the third module 1130, and the fourth module 1140 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding function.
[0127] According to an embodiment of the present disclosure, the data processing device 1100 may further include a fifth module 1150 and a sixth module 1160.
[0128] The fifth module 1150 may execute the implementation details of the foregoing embodiments, and is configured to perform a first process on each first data frame in the first multimedia data to obtain a corresponding second data frame, and store the second data frame in the buffer pool, so that the target model processes based on the data frames in the buffer pool.
[0129] The sixth module 1160 may execute the implementation details of the foregoing embodiments, and is configured to perform a second process on each first data frame in the first multimedia data to obtain a corresponding fourth data frame.
[0130] For example, any combination of the first module 1110, the second module 1120, the third module 1130, the fourth module 1140, the fifth module 1150, and the sixth module 1160 may be combined and implemented in one module, or any one of the modules may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the first module 1110, the second module 1120, the third module 1130, the fourth module 1140, the fifth module 1150, and the sixth module 1160 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable manner that can be achieved by integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of them. Alternatively, at least one of the first module 1110, the second module 1120, the third module 1130, the fourth module 1140, the fifth module 1150, and the sixth module 1160 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0131] It should be understood that the data processing device in the embodiments of the present disclosure corresponds to the data processing method part in the embodiments of the present disclosure, and their specific implementation details are the same, and will not be repeated here.
[0132] It should be noted that in the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs. In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the user's authorization or consent has been obtained.
[0133] Embodiments of the present disclosure also provide an electronic device, which may include a memory, at least one processor, and target hardware. Among them, the target hardware is used to collect first multimedia data; the memory is used to store computer instructions and a target model; the processor is used to load the target model to process the first multimedia data to obtain corresponding first information; the processor is also used to load the computer instructions to implement: processing the corresponding first multimedia data based on the first information to obtain second multimedia data, and outputting the second multimedia data to a target application based on the second multimedia data.
[0134] Figure 12 Schematically shows a block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure.
[0135] As Figure 12 shown, the electronic device 1200 according to an embodiment of the present disclosure includes a processor 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage section 1208 into a random access memory (RAM) 1203. The processor 1201 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1201 may also include on-board memory for caching purposes. The processor 1201 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0136] In the RAM 1203, various programs and data required for the operation of the electronic device 1200 are stored. The processor 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. The processor 1201 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 1202 and / or the RAM 1203. It should be noted that the program may also be stored in one or more memories other than the ROM 1202 and the RAM 1203. The processor 1201 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0137] According to an embodiment of the present disclosure, the electronic device 1200 may further include an input / output (I / O) interface 1205, and the input / output (I / O) interface 1205 is also connected to the bus 1204. The electronic device 1200 may further include one or more of the following components connected to the I / O interface 1205: an input portion 1206 including target hardware, etc.; an output portion 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1208 including a hard disk, etc.; and a communication portion 1209 including a network interface card such as a LAN card, a modem, etc. The communication portion 1209 performs communication processing via a network such as the Internet. The drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read therefrom can be installed into the storage portion 1208 as needed.
[0138] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0139] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 1202 and / or RAM 1203 and / or one or more memories other than ROM 1202 and RAM 1203.
[0140] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program includes program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the processing method provided by the embodiments of the present disclosure.
[0141] When the computer program is executed by the processor 1201, the above functions defined in the system / apparatus of the embodiments of the present disclosure are performed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0142] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 1209, and / or be installed from the removable medium 1211. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0143] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1209, and / or be installed from the removable medium 1211. When the computer program is executed by the processor 1201, the above functions defined in the system of the embodiments of the present disclosure are performed. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0144] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include but are not limited to, such as Java, C++, python, the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0146] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0147] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A data processing method, comprising: Obtaining first multimedia data in response to a target instruction; The target instruction is generated by a target application, and the first multimedia data is collected by target hardware; Processing the first multimedia data based on a target model to obtain corresponding first information; Processing the corresponding first multimedia data based on the first information to obtain second multimedia data; Outputting based on the second multimedia data to the target application.
2. The method according to claim 1, the method further comprising: Performing first processing on each first data frame in the first multimedia data to obtain a corresponding second data frame, so that the target model processes the second data frame; The first processing is used to make the first data frame meet the model input conditions; The processing the corresponding first multimedia data based on the first information to obtain second multimedia data includes: For each first data frame, superimposing the first information onto the first data frame to obtain a corresponding third data frame.
3. The method according to claim 1, the method further comprising: For each first data frame in the first multimedia data, during output, in response to not obtaining the corresponding first information, directly outputting the first data frame to the target application or outputting a fourth data frame that has undergone second processing to the target application; the second processing is used to make the first data frame meet the application input conditions.
4. The method according to claim 1, for each first data frame in the first multimedia data, the method further comprising: Performing first processing to obtain a corresponding second data frame, and storing the second data frame in a buffer pool, so that the target model processes based on the data frames in the buffer pool; Performing second processing to obtain a corresponding fourth data frame; During output, in response to not obtaining the first information of the output frame, directly outputting the fourth data frame to the target application; During output, in response to obtaining the first information of the output frame, processing the fourth data frame of the output frame based on the first information, and outputting a fifth data frame to the target application.
5. The method according to claim 4, processing the first multimedia data based on a target model includes: Determining the number of data frames in the buffer pool that have not been processed by the model; In response to the number of data frames being less than a threshold, controlling the target model to process the data frames that have not been processed by the model based on a first strategy, so that the target model meets the processing accuracy conditions; In response to the number of data frames being greater than the threshold, controlling the target model to process the data frames that have not been processed by the model based on a second strategy, so that the target model meets the processing efficiency conditions.
6. The method according to claim 4, the method further comprising: Obtaining a status identifier of an output frame based on a communication connection corresponding to the target model thread; In response to the status identifier indicating that the first data frame corresponding to the output frame has not been processed by the target model or the processing of the target model is not completed, determining that the first information of the output frame has not been obtained; In response to not obtaining the status identifier, determining that the first information of the output frame has not been obtained.
7. The method according to claim 4, wherein the method further comprises: For the first process, calling a first target function from an acceleration library to perform intelligent acceleration on the first process; For the second process, calling a second target function from the acceleration library to perform intelligent acceleration on the second process.
8. A data processing device, comprising: A first module, configured to obtain first multimedia data in response to a target instruction; The target instruction is generated by a target application, and the first multimedia data is collected by target hardware; A second module, configured to process the first multimedia data based on a target model to obtain corresponding first information; A third module, configured to process the corresponding first multimedia data based on the first information to obtain second multimedia data; A fourth module, configured to output the second multimedia data to the target application.
9. The device according to claim 8, further comprising: A fifth module, configured to perform a first process on each first data frame in the first multimedia data to obtain a corresponding second data frame, and store the second data frame in a buffer pool, so that the target model processes based on the data frames in the buffer pool; A sixth module, configured to perform a second process on each first data frame in the first multimedia data to obtain a corresponding third data frame.
10. An electronic device, a memory, at least one processor, and target hardware, wherein: The target hardware is configured to collect first multimedia data; The memory is configured to store computer instructions and the target model; The processor is configured to load the target model to process the first multimedia data to obtain corresponding first information; The processor is further configured to load the computer instructions to: process the corresponding first multimedia data based on the first information to obtain second multimedia data, and output the second multimedia data to the target application.