Picture reasoning method, system and equipment based on Pipeline and medium
By adopting a Pipeline-based method in the picture inference system, dynamically loading the hardware accelerated decoding plug-in, and combining the yolov series algorithms, the problem of traditional image parsing is solved, and fast analysis and efficient inference of different image formats are achieved.
Patent Information
- Application Number
- CN202510157905.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-16
AI Technical Summary
The traditional image analysis method has problems such as slow resolution, low efficiency, and single image format, making it difficult to achieve fast dynamic analysis of different image formats, resulting in inefficient inference of single image parsing.
Pipeline-based picture inference method is adopted, by obtaining the original binary data of the picture, using the decodebin plug-in to dynamically load the decoder and demultiplexer, the hardware accelerated decoding plug-in is preferred, the image decoding and color space conversion is completed, and the yolov series object detection methods, object classification algorithms and feature extraction methods are used for inference processing, and the results are finally integrated into a custom protocol format and returned to the http server for visual display.
Through this method, rapid dynamic analysis of different image formats is achieved, which significantly improves the efficiency of single image parsing and inference. Compared with traditional technology, the optimization time is reduced by 50%.
Smart Images

Figure CN120014415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent visual analysis of artificial intelligence, and specifically to a Pipeline-based image reasoning method, system, device and medium. Background Art
[0002] With the development of mobile Internet, smart phones and social networks, massive amounts of image information have been generated. Images, which are not restricted by region or language, have gradually replaced cumbersome and subtle texts and become the main medium for conveying ideas. At the same time, with the rapid development of artificial intelligence and the emphasis on the acquisition and analysis of image information, image recognition technology and image recognition efficiency have become particularly important.
[0003] Image recognition technology is a technology that uses intelligent means to decode, process, and analyze images to identify targets and objects in various modes. The image recognition process includes image decoding, preprocessing, image segmentation, feature vector extraction, and judgment matching. In layman's terms, it is to use computers to understand the content of images like humans, and may generate more intelligent information. However, the traditional image parsing method is the opencv decoding method, which has problems such as slow parsing, low efficiency, and single parsed image format.
[0004] Therefore, how to achieve fast dynamic parsing of different image formats and improve the efficiency of single image parsing and reasoning is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The technical task of the present invention is to provide a Pipeline-based image reasoning method, system, device and medium to solve the problem of how to achieve fast dynamic parsing of different image formats and improve the efficiency of single image parsing and reasoning.
[0006] The technical task of the present invention is achieved in the following way: a Pipeline-based image reasoning method, which is specifically as follows:
[0007] Get the image source and convert it into the original image binary format;
[0008] According to the acquired original data source format, the decodebin plug-in is used to dynamically load the available decoders and demultiplexers, and the hardware accelerated decoding plug-in is given priority to complete the decoding of the image to be processed;
[0009] Normalize the decoded image;
[0010] Obtain the inference result using the corresponding inference method according to the loaded inference model;
[0011] The inference results are integrated into a custom protocol format and returned to the http server for visual display.
[0012] As a preferred method, the image source is obtained and converted into the original image binary format as follows:
[0013] According to the given picture URI, use the REST service to obtain the original information of the picture, that is, the picture information in binary format;
[0014] Pass the binary image information to the starting image source plugin of the Pipeline pipeline.
[0015] Preferably, the decodebin plug-in is used to dynamically load the available decoders and demultiplexers according to the obtained original data source format, and the hardware accelerated decoding plug-in is given priority to obtain the image to be processed, and the decoding of the image to be processed is completed as follows:
[0016] The decodebin plug-in automatically connects different decoders to the pipeline to complete image decoding based on the original image information;
[0017] Pass the decoded data from the upstream Pipeline plugin to the downstream plugin via the hardware buffer ID;
[0018] At the same time, videoconvert is used to convert and adapt the corresponding image color space. Since its upstream and downstream plug-ins already understand each other, when image color space conversion is not required, it runs in pass-through mode to reduce the impact on performance.
[0019] Preferably, the inference model integrates the YOLO series of target detection methods, target classification algorithms and feature extraction methods.
[0020] Preferably, obtaining the inference result using the corresponding inference method according to the loaded inference model includes two stages: primary inference and secondary inference;
[0021] Among them, one inference is used to detect or classify targets in the image, specifically: according to the trained model, the target information in the image is detected or identified from the decoded and normalized image information, and the integrated general yolov4 model inference method and yolov5 model inference method are used as needed according to different models to achieve target reasoning detection, such as detecting faces, pedestrians, vehicles and other targets in the image; wherein, the target information in the image includes the target name, target confidence and target coordinate position;
[0022] Secondary reasoning is used to identify target attributes and extract target features based on the results obtained from the first reasoning. For example, faces can be identified by age, gender, whether they are wearing masks, glasses, hats, and other attributes, and face feature vectors can be obtained to facilitate face comparison, image search, and other functions. In addition to facial attributes, pedestrians can also be identified by clothing, hairstyle and other attributes. Vehicle attributes include vehicle color, model, style, license plate information, etc. Secondary reasoning is not a mandatory option. When the model only needs to detect no features or attribute analysis, secondary reasoning can be omitted.
[0023] Preferably, the inference results are integrated into a custom protocol format and returned to the http server for visualization as follows:
[0024] Convert the inference model detection results to the corresponding size so that the detection target coordinates correspond to the original image pixel size;
[0025] The detected target information results are integrated into the custom protocol and combined into a message queue and returned to the http server for visual display.
[0026] A Pipeline-based image reasoning system, the system comprising:
[0027] The image source acquisition module is used to obtain the original information of the image, that is, the image information in binary format, based on the given image URI using the REST service, and pass the binary image information to the starting image source plug-in of the Pipeline pipeline;
[0028] The decoding module is used to automatically connect different decoders to the Pipeline to complete image decoding through the decodebin plug-in according to the original information of the image, and transfer the decoded data from the upstream plug-in of the pipeline to the downstream plug-in through the hardware buffer ID; at the same time, videoconvert is used to convert and adapt the corresponding image color space;
[0029] The preprocessing module is used to normalize the decoded images to adapt to the inference model recognition;
[0030] The reasoning module is used to obtain the reasoning result using the corresponding reasoning method according to the loaded reasoning model;
[0031] The post-processing and result return module is used to integrate the inference results into a custom protocol format and return them to the http server for visual display.
[0032] Preferably, the reasoning module includes a primary reasoning submodule and a secondary reasoning submodule;
[0033] Among them, the primary reasoning submodule is used to detect or classify targets in the image, specifically: according to the trained model, the target information in the image is detected or identified from the decoded and normalized image information, and the integrated general yolov4 model reasoning method and yolov5 model reasoning method are used as needed according to different models to achieve the reasoning detection of the target, such as detecting faces, pedestrians, vehicles and other targets in the image; wherein, the target information in the image includes the target name, target confidence and target coordinate position;
[0034] The secondary reasoning submodule is used to identify target attributes and extract target features based on the results obtained from the first reasoning. For example, faces can be identified by age, gender, whether they are wearing masks, glasses, hats, and other attributes, and face feature vectors can be obtained to facilitate face comparison, image search, and other functions. In addition to facial attributes, pedestrians can also be identified by clothing, hairstyle and other attributes. Vehicle attributes include vehicle color, model, style, license plate information, etc. Secondary reasoning is not a must. When the model only needs to detect no features or attribute analysis, secondary reasoning can be omitted.
[0035] An electronic device comprising: a memory and at least one processor;
[0036] Wherein, the memory stores computer-executable instructions;
[0037] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the Pipeline-based image reasoning method as described above.
[0038] A computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the above-mentioned Pipeline-based image reasoning method is implemented.
[0039] The Pipeline-based image reasoning method, system, device, and medium of the present invention have the following advantages:
[0040] (I) The present invention obtains the original binary data of the image through the image URI, passes the original binary data into the pipeline, completes the image decoding and image color space conversion in the pipeline through the decodebin+videocovert plug-in to facilitate the mutual understanding of the upstream and downstream plug-ins to improve performance, and then adopts the yolov series of reasoning methods to perform target detection, target attribute recognition and target feature vector extraction, and integrates the reasoning results into a custom protocol format and returns them to the server for visual display. Each plug-in is independent and isolated from each other in the pipeline, and each plug-in is responsible for its own function and realizes the transmission of data streams through the cache area; the present invention abandons the traditional opencv decoding method, makes full use of gpu hard decoding, and uses the underlying hardware acceleration plug-in when applicable to provide the best performance; compared with the decoding of opencv and other technologies in the prior art, it has more efficient processing capabilities, and the optimization time of a picture is reduced by 50% compared with the traditional technology;
[0041] (ii) The present invention combines different processing function modules into a pipeline by plug-in, realizes image decoding, target detection, target attribute recognition and feature extraction, and makes full use of the GPU hard decoding function. When applicable, the underlying hardware acceleration plug-in is used to improve the efficiency of single image reasoning, which has good promotion and application value;
[0042] (III) The present invention uses a GPU with high-speed computing capabilities and, when applicable, a hardware acceleration plug-in. The decodebin plug-in is used to autonomously load different decoders according to the image format to complete image decoding, abandoning the traditional OpenCV decoding method. The original image data is directly pushed into the Pipeline for decoding, preprocessing, image reasoning, image secondary reasoning, and post-processing. Finally, the result is returned through HTTP, which greatly improves the performance.
[0043] (IV) The image reasoning method based on Pipeline in the present invention solves the problem of slow traditional image parsing. It uses the decodebin plug-in to autonomously load the decoder, target detection and recognition, attribute and feature extraction, and post-processing to realize the process arrangement of image reasoning. It uses hard decoding to improve decoding efficiency and speed up the entire pipeline reasoning process.
[0044] (V) The present invention realizes the process arrangement of image reasoning by combining plug-ins into a pipeline to perform hardware decoding and reasoning of images, and integrating the results into a custom protocol format and returning them to the server for visual display; the present invention uses the underlying hardware acceleration plug-in to improve the reasoning efficiency when applicable, and the entire reasoning process uses different modular plug-ins according to different functions, so that the algorithm service is more user-friendly;
[0045] (VI) The present invention uses decodebin decoding, preprocessing, primary reasoning, secondary reasoning, post-processing and other processes to complete the detection and recognition of image targets, which greatly improves the reasoning efficiency;
[0046] (VII) The present invention obtains the original binary data of the image, uses the decodebin plug-in to load the decoder autonomously, and uses the videoconvert plug-in to convert the image color space to facilitate mutual understanding between upstream and downstream plug-ins. One-time reasoning realizes image target detection, mainly including target name and target coordinate position. Secondary reasoning completes target attribute recognition and feature extraction. The post-processing module integrates the reasoning results into the custom protocol to complete the visualization display of the entire reasoning target result. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The present invention is further described below in conjunction with the accompanying drawings.
[0048] Attached Figure 1 This is a flowchart of the Pipeline-based image reasoning method. DETAILED DESCRIPTION
[0049] The Pipeline-based image reasoning method, system, device and medium of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments of the specification.
[0050] Embodiment 1:
[0051] As attached Figure 1 The present embodiment provides a Pipeline-based image reasoning method, which is specifically as follows:
[0052] S1. Get the image source and convert it into the original image binary format;
[0053] S2. According to the acquired original data source format, the decodebin plug-in is used to dynamically load the available decoders and demultiplexers, and the hardware accelerated decoding plug-in is given priority to complete the decoding of the image to be processed;
[0054] S3, normalizing the decoded image;
[0055] S4. Obtain inference results using the corresponding inference method according to the loaded inference model;
[0056] S5, integrating the inference results into a custom protocol format and returning it to the http server for visual display. The steps of obtaining the image source and converting it into the original image binary format in step S1 of this embodiment are as follows:
[0057] S101, according to the given picture URI, using the REST service to obtain the original information of the picture, that is, the picture information in binary format;
[0058] S102: The image information in binary format is passed to the starting image source plug-in of the Pipeline pipeline.
[0059] In step S2 of this embodiment, the decodebin plug-in is used to dynamically load the available decoder and demultiplexer according to the acquired original data source format, and the hardware accelerated decoding plug-in is preferentially selected to obtain the image to be processed, and the decoding of the image to be processed is completed as follows:
[0060] S201, using the decodebin plug-in to automatically connect different decoders to the pipeline according to the original information of the image to complete image decoding;
[0061] S202, transferring the decoded data from the Pipeline upstream plug-in to the downstream plug-in through the hardware buffer ID;
[0062] S203. At the same time, videoconvert is used to convert and adapt the color space of the corresponding image. Since its upstream and downstream plug-ins already understand each other, when the image color space conversion is not required, it runs in a pass-through mode to reduce the impact on performance.
[0063] The reasoning model in step S4 of this embodiment integrates the YOLO series target detection method, target classification algorithm and feature extraction method.
[0064] In step S4 of this embodiment, obtaining the inference result by using the corresponding inference method according to the loaded inference model includes two stages: primary inference and secondary inference;
[0065] Among them, one inference is used to detect or classify targets in the image, specifically: according to the trained model, the target information in the image is detected or identified from the decoded and normalized image information, and the integrated general yolov4 model inference method and yolov5 model inference method are used as needed according to different models to achieve target reasoning detection, such as detecting faces, pedestrians, vehicles and other targets in the image; wherein, the target information in the image includes the target name, target confidence and target coordinate position;
[0066] Secondary reasoning is used to identify target attributes and extract target features based on the results obtained from the first reasoning. For example, faces can be identified by age, gender, whether they are wearing masks, glasses, hats, and other attributes, and face feature vectors can be obtained to facilitate face comparison, image search, and other functions. In addition to facial attributes, pedestrians can also be identified by clothing, hairstyle and other attributes. Vehicle attributes include vehicle color, model, style, license plate information, etc. Secondary reasoning is not a mandatory option. When the model only needs to detect no features or attribute analysis, secondary reasoning can be omitted.
[0067] In this embodiment, the inference results are integrated into a custom protocol format and returned to the http server for visual display as follows:
[0068] ① Convert the detection results of the inference model to the corresponding size so that the detection target coordinates correspond to the pixel size of the original image;
[0069] ②Integrate the detected target information results into a custom protocol, combine them into a message queue and return them to the http server for visual display.
[0070] Embodiment 2:
[0071] This embodiment provides a Pipeline-based image reasoning system, which includes:
[0072] The image source acquisition module is used to obtain the original information of the image, that is, the image information in binary format, based on the given image URI using the REST service, and pass the binary image information to the starting image source plug-in of the Pipeline pipeline;
[0073] The decoding module is used to automatically connect different decoders to the Pipeline to complete image decoding through the decodebin plug-in according to the original information of the image, and transfer the decoded data from the upstream plug-in of the pipeline to the downstream plug-in through the hardware buffer ID; at the same time, videoconvert is used to convert and adapt the corresponding image color space;
[0074] The preprocessing module is used to normalize the decoded images to adapt to the inference model recognition;
[0075] The reasoning module is used to obtain the reasoning result using the corresponding reasoning method according to the loaded reasoning model;
[0076] The post-processing and result return module is used to integrate the inference results into a custom protocol format and return them to the http server for visual display.
[0077] The reasoning module in this embodiment includes a primary reasoning submodule and a secondary reasoning submodule;
[0078] Among them, the primary reasoning submodule is used to detect or classify targets in the image, specifically: according to the trained model, the target information in the image is detected or identified from the decoded and normalized image information, and the integrated general yolov4 model reasoning method and yolov5 model reasoning method are used as needed according to different models to achieve the reasoning detection of the target, such as detecting faces, pedestrians, vehicles and other targets in the image; wherein, the target information in the image includes the target name, target confidence and target coordinate position;
[0079] The secondary reasoning submodule is used to identify target attributes and extract target features based on the results obtained from the first reasoning. For example, faces can be identified by age, gender, whether they are wearing masks, glasses, hats, and other attributes, and face feature vectors can be obtained to facilitate face comparison, image search, and other functions. In addition to facial attributes, pedestrians can also be identified by clothing, hairstyle and other attributes. Vehicle attributes include vehicle color, model, style, license plate information, etc. Secondary reasoning is not a must. When the model only needs to detect no features or attribute analysis, secondary reasoning can be omitted.
[0080] This embodiment makes the algorithm service lightweight through plug-in module connection, which is convenient for process orchestration. At the same time, when applicable, the overall reasoning efficiency is improved by adopting the underlying hardware acceleration plug-in and hardware decoding method.
[0081] Embodiment 3:
[0082] This embodiment also provides an electronic device, including: a memory and at least one processor;
[0083] Wherein, the memory stores computer-executable instructions;
[0084] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the Pipeline-based image reasoning method described in any one of the present inventions.
[0085] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.
[0086] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.
[0087] Embodiment 4:
[0088] This embodiment also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are loaded by a processor to enable the processor to execute the Pipeline-based image reasoning method in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0089] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0090] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.
[0091] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0092] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A Pipeline-based image reasoning method, characterized in that: The method is as follows: Get the image source and convert it into the original image binary format; According to the acquired original data source format, the decodebin plug-in is used to dynamically load the available decoders and demultiplexers, and the hardware accelerated decoding plug-in is given priority to complete the decoding of the image to be processed; Normalize the decoded image; Obtain the inference result using the corresponding inference method according to the loaded inference model; The inference results are integrated into a custom protocol format and returned to the http server for visual display.
2. The Pipeline-based image reasoning method according to claim 1, characterized in that: Get the image source and convert it into the original image binary format as follows: According to the given picture URI, use the REST service to obtain the original information of the picture, that is, the picture information in binary format; Pass the binary image information to the starting image source plugin of the Pipeline pipeline.
3. The Pipeline-based image reasoning method according to claim 1 or 2, characterized in that: According to the obtained original data source format, the decodebin plug-in is used to dynamically load the available decoders and demultiplexers, and the hardware accelerated decoding plug-in is given priority to obtain the image to be processed and complete the decoding of the image to be processed as follows: The decodebin plug-in automatically connects different decoders to the pipeline to complete image decoding based on the original image information; Pass the decoded data from the upstream Pipeline plugin to the downstream plugin via the hardware buffer ID; At the same time, videoconvert is used to convert and adapt the corresponding image color space. When image color space conversion is not required, it runs in pass-through mode.
4. The Pipeline-based image reasoning method according to claim 3, characterized in that: The inference model integrates the YOLO series target detection method, target classification algorithm and feature extraction method.
5. The Pipeline-based image reasoning method according to claim 4, characterized in that: Obtaining the inference results using the corresponding inference method according to the loaded inference model includes two stages: primary inference and secondary inference; Among them, one inference is used to detect or classify the target in the image, specifically: according to the trained model, the target information in the image is detected or identified from the decoded and normalized image information, and the integrated general yolov4 model inference method and yolov5 model inference method are used as needed according to different models to realize the inference detection of the target; among them, the target information in the image includes the target name, target confidence and target coordinate position; Secondary reasoning is used to identify target attributes and extract target features based on the results obtained from primary reasoning.
6. The Pipeline-based image reasoning method according to claim 5, characterized in that: The inference results are integrated into a custom protocol format and returned to the http server for visualization as follows: Convert the inference model detection results to the corresponding size so that the detection target coordinates correspond to the original image pixel size; The detected target information results are integrated into the custom protocol and combined into a message queue and returned to the http server for visual display.
7. A Pipeline-based image reasoning system, characterized in that: The system includes: The image source acquisition module is used to obtain the original information of the image, that is, the image information in binary format, based on the given image URI using the REST service, and pass the binary image information to the starting image source plug-in of the Pipeline pipeline; The decoding module is used to automatically connect different decoders to the Pipeline to complete image decoding through the decodebin plug-in according to the original information of the image, and transfer the decoded data from the upstream plug-in of the pipeline to the downstream plug-in through the hardware buffer ID; at the same time, videoconvert is used to convert and adapt the corresponding image color space; The preprocessing module is used to normalize the decoded images to adapt to the inference model recognition; The reasoning module is used to obtain the reasoning result using the corresponding reasoning method according to the loaded reasoning model; The post-processing and result return module is used to integrate the inference results into a custom protocol format and return them to the http server for visual display.
8. The Pipeline-based image reasoning system according to claim 7, characterized in that: The reasoning module includes a primary reasoning submodule and a secondary reasoning submodule; Among them, the primary reasoning submodule is used to detect or classify targets in the image, specifically: according to the trained model, the target information in the image is detected or identified from the decoded and normalized image information, and the integrated general yolov4 model reasoning method and yolov5 model reasoning method are used as needed according to different models to achieve the reasoning detection of the target; wherein, the target information in the image includes the target name, target confidence and target coordinate position; The secondary reasoning submodule is used to identify target attributes and extract target features based on the results obtained from the primary reasoning.
9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the Pipeline-based image reasoning method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the Pipeline-based image reasoning method as described in any one of claims 1 to 6 is implemented.