Image processing method and device, electronic equipment and storage medium

By using the prediction results output from the pre-trained target recognition model and the external polygon box in image processing, road marking elements are generated, which solves the problem of low image sample data processing accuracy in the prior art, and improves the training accuracy of the image recognition model on the autonomous driving vehicle side.

CN120236256APending Publication Date: 2025-07-01MUSHROOM CHELIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510356866.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the image sample data processing accuracy of the image recognition model used to train the autonomous driving vehicle end is low, which affects the model training process.

Method used

By obtaining the image to be processed, inputting a pre-trained target recognition model, adding an external polygon box based on the output prediction results, and generating road marking elements based on the positioning information in the box and the prediction results to improve the recognition accuracy.

Benefits of technology

By adding positioning information of external polygon boxes and prediction results in the image processing, the location of road marking elements in the image to be processed can be accurately displayed, thereby improving the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236256A_ABST
    Figure CN120236256A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining a to-be-processed image; inputting the to-be-processed image into a pre-trained target recognition model; adding a corresponding external polygon frame according to a prediction result output by the target identification model; and according to the circumscribed polygon frame and the positioning information of the target in the prediction result, generating a road marking element in the to-be-processed image. According to the invention, the image post-processing process is improved, so that the precision of the image sample used for training is higher, and the accuracy of the trained model is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of image processing and image post-processing, and particularly to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Image post-processing refers to further processing and optimizing the acquired image during the image processing to improve the image quality, enhance its visual effect, or meet specific application requirements.

[0003] In the field of autonomous driving, the post-processing of road signs is beneficial to the training of the in-vehicle image recognition model. However, the current processing accuracy of the image sample data for training is low, thus affecting the subsequent model training process. Summary of the Invention

[0004] Embodiments of this application provide an image processing method, apparatus, electronic device, and storage medium to improve the recognition accuracy.

[0005] Embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, embodiments of this application provide an image processing method, where the processing method includes:

[0007] Obtain an image to be processed;

[0008] Input the image to be processed into a pre-trained target recognition model;

[0009] Add a corresponding circumscribed polygon frame according to the prediction result output by the target recognition model;

[0010] Generate road sign elements in the image to be processed according to the circumscribed polygon frame and the positioning information of the target in the prediction result.

[0011] In some embodiments, the adding a corresponding circumscribed polygon frame according to the prediction result output by the target recognition model includes:

[0012] Add a corresponding circumscribed rectangle frame to each target according to the multiple targets in the prediction result output by the target recognition model;

[0013] If the circumscribed rectangle frame meets the requirements, save it in the image to be processed;

[0014] If the circumscribed rectangle frame does not meet the requirements, delete it or re-obtain the prediction result.

[0015] In some embodiments, the target recognition model includes:

[0016] An encoder, which is used to perform downsampling operations;

[0017] A decoder, which performs upsampling operations;

[0018] A concat layer, which concatenates multiple tensors in corresponding dimensions; and

[0019] An average loss module is used to constrain the position, a multi-loss module is used to optimize the convergence effect, and corresponding loss values loss are output respectively according to the feature map information of the network.

[0020] In some embodiments, the target recognition model includes:

[0021] A network structure including five layers is adopted and used as the encoder and the decoder respectively;

[0022] The loss value loss of the input part of the encoder and the output part of the decoder is compared as the first loss value loss;

[0023] The difference results between the loss values loss in the second layer, the third layer, the fourth layer and the fifth layer of the network and the loss value loss of the label are compared respectively as the second loss value loss, the third loss value loss, the fourth loss value loss, and the fifth loss value loss;

[0024] The first loss value loss, the second loss value loss, the third loss value loss, the fourth loss value loss and the fifth loss value loss are fused to obtain a fused loss value loss, and then the final loss value loss is obtained through backpropagation optimization.

[0025] In some embodiments, the method further includes:

[0026] A loss value loss is obtained by comparing the feature differences between the encoder and the decoder in the network;

[0027] For each layer, a loss function is used to constrain and optimize the loss value loss to obtain the prediction result output by the target recognition model.

[0028] In some embodiments, generating road marking elements in the to-be-processed image according to the circumscribed polygon box and the positioning information of the target in the prediction result includes:

[0029] According to the circumscribed polygon box, the pixel coordinate information of the four vertices is determined;

[0030] According to the pixel coordinate information of the four vertices, calculate the adjacent distances between the four coordinate points in sequence, and use the two sides with the closest adjacent distance as the start and end of the arrow to obtain the arrow direction;

[0031] Display the arrow direction and the positioning information of the target in the prediction result in the map data of the third-party platform to generate a road sign element.

[0032] In some embodiments, the method further includes:

[0033] Combine the road sign image data and the label data corresponding to the image data to generate sample data;

[0034] Randomly divide the sample data into test data and training data, and save the test data and the training data into the mdb database respectively;

[0035] Read the data in the mdb database and parse it into a target matrix to input into the network for training to obtain the trained target recognition model.

[0036] In a second aspect, an embodiment of the present application further provides an image processing device, where the device includes:

[0037] An acquisition module, configured to acquire an image to be processed;

[0038] An input module, configured to input the image to be processed into a pre-trained target recognition model;

[0039] An adding module, configured to add a corresponding circumscribed polygon frame according to the prediction result output by the target recognition model;

[0040] A generating module, configured to generate a road sign element in the image to be processed according to the circumscribed polygon frame and the positioning information of the target in the prediction result.

[0041] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, where the executable instructions, when executed, cause the processor to execute the above method.

[0042] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute the above method.

[0043] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects: obtaining an image to be processed; inputting the image to be processed into a pre-trained target recognition model; adding a corresponding circumscribed polygon frame according to the prediction result output by the target recognition model; generating road marking elements in the image to be processed according to the circumscribed polygon frame and the positioning information of the target in the prediction result. By the above method, adding a circumscribed polygon frame on the basis of the prediction result can accurately display the position where the road marking elements are generated in the image to be processed. Description of the Drawings

[0044] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0045] Figure 1 It is a schematic flowchart of the image processing method in the embodiments of the present application;

[0046] Figure 2 It is a schematic structural diagram of the target recognition model of the image processing method in the embodiments of the present application;

[0047] Figure 3 It is a schematic structural diagram of the image processing device in the embodiments of the present application;

[0048] Figure 4 It is a schematic diagram of the effect of the image processing method in the embodiments of the present application;

[0049] Figure 5 It is a schematic structural diagram of an electronic device in the embodiments of the present application. Detailed Description of the Embodiments

[0050] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments and corresponding drawings of the present application. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0051] The following will describe in detail the technical solutions provided by the embodiments of the present application in conjunction with the drawings.

[0052] The embodiments of the present application provide an image processing method. As Figure 1 shown, a schematic flowchart of the image processing method in the embodiments of the present application is provided. The method at least includes the following steps S110 to step S140:

[0053] Step S110, obtaining an image to be processed.

[0054] The image to be processed includes at least an object to be classified.

[0055] Step S120: Input the image to be processed into a pre-trained object recognition model.

[0056] After relevant preprocessing, the image to be processed is input into a pre-trained object recognition model, and the output result.

[0057] Step S130: Add a corresponding circumscribed polygon box according to the prediction result output by the object recognition model.

[0058] Add a circumscribed polygon box at the corresponding position in the prediction result output by the object recognition model.

[0059] Step S140: Generate a road sign element in the image to be processed according to the circumscribed polygon box and the positioning information of the object in the prediction result.

[0060] Specifically, the added circumscribed polygon box and the pixel coordinate information of the object in the prediction result can be used as positioning information, and a road sign element can be generated and displayed on a relevant platform, so as to be used for subsequent classification tasks.

[0061] In an embodiment of the present application, adding a corresponding circumscribed polygon box according to the prediction result output by the object recognition model includes: adding a corresponding circumscribed rectangle box to each of multiple objects in the prediction result output by the object recognition model; if the circumscribed rectangle box meets the requirements, save it in the image to be processed; if the circumscribed rectangle box does not meet the requirements, delete it or re-obtain the prediction result.

[0062] If multiple objects are included in the prediction result output by the object recognition model, a corresponding circumscribed rectangle box needs to be added to each object. At the same time, check whether the circumscribed rectangle box meets the requirements, save the circumscribed rectangle box that meets the requirements in the image to be processed, and those that do not meet the requirements do not need to be saved and need to delete or re-obtain the prediction result in the object recognition model.

[0063] It should be noted that meeting the requirements means whether the size of the circumscribed rectangle box fits and whether the direction is consistent with the arrow direction, etc.

[0064] In one embodiment of the present application, the target recognition model includes: an encoder for performing downsampling operations; a decoder for performing upsampling operations; a concat layer for concatenating multiple tensors in corresponding dimensions; and an average loss module for constraining positions, a multi-loss module for optimizing the convergence effect, and respectively outputting corresponding loss values loss according to the feature map information of the network.

[0065] As Figure 2 shown, it mainly includes: encoder: encoder, decoder: decoder, concat: concatenation, input: input, output: output.

[0066] Specifically, the encoder mainly performs downsampling operations. In image processing, downsampling usually refers to reducing the resolution or size of an image. By calculating the statistical quantity of pixel values within each small block to represent the pixel values of the entire small block, the image is thus shrunk. Specifically, the decoder mainly performs upsampling operations, which mainly convert a low-resolution image into a high-resolution image.

[0067] Preferably, the average loss module in Figure 2 is used to constrain positions, the multi-loss module is used to optimize the convergence effect, and at the same time, corresponding loss values loss are respectively output according to the feature map information of the network.

[0068] In addition, both the average loss module and the multi-loss module take the feature map information of the network as input. For example, the average loss module uses operations including but not limited to mse, convolutional feature extraction, pooling, etc. The multi-loss module uses operations including but not limited to cross-entropy, convolutional feature extraction, weight weighting, etc.

[0069] In an embodiment of the present application, the target recognition model includes: adopting a network structure with five layers, and serving as an encoder and a decoder respectively; comparing the loss value between the input part of the encoder and the output part of the decoder as the first loss value; respectively comparing the difference results between the loss values in the second, third, fourth, and fifth layers of the network and the loss value of the label as the second loss value, the third loss value, the fourth loss value, and the fifth loss value; fusing the first loss value, the second loss value, the third loss value, the fourth loss value, and the fifth loss value to obtain a fused loss value, and then optimizing through backpropagation to obtain the final loss value.

[0070] As Figure 2 shown, first, a basic network structure with five layers is established, divided into an encoder part and a decoder part. A connection relationship is established through a concat splicing layer between the encoder and the decoder. Then, the loss between the input part of the encoder and the output part of the decoder is compared (one loss). Further, the loss differences between the second, third, fourth, and fifth layers of the network and the label are compared (four losses). The five obtained losses are fused and optimized through backpropagation. Finally, the prediction result of the network is obtained, and the minimum bounding rectangle is used on the prediction result. Then, the four vertices of the bounding rectangle are obtained, so as to determine the direction of the arrow in the image. Preferably, after obtaining the arrow direction and the positioning information of the image, it is displayed on the map of QGIS, so as to generate the ground arrow element required by the production group. That is to say, the pixel coordinates of four points are obtained using the minimum bounding matrix, and the physical coordinates of these four points can be obtained by combining the coordinates of these four points with gps positioning, and these four physical coordinates are then used for production.

[0071] In an embodiment of the present application, the method further includes: obtaining a loss value by comparing the feature differences between the encoder and the decoder in the network; using a loss function to constrain and optimize the loss value for each layer to obtain the prediction result output by the target recognition model.

[0072] Obtaining a loss value by comparing the feature differences between the encoder and the decoder in the network, and then using a loss function to constrain and optimize the loss value for each layer to obtain the prediction result output by the target recognition model.

[0073] In one embodiment of the present application, generating a road sign element in the image to be processed according to the circumscribed polygon frame and the positioning information of the target in the prediction result includes: determining the pixel coordinate information of the four vertices according to the circumscribed polygon frame; calculating the adjacent distances between the four coordinate points in sequence according to the pixel coordinate information of the four vertices, and using the two sides with the closest adjacent distances as the start and end of the arrow to obtain the arrow direction; displaying the arrow direction and the positioning information of the target in the prediction result in the map data of a third-party platform to generate a road sign element.

[0074] The minimum circumscribed rectangle is used to obtain the four vertices, and then the adjacent distances between the four coordinate points are calculated in sequence according to the pixel coordinate information of the four vertices, and the two sides with the closest adjacent distances are used as the start and end of the arrow to obtain the arrow direction.

[0075] The arrow direction in the result obtained through the added circumscribed polygon and the positioning information of the target in the prediction result are displayed in the map data of a third-party platform (such as QGIS) to generate a road sign element. As Figure 4 shown, the left is the original image, the middle is the display of the recognition result on the original image, and the right is the display of the recognition result + the minimum circumscribed rectangle frame on the original image.

[0076] In one embodiment of the present application, the method further includes: combining the road sign image data and the label data corresponding to the image data to generate sample data; randomly grouping the sample data into test data and training data, and separately saving the test data and the training data into an mdb database; reading the data in the mdb database and parsing it into a target matrix to be input into the network for training to obtain the trained target recognition model.

[0077] For the training stage of the target recognition model, the obtained road surface arrow image data and json data are combined to generate the required sample data. At the same time, the data that does not meet the specifications is processed to obtain data that meets the specifications.

[0078] Then, the obtained data is randomly grouped into test data and training data, and these two parts of data are separately saved into the mdb database using a program. Using the mdb database method will be more efficient than directly reading the image data. The mdb database is mainly used to create and manage small database systems. The MDB file contains various database objects such as data tables, queries, forms, reports, macros, and modules.

[0079] Then, the mdb data is read and parsed into a matrix of 1080*1920*3 and input into the network for training to obtain the trained model.

[0080] Finally, for the verification stage of the target recognition model, the obtained training model is used for prediction, and the prediction results are compared with the true image labels. It is found through comparison that the recognition accuracy has been significantly improved, and tests are carried out on a larger range of actual data, and the detection results have all passed the tests.

[0081] The embodiment of the present application also provides an image processing device 300, as Figure 3 shown, which provides a schematic structural diagram of the image processing device in the embodiment of the present application. The image processing device 300 at least includes: an acquisition module 310, an input module 320, an addition module 330, and a generation module 340, where:

[0082] In an embodiment of the present application, the acquisition module 310 is specifically configured to: acquire an image to be processed.

[0083] The image to be processed at least includes an object to be classified.

[0084] In an embodiment of the present application, the input module 320 is specifically configured to: input the image to be processed into a pre-trained target recognition model.

[0085] After relevant preprocessing, the image to be processed is input into a pre-trained target recognition model, and the output result.

[0086] In an embodiment of the present application, the addition module 330 is specifically configured to: add a corresponding circumscribed polygon frame according to the prediction result output by the target recognition model.

[0087] Add a circumscribed polygon frame at the corresponding position in the prediction result output by the target recognition model.

[0088] In an embodiment of the present application, the generation module 340 is specifically configured to: generate a road sign element in the image to be processed according to the circumscribed polygon frame and the positioning information of the target in the prediction result.

[0089] According to the added circumscribed polygon frame and the pixel coordinate information of the target in the prediction result, a road sign element can be generated and displayed on a relevant platform, so that it can be used for subsequent classification tasks.

[0090] In an embodiment of the present application, the addition module 330 is further configured to:

[0091] According to multiple targets in the prediction result output by the target recognition model, add a corresponding circumscribed rectangle frame to each target;

[0092] If the circumscribed rectangle frame meets the requirements, save it in the image to be processed;

[0093] If the external rectangular box does not meet the requirements, delete or re-obtain the prediction result.

[0094] In an embodiment of the present application, the target recognition model includes:

[0095] An encoder, which is used to perform downsampling operations;

[0096] A decoder, which performs upsampling operations;

[0097] A concat layer, which concatenates multiple tensors in corresponding dimensions; and

[0098] An averageloss module is used to constrain the position, a multi-loss module is used to optimize the convergence effect, and corresponding loss values loss are output respectively according to the feature map information of the network.

[0099] In an embodiment of the present application, the target recognition model includes:

[0100] A network structure including five layers is adopted and used as the encoder and decoder respectively;

[0101] The loss value loss of the input part of the encoder and the output part of the decoder is compared as the first loss value loss;

[0102] The difference results between the loss values loss of the network in the second layer, the third layer, the fourth layer and the fifth layer and the loss value loss of the label are respectively compared as the second loss value loss, the third loss value loss, the fourth loss value loss, and the fifth loss value loss;

[0103] The first loss value loss, the second loss value loss, the third loss value loss, the fourth loss value loss and the fifth loss value loss are fused to obtain a fused loss value loss, and then the final loss value loss is optimized through backpropagation.

[0104] In an embodiment of the present application, it further includes: a comparison module, which is used for:

[0105] Obtaining a loss value loss by comparing the feature differences between the encoder and the decoder in the network;

[0106] For each layer, a loss function is used to constrain and optimize the loss value loss to obtain the prediction result output by the target recognition model.

[0107] In an embodiment of the present application, the generation module 340 is further used for:

[0108] Determine the pixel coordinate information of the four vertices according to the circumscribed polygon frame;

[0109] According to the pixel coordinate information of the four vertices, calculate the adjacent distances between the four coordinate points in sequence, and use the two sides with the closest adjacent distance as the start and end of the arrow to obtain the arrow direction;

[0110] Display the arrow direction and the positioning information of the target in the prediction result in the map data of the third-party platform to generate a road sign element.

[0111] In an embodiment of the present application, the method further includes:

[0112] Combine the road sign image data and the label data corresponding to the image data to generate sample data;

[0113] Randomly group the sample data into test data and training data, and save the test data and the training data into the mdb database respectively;

[0114] Read the data in the mdb database and parse it into a target matrix and input it into the network for training to obtain the trained target recognition model.

[0115] It can be understood that the above image processing device can implement each step of the image processing method provided in the foregoing embodiment. The relevant explanations about the image processing method are applicable to the image processing device and will not be elaborated here.

[0116] Figure 5 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 5 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0117] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a bidirectional arrow is used in

[0118] Memory, which is used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0119] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming an image processing device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0120] Obtain the image to be processed;

[0121] Input the image to be processed into a pre-trained target recognition model;

[0122] Add a corresponding circumscribed polygon box according to the prediction result output by the target recognition model;

[0123] Generate road marking elements in the image to be processed according to the circumscribed polygon box and the positioning information of the target in the prediction result.

[0124] The above as in this application Figure 1The method executed by the image processing apparatus disclosed in the illustrated embodiment can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in hardware in the processor or by instructions in software form. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0125] The electronic device can also execute Figure 1 the method executed by the image processing apparatus in Figure 1 the illustrated embodiment and implement the functions of the image processing apparatus in

[0126] Embodiments of the present application also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 1 the method executed by the image processing apparatus in the illustrated embodiment, and specifically used to execute:

[0127] Obtain the image to be processed;

[0128] Input the image to be processed into a pre-trained target recognition model;

[0129] Add a corresponding circumscribed polygon frame according to the prediction result output by the target recognition model;

[0130] Generate road sign elements in the image to be processed according to the circumscribed polygon frame and the positioning information of the target in the prediction result.

[0131] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0133] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0135] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0136] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0137] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0138] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0139] Those skilled in the art will appreciate that the embodiments of the present application may be provided as a method, system, or computer program product. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0140] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An image processing method, wherein: The processing method comprises: Get the image to be processed; Inputting the image to be processed into a pre-trained object recognition model; According to the prediction result output by the target recognition model, a corresponding circumscribed polygonal box is added; A road sign element is generated in the image to be processed according to the circumscribed polygonal frame and the positioning information of the target in the prediction result.

2. The method of claim 1, wherein: The adding of a corresponding circumscribed polygonal frame according to the prediction result output by the target recognition model comprises: According to multiple targets in the prediction results output by the target recognition model, adding a corresponding circumscribed rectangular frame to each target; If the circumscribed rectangular frame meets the requirements, then it is saved in the image to be processed; If the bounding rectangle does not meet the requirements, the prediction result is deleted or re-acquired.

3. The method of claim 2, wherein: The target recognition model comprises: Encoder, used for downsampling operation; Decoder decoder, performs upsampling operation; The concat layer concatenates multiple tensors in corresponding dimensions; and The averageloss module is used to constrain the position, the multi-loss module is used to optimize the convergence effect, and the corresponding loss value loss is output according to the feature map information of the network.

4. The method of claim 3, wherein: The target recognition model comprises: A five-layer network structure is used, which serves as encoder and decoder respectively; Comparing the loss value loss of the encoder input part and the decoder output part as the first loss value loss; Compare the difference between the loss value loss of the network in the second layer, the third layer, the fourth layer and the fifth layer and the loss value loss of the label as the second loss value loss, the third loss value loss, the fourth loss value loss and the fifth loss value loss; The first loss value loss, the second loss value loss, the third loss value loss, the fourth loss value loss and the fifth loss value loss are fused to obtain a fused loss value loss, and then optimized through back propagation to obtain a final loss value loss.

5. The method of claim 4, wherein: The method further comprises: Obtaining a loss value loss by comparing the feature differences between the encoder and the decoder in the network; For each layer, the loss value loss is optimized using a loss function constraint to obtain a prediction result output by the target recognition model.

6. The method of claim 1, wherein: The step of generating a road sign element in the image to be processed according to the circumscribed polygonal frame and the positioning information of the target in the prediction result includes: Determine pixel coordinate information of four vertices according to the circumscribed polygonal frame; According to the pixel coordinate information of the four vertices, the adjacent distances between the four coordinate points are calculated in sequence, and the two sides with the closest adjacent distances are used as the beginning and the end of the arrow to obtain the direction of the arrow; The arrow direction and the positioning information of the target in the prediction result are displayed in the map data of the third-party platform to generate a road sign element.

7. The method according to any one of claims 1 to 6, wherein: The method further comprises: Combining the road sign image data and the label data corresponding to the image data to generate sample data; The sample data is divided into test data and training data by random grouping, and the test data and the training data are respectively saved in the mdb database; The data in the mdb database is read and parsed into a target matrix, which is input into the network for training to obtain the trained target recognition model.

8. An image processing device, wherein: The device comprises: An acquisition module, used for acquiring an image to be processed; An input module, used for inputting the image to be processed into a pre-trained object recognition model; An adding module, used for adding a corresponding circumscribed polygonal box according to the prediction result output by the target recognition model; A generation module is used to generate road sign elements in the image to be processed according to the circumscribed polygonal box and the positioning information of the target in the prediction result.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.