Image processing method and vehicle

By using automated image processing methods to extract target objects from background templates and add descriptive information, the problem of low efficiency in manual operations in existing technologies is solved, achieving efficient and standardized image processing and information integrity.

CN121937549APending Publication Date: 2026-04-28GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREAT WALL MOTOR CO LTD
Filing Date
2026-01-13
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, image processing of automotive parts photos relies on manual operation, resulting in low processing efficiency and unstable accuracy.

Method used

By using automated image processing methods, target objects are extracted from background templates and descriptive information is added, integrating these into an automated process to achieve efficient generation of target images.

Benefits of technology

It improves image processing efficiency, ensures the style consistency and standardization of image synthesis, and enhances the information integrity of the target image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937549A_ABST
    Figure CN121937549A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and a vehicle, and relates to the technical field of image processing. The method comprises the following steps: acquiring at least one image, and determining a corresponding background template; the at least one image at least comprises an overall image of a target object; processing the overall image so as to extract the target object from the overall image to obtain a cutout; obtaining a draft image according to the cutout and the background template; according to the description information corresponding to the draft image, the description information is added to the draft image to obtain the target image, automatic generation of the target image is achieved, manual operation is not needed, and the image processing efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more particularly to an image processing method and a vehicle. Background Technology

[0002] The applications of automotive parts photos span multiple fields, including parts ordering and after-sales service. In the parts ordering process, photos are used for comparison and verification to increase order accuracy. In after-sales service companies, their applications include on-site inspection, repackaging, and other after-sales service.

[0003] However, due to the diverse types and shapes of parts, photos need to be edited to standardize their format before use. Related technologies primarily rely on manual operation, involving taking photos with a camera and then processing them using image editing tools to obtain the desired format. This process is cumbersome, time-consuming, and inefficient. Summary of the Invention

[0004] This application provides an image processing method and a vehicle to solve the problem of low processing efficiency when processing images manually in related technologies.

[0005] In a first aspect, embodiments of this application provide an image processing method, including: Acquire at least one image and determine the corresponding background template; the at least one image must contain at least one complete image of the target object; The entire image is processed to extract the target object from the entire image, resulting in a cutout image; Based on the cutout and background template, a draft image is obtained; Based on the description information corresponding to the draft image, add description information to the draft image to obtain the target image.

[0006] Based on the above technical content, this embodiment, after acquiring at least one image, can determine the corresponding background template based on the image, and then process the overall image to extract the target object from the overall image, obtaining a cutout. Based on the cutout and the background template, a draft image is obtained. Finally, based on the description information corresponding to the draft image, the description information is added to the draft image to obtain the target image. This embodiment integrates multiple image processing workflows, such as determining the background template, extracting the target object, generating the draft image, and adding description information, into an automated process. It has a high degree of automation, achieving target image output without manual operation, thereby effectively improving image processing efficiency. Furthermore, since this application obtains the target image based on a unified background template, it ensures the style consistency and standardization of different image synthesis. Simultaneously, this embodiment also adds corresponding description information, enabling the target image to not only contain visual information but also possess semantic information communication capabilities, improving the information integrity of the target image.

[0007] In one possible implementation, at least one image also contains at least one detailed image of a local region of the target object; Based on the cutout and background template, a draft image is obtained, including: Remove redundant regions from the detail image, except for local regions, to obtain the detail region image; Based on the cutout, detailed area images, and background template, a draft image is obtained.

[0008] In this embodiment, by introducing detailed images of local areas of the target object based on the overall image, the local features of the target object can be better represented, improving the accuracy of the target object presentation. Furthermore, by removing redundant areas from the detailed images, the images of these areas focus on core features, avoiding redundant areas occupying space in the background template or interfering with visual focus. This strengthens the information concentration of the draft image and improves the display effect.

[0009] In one possible implementation, a draft image is obtained based on the cutout, the detailed region image, and the background template, including: Based on the area in the background template used to draw the overall image, determine the position and orientation parameters of the cutout within the background template; Determine the position parameters of the detail area image in the background template based on the area used to draw the detail image in the background template; Based on the position and orientation parameters of the cutout image in the background template, the cutout image is drawn onto the background template. Then, based on the position parameters of the detail area image in the background template, the detail area image is drawn onto the background template to obtain the draft image.

[0010] Here, the cutout image is drawn onto the background template using its position parameters within the template, and the detail area image is drawn onto the background template using its position parameters within the template, achieving precise positioning of both the cutout and the detail area image within the background template. Furthermore, drawing the cutout onto the background template based on orientation parameters corrects the tilt of the target object in the overall image, making the target object appear more three-dimensional.

[0011] In one possible implementation, the position and orientation parameters of the cutout within the background template are determined based on the area in the background template used to draw the overall image, including: Based on the mask corresponding to the overall image, the smallest rectangle surrounding the target object and the rotation angle of the smallest rectangle are obtained, and the rotation angle is used as the direction parameter; wherein, the mask is obtained based on the overall image, and the size of the mask is the same as the size of the overall image. On the mask, the pixel value of the area where the target object is located is a first preset value, and the pixel value of the area where the non-target object is located is a second preset value. The first scaling factor is determined based on the area in the background template used to draw the overall image, as well as the width and height of the smallest rectangle; The position parameters of the cutout in the background template are determined based on the width and height of the smallest rectangle and the first scaling factor.

[0012] Specifically, the first scaling factor is determined by the area in the background template used to draw the overall image, as well as the width and height of the smallest rectangle. The position parameters of the cutout in the background template are then determined based on the first scaling factor, ensuring that the size of the cutout is accurately matched with the area of ​​the overall image to be drawn, and avoiding layout imbalance caused by the cutout size being too large or too small.

[0013] In one possible implementation, the cutout is drawn onto the background template based on its position and orientation parameters within the background template, including: Rotate the cutout and mask according to the orientation parameters; Based on the smallest rectangle, the rotated cutout and mask are cropped to obtain the cropped cutout and mask. Based on the first scaling factor, the cropped cutout and mask are scaled to obtain the scaled cutout and mask. Based on the scaled cutout, the scaled mask, the background template, and the position parameters of the cutout in the background template, the scaled cutout is drawn onto the background template.

[0014] This embodiment of the application corrects the tilt of the target object in the overall image by rotating the cutout and mask using direction parameters, making the target object appear more three-dimensional. Then, the rotated cutout and mask are cropped to remove edge areas unrelated to the target object. Finally, the cropped cutout and mask are scaled using a first scaling factor to ensure that the cutout fits the size of the background template.

[0015] In one possible implementation, redundant regions other than local regions are removed from the detail image to obtain a detail region image, including: In detailed images, determine the location information of local regions; Based on the location information of the local region, the local region is cropped from the detail image to obtain the detail region image.

[0016] Based on the above technical content, by cropping out local areas from detailed images, the resulting detailed area images do not contain redundant areas other than the local areas in the detailed images, thus improving the clarity and recognizability of the local areas and avoiding interference from redundant areas.

[0017] In one possible implementation, determining the location information of local regions in a detail image includes: Classify the detail images to determine their types; If the type of the detail image is a preset type, then the corresponding target detection model is determined according to the preset type; The detailed image is input into the object detection model to obtain the location information of the local region; the object detection model is used to determine the location information of the local region in a detailed image of a preset type.

[0018] In this embodiment, different preset types of detail images correspond to different detection models. By using the target detection model corresponding to the detail image type, the location information of the local area is determined, thereby improving the positioning accuracy of the local area in different types of detail images, that is, improving the accuracy of the location information.

[0019] In one possible implementation, the detail images are classified to obtain the types of detail images, including: Perform text detection on detailed images; If the detail image contains text, the type of the detail image is determined based on the number of characters contained in the text; If the detail image does not contain text, perform image segmentation on the detail image to obtain the image segmentation result, and determine the type of the detail image based on the image segmentation result.

[0020] Among them, multiple methods such as text detection, text recognition, and image segmentation are used in a coordinated manner to determine the type of detailed images. The classification logic is clear and improves the accuracy of determining the type of detailed images.

[0021] In one possible implementation, the background template includes a first background template and a second background template; Determine the corresponding background template, including: If at least one image contains a detailed image of the target object, then the background template is determined to be the first background template; If at least one image does not contain a detailed image of the target object, then the background template is determined to be the second background template.

[0022] Here, based on whether at least one image contains a detailed image of the target object, a corresponding background template is determined, so that the background template is accurately matched with the type of the input image, avoiding waste or insufficiency of background template space and improving the layout rationality of the target image.

[0023] Secondly, embodiments of this application provide an image processing apparatus, including: The processing module is used to acquire at least one image and determine the corresponding background template; the at least one image contains at least one complete image of the target object; The processing module is also used to process the overall image to extract the target object from the overall image and obtain the cutout image; The processing module is also used to obtain a draft image based on the cutout and background template; The processing module is also used to add descriptive information to the draft image based on the descriptive information corresponding to the draft image, so as to obtain the target image.

[0024] Thirdly, embodiments of this application provide a vehicle including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the computer program, it implements the image processing method as described in any of the first aspects.

[0025] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image processing method as described in any of the first aspects.

[0026] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application; Figure 2 This is a schematic flowchart of an image processing method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a first background template provided in an embodiment of this application; Figure 4 This is a schematic diagram of a second background template provided in one embodiment of this application; Figure 5 This is a schematic diagram illustrating the process of obtaining image matting based on the overall image according to an embodiment of this application; Figure 6 This is a schematic flowchart of an image processing method provided in another embodiment of this application; Figure 7 This is a schematic diagram of an image processing procedure provided in an embodiment of this application; Figure 8 This is a schematic diagram of a button diagram provided in an embodiment of this application; Figure 9 This is a schematic diagram of a weaving pattern provided in one embodiment of this application; Figure 10 This is a schematic diagram of an encoding diagram provided in an embodiment of this application; Figure 11 This is a schematic diagram of other types of images provided in one embodiment of this application; Figure 12 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application; Figure 13 This is a schematic diagram of the structure of a vehicle provided in one embodiment of this application. Detailed Implementation

[0030] The present application will be described more clearly below with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the function of the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.

[0031] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0032] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0033] In the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0034] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0035] Furthermore, the term "multiple" mentioned in the embodiments of this application should be interpreted as two or more.

[0036] In related technologies, the main approach relies on manual processing of captured images to obtain the desired result. This process is cumbersome and time-consuming, leading to low efficiency. Furthermore, manual processing depends on the operator's experience and subjective judgment; the accuracy of image processing can vary between different individuals or even the same individual at different times, compromising the final image quality.

[0037] The applicant has discovered that, in order to improve image processing efficiency, it is necessary to consider a new method to automate image processing.

[0038] To improve processing efficiency through automated image processing, this application's implementation determines a corresponding background template after acquiring at least one image. Since at least one image contains at least one overall image of the target object, the overall image can be processed to extract the target object, resulting in a cutout. Based on the cutout and the background template, a draft image is obtained. Adding the corresponding descriptive information to the draft image yields the target image. This embodiment integrates multiple image processing workflows—including determining the background template, extracting the target object, generating the draft image, and adding descriptive information—into a single automated process. This high degree of automation eliminates the need for manual operation, effectively improving image processing efficiency, ensuring image processing accuracy, and ultimately guaranteeing the quality of the target image. Furthermore, since this application obtains the target image based on a unified background template, it ensures stylistic consistency and standardization in the synthesis of different images. Additionally, this embodiment adds corresponding descriptive information, enabling the target image to not only contain visual information but also semantic information, thus enhancing the information integrity of the target image.

[0039] First refer to Figure 1 , Figure 1 The illustration schematically depicts an application scenario provided according to an embodiment of this application. The device involved in this application scenario includes a server. The server can process received overall and detail images to generate a target image, or it can generate a target image based solely on the overall image. The overall image includes at least a frontal image of the target object, and may also include a back image.

[0040] The following is combined with Figure 1 Application scenarios, refer to Figure 2 and Figure 6 This application describes an image processing method provided according to exemplary embodiments. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application can be applied to any applicable scenario.

[0041] It should be noted that the embodiments of this application can be applied to a server or a vehicle host, that is, the image processing method provided by the exemplary embodiments of this application can be executed on the server or the vehicle host.

[0042] The server can be a monolithic server or a distributed server spanning multiple computers or computer data centers. Servers can also be of various categories, such as, but not limited to, web servers, application servers, database servers, or proxy servers.

[0043] Optionally, a server may include hardware, software, or embedded logic components for performing suitable functions supported or implemented by the server, or a combination of two or more such components. For example, a server may be a blade server, a cloud server, or a server group consisting of multiple servers, which may include one or more of the above-mentioned categories of servers, etc.

[0044] It should be noted that the image processing method provided according to the exemplary embodiments of this application can be executed on the same device or on different devices.

[0045] refer to Figure 2 , Figure 2 This is a schematic flowchart of an image processing method provided in an embodiment of this application. Figure 2 As shown, the method in the embodiments of this application may include: Step 201: Obtain at least one image and determine the corresponding background template; the at least one image contains at least one overall image of the target object.

[0046] The target object can be a component or other product; this application does not limit the specific type of the target object.

[0047] At least one image contains at least one overall image of the target object, wherein the overall image can display the overall information of the target object.

[0048] The overall image can be a frontal image or a back image of the target object, and the number of images can be one or more.

[0049] In another implementation scenario, at least one image may also contain a detail image showing a local area of ​​the target object. The detail image is used to display key image information of the details to supplement the fine features of the local area of ​​the target object.

[0050] Background templates are preset image files with specific layouts and styles that can be used to place images of target objects.

[0051] In some embodiments, since the overall image and the detail image describe different features of the target object, and at least one image includes at least one overall image, the corresponding background template can be determined based on whether the at least one image contains a detail image.

[0052] In one possible implementation, the background template includes a first background template and a second background template; determining the corresponding background template includes: if at least one image also contains a detail image of the target object, then the background template is determined to be the first background template; if at least one image does not contain a detail image of the target object, then the background template is determined to be the second background template.

[0053] Figure 3 This is a schematic diagram of a first background template provided in an embodiment of this application, by Figure 3 It can be seen that the first background template includes an area for drawing the overall image, an area for drawing detailed images, and an area for drawing descriptive information. Figure 4 This is a schematic diagram of a second background template provided in an embodiment of this application, by Figure 4 It can be seen that the second background template includes an area for drawing the overall image and an area for drawing descriptive information.

[0054] Here, based on whether at least one image contains a detailed image of the target object, a corresponding background template is determined, so that the background template is accurately matched with the type of the input image, avoiding waste or insufficiency of background template space and improving the layout rationality of the target image.

[0055] Here, each image in at least one image needs to be classified to determine whether at least one image contains a detail image. Optionally, the images can be classified using a multimodal large-scale model that generates text from images. Specifically, the image and prompt words can be used as input to the multimodal large-scale model, and the model can be called to obtain the classification results of the images.

[0056] The prompt can be "This is a photo of a car part. Please judge whether it fully shows the whole part. Answer yes or no." Alternatively, the prompt can be any other description, as long as the multimodal large model can obtain an image that is either a whole image or a detailed image. This application does not limit the specific content of the description.

[0057] In another implementation scenario, determining whether an image is a whole image or a detail image can be done through a constraint parameter interface. Specifically, parameters of the image to be processed can be set, and the whole image and detail image can be transmitted in different fields. For example, the image type can be constrained using JSON format, as shown below: {"Overall Image": ["Overall Image 1.png", "Overall Image 2.png"], "Detail Image": ["Detail Image 1.png", "Detail Image 2.png"]} Here, since the type of each image has been constrained through different fields, that field can indicate the type of the image.

[0058] In another implementation scenario, when determining whether an image is a whole image or a detail image, a sufficient number of whole images and detail images can be collected, and then the whole images and detail images can be labeled to obtain whether they are whole images. Then, the labeled whole images and detail images can be used to train an image classification model, and the trained image classification model can be used to classify whether an image is a whole image.

[0059] Step 202: Process the overall image to extract the target object from the overall image and obtain the cutout image.

[0060] In one implementation scenario, since the overall image usually includes non-target objects in addition to the target object, in order to remove the non-target objects and retain only the target object, an image containing only the target object can be obtained from the overall image through image segmentation, i.e., image matting.

[0061] Image segmentation includes, but is not limited to, semantic segmentation and instance segmentation. Taking semantic segmentation as an example, the entire image is input into an algorithm model used for image segmentation and matting tasks to obtain a mask for the target object. This mask is then used to extract the target object from the original image, resulting in the matted image. Here, a steering wheel is used as an example; the specific process can be found in [reference needed]. Figure 5 As shown, Figure 5 This is a schematic diagram of the process of obtaining image cutout based on the whole image according to an embodiment of this application.

[0062] The algorithm models include, but are not limited to, U-Net and BiRefNet. The mask is an image of the same size as the overall image, where the pixel value of the target object is 1, and the pixel value of non-target object areas is 0.

[0063] In another implementation scenario, if the image is matted through instance segmentation, the basic processing flow is similar to that of semantic segmentation. The difference is that instance segmentation will produce multiple masks, most of which are masks of non-target objects. Therefore, it is necessary to filter the multiple masks, such as selecting the mask that is relatively centered in the image, to select the mask of the target object.

[0064] In another implementation scenario, in addition to image segmentation, a background removal model can be used to obtain the mask of the target object. The background removal model can be a background removal model (RMBG) or other models that can remove the background.

[0065] In some embodiments, after obtaining the cutout, noise reduction algorithms such as Gaussian filtering can be used to process the cutout, thereby reducing noise in the target object area of ​​the cutout and improving the cutout quality.

[0066] In some embodiments, after processing the overall image to obtain the mask of the target object, edge extraction algorithms such as the Canny algorithm can be used to process the mask to obtain the mask of the target object's edge. The edge map of the target object can be obtained from the overall image using the mask of the target object's edge. Then, image processing algorithms such as mean filtering are used to process the edge map, making the edges of the component area smooth and also providing anti-aliasing functionality.

[0067] Step 203: Obtain the draft image based on the cutout and background template.

[0068] In some embodiments, the background template contains an area for drawing the overall image. After obtaining the cutout corresponding to the overall image, the cutout can be drawn onto this area to obtain the corresponding draft image.

[0069] In one implementation scenario, if there is only one image, when drawing the cutout onto that area, the center of the cutout can be set to coincide with the center of the corresponding area, or the top left corner of the cutout can coincide with the top left corner of the area, and so on.

[0070] In another implementation scenario, if there are multiple overall images, the area in the background template used to draw the overall image can be divided into multiple sub-regions, and each sub-region can be used to place one overall image.

[0071] In another implementation scenario, if there are multiple overall images, the current overall image can be drawn sequentially based on the previous overall images.

[0072] In another implementation scenario, if at least one image contains a detail image, the detail image also needs to be drawn onto the background template.

[0073] Step 204: Add descriptive information to the draft image based on the descriptive information corresponding to the draft image to obtain the target image.

[0074] The description information describes each image in the draft image. For example, for the cutout in the draft image, since the cutout is obtained based on the overall image, the description information of the cutout can be to show the overall outline of the target object.

[0075] Based on this, since there may be multiple images in the target image, the description information can include not only the text information mentioned above, but also the image number information, so as to achieve accurate description of multiple images.

[0076] In one implementation scenario, if the draft image contains a detail image, the descriptive information corresponding to that detail image can be something like showing local details of the target object.

[0077] In some embodiments, the descriptive information corresponding to the draft image can be determined using constraint text or a large language model.

[0078] The background template contains areas for drawing descriptive information. After obtaining this descriptive information, it can be added to the draft image based on these areas. This addition of descriptive information to the draft image can be achieved using OpenCV.

[0079] In some embodiments, after obtaining the target image, a corresponding watermark can be added to the target image as needed, which can be achieved using a watermarking tool.

[0080] In this embodiment, after acquiring at least one image, a corresponding background template can be determined based on the image. Then, the overall image is processed to extract the target object from the overall image, resulting in a cutout. Based on the cutout and the background template, a draft image is obtained. Finally, based on the description information corresponding to the draft image, the description information is added to the draft image to obtain the target image. This embodiment integrates multiple image processing workflows, such as determining the background template, extracting the target object, generating the draft image, and adding description information, into an automated process. This highly automated process eliminates the need for manual operation, effectively improving image processing efficiency. Furthermore, since this application obtains the target image based on a unified background template, it ensures the consistency and standardization of the style of different image compositions. Simultaneously, this embodiment also adds corresponding description information, enabling the target image to not only contain visual information but also possess semantic information communication capabilities, thus enhancing the information integrity of the target image.

[0081] In addition, when obtaining a draft image based on the cutout and background template, this application embodiment also needs to consider that if at least one image also contains a detail image, the detail image is drawn onto the background template to obtain a draft image containing the overall image and the detail image, so as to more accurately display the features of the target object. Figure 6 This is a schematic flowchart of an image processing method provided in another embodiment of this application. Figure 7 This is a schematic diagram of an image processing process provided in an embodiment of this application, combined with... Figure 6 and Figure 7 As shown, the method includes: Step 601: Obtain at least one image and determine the corresponding background template; the at least one image contains at least one overall image of the target object.

[0082] Step 602: Process the overall image to extract the target object from the overall image to obtain the cutout image.

[0083] Here, the implementation methods for steps 601 and 602 are described in [reference needed]. Figure 2 The relevant descriptions in the embodiments will not be repeated here.

[0084] Step 603: If at least one image also contains at least one detail image of a local area of ​​the target object; remove redundant areas other than the local area from the detail image to obtain a detail area image; obtain a draft image based on the cutout, the detail area image, and the background template.

[0085] By introducing detailed images of local areas of the target object on top of the overall image, the local features of the target object can be better represented, thus improving the accuracy of the target object presentation.

[0086] In addition, due to shooting issues, there are usually a lot of redundant areas in the detailed images. By removing the redundant areas of the detailed images, the detailed images can focus on the core features, avoiding the redundant areas occupying the space of the background template or interfering with the visual focus. This strengthens the information concentration of the draft image and improves the display effect.

[0087] In some embodiments, removing redundant regions other than local regions from a detail image to obtain a detail region image includes: determining the location information of local regions in the detail image; and cropping local regions from the detail image based on the location information of local regions to obtain a detail region image.

[0088] Here, by cropping out local regions from the detail image, the resulting detail region image does not contain redundant regions other than the local region in the detail image, thus improving the clarity and recognizability of the local region and avoiding interference from redundant regions.

[0089] In one implementation scenario, the detailed image region is a rectangular image.

[0090] In a detail image, determining the location information of a local region includes: classifying the detail image to obtain its type; if the detail image type is a preset type, determining the corresponding target detection model based on the preset type; inputting the detail image into the target detection model to obtain the location information of the local region; wherein, the target detection model is used to determine the location information of the local region in a detail image of a preset type.

[0091] Taking vehicle parts as an example, the types of detailed images include, but are not limited to, button diagrams, weave diagrams, coding diagrams, and other types of images. For example, button diagrams can be found here. Figure 8 As shown, Figure 8 This shows the area corresponding to a steering wheel function button, along with the top left corner, width, and height of that area; a diagram can be referenced. Figure 9 As shown, the encoding diagram can be referenced. Figure 10As shown. Here, other types of images are mainly used to show details of other special areas of the component, such as the edges of the component, devices at the ports of pipes or cables, etc. For details, please refer to... Figure 11 As shown, here, due to the different shapes of the devices at the ports of pipes or cables, Figure 11 This is just an example using a circle.

[0092] Since different types of detail images determine local region location information in different ways, it is necessary to determine the type of detail image. Optionally, the detail image can be classified to obtain its type, including: performing text detection on the detail image; if the detail image contains text, determining the type of detail image based on the number of characters contained in the text; if the detail image does not contain text, performing image segmentation on the detail image to obtain the image segmentation result, and determining the type of detail image based on the image segmentation result.

[0093] The process involves text detection in the detail image. If text is present, the number of characters in the text is identified. If the text consists of only one or two characters, the detail image is classified as a key image. If the text contains multiple characters (i.e., a string of characters), the detail image is classified as an encoded image. If no text is detected in the detail image, image segmentation is performed. If the segmentation result does not yield a corresponding mask, the detail image is classified as a woven image. If a corresponding mask can be obtained, the detail image is classified as another type of image.

[0094] Here, multiple methods such as text detection, text recognition, and image segmentation are used in a coordinated manner to determine the type of detailed images. The classification logic is clear and improves the accuracy of determining the type of detailed images.

[0095] In another implementation scenario, when classifying detailed images and determining their types, a multimodal large model for generating text from images can be used. Specifically, the detailed image and its corresponding prompt can be input into the multimodal large model, which can then determine the type of the detailed image.

[0096] At this point, the prompt could be something like, "This is a detailed image of a local area of ​​a car part. Please determine if there is a button area, a knitted area, or a text area in this image. If there is a button, please answer 1; if there is a knitted area, please answer 2; if there is a text area, please answer 3; if there are none, please answer 4." Alternatively, the prompt could be any other relevant description to help the multimodal large model determine the type of the detailed image.

[0097] After determining the type of the detailed image, it can be determined whether that type is a preset type. In one implementation scenario, preset types include, but are not limited to, button images, weaving images, and coding images.

[0098] For keypad images, a target detection model can be trained to detect keypad areas. This requires collecting a sufficient amount of keypad image data, then labeling the data to obtain a tag file containing the locations of keypad areas within the keypad image. For example, a tag file format might look like this: 324,413,387,499,0 215,297,380,463,0 The first two columns represent the horizontal and vertical coordinates of the top left corner of the button area, the third and fourth columns represent the width and height of the button area, and the last column represents the button type.

[0099] It should be noted that the above is only an example of two lines of data in the label file, and this application does not limit the amount of data in the label file.

[0100] After obtaining the label file, the button image and label file can be used to train the object detection model. After training, the trained object detection model can be used to detect the button image and obtain the position information of the button area.

[0101] Here, the object detection model can be YOLO or other models. This application does not specifically limit the type of object detection model.

[0102] For woven patterns, the training method for the corresponding object detection model is basically the same as that for button patterns, so it will not be elaborated here.

[0103] For an encoded image, the corresponding object detection models can include text detection models and text orientation detection models. Specifically, a text detection model can be used to detect the regions containing text in the encoded image, and then a text orientation detection model can be used to obtain the orientation of the text. Based on this orientation, an affine transformation is applied to orthogonalize the text regions. Text detection models include, but are not limited to, PSENet, DBNet, and Paddle, while text orientation detection models can be convolutional neural networks (CNNs).

[0104] It should be noted that the text detection model is also trained, and its training method is similar to that of the key image corresponding to the target detection model. The difference is that the content of the label file is not coordinate information, but the angle between the text direction and the horizontal direction, and the model is a regression model.

[0105] In another implementation scenario, if the detail image is of a different type than the preset type, the method for removing redundant regions is similar to image segmentation of the overall image. The difference is that after obtaining the corresponding mask using image segmentation, it is necessary to calculate the coordinates of the bounding box surrounding the component in the detail image. Then, based on the coordinates of the bounding box, the local area of ​​the target object is cropped from the detail image to obtain the detail region image.

[0106] The bounding box is calculated as follows: obtain the set of pixel coordinates {(x,y)} in the mask where the pixel value is not 0. Since the pixel value of the area where the target object is located in the mask is 1 and the pixel value of the area where the non-target object is located is 0, the set of pixel coordinates with non-zero pixel values ​​extracted here covers the area where the target object is located. This set can represent the bounding box. Therefore, the left edge coordinate of the bounding box is min({x}), the right edge coordinate is max({x}), the top edge coordinate is min({y}), and the bottom edge coordinate is max({y}).

[0107] In this embodiment, different preset types of detail images correspond to different detection models. By using the target detection model corresponding to the detail image type, the location information of the local area is determined, thereby improving the positioning accuracy of the local area in different types of detail images, that is, improving the accuracy of the location information.

[0108] In some embodiments, obtaining a draft image based on the cutout, the detail region image, and the background template includes: determining the position and orientation parameters of the cutout in the background template based on the area in the background template used to draw the overall image; determining the position parameters of the detail region image in the background template based on the area in the background template used to draw the detail image; drawing the cutout onto the background template based on the position and orientation parameters of the cutout, and drawing the detail region image onto the background template based on the position parameters of the detail region image, thereby obtaining the draft image.

[0109] Here, the cutout image is drawn onto the background template using its position parameters within the template, and the detail area image is drawn onto the background template using its position parameters within the template, achieving precise positioning of both the cutout and the detail area image within the background template. Furthermore, drawing the cutout onto the background template based on orientation parameters corrects the tilt of the target object in the overall image, making the target object appear more three-dimensional.

[0110] In order to determine the position parameters of the cutout and detail area images in the background template, it is necessary to determine the corresponding areas in the background template for drawing the overall image and the detail area images. The specific method is to manually set the corresponding areas for each background template. For example, the area for drawing the overall image is {X1,Y1,W1,H1}, and the area for drawing the detail area image is {X2,Y2,W2,H2}.

[0111] Where X1 and Y1 are the horizontal and vertical coordinates of the top left corner of the overall image area, respectively, and W1 and H1 are the width and height of the area. Similarly, X2 and Y2 are the horizontal and vertical coordinates of the top left corner of the detailed image area, respectively, and W2 and H2 are the width and height of the area.

[0112] Since the overall image can be a single image or multiple images, in one implementation scenario, if there is only one overall image, the position and direction parameters of the cutout in the background template are determined based on the area used to draw the overall image in the background template. This includes: obtaining the smallest rectangle surrounding the target object and its rotation angle based on the mask corresponding to the overall image, and using the rotation angle as the direction parameter; wherein, the mask is obtained based on the overall image, and the size of the mask is the same as the size of the overall image; on the mask, the pixel value of the area where the target object is located is a first preset value, and the pixel value of the area where the non-target object is located is a second preset value; a first scaling factor is determined based on the area used to draw the overall image in the background template, and the width and height of the smallest rectangle; and the position parameters of the cutout in the background template are determined based on the width and height of the smallest rectangle, and the first scaling factor.

[0113] In this context, the pixel values ​​of the area where the target object is located on the mask are first preset values, and the pixel values ​​of the area where the non-target object is located are second preset values. Optionally, the first preset value can be 1 and the second preset value can be 0.

[0114] In one implementation scenario, a convex hull rotation caliper algorithm can be applied to the mask corresponding to the overall image to obtain the smallest rectangle surrounding the target object. In this process, the rotation angle of the smallest rectangle can be obtained, which is the direction parameter. When drawing the cutout, this direction parameter can be used to place the target object in a more three-dimensional way.

[0115] Let W and H be the width and height of the smallest rectangle, respectively. Then, based on the width and height of the smallest rectangle and the width and height of the area in the background template used to draw the overall image, calculate the first scaling factor to ensure that the image neither exceeds the area in the background template used to draw the overall image nor is too small.

[0116] Alternatively, one method for calculating the first scaling factor is as follows:

[0117] Where scale represents the first scaling factor, H1 and W1 are the height and width of the region used to draw the overall image in the background template, H and W are the height and width of the minimum rectangle, and factor is a factor used to adjust the size of the overall image in the final target image. The larger the value, the larger the proportion of the overall image in the final image. Usually, factor is greater than 0 and less than 1.

[0118] Based on the first scaling factor, obtain the width w and height h of the smallest rectangle after scaling. The position parameters of the cutout image within the background template can then be determined using the following method: x = W1 / 2 - w / 2 + X1 y = H1 / 2 - h / 2 + Y1 Where x and y are the horizontal and vertical coordinates of the top left corner of the overall image area drawn by the background template, respectively; X1 and Y1 are the horizontal and vertical coordinates of the top left corner of the overall image area drawn by the background template, respectively; W1 and H1 are the width and height of the area, respectively; and w and h are the width and height of the smallest rectangle after scaling, respectively.

[0119] It should be noted that the above calculation of the position parameters of the cutout in the background template is based on the assumption that the center of the cutout coincides with the center of the overall image area drawn by the background template, that is, the cutout is placed at the center of the overall image area drawn by the background template. However, the cutout can also be placed in other ways within the overall image area drawn by the background template, and this application does not limit this. In addition, the above uses the horizontal rightward direction as the positive x-axis and the vertical downward direction as the positive y-axis.

[0120] In another implementation scenario, if there are multiple overall images, before calculating the orientation and position parameters, the area in the background template where the overall images are drawn can be divided into multiple sub-regions based on the number of overall images. Each sub-region can hold one overall image. The process of calculating the position and orientation parameters of any given overall image in a sub-region is the same as that described above when there is only one overall image, and will not be elaborated further here.

[0121] Here, there are multiple ways to divide the image into sub-regions. Taking two whole images as an example, one way to divide the image into sub-regions is to calculate the ratio of the width to the height and the ratio of the height to the width of each whole image. If the minimum ratio is less than 0.5, i.e., min(W / H, H / W) < 0.5, then the area where the background template is used to draw the whole image is divided into two equal parts, upper and lower, as two sub-regions. Otherwise, it is divided into two equal parts, left and right.

[0122] In addition, if the area for drawing the overall image is divided into upper and lower parts, before determining the position and orientation parameters of each cutout in the background template, if the width of the cutout is less than its height, then the cutout is rotated by 90°, and then the position and orientation parameters of the rotated cutout in the background template are calculated.

[0123] In this embodiment, a first scaling factor is determined by the area in the background template used to draw the overall image, as well as the width and height of the smallest rectangle. The position parameters of the cutout in the background template are determined based on the first scaling factor, ensuring that the size of the cutout is accurately matched with the area of ​​the overall image to be drawn, and avoiding layout imbalance caused by the cutout size being too large or too small.

[0124] In this embodiment, when determining the position parameters of the detail region image in the background template based on the region {X2,Y2,W2,H2} used to draw the detail image, since the detail region image is rectangular, its width and height are set to width_i and height_i, respectively, where i represents the i-th detail region image. Before calculating the position parameters of the detail region image in the background template, the scaling factor of each detail region image can be calculated based on the height H2 of the region used to draw the detail region image in the background template: scale_i = H2 / height_i. The position parameters of each detail region image can be determined based on the positions of the previous few detail region images.

[0125] The location parameters of multiple detailed region images can be referenced as follows: Position parameters of the first detailed region image:

[0126] Positional parameters of the second detailed region image:

[0127] The positional parameters of the third detailed region image are:

[0128] If there are other detailed area images subsequently, they can be determined sequentially based on the positions of the previous detailed area images. It should be noted that the position parameters of each detailed area image are represented by the horizontal and vertical coordinates of the upper left corner, the width, and the height. Since the area where the detailed area image is drawn is {X2,Y2,W2,H2}, the above example uses the upper left corner of the first detailed area image coinciding with the upper left corner of the area where the detailed area image is drawn. However, other placement methods are also possible, and this application does not limit them.

[0129] After determining the position and orientation parameters of the cutout image within the background template, the cutout image is drawn onto the background template based on these parameters. This includes: rotating the cutout image and mask according to the orientation parameters; cropping the rotated cutout image and mask according to the minimum rectangle to obtain the cropped cutout image and mask; scaling the cropped cutout image and mask according to the first scaling factor to obtain the scaled cutout image and mask; and drawing the scaled cutout image onto the background template based on the scaled cutout image, the scaled mask, the background template, and the position parameters of the cutout image within the background template.

[0130] When rotating the cutout and mask according to the direction parameters, affine transformation or rigid transformation can be used to rotate the cutout and mask.

[0131] In some embodiments, when scaling the cropped cutout and mask according to the first scaling factor, an interpolation algorithm can be used. The interpolation algorithm can be a bilinear interpolation algorithm or other interpolation algorithms; this application does not limit the type of interpolation algorithm.

[0132] In some embodiments, when drawing the scaled-down cutout onto the background template based on the scaled-down cutout, the scaled-down mask, the background template, and the position parameters of the cutout in the background template, a reference image with the same size as the background template and a pixel value of 0 can be created. Based on the position parameters of the cutout in the background template, the scaled-down mask is drawn onto the reference image to obtain the target mask. Then, the pixel values ​​in the target mask are flipped (pixel values ​​of 1 are converted to 0, and pixel values ​​of 0 are converted to 1). A Hadamard product is then calculated between the target mask and the original background template to obtain a new image. Finally, the scaled-down cutout is added to the corresponding area of ​​this new image as a matrix sum. This corresponding area is the area corresponding to the position parameters of the cutout in the background template. At this point, the scaled-down cutout has been drawn onto the background template.

[0133] It should be noted that, as mentioned above, when determining the position parameters of the cutout in the background template, these position parameters are obtained by scaling the width w and height h of the smallest rectangle after scaling by the first scaling factor. The scaled cutout is also obtained by scaling the cropped cutout according to the first scaling factor. The cropped cutout is obtained by cropping the smallest rectangle from the rotated cutout, meaning that the cropped cutout has the same size as the smallest rectangle. Therefore, in this embodiment, the position parameters of the cutout in the background template are the same as the position parameters of the scaled cutout in the background template. Based on the position parameters of the cutout in the background template, the scaled cutout can be drawn onto the background template.

[0134] In this embodiment, for the detailed area image, the process of drawing the detailed area image onto the background template according to the position parameters of the detailed area image in the background template is similar to the process of drawing the scaled cutout onto the background template as described above. However, since the detailed area image is rectangular, the pixel value of the corresponding area in the background template can be set to 0 directly according to the position parameters of the detailed area image, and then the detailed area image can be added to the corresponding area in the form of matrix summation.

[0135] This embodiment of the application corrects the tilt of the target object in the overall image by rotating the cutout and mask using direction parameters, making the target object appear more three-dimensional. Then, the rotated cutout and mask are cropped to remove edge areas unrelated to the target object. Finally, the cropped cutout and mask are scaled using a first scaling factor to ensure that the cutout fits the size of the background template.

[0136] Step 604: Add descriptive information to the draft image based on the descriptive information corresponding to the draft image to obtain the target image.

[0137] As can be seen from the above embodiments, the descriptive information may include text information and numbering information. Optionally, the descriptive information corresponding to the draft image can be determined using constraint text or a language model. Here, we will take the determination of the descriptive information corresponding to the draft image using constraint text as an example. The text format of the constraint text can be user-defined, such as "1,2 shows the overall outline; 3,4 shows the component code; 5,6,7 shows the local details of the component".

[0138] To do this, each image in the draft image needs to be numbered. If there are two overall images, the corresponding scaled-down images of the overall images are numbered 1 and 2, and the detail area images are numbered 3, 4, 5, and so on. If there is only one overall image, then number 2 is used on the detail area images, and the detail area images are numbered 2, 3, 4, and so on.

[0139] Here, since the encoded information is a relatively important type of information that can better represent the model of the target object, after determining the number of the overall image, the encoded image can be numbered first, and other detailed area images can be numbered in the same way. After determining the number of each image, the corresponding string text can be obtained directly.

[0140] In addition, the placement of the number can be determined based on the position parameters of the cutout image and the position parameters of the detailed area image.

[0141] In one implementation scenario, the number corresponding to the overall image can be located at the top center of the overall image. In this case, the position of the number is calculated according to the position parameters of the cutout as x0=W1 / 2+X1, y0=H1 / 2-h / 2+Y1, where x0 and y0 are the horizontal and vertical coordinates of the number, X1 and Y1 are the horizontal and vertical coordinates of the upper left corner of the area used to draw the overall image, W1 and H1 are the width and height of the area, and h is the height of the smallest rectangle after scaling, that is, the height of the cutout after scaling.

[0142] Alternatively, the number corresponding to the overall image can be located in other positions, such as the upper left corner, the lower center, or the lower right corner of the overall image. This application does not limit this.

[0143] For the detail image, the number can be drawn inside the detail image, that is, inside the detail region image. Its starting coordinates can be the same as the upper left corner coordinates of the detail region image, or it can be located at the upper center of the detail region image or other positions, etc. This application does not limit this.

[0144] The size and format of the overall image and detail image numbering and text information can be set according to actual needs. In addition, the drawing of numbering and text information can be implemented using OpenCV.

[0145] Here, when adding description information, since the background template contains an area {X3,Y3,W3,H3} for drawing description information, where X3 and Y3 are the coordinates of the upper left corner of the area, and W3 and H3 are the width and height of the area, the description information can be drawn in the area based on the location information of the area.

[0146] In some embodiments, the process of generating a target image based on at least one image can also be uniformly implemented using an image generation algorithm based on neural networks and guided by prompt words, which has a high degree of automation.

[0147] In this embodiment, after acquiring at least one image, a corresponding background template can be determined. The at least one image contains at least one overall image of the target object. The overall image is processed to extract the target object from it, resulting in a cutout. If the at least one image also contains at least one detail image of a local area of ​​the target object, redundant areas other than the local area are removed from the detail image to obtain a detail area image. Based on the cutout, the detail area image, and the background template, a corresponding draft image is obtained. Here, by introducing a detail image of a local area of ​​the target object on top of the overall image, the local features of the target object can be better represented, improving the accuracy of the target object presentation. After obtaining the draft image, the corresponding descriptive information is added to the draft image to obtain the target image. This ensures that the target image not only contains visual information but also has semantic information sensing capabilities, improving the information completeness of the target image. Since this embodiment integrates multiple image processing workflows such as determining the background template, extracting the target object, generating the draft image, and adding descriptive information into an automated process, the degree of automation is high. The target image can be output without manual operation, thereby effectively improving image processing efficiency. In addition, since this application obtains the target image based on a unified background template, it ensures the style consistency and standardization of different image synthesis.

[0148] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0149] Figure 12 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application. For example... Figure 12 As shown, the image processing apparatus provided in this embodiment may include: a processing module 1201.

[0150] The processing module 1201 is used to acquire at least one image and determine the corresponding background template; the at least one image contains at least one overall image of the target object; The processing module 1201 is also used to process the overall image to extract the target object from the overall image and obtain the cutout image; The processing module 1201 is also used to obtain a draft image based on the cutout and background template; The processing module 1201 is also used to add descriptive information to the draft image based on the descriptive information corresponding to the draft image, so as to obtain the target image.

[0151] In one possible implementation, at least one image further includes at least one detail image of a local region of the target object; the processing module 1201 is specifically used for: Remove redundant regions from the detail image, except for local regions, to obtain the detail region image; Based on the cutout, detailed area images, and background template, a draft image is obtained.

[0152] In one possible implementation, the processing module 1201 is specifically used for: Based on the area in the background template used to draw the overall image, determine the position and orientation parameters of the cutout within the background template; Determine the position parameters of the detail area image in the background template based on the area used to draw the detail image in the background template; Based on the position and orientation parameters of the cutout image in the background template, the cutout image is drawn onto the background template. Then, based on the position parameters of the detail area image in the background template, the detail area image is drawn onto the background template to obtain the draft image.

[0153] In one possible implementation, the processing module 1201 is specifically used for: Based on the mask corresponding to the overall image, the smallest rectangle surrounding the target object and the rotation angle of the smallest rectangle are obtained, and the rotation angle is used as the direction parameter; wherein, the mask is obtained based on the overall image, and the size of the mask is the same as the size of the overall image. On the mask, the pixel value of the area where the target object is located is a first preset value, and the pixel value of the area where the non-target object is located is a second preset value. The first scaling factor is determined based on the area in the background template used to draw the overall image, as well as the width and height of the smallest rectangle; The position parameters of the cutout in the background template are determined based on the width and height of the smallest rectangle and the first scaling factor.

[0154] In one possible implementation, the processing module 1201 is specifically used for: Rotate the cutout and mask according to the orientation parameters; Based on the smallest rectangle, the rotated cutout and mask are cropped to obtain the cropped cutout and mask. Based on the first scaling factor, the cropped cutout and mask are scaled to obtain the scaled cutout and mask. Based on the scaled cutout, the scaled mask, the background template, and the position parameters of the cutout in the background template, the scaled cutout is drawn onto the background template.

[0155] In one possible implementation, the processing module 1201 is specifically used for: In detailed images, determine the location information of local regions; Based on the location information of the local region, the local region is cropped from the detail image to obtain the detail region image.

[0156] In one possible implementation, the processing module 1201 is specifically used for: Classify the detail images to determine their types; If the type of the detail image is a preset type, then the corresponding target detection model is determined according to the preset type; The detailed image is input into the object detection model to obtain the location information of the local region; the object detection model is used to determine the location information of the local region in a detailed image of a preset type.

[0157] In one possible implementation, the processing module 1201 is specifically used for: Perform text detection on detailed images; If the detail image contains text, the type of the detail image is determined based on the number of characters contained in the text; If the detail image does not contain text, perform image segmentation on the detail image to obtain the image segmentation result, and determine the type of the detail image based on the image segmentation result.

[0158] In one possible implementation, the background template includes a first background template and a second background template; the processing module 1201 is specifically used for: If at least one image contains a detailed image of the target object, then the background template is determined to be the first background template; If at least one image does not contain a detailed image of the target object, then the background template is determined to be the second background template.

[0159] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0160] Figure 13 This is a schematic diagram of the structure of a vehicle provided in one embodiment of this application. Figure 13 As shown, the vehicle 1300 of this embodiment includes a processor 1301 and a memory 1302. The memory 1302 stores a computer program 1303 that can run on the processor 1301. When the processor 1301 executes the computer program 1303, it implements the steps in any of the above-described method embodiments, for example... Figure 2 Steps 201 to 204 are shown. Alternatively, when processor 1301 executes computer program 1303, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 12 The function of module 1201 shown.

[0161] For example, computer program 1303 may be divided into one or more modules / units, one or more of which are stored in memory 1302 and executed by processor 1301 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of computer program 1303 in vehicle 1300.

[0162] Those skilled in the art will understand that Figure 13 This is merely an example of a vehicle and does not constitute a limitation on the vehicle. It may include more or fewer components than shown, or combinations of certain components, or different components, such as input / output devices, network access devices, buses, etc.

[0163] Processor 1301 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0164] The memory 1302 can be an internal storage unit of the vehicle, such as a hard drive or memory, or an external storage device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc. The memory 1302 can also include both internal and external storage devices. The memory 1302 is used to store computer programs and other programs and data required by the vehicle. The memory 1302 can also be used to temporarily store data that has been output or will be output.

[0165] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0166] An embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image processing method.

[0167] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0168] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0169] In the embodiments provided in this application, it should be understood that the disclosed devices / vehicles and methods can be implemented in other ways. For example, the device / vehicle embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0170] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0171] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0172] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0173] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An image processing method, characterized in that, include: Obtain at least one image and determine the corresponding background template; The at least one image contains at least one complete image of the target object; The overall image is processed to extract the target object from the overall image, resulting in a cutout image; Based on the cutout and the background template, a draft image is obtained; Based on the description information corresponding to the draft image, the description information is added to the draft image to obtain the target image.

2. The image processing method according to claim 1, characterized in that, The at least one image also includes at least one detailed image of a local region of the target object; The process of obtaining a draft image based on the cutout and the background template includes: Remove redundant regions from the detailed image, excluding the local region, to obtain the detailed region image; A draft image is obtained based on the cutout, the detailed area image, and the background template.

3. The image processing method according to claim 2, characterized in that, The process of obtaining a draft image based on the cutout, the detailed region image, and the background template includes: Based on the area in the background template used to draw the overall image, determine the position and direction parameters of the cutout in the background template; Based on the area in the background template used for drawing the detail image, determine the position parameters of the detail area image in the background template; Based on the position and direction parameters of the cutout in the background template, the cutout is drawn onto the background template, and based on the position parameters of the detail area image in the background template, the detail area image is drawn onto the background template to obtain a draft image.

4. The image processing method according to claim 3, characterized in that, The step of determining the position and orientation parameters of the cutout image within the background template based on the area used to draw the overall image in the background template includes: Based on the mask corresponding to the overall image, the smallest rectangle surrounding the target object and the rotation angle of the smallest rectangle are obtained, and the rotation angle is used as the direction parameter; wherein, the mask is obtained based on the overall image, and the size of the mask is the same as the size of the overall image; on the mask, the pixel value of the area where the target object is located is a first preset value, and the pixel value of the area where the non-target object is located is a second preset value; The first scaling factor is determined based on the area in the background template used to draw the overall image, and the width and height of the minimum rectangle; The position parameters of the cutout in the background template are determined based on the width and height of the minimum rectangle and the first scaling factor.

5. The image processing method according to claim 4, characterized in that, The step of drawing the cutout onto the background template based on the position and orientation parameters of the cutout in the background template includes: The cutout and the mask are rotated according to the direction parameters; Based on the minimum rectangle, the rotated cutout and mask are cropped to obtain the cropped cutout and mask. Based on the first scaling factor, the cropped cutout and mask are scaled to obtain the scaled cutout and mask. Based on the scaled cutout, the scaled mask, the background template, and the position parameters of the cutout in the background template, the scaled cutout is drawn onto the background template.

6. The image processing method according to any one of claims 2 to 5, characterized in that, The step of removing redundant regions other than the local region from the detail image to obtain a detail region image includes: In the detailed image, determine the location information of the local region; Based on the location information of the local region, the local region is cropped from the detail image to obtain the detail region image.

7. The image processing method according to claim 6, characterized in that, Determining the location information of the local region in the detailed image includes: The detailed images are classified to obtain their types; If the type of the detailed image is a preset type, then the corresponding target detection model is determined according to the preset type; The detailed image is input into the target detection model to obtain the location information of the local region; wherein, the target detection model is used to determine the location information of the local region in the detailed image of the preset type.

8. The image processing method according to claim 7, characterized in that, The process of classifying the detail images to obtain their types includes: Perform text detection on the detailed image; If the detailed image contains text, the type of the detailed image is determined based on the number of characters contained in the text; If the detail image does not contain text, perform image segmentation on the detail image to obtain the image segmentation result, and determine the type of the detail image based on the image segmentation result.

9. The image processing method according to any one of claims 1 to 5, characterized in that, The background template includes a first background template and a second background template; Determining the corresponding background template includes: If the at least one image also contains a detailed image of the target object, then the background template is determined to be the first background template; If at least one image does not contain a detailed image of the target object, then the background template is determined to be the second background template.

10. A vehicle comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the image processing method as described in any one of claims 1 to 9.