Image Processing Method, Apparatus and Medium

By splicing and fusing cabin images from multiple perspectives, and combining computer vision technology to identify and build virtual images, the problem of difficult monitoring of items in concealed sight lines is solved, and more comprehensive cabin perception and monitoring is achieved.

CN114332830BActive Publication Date: 2025-07-18CHINA AUTOMOTIVE INNOVATION CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111535288.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-07-18
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

In the prior art, computer vision in the cabin of the car is difficult to effectively monitor items in concealed areas of sight, resulting in poor perception effects.

Method used

By obtaining multiple original cabin images from different perspectives, stitching and fusion, using computer vision perception technology to identify the target object, and constructing a layer of the target virtual image, and finally superimposing the cabin panoramic image with the target layer to highlight the target virtual image.

Benefits of technology

It significantly improves the perception effect of the cabin, and can identify and highlight objects in hidden sight areas, making it easier for users to monitor the internal situation of the cabin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332830B_ABST
    Figure CN114332830B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, apparatus and medium, relating to the technical field of image processing. The method includes: obtaining a plurality of original cabin images from different perspectives; splicing the plurality of original cabin images to obtain a panoramic cabin image; determining a target object in the plurality of original cabin images based on computer vision perception; determining a target virtual image corresponding to the target object and obtaining a target layer including the target virtual image; and superimposing the panoramic cabin image and the target layer to obtain a target panoramic image including the target virtual image. The present application uses computer vision technology to perceive the target object in the cabin image, and uses the target virtual image to represent the detected target object, so that the synthesized target panoramic image can effectively improve the perception effect of the cabin.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and specifically to an image processing method, apparatus, and medium. Background Art

[0002] At present, computer vision applications in the automotive cockpit have become increasingly widespread, and static cockpit monitoring technologies are also constantly evolving and updating. In related technologies, based on real in-cabin images, remote monitoring of drivers, passengers, in-cabin items, etc. can be carried out, but there are problems in detecting items in hidden areas of sight. Summary of the Invention

[0003] In order to improve the perception effect of the vehicle cabin, this application provides an image processing method, apparatus, and medium. The technical solutions are as follows:

[0004] In a first aspect, this application provides an image processing method, the method includes:

[0005] Obtain multiple original cabin images from different perspectives;

[0006] Stitch the multiple original cabin images to obtain a cabin panoramic image;

[0007] Based on computer vision perception, determine the target object in the multiple original cabin images;

[0008] Determine the target virtual image corresponding to the target object, and obtain a target layer including the target virtual image;

[0009] Overlay the cabin panoramic image and the target layer to obtain a target panoramic image including the target virtual image.

[0010] Optionally, the step of stitching the multiple original cabin images to obtain a cabin panoramic image includes:

[0011] Perform feature point detection on each of the multiple original cabin images to obtain the feature points of each original cabin image and the description information of the feature points;

[0012] Match the feature points of each original cabin image according to the description information of the feature points to obtain feature point pairs;

[0013] Determine the association relationship between each of the original cabin images based on the feature point pairs;

[0014] According to the association relationship between each of the original cabin images, stitch and fuse the multiple original cabin images to obtain the cabin panoramic image.

[0015] Optionally, determining the target object in the multiple original cabin images based on computer vision perception includes:

[0016] Inputting the multiple original cabin images into a trained computer vision perception model respectively for visual perception processing to obtain a perception result, where the perception result indicates the target object in the multiple original cabin images and the attribute information of the target object.

[0017] Further, the attribute information includes the target object category. Determining the target virtual image corresponding to the target object and obtaining the target layer containing the target virtual image includes:

[0018] Determining the target layer type corresponding to the target object category based on a preset mapping relationship; the preset mapping relationship represents the corresponding relationship between the object category and the layer type;

[0019] Determining the target virtual image corresponding to the target object according to the target object category and the target layer type;

[0020] Generating the target layer containing the target virtual image according to the target layer type and the target virtual image.

[0021] Further, the attribute information further includes the first coordinate information of the target object in the camera coordinate system. Obtaining the target layer containing the target virtual image further includes:

[0022] Determining the coordinate conversion relationship between the camera coordinate system and a pre-calibrated cabin coordinate system;

[0023] Obtaining the second coordinate information of the target object in the cabin coordinate system according to the first coordinate information and the coordinate conversion relationship;

[0024] Determining the display area information of the target virtual image in the target layer according to the second coordinate information.

[0025] Optionally, superimposing the cabin panoramic image and the target layer to obtain a target panoramic image containing the target virtual image includes:

[0026] Performing digital image synthesis on the cabin panoramic image and the target layer with the same scale in a preset cabin coordinate system to obtain the target panoramic image in the cabin coordinate system, and representing the target object with the target virtual image in the target panoramic image.

[0027] Optionally, the method further includes:

[0028] Send the target panoramic image to the user terminal so that the target panoramic image can be correspondingly displayed on the user terminal according to the user's perspective;

[0029] And display the target virtual image in a preset display style in the target panoramic image.

[0030] Further, the attribute information further includes the state information of the target object in the vehicle cabin, and the method further includes:

[0031] Generate interaction information corresponding to the target virtual image according to the state information;

[0032] When displaying the target virtual image, display the corresponding interaction information, and the interaction information includes status prompt information or operation instruction information.

[0033] In a second aspect, the present application provides an image processing apparatus, and the apparatus includes:

[0034] An acquisition module, configured to acquire a plurality of original vehicle cabin images from different perspectives;

[0035] A stitching module, configured to stitch the plurality of original vehicle cabin images to obtain a vehicle cabin panoramic image;

[0036] A perception module, configured to determine a target object in the plurality of original vehicle cabin images based on computer vision perception;

[0037] A virtualization module, configured to determine a target virtual image corresponding to the target object and obtain a target layer including the target virtual image;

[0038] An overlay module, configured to overlay the vehicle cabin panoramic image and the target layer to obtain a target panoramic image including the target virtual image.

[0039] In a third aspect, the present application provides a computer-readable storage medium, in which at least one instruction or at least one segment of program is stored, and the at least one instruction or at least one segment of program is loaded and executed by a processor to implement an image processing method as described in the first aspect.

[0040] In a fourth aspect, the present application provides a computer device, and the computer device includes a processor and a memory, and at least one instruction or at least one segment of program is stored in the memory, and the at least one instruction or at least one segment of program is loaded and executed by the processor to implement an image processing method as described in the first aspect.

[0041] Fifth aspect, the present application provides a computer program product, which includes computer instructions that, when executed by a processor, implement an image processing method as described in the first aspect.

[0042] The image processing method, device, and medium provided by the present application have the following technical effects:

[0043] On the one hand, the solution provided by the present application stitches multiple original cabin images from different perspectives into a panoramic cabin image. On the other hand, based on computer vision perception technology, it identifies the target objects in each original cabin image and constructs a target layer containing the target virtual image corresponding to the target object. Thus, the panoramic cabin image in the real scene can be superimposed with the target layer to obtain a target panoramic image combining reality and virtuality, and the target object is represented by the target virtual image in the target panoramic image, and the target virtual image can also be highlighted. The technical solution provided by the present application can significantly improve the perception effect of the cabin, enabling the observer to understand the real situation inside the cabin, effectively avoiding the problem that it is difficult for the human eye to detect objects in hidden areas of sight, and facilitating the user to monitor the cabin.

[0044] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings

[0045] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0046] Figure 1 is a flowchart showing an image processing method provided by an embodiment of the present application;

[0047] Figure 2 is a flowchart showing a process of stitching multiple original cabin images provided by an embodiment of the present application;

[0048] Figure 3 is a flowchart showing a process of generating a target layer containing a target virtual image provided by an embodiment of the present application;

[0049] Figure 4 is a flowchart showing a process of determining the display area of the target virtual image in the target layer provided by an embodiment of the present application;

[0050] Figure 5 is a schematic diagram of an image processing device provided by an embodiment of the present application;

[0051] Figure 6 It is a schematic diagram of the hardware structure of a device provided by an embodiment of the present application for implementing an image processing method. Detailed implementation manners

[0052] To improve the perception effect of the vehicle cabin, embodiments of the present application provide an image processing method, apparatus and medium. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout.

[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0054] To facilitate the understanding of the technical solutions described in the embodiments of the present application and the technical effects produced thereby, the embodiments of the present application explain the relevant professional terms involved:

[0055] Image stitching: Image Stitching, is a technology that uses real-scene images to form a panoramic space. It stitches multiple images into a large-scale image or a 360-degree panoramic image. Image stitching technology involves technologies such as computer vision, computer graphics, digital image processing, and some mathematical tools.

[0056] Computer vision: Computer Vision, CV, is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking and measurement on targets, and further performing graphics processing to make the computer processing results become images more suitable for human eyes to observe or be transmitted to instruments for detection.

[0057] Please refer toFigure 1 , Figure 1 is a flowchart of an image processing method provided by an embodiment of the present application. The present application provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, more or fewer operation steps may be included. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order of the method shown in the embodiment or the accompanying drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). As Figure 1 shown, an image processing method provided by an embodiment of the present application may include the following steps:

[0058] S100: Obtain a plurality of original cabin images from different perspectives.

[0059] In the embodiment of the present application, a plurality of cameras installed at different positions inside the cabin are used for shooting to obtain a plurality of original cabin images from different perspectives. Preferably, at least 3 cameras are arranged inside the cabin, for example, they can be respectively deployed at positions such as the interior rearview mirror, the left and right B-pillars, etc., so that the field of view of the cameras covers the interior of the cabin and there is no occlusion in the stitched cabin panoramic image.

[0060] S200: Stitch the plurality of original cabin images to obtain a cabin panoramic image.

[0061] In the embodiment of the present application, based on the image stitching technology, the two-dimensional original cabin images can be stitched to obtain a three-dimensional cabin panoramic image. Specifically, feature extraction and matching are performed on the plurality of original cabin images, and the plurality of original cabin images are registered, stitched, and fused based on the matching results of the features. Image registration is used to determine the overlapping area and overlapping position between the plurality of original cabin images to be stitched. Further, after image stitching and fusion, brightness, color, etc. can also be equalized.

[0062] In an embodiment of the present application, due to differences in the positions, postures, imaging performances, etc. of the plurality of cameras, the scales, resolutions, calibrated image coordinate systems, etc. of the plurality of original cabin images taken may also be different. Therefore, it is necessary to project the plurality of original cabin images onto an image model in the same coordinate system before image stitching. Generally, there are methods such as cylindrical projection, cube projection, and spherical projection. The cabin panoramic image obtained after spherical projection, stitching, and fusion can be a 360-degree spherical image. It can be understood that spherical projection simulates the characteristics of human eye observation and projects the original cabin image into the spherical model through perspective transformation to form an observation sphere.

[0063] S300: Based on computer vision perception, determine the target objects in the multiple original cabin images.

[0064] In one embodiment of the present application, the multiple original cabin images are respectively input into a trained computer vision perception model for visual perception processing to obtain a perception result, which indicates the target objects in the multiple original cabin images and the attribute information of the target objects. Among them, the attribute information may include the object category of the target object, the position information and size information of the target object in a certain coordinate system, the status information of the target object, etc. Exemplarily, a deep neural network model is used to perform face recognition and object detection on each original cabin image to obtain the faces and objects existing in each original cabin image, and identify the user identity corresponding to the face, the category corresponding to the object, the coordinate information of the object in the image coordinate system or camera coordinate system, the motion state of the object, etc.

[0065] Using computer vision technology based on machine learning, the original cabin images can be processed for target detection, recognition, classification, tracking, etc., so as to detect the target objects in the multiple original cabin images. The target objects can be faces, human bodies, pets (such as cats, dogs, etc.), objects (such as wallets, keys, computers, tablets, mobile phones, etc.), etc. At the same time, the object category, image area, depth of field, external static features, vital signs, motion trajectory, etc. of the target objects can also be identified. The target objects identified in the multiple original cabin images can be used as the augmented reality content to be enhanced in the above-mentioned cabin panoramic image and highlighted. By using the powerful perception ability of computer vision, objects that may not be immediately recognized by the human eye can be identified, so as to effectively improve the target detection effect of the original cabin images and further facilitate the monitoring of the interior of the cabin.

[0066] S400: Determine the target virtual image corresponding to the target object and obtain a target layer containing the target virtual image.

[0067] In the embodiment of the present application, the target virtual image refers to a virtual image that can represent the target object, corresponding to the real image of the target object in the real world, but more prominent and eye-catching than the real image. The target virtual image can be an abstract and general representation of the category to which the target object belongs, or an extraction or rendering of the real image of the target object. Further, the same target object can have multiple target virtual images with different image styles, such as a simple style of black and white lines, a cartoon image style, etc., for the user to set or switch according to personal style preferences. Exemplarily, for a detected mobile phone, its corresponding virtual image can be a cartoon mobile phone image with black and white lines, or a standard mobile phone image consistent with the model of the mobile phone itself.

[0068] In an embodiment of the present application, in order to superimpose a target virtual image onto a panoramic image of a real vehicle cabin, one or more layers are set up to place the target virtual images of each target object. Specifically, the target virtual image is displayed in the layer by loading the resource file corresponding to the target virtual image. The layer can accurately position the target virtual image so that in the subsequent image superimposition process, the target virtual image in the target layer can cover the target object in the panoramic image of the vehicle cabin. The positioning of the target virtual image in the layer is related to the position and size of the target object in the vehicle cabin, the scale of the image model corresponding to the layer, etc., and its position in the layer corresponds to the content area to be enhanced in the panoramic image of the vehicle cabin. Preferably, the scale of the set layer is consistent with the panoramic image of the vehicle cabin.

[0069] It can be understood that in a real original vehicle cabin image or a panoramic image of the vehicle cabin, due to reasons such as seat occlusion and insufficient light, the human eye may not be able to quickly find the object to be searched. Therefore, in the embodiment of the present application, the target objects identified in multiple original vehicle cabin images are used as the augmented reality content to be enhanced in the above-mentioned panoramic image of the vehicle cabin, and the target virtual image is prominently displayed in the image area to be enhanced, so that the user can quickly find and locate the target object when monitoring the vehicle cabin.

[0070] S500: Superimpose the panoramic image of the vehicle cabin and the target layer to obtain a target panoramic image containing the target virtual image.

[0071] In an embodiment of the present application, in a preset vehicle cabin coordinate system, the panoramic image of the vehicle cabin and the target layer with the same scale are digitally image - synthesized to obtain the target panoramic image in the vehicle cabin coordinate system, and in the target panoramic image, the target object is represented by the target virtual image. Here, the same scale means that the pixel sizes are the same. Specifically, the target layer and the panoramic image of the vehicle cabin can be directly superimposed, or image fusion can be performed according to different weight coefficients. Further, brightness and color equalization processing are performed on the fused image.

[0072] In a practical application, the collected original vehicle cabin images are sent to a vehicle cabin monitoring system in the cloud through a vehicle - mounted communication device, and a target panoramic image is generated in the cloud. Then, the generated target panoramic image can be sent to a user terminal so that the target panoramic image is correspondingly displayed on the user terminal according to the user's perspective, and the target virtual image is displayed in a preset display style in the target panoramic image. That is, the target panoramic image can be displayed on the terminal screen, and the local or all areas in the target panoramic image can be correspondingly displayed according to the perspective adjusted by the user through the terminal screen, and the target virtual image can be displayed on the terminal screen, such as using dynamic display styles such as flashing and magnifying to make it have a prominent and eye - catching display effect.

[0073] Further, based on the foregoing embodiments, the attribute information of the target object obtained through computer vision perception may further include the state information of the target object in the vehicle cabin, such as the vital signs of an animal and the degree of user fatigue recognized through face recognition. Therefore, the interaction information corresponding to the target virtual image may further be generated according to the state information, and the corresponding interaction information may be synchronously displayed when the target virtual image is displayed, where the interaction information may include status prompt information or operation instruction information. Exemplarily, the interaction information may be a text element in the target layer and is displayed together with the corresponding target virtual image, so as to prompt through text that there is an animal here, the driver is unconscious, etc., or prompt the monitor that the door can be remotely opened, the vehicle driving can be remotely controlled, etc.

[0074] The image processing method provided in the foregoing embodiments, on the one hand, stitches a plurality of original vehicle cabin images from different perspectives into a panoramic vehicle cabin image, and on the other hand, based on computer vision perception technology, identifies the target objects in each original vehicle cabin image and constructs a target layer including the target virtual image corresponding to the target object, so that the panoramic vehicle cabin image in the real scene can be superimposed on the target layer to obtain a target panoramic image combining the real and the virtual, and the target object is represented by the target virtual image in the target panoramic image, and the target virtual image can also be highlighted. The technical solution provided in this application can significantly improve the perception effect of the vehicle cabin, enable the observer to understand the real situation inside the cabin, effectively avoid the problem that it is difficult for the human eye to detect objects in the hidden line of sight, and facilitate the user to monitor the vehicle cabin at the same time.

[0075] Please refer to Figure 2 , Figure 2 which is a schematic flow chart of stitching a plurality of original vehicle cabin images provided by an embodiment of the present application. As Figure 2 shown, the above panoramic vehicle cabin image can be obtained through the following steps:

[0076] S210: Perform feature point detection on the plurality of original vehicle cabin images respectively to obtain the feature points of each original vehicle cabin image and the description information of the feature points.

[0077] Specifically, the Harris algorithm (a corner detection algorithm) or the SIFT (Scale-invariant feature transform) algorithm is used to detect the local features in the original vehicle cabin image. Feature points, that is, key points, can be corner points, edge points, bright points in dark areas, dark points in bright areas, etc. in the image. The description information of the feature points is also called a descriptor, which is the local feature information of the feature points and may include but is not limited to position, scale, orientation, etc. Generally, the local feature information can be expressed by a spatial vector.

[0078] S220: Match the feature points of each of the original cabin images according to the description information of the feature points to obtain feature point pairs.

[0079] It can be understood that in order to be able to splice into a panoramic cabin image, there are overlapping cabin scenes among the multiple original cabin images taken. The features of the feature points in the overlapping parts of different images are similar, and they can be matched to obtain paired feature points.

[0080] Specifically, an exhaustive matching method is used for feature point matching. For two original cabin images with overlapping scenes, respective sets of feature point description information are established, the feature point description information in the two sets is traversed, and the similarity between each pair is calculated, such as calculating the Euclidean distance, and two feature points with a similarity meeting a preset threshold are used as a matched feature point pair. In addition, a search method based on a binary tree can also be used to determine the nearest neighbor feature points of each feature point in each original cabin image, so as to form feature point pairs.

[0081] S230: Determine the association relationship between each of the original cabin images based on the feature point pairs.

[0082] Based on the feature point pairs between each of the original cabin images, the association relationship between each pair of the original cabin images can be determined. The association relationship can include the spatial order, scale transformation relationship, etc. of the original cabin images.

[0083] S240: Splice and fuse the multiple original cabin images according to the association relationship between each of the original cabin images to obtain the panoramic cabin image.

[0084] Feasibly, according to the spatial order and scale transformation relationship of the original cabin images indicated by the association relationship, map the multiple original cabin images onto the same panoramic image model for splicing and fusing to obtain the above-mentioned panoramic cabin image. Preferably, this panoramic image model can be a spherical model set in a pre-calibrated cabin coordinate system.

[0085] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a process for generating a target layer including a target virtual image provided by an embodiment of the present application. As Figure 3 shown, on the basis of the foregoing embodiment, if the attribute information of the target object includes the target object category of the target object, the above-mentioned target layer can be obtained through the following steps, which can specifically include:

[0086] S411: Determine the target layer type corresponding to the target object category based on a preset mapping relationship; the preset mapping relationship represents the corresponding relationship between the object category and the layer type.

[0087] That is, the layers are divided according to the target object category. Exemplarily, four types of layers can be set, corresponding to human faces, vehicle bodies, animals, and static objects respectively. In addition, the layer type can also indicate the image style.

[0088] S412: Determine the target virtual image corresponding to the target object according to the target object category and the target layer type.

[0089] For the same target object category, there can be virtual images with different styles. According to the image style indicated by the layer type, the suitable target virtual image is determined.

[0090] S413: Generate the target layer containing the target virtual image according to the target layer type and the target virtual image.

[0091] Specifically, the target virtual image is displayed in the corresponding layer by loading the resource file corresponding to the target virtual image, and the target layer containing the target virtual image is obtained.

[0092] In the above embodiments, the layer type is selected according to the target object category, and then the suitable target virtual image is matched, so that the styles and categories of the target virtual images in the same target layer are unified, which is convenient for management and display.

[0093] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a process for determining the display area of a target virtual image in a target layer provided by an embodiment of the present application. As Figure 4 shown, on the basis of the foregoing embodiments, the attribute information of the target object further includes the first coordinate information of the target object in the camera coordinate system. Then, the display area information can be determined through the following steps, which may specifically include:

[0094] S421: Determine the coordinate conversion relationship between the camera coordinate system and the pre-calibrated vehicle cabin coordinate system.

[0095] It can be understood that the camera coordinate system is a three-dimensional rectangular coordinate system established with the focus center of the camera as the origin and the optical axis as the Z axis, and the Z axis is perpendicular to the image plane. The pre-calibrated vehicle cabin coordinate system is a three-dimensional rectangular coordinate system established with the pre-calibrated origin and direction. Based on the installation position and posture of the camera, the positioning of the camera in the vehicle cabin coordinate system can be determined, and further the coordinate conversion relationship between the camera coordinate system and the vehicle cabin coordinate system can be determined, and this coordinate conversion relationship can be represented by a conversion matrix.

[0096] S422: Obtain the second coordinate information of the target object in the vehicle cabin coordinate system according to the first coordinate information and the coordinate conversion relationship.

[0097] Based on computer vision technology, parameters such as the relative size and depth of field of the target object can be perceived from the image plane, and then converted into the first coordinate information in the camera coordinate system. The first coordinate information is three-dimensional coordinates. The first coordinate information can be the three-dimensional coordinates of the edge detection points of the target object, which are used to delineate the area of the target object in the camera coordinate system. Based on the above conversion matrix, the first coordinate information can be converted into the second coordinate information in the vehicle cabin coordinate system, and the second coordinate information is also three-dimensional coordinates.

[0098] S423: Determine the display area information of the target virtual image in the target layer according to the second coordinate information.

[0099] According to the second coordinate information of the target object in the vehicle cabin coordinate system, the area of the target object in the vehicle cabin panoramic image can be mapped. The area of the target object in the vehicle cabin panoramic image is also the area corresponding to the target virtual image in the target layer. Precise positioning of the target virtual image in the target layer can ensure that the target virtual image in the target layer can cover the target object in the vehicle cabin panoramic image during the image overlay process.

[0100] An embodiment of the present application also provides an image processing device 500, as Figure 5 shown. The device 500 may include:

[0101] An acquisition module 510, configured to acquire multiple original vehicle cabin images from different perspectives;

[0102] A stitching module 520, configured to stitch the multiple original vehicle cabin images to obtain a vehicle cabin panoramic image;

[0103] A perception module 530, configured to determine the target object in the multiple original vehicle cabin images based on computer vision perception;

[0104] A virtualization module 540, configured to determine the target virtual image corresponding to the target object and obtain a target layer including the target virtual image;

[0105] An overlay module 550, configured to overlay the vehicle cabin panoramic image and the target layer to obtain a target panoramic image including the target virtual image.

[0106] In an embodiment of the present application, the stitching module 520 may include:

[0107] A feature detection unit, configured to perform feature point detection on each of the multiple original vehicle cabin images to obtain the feature points of each original vehicle cabin image and the description information of the feature points;

[0108] A feature matching unit, configured to match the feature points of each of the original cabin images according to the description information of the feature points, so as to obtain feature point pairs;

[0109] An image association relationship determination unit, configured to determine the association relationship between each of the original cabin images based on the feature point pairs;

[0110] An image stitching unit, configured to stitch and fuse the multiple original cabin images according to the association relationship between each of the original cabin images, so as to obtain the panoramic cabin image.

[0111] In an embodiment of the present application, the perception module 530 may include:

[0112] A neural network perception unit, configured to respectively input the multiple original cabin images into a trained computer vision perception model for visual perception processing, so as to obtain a perception result, where the perception result indicates a target object in the multiple original cabin images and attribute information of the target object.

[0113] In an embodiment of the present application, the virtualization module 540 may include:

[0114] A layer type selection unit, configured to determine a target layer type corresponding to the target object category based on a preset mapping relationship; the preset mapping relationship represents the corresponding relationship between the object category and the layer type;

[0115] A virtual image determination unit, configured to determine a target virtual image corresponding to the target object according to the target object category and the target layer type;

[0116] A layer generation unit, configured to generate the target layer including the target virtual image according to the target layer type and the target virtual image.

[0117] In an embodiment of the present application, the virtualization module 540 may further include:

[0118] A coordinate transformation relationship determination unit, configured to determine a coordinate transformation relationship between the camera coordinate system and a pre-calibrated cabin coordinate system;

[0119] A coordinate transformation unit, configured to obtain second coordinate information of the target object in the cabin coordinate system according to the first coordinate information and the coordinate transformation relationship;

[0120] A position determination unit, configured to determine display area information of the target virtual image in the target layer according to the second coordinate information.

[0121] In an embodiment of the present application, the overlay module 550 may include:

[0122] A digital image synthesis unit, configured to perform digital image synthesis on the panoramic cabin image and the target layer of the same scale in a preset cabin coordinate system to obtain the target panoramic image in the cabin coordinate system, and represent the target object with the target virtual image in the target panoramic image.

[0123] In an embodiment of the present application, the device 500 may further include:

[0124] A first display unit, configured to send the target panoramic image to a user terminal so that the target panoramic image is correspondingly displayed on the user terminal according to the user's perspective;

[0125] A second display unit, configured to display the target virtual image in a preset display style in the target panoramic image.

[0126] In another embodiment of the present application, the device 500 may further include:

[0127] An interaction information determination unit, configured to generate interaction information corresponding to the target virtual image according to the status information of the target object in the cabin;

[0128] A third display unit, configured to display the corresponding interaction information when displaying the target virtual image, and the interaction information includes status prompt information or operation instruction information.

[0129] It should be noted that when the device provided in the above embodiment realizes its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions may be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment and will not be repeated here.

[0130] An embodiment of the present application provides a computer device, which includes a processor and a memory. At least one instruction or at least one segment of program is stored in the memory, and the at least one instruction or the at least one segment of program is loaded and executed by the processor to implement an image processing method as provided in the above method embodiment.

[0131] Figure 6 A hardware structure diagram of a device for implementing an image processing method provided in an embodiment of the present application is shown. The device may participate in forming or include the device or system provided in an embodiment of the present application. As Figure 6As shown, the device 10 may include one or more processors 1002 (illustrated as 1002a, 1002b, ……, 1002n in the figure) (the processor 1002 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 6 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the device 10 may further include more or fewer components than Figure 6 shown in, or have a different configuration from Figure 6 that shown.

[0132] It should be noted that the above one or more processors 1002 and / or other data processing circuits can generally be referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the device 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0133] The memory 1004 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the methods described in the embodiments of the present application. The processor 1002 executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, that is, implements the above-mentioned image processing method. The memory 1004 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 1004 may further include a memory remotely set relative to the processor 1002, and these remote memories can be connected to the device 10 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0134] The transmission device 1006 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the device 10. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 1006 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0135] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the device 10 (or mobile device).

[0136] An embodiment of the present application also provides a computer-readable storage medium, which can be disposed in a server to store at least one instruction or at least one segment of a program related to an image processing method in a method embodiment. The at least one instruction or the at least one segment of the program is loaded and executed by the processor to implement the image processing method provided in the above method embodiment.

[0137] Optionally, in this embodiment, the above storage medium can be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0138] An embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes an image processing method provided in the above various optional implementation manners.

[0139] As can be seen from the embodiments of the image processing method, apparatus, and medium provided by the present application above, on the one hand, the solution provided by the present application stitches multiple original cabin images from different perspectives into a panoramic cabin image. On the other hand, based on computer vision perception technology, it identifies target objects in each original cabin image and constructs a target layer containing the target virtual image corresponding to the target object. Thus, the panoramic cabin image in the real scene can be superimposed on the target layer to obtain a target panoramic image combining reality and virtuality, and the target object is represented by the target virtual image in the target panoramic image. Moreover, the target virtual image can be prominently displayed. The technical solution provided by the present application can significantly improve the perception effect of the cabin, enabling the observer to understand the real situation inside the cabin, effectively avoiding the problem that it is difficult for the human eye to detect objects in hidden areas of the line of sight, and facilitating the user's monitoring of the cabin at the same time.

[0140] It should be noted that the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0141] Each embodiment in the present application is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the apparatus, device, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0142] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.

[0143] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining a plurality of original cabin images from different perspectives, where the fields of view of the plurality of original cabin images cover the interior of the cabin; Stitching the plurality of original cabin images to obtain a panoramic cabin image; Based on computer vision perception, determining target objects in the plurality of original cabin images and attribute information of the target objects, where the attribute information includes the target object category and the status information of the target object inside the cabin; Based on a preset mapping relationship, determining a target layer type corresponding to the target object category; the preset mapping relationship represents the corresponding relationship between the object category and the layer type; Determining a target virtual image adapted to the target object according to the image style indicated by the target layer type; generating interaction information corresponding to the target virtual image according to the status information; Generating a target layer including the target virtual image according to the target layer type and the target virtual image; Overlaying the panoramic cabin image and the target layer to obtain a target panoramic image, so as to synchronously display the target virtual image and the interaction information in the target panoramic image.

2. The method according to claim 1, wherein The step of stitching the plurality of original cabin images to obtain a panoramic cabin image includes: Performing feature point detection on each of the plurality of original cabin images to obtain the feature points of each original cabin image and the description information of the feature points; Matching the feature points of each original cabin image according to the description information of the feature points to obtain feature point pairs; Determining the association relationship between each of the original cabin images based on the feature point pairs; Stitching and fusing the plurality of original cabin images according to the association relationship between each of the original cabin images to obtain the panoramic cabin image.

3. The method according to claim 1, characterized in that, The step of determining target objects in the plurality of original cabin images based on computer vision perception includes: Inputting each of the plurality of original cabin images into a trained computer vision perception model for visual perception processing to obtain a perception result, where the perception result indicates the target objects in the plurality of original cabin images and the attribute information of the target objects.

4. The method according to claim 3, wherein The attribute information further includes first coordinate information of the target object in the camera coordinate system, and the step of generating a target layer including the target virtual image further includes: Determining the coordinate conversion relationship between the camera coordinate system and a pre-calibrated cabin coordinate system; Obtaining second coordinate information of the target object in the cabin coordinate system according to the first coordinate information and the coordinate conversion relationship; Determining the display area information of the target virtual image in the target layer according to the second coordinate information.

5. The method according to claim 1, characterized in that, The step of overlaying the panoramic cabin image and the target layer to obtain a target panoramic image includes: Under a preset cabin coordinate system, performing digital image synthesis on the panoramic cabin image and the target layer with the same scale to obtain the target panoramic image in the cabin coordinate system, and representing the target object with the target virtual image in the target panoramic image.

6. The method according to claim 1, wherein The method further includes: Send the target panoramic image to the user terminal so that the target panoramic image can be correspondingly displayed on the user terminal according to the user's perspective; And display the target virtual image in a preset display style in the target panoramic image.

7. The method according to claim 3, wherein The method further includes: When displaying the target virtual image, display the corresponding interaction information, where the interaction information includes status prompt information or operation instruction information.

8. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to acquire a plurality of original cabin images from different perspectives, and the fields of view of the plurality of original cabin images cover the interior of the cabin; A splicing module, configured to splice the plurality of original cabin images to obtain a cabin panoramic image; A perception module, configured to determine a target object and attribute information of the target object in the plurality of original cabin images based on computer vision perception, where the attribute information includes the target object category and the status information of the target object in the cabin; A layer type selection module, configured to determine a target layer type corresponding to the target object category based on a preset mapping relationship; the preset mapping relationship represents the corresponding relationship between the object category and the layer type; A virtual image determination module, configured to determine a target virtual image adapted to the target object according to the image style indicated by the target layer type; An interaction information determination module, configured to generate interaction information corresponding to the target virtual image according to the status information; A layer generation module, configured to generate a target layer including the target virtual image according to the target layer type and the target virtual image; An overlay module, configured to overlay the cabin panoramic image and the target layer to obtain a target panoramic image, so as to synchronously display the target virtual image and the interaction information in the target panoramic image.

9. A computer-readable storage medium, characterized in that, At least one instruction or at least one program is stored in the computer-readable storage medium, and the at least one instruction or the at least one program is loaded and executed by a processor to implement an image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image display method and device, display equipment and computer readable storage medium

    CN112037314A

  • Airport surface operation management system and method based on digital twinning

    CN112053085A

  • Method and device for determining position and posture of panoramic image in three-dimensional model

    CN112184815A