Image processing method, electronic device, and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2026-08-11
AI Technical Summary
智能头戴设备具有三维内容显示与交互功能,但目前三维内容的生产还主要依靠专业的设备和专业的人员
Smart Images

Figure CN116468917B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image processing method, electronic device and storage medium. Background Technology
[0002] Smart head-mounted devices (such as AR glasses, VR glasses, or MR glasses) are developing rapidly and their applications are becoming increasingly widespread. While smart head-mounted devices offer 3D content display and interaction capabilities, the production of 3D content currently relies primarily on specialized equipment and personnel. Capturing and sharing 3D images is not as convenient as capturing 2D images. Smart head-mounted devices equipped with cameras offer a more natural and convenient way to capture images compared to handheld devices (such as mobile phones). Summary of the Invention
[0003] In a first aspect, embodiments of this application provide an image processing method applied to a smart wearable device, comprising:
[0004] Acquire multiple image frames at equal time intervals;
[0005] For each of the image frames, feature extraction is performed to determine the feature points of the current image frame, wherein the feature point data of the feature points includes a descriptor describing the image region surrounding the feature points;
[0006] Each of the aforementioned feature points is matched with all feature points of the previous image frame to determine the feature points of the previous image frame that match the aforementioned feature points.
[0007] The matching relationship data and image data are sent to the terminal, so that the terminal determines the three-dimensional model of the target based on the matching relationship data and the image data. The matching relationship data includes the matched feature points, the feature point ID and feature point data of the feature points of the previous image frame, and the image data includes the image data of the current image frame.
[0008] In some embodiments, it also includes:
[0009] For each of the aforementioned feature points, feature matching is performed with all feature points of the previous image frame to determine the Euclidean distance between each of the aforementioned feature points and all feature points of the previous image frame;
[0010] The feature points of the previous image frame are determined based on the Euclidean distances and the feature point matching.
[0011] Alternatively, based on the fast nearest neighbor matching algorithm, feature matching is performed on each of the aforementioned feature points and all feature points of the previous image frame to determine the feature points of the previous image frame that match the aforementioned feature points.
[0012] In some embodiments, it also includes:
[0013] If the feature point is determined to be mismatched with all feature points in the previous image frame, the feature point is discarded.
[0014] If each feature point is determined to be mismatched with all feature points in the previous image frame, the current image frame is discarded.
[0015] In some embodiments, it also includes:
[0016] If the feature point is determined to be mismatched with all feature points in the previous image frame, the feature point is retained.
[0017] If each feature point is determined to be mismatched with all feature points of the previous image frame, the current image frame is retained, and the previous image frame is discarded.
[0018] Secondly, embodiments of this application provide an image processing method applied to a terminal, comprising:
[0019] The system receives matching relationship data and image data sent by a smart wearable device. The matching relationship data is determined by the smart wearable device based on feature matching of each feature point in the current image frame and all feature points in the previous image frame. The matching relationship data includes the matched feature points, the feature point IDs of the feature points in the previous image frame, and the feature point data. The image data includes the image data of the current image frame, and the feature point data of the feature points includes a descriptor describing the image region surrounding the feature point.
[0020] Based on the matching relationship data, the current image frame, and the previous image frame, the camera pose and three-dimensional spatial points of the current image frame are determined;
[0021] Perform BA optimization on all the determined camera poses and the three-dimensional space points;
[0022] If no new image frame is received, perform BA optimization on all the determined camera poses and 3D spatial points to obtain the final camera poses and 3D spatial points.
[0023] Based on all received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames, a 3D model is determined.
[0024] In some embodiments, it also includes:
[0025] The camera pose of the current image frame is set as an identity matrix, and the target matrix is determined based on the identity matrix and the matching relationship data.
[0026] The target matrix is decomposed to obtain the camera pose of the current image frame;
[0027] The three-dimensional spatial points are determined by triangulation based on the camera pose of the current image frame, the camera pose of the previous image frame, and the matching relationship data.
[0028] In some embodiments, it also includes:
[0029] All received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames are input into the target algorithm to output the 3D model. The target algorithm is either the MVS algorithm or the NeRF algorithm.
[0030] In some embodiments, it also includes:
[0031] Based on the 3D model, output multi-view rendered videos or 3D data files.
[0032] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the image processing methods described above.
[0033] Fourthly, embodiments of this application also provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image processing method as described above.
[0034] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements any of the image processing methods described above. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This is a schematic diagram of the structure of a smart wearable device provided in one embodiment of this application.
[0037] Figure 2 This is one of the schematic flowcharts of an image processing method provided in an embodiment of this application;
[0038] Figure 3 This is a second schematic flowchart of an image processing method provided in one embodiment of this application;
[0039] Figure 4This is the third schematic flowchart of an image processing method provided in one embodiment of this application;
[0040] Figure 5 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0041] To make the technical solutions and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0042] Some image processing methods provided in this application can be applied to smart wearable devices such as wearable devices, augmented reality (AR) / virtual reality (VR) devices, and this application does not impose any restrictions on the specific type of smart wearable device.
[0043] For example, Figure 1 This is a schematic diagram of the structure of a smart wearable device provided in one embodiment of this application, as shown below. Figure 1 As shown, the smart wearable device 100 may include a processor 110, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, a wireless communication module 160, a sensor module 180, a button 190, a light emitting diode (LED) lamp 191, a camera 193, a display component 194, and an optical engine 195, etc.; wherein, the sensor module 180 includes a touch sensor 180K; the optical engine 195 includes a lens and a display screen.
[0044] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the smart wearable device 100. In other embodiments of this application, the smart wearable device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0045] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.
[0046] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0047] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0048] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (DCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the smart wearable device 100.
[0049] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality.
[0050] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display unit 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the shooting function of the smart wearable device 100. The processor 110 and the display unit 194 communicate via the DSI interface to enable the display function of the smart wearable device 100.
[0051] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display component 194, a wireless communication module 160, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0052] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, or USB Type-C port. USB port 130 can be used to connect a charger to charge the smart wearable device 100, and can also be used for data transfer between the smart wearable device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other smart wearable devices, such as smart bracelets, smart rings, or smartphones.
[0053] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the smart wearable device 100. In other embodiments of this application, the smart wearable device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0054] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the smart wearable device 100. While charging the battery 142, the charging management module 140 can also supply power to the smart wearable device 100 via the power management module 141.
[0055] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display unit 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0056] The wireless communication function of the smart wearable device 100 can be implemented through the antenna 1, the wireless communication module 160, the modem processor, and the baseband processor.
[0057] Antenna 1 is used to transmit and receive electromagnetic wave signals. Each antenna in the smart wearable device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in conjunction with a tuning switch.
[0058] The wireless communication module 160 can provide solutions for wireless communication applications on the smart wearable device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 1.
[0059] In some embodiments, the antenna 1 of the smart wearable device 100 is coupled to the wireless communication module 160, enabling the smart wearable device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0060] The smart wearable device 100 implements display functions through a GPU, a display component 194, and an application processor. The GPU is a microprocessor for image processing, connecting the display component 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0061] Display component 194 is used to display images, videos, etc. Display component 194 may include a display lens or a display mask, and may also include a display screen. The display lens or display mask can be an optical waveguide (e.g., a diffractive waveguide or geometric waveguide), a freeform prism, or free space, etc.; the display lens or display mask serves as the propagation path for the imaging light, transmitting the virtual image to the human eye. To enable AR glasses to simultaneously see real and virtual images, a waveguide can be used to transmit the light from the virtual image into the human eye.
[0062] The aforementioned display screen can be a display panel, which can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.
[0063] In some embodiments, the smart wearable device 100 may include one or N display components 194, where N is a positive integer greater than 1.
[0064] The smart wearable device 100 can achieve shooting functions through ISP, camera 193, video codec, GPU, display component 194 and application processor.
[0065] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0066] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the smart wearable device 100 may include one or N cameras 193, where N is a positive integer greater than 1. In this embodiment, camera 193 may include at least one infrared camera.
[0067] For example, the camera 193 can be used to acquire multiple image frames, and to acquire multiple image frames at equal time intervals from the multiple image frames for performing the image processing method of the embodiments of this application.
[0068] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when the smart wearable device 100 selects a frequency, the DSP performs Fourier transforms on the frequency energy.
[0069] Video codecs are used to compress or decompress digital video. The smart wearable device 100 may support one or more video codecs. Thus, the smart wearable device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0070] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent wearable devices to perform applications such as image recognition, facial recognition, speech recognition, and text understanding.
[0071] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area may store data created during the use of the smart wearable device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of the smart wearable device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located within the processor.
[0072] Touch sensor 180K, also known as a "touch device," can be disposed on display component 194. The touch sensor 180K and display component 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K is used to detect touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display component 194. In other embodiments, touch sensor 180K may also be disposed on the surface of smart wearable device 100, in a different location than display component 194.
[0073] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. The smart wearable device 100 can receive button input and generate key signal inputs related to user settings and function control of the smart wearable device 100.
[0074] The optical engine 195 is mainly used for imaging, including lenses and displays. The lenses here can be optical components.
[0075] Other image processing methods provided in this application can be applied to terminal devices such as mobile phones, tablets, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). They can also be applied to databases, servers, and service response systems based on terminal artificial intelligence. Generally, the processing power of the terminal devices described in this application is stronger than that of smart wearable devices; for example, the terminal devices have graphics processing units with more cores. This application does not limit the specific type of terminal device. The mobile terminal (terminal device) in this application includes various handheld devices, in-vehicle devices, computing devices, or other processing devices connected to a wireless modem with wireless communication capabilities, such as mobile phones, tablets, desktop laptops, and smart devices capable of running applications, including the central control console of a smart car. Specifically, it can refer to user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device. Terminal devices can also be satellite phones, cellular phones, smartphones, wireless data cards, wireless modems, machine-type communication devices, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical care, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, and terminal devices in 5G networks or future communication networks. Mobile terminals can be battery-powered or attached to and powered by the power system of a vehicle or vessel.The power system of a vehicle or ship can also charge the battery of a mobile terminal to extend the communication time of the mobile terminal.
[0076] Reference Figure 2 One embodiment of this application also provides an image processing method applied to a smart wearable device, including but not limited to the following steps:
[0077] Step 201: Acquire multiple image frames at equal time intervals;
[0078] Step 202: Perform feature extraction on each of the image frames to determine the feature points of the current image frame, wherein the feature point data of the feature points includes a descriptor describing the image region surrounding the feature points;
[0079] Step 203: Perform feature matching between each feature point and all feature points of the previous image frame to determine the feature points of the previous image frame that match the feature points.
[0080] Step 204: Send the matching relationship data and image data to the terminal, so that the terminal determines the three-dimensional model of the target based on the matching relationship data and the image data, wherein the matching relationship data includes the matched feature points, the feature point ID and feature point data of the feature points of the previous image frame, and the image data includes the image data of the current image frame.
[0081] In step 201 above, the smart wearable device can receive user input commands, activate its camera to capture video segments, and select multiple image frames at equal time intervals from the video segments in real time using an equally spaced strategy. Alternatively, the smart wearable device can also acquire multiple image frames at equal time intervals by capturing images at equal time intervals using its camera; this is not a limitation of this application. The smart wearable device can be AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, etc. For example, the time interval can be set to 500ms, meaning the acquisition interval for each image frame is 500ms. The above time interval can also be adjusted according to specific circumstances.
[0082] It is understood that each of the above image frames may include one or more targets, or may not include any targets. The targets between each image frame may be the same target or different targets. For example, an image frame may include target I and target II, and the previous image frame of an image frame may include target I, target II and target III. This is not a limitation of this application.
[0083] In step 202 above, feature extraction is performed on each image frame to obtain the feature points of each image frame, and the feature points of the current image frame are determined from all the feature points.
[0084] In this embodiment, the feature used for matching can be a SIFT feature.
[0085] SIFT (Scale Invariant Feature Transform) extracts local features from an image, finds extrema in scale space, and extracts their location, scale, and orientation information. Applications of SIFT include object recognition, robot mapping and navigation, image stitching, 3D model building, gesture recognition, and image tracking.
[0086] The characteristics of SIFT features are as follows:
[0087] 1. It remains invariant to rotation, scaling, and brightness changes, and also exhibits a certain degree of stability to changes in viewing angle and noise.
[0088] 2. Uniqueness and rich information content make it suitable for fast and accurate matching in massive feature data;
[0089] 3. Abundance: Even a few objects can generate a large number of SIFT feature vectors;
[0090] 4. Scalability: It can be easily combined with other forms of feature vectors;
[0091] The essence of the SIFT algorithm is to find key points (feature points) in different scale spaces, calculate the size, orientation, and scale information of the key points, and use this information to construct key points to describe the feature points.
[0092] It should be noted that the feature point data in this embodiment includes a descriptor describing the image region surrounding the feature point. In this embodiment, the descriptor can be a 128-dimensional vector describing the image information of a 16*16 pixel region surrounding the feature point. That is, when matching feature points in two adjacent image frames, the descriptors of each feature point in the two image frames can be matched. When the descriptors match, it means that the two feature points have been successfully matched.
[0093] In step 203 above, feature matching is performed between each feature point of the current image frame and all feature points of the previous image frame to determine the feature points of the previous image frame that have successfully matched.
[0094] It should be noted that the feature matching method in this embodiment can perform matching based on the similarity of feature points or based on a fast nearest neighbor matching algorithm. In step 204 above, the matching relationship data and image data are sent to the terminal, enabling the terminal to determine the three-dimensional model of the target based on the matching relationship data and image data.
[0095] It should be noted that the matching relationship data includes the matched feature points, the feature point IDs of the feature points in the previous image frame, and the feature point data, while the image data includes the image data of the current image frame.
[0096] The feature matching of each target feature and the feature points of the previous frame to be processed is performed to obtain the matching relationship data. The matching relationship data and the image data of the current frame to be processed are sent to the terminal. The terminal receives the matching relationship data and the image data of the current frame to be processed and can use the terminal's computing resources to perform calculations to determine the three-dimensional model of the target.
[0097] The image processing method provided in this application acquires multiple image frames at equal time intervals, extracts features from each image frame to determine the feature points of the current image frame, and performs feature matching between each feature point and all feature points of the previous image frame to determine the feature points of the previous image frame that match the feature points. Finally, the matching relationship data and image data are sent to the terminal, enabling the terminal to determine the 3D model of the target based on the matching relationship data and image data. This application embodiment acquires target images through a smart wearable device, performs calculations using computing resources, and then sends the results to the terminal for further calculations. By having the terminal and the smart wearable device jointly complete the calculations, the limitations of computing power and power consumption on the smart wearable device can be effectively overcome, resulting in a highly detailed 3D model.
[0098] In some embodiments, it also includes:
[0099] For each of the aforementioned feature points, feature matching is performed with all feature points of the previous image frame to determine the Euclidean distance between each of the aforementioned feature points and all feature points of the previous image frame;
[0100] The feature points of the previous image frame are determined based on the Euclidean distances and the feature point matching.
[0101] Alternatively, based on the fast nearest neighbor matching algorithm, feature matching is performed on each of the aforementioned feature points and all feature points of the previous image frame to determine the feature points of the previous image frame that match the aforementioned feature points.
[0102] It is understood that the Euclidean distance between each of the above feature points and all feature points in the previous image frame, or when performing feature matching between each of the above feature points and all feature points in the previous image frame based on the fast nearest neighbor matching algorithm, can be represented by a descriptor.
[0103] This embodiment provides a specific method for feature matching.
[0104] During feature matching, it is necessary to determine the Euclidean distance between each feature point and all feature points in the previous image frame. We can denote the Euclidean distances from the current feature point 1 to feature points A, B, C, etc., in the previous image frame as x1, x2, x3, and so on. Then, we determine the similarity to the current feature point 1 based on these Euclidean distances. As mentioned above, if x1 < x2 < x3, then the Euclidean distance between feature point A and the current feature point 1 is the smallest, meaning the similarity is the largest. In other words, feature point A in the previous image frame is a successfully matched feature point with the current feature point 1.
[0105] In some scenarios, when the descriptor corresponding to a feature point has a high dimensionality, coupled with an increased number of feature points due to complex scenarios, a fast nearest neighbor matching algorithm can be used. Examples include randomized Kd-trees and K-means trees with priority search. In some embodiments, the algorithm may also include:
[0106] If the feature point is determined to be mismatched with all feature points in the previous image frame, the feature point is discarded.
[0107] If each feature point is determined to be mismatched with all feature points in the previous image frame, the current image frame is discarded.
[0108] Specifically, in this embodiment, during the feature matching process, there is a phenomenon where the feature points in the current image frame do not match all the feature points in the previous image frame. For example, if the Euclidean distance between each feature point and all the feature points in the previous image frame is greater than a preset value, that is, the similarity does not meet the preset value, then it can be determined that the feature point does not match all the feature points in the previous image frame, and the feature point can be discarded.
[0109] If every feature point in the current image frame does not match every feature point in the previous image frame, then the current image frame can be discarded.
[0110] It is understood that in this embodiment, choosing to discard the current image frame when discarding the previous image frame can be understood as using the next image frame as the current image frame and matching the feature points with the previous image frame.
[0111] In some embodiments, it also includes:
[0112] If the feature point is determined to be mismatched with all feature points in the previous image frame, the feature point is retained.
[0113] If each feature point is determined to be mismatched with all feature points of the previous image frame, the current image frame is retained, and the previous image frame is discarded.
[0114] Specifically, this embodiment differs from the above embodiments in that, when it is determined that the feature points of the current image frame do not match all feature points of the previous image frame, this embodiment retains the feature point and discards the feature points of the previous image frame.
[0115] Iterate through each feature point of the current image frame and match each feature point with the previous image frame. If no feature point of the current image frame matches any feature point of the previous image frame, discard the previous image frame.
[0116] It is understood that in this embodiment, choosing to discard the previous image frame when discarding the current image frame can be understood as matching feature points between the current processing frame and the image frame before the previous image frame.
[0117] Reference Figure 3 This application provides an image processing method according to one embodiment, applied to a terminal, including but not limited to the following steps:
[0118] Step 301: Receive matching relationship data and image data sent by the smart wearable device. The matching relationship data is determined by the smart wearable device based on feature matching of each feature point in the current image frame and all feature points in the previous image frame. The matching relationship data includes the matched feature points, the feature point IDs and feature point data of the feature points in the previous image frame, and the image data includes the image data of the current image frame. The feature point data of the feature points includes a descriptor describing the image region surrounding the feature point.
[0119] Step 302: Based on the matching relationship data, the current image frame, and the previous image frame, determine the camera pose and three-dimensional spatial points of the current image frame;
[0120] Step 303: Perform BA optimization on all the determined camera poses and 3D spatial points;
[0121] Step 304: Determine that no new image frame has been received, perform BA optimization on all the determined camera poses and 3D spatial points to obtain the final camera poses and 3D spatial points.
[0122] Step 305: Based on all received image frames, the final camera pose, camera intrinsic parameters, and three-dimensional spatial points of all image frames, determine the three-dimensional model.
[0123] First, it should be noted that the image processing method provided in this embodiment is applied to a terminal. The terminal and the smart wearable device can establish communication and transmit data. The data processing of the terminal and the smart wearable device can be carried out in parallel to improve the overall generation efficiency of the stereoscopic image. The above communication can be established through the aforementioned wireless communication technology or wired communication technology, which is not a limitation of this application.
[0124] In step 301 above, the terminal receives image data of all frames to be processed sent by the smart wearable device, as well as matching relationship data between the current image frame and the previous image frame. It should be noted that the image frame data received by the terminal from the smart wearable device can be sent directly by the smart wearable device or relayed, for example, first sent by the smart wearable device to a cloud server, and then sent to the terminal via the cloud server. It should be explained that in this embodiment, the terminal receives multiple consecutive image frames sent by the smart wearable device, each image frame corresponding to one set of image data. The image data of all image frames can be understood as follows: each image frame is sequentially used as the current frame to be processed, and the image data corresponding to each frame is sent to the terminal for processing. In this way, the terminal can receive image data of consecutive frames and thus build a 3D model.
[0125] In this embodiment, the current frame to be processed is selected and determined in real time by the smart wearable device based on the aforementioned equal interval strategy. The previous frame to be processed refers to the frame preceding the current frame to be processed, which will not be elaborated here.
[0126] The matching relationship data is determined by the smart wearable device through feature matching between each feature point in the current image frame and all feature points in the previous image frame. The matching relationship data includes matched feature points, feature point IDs and feature point data from the previous image frame, and image data including the image data of the current image frame. The feature point data includes a descriptor describing the image region surrounding the feature point. In this embodiment, the descriptor can be a 128-dimensional vector describing image information of a 16*16 pixel region surrounding the feature point.
[0127] In step 302 above, the terminal uses computing resources, specifically the computing resources of the mobile phone, to determine the camera pose and three-dimensional spatial points of the current image frame based on the matching relationship data sent by the smart wearable device, the current image frame, and the previous image frame.
[0128] It should be noted that the camera pose of the current image frame is the camera's pose matrix in the current state, and the three-dimensional spatial point is the three-dimensional spatial point generated by triangulation of the camera pose of the current image frame and the camera pose of the previous image frame.
[0129] Secondly, in step 303 above, BA (Bundle Adjustment) optimization is performed on the camera pose and 3D spatial points. BA aims to improve the accuracy of the camera pose and 3D spatial points to obtain the optimal camera pose and 3D spatial points for generating the final 3D image. It refers to extracting the optimal 3D model and camera parameters from the visual image as optimization data for each frame to be processed. Further, in step 304 above, when no new image frames are received, i.e., when the image acquisition by the smart wearable device ends, global optimization is performed on all camera poses and 3D spatial points to obtain the final camera pose and 3D spatial points.
[0130] In some examples, step 303 also includes:
[0131] Alternatively, perform local BA optimization on the camera pose and the three-dimensional spatial points of the image frames within the first frame sequence;
[0132] The first frame sequence mentioned above is a frame sequence that includes at least the current image frame and the previous image frame. It can be understood that the first frame sequence may also include image frames before the previous image frame, such as two or three consecutive image frames before the previous image frame. The number of image frames in the first frame sequence is not a limitation of this application.
[0133] In some examples, step 304 also includes:
[0134] Alternatively, global BA optimization can be performed on the camera pose and the three-dimensional spatial points of the image frames in the second frame sequence to obtain the final camera pose and three-dimensional spatial points.
[0135] For example, a keyframe is determined from each first frame sequence. The second frame sequence can be a frame sequence composed of all keyframes. After determining that no new image frame has been received, global BA optimization is performed on all keyframes in the second frame sequence to determine the camera pose and 3D spatial points after global BA optimization of the keyframes. These, together with the camera pose and 3D spatial points of the image frames that are not keyframes, constitute the final camera pose and 3D spatial points.
[0136] It should be noted that local optimization usually includes optimizing the frame sequence, including the current image frame, while global optimization optimizes the key frames of each frame sequence to obtain the final camera pose and 3D spatial points. In this embodiment, the final 3D spatial points of all image frames form a 3D sparse point cloud.
[0137] It is understood that, in the embodiments of this application, BA optimization can be performed on all image frames with generated 3D points after receiving the current image frame, or local BA optimization can be performed on some image frames (i.e., image frames containing the first frame sequence content) including the current image frame after receiving the current image frame.
[0138] In this embodiment of the application, when it is determined that no new image frame has been received, BA optimization is performed on all image frames with generated 3D points. Global BA optimization can also be performed on image frames in the second frame sequence composed of key frames determined by the first frame sequence, so as to obtain the final camera pose and 3D spatial points of all image frames.
[0139] Finally, based on all received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames, the 3D model is determined, thus obtaining the final 3D image.
[0140] In this embodiment, a three-dimensional image can be obtained by performing three-dimensional imaging based on the image data of the image frame.
[0141] Optionally, multi-view rendered videos or 3D data files can be output based on the 3D model.
[0142] In some embodiments, the method further includes: inputting all received image frames, the final camera pose, camera intrinsic parameters, and three-dimensional spatial points of all image frames into a target algorithm, and outputting the three-dimensional model, wherein the target algorithm is an MVS algorithm or a NeRF algorithm.
[0143] Specifically, all received image frames, their final camera poses, camera intrinsic parameters, and 3D spatial points are input into the MVS or NeRF algorithm to obtain a 3D model. In this embodiment, the 3D spatial points are a 3D sparse point cloud. By using the MVS or NeRF algorithm, a 3D dense point cloud can be output, thus forming a 3D image.
[0144] It's worth noting that the MVS (Multi-view Stereo) algorithm can construct highly detailed 3D models from images independently, acquiring a large image dataset to build a 3D geometric model for image parsing. The NeRF (Neural Radiation Field) algorithm can generate new views of complex 3D scenes based on a partial 2D image set. In the NeRF neural network model, the input view of the scene is trained using a rendering loss to recreate the scene. A complete scene is rendered by acquiring representative input images of the scene and interpolating between them. Therefore, the NeRF algorithm is an efficient method for generating images from synthetic data.
[0145] In some embodiments, it also includes:
[0146] The camera pose of the current image frame is set as an identity matrix, and the target matrix is determined based on the identity matrix and the matching relationship data.
[0147] The target matrix is decomposed to obtain the camera pose of the current image frame;
[0148] The three-dimensional spatial points are determined by triangulation based on the camera pose of the current image frame, the camera pose of the previous image frame, and the matching relationship data.
[0149] It is understood that this embodiment describes the process of generating points in three-dimensional space.
[0150] Specifically, this embodiment requires at least two temporally adjacent matching frames, namely the current image frame and the previous image frame. The camera pose of the first frame (i.e., the previous image frame) is set as the identity matrix. Then, the target matrix is estimated using the matching feature points between the first and second frames (the current image frame). The camera pose of the second frame image is obtained by decomposing the target matrix.
[0151] After estimating the poses of the first two frames, three-dimensional spatial points are generated through triangulation.
[0152] In this embodiment, for the image data of the current image frame that is continuously added later, 2D-2D matching is performed between the image data of the current image frame and the image data of the previous image frame to obtain the 2D-3D matching between the two. Then, RANSAC-PNP is used to calculate the camera pose of the current frame. Then, triangulation is used to generate three-dimensional spatial points in the current frame, and the newly added three-dimensional spatial points are added to the historical data.
[0153] For example, if the previous image frame is denoted as Frame1, the current image frame as Frame2, the next image frame as Frame3, and so on.
[0154] First, perform 2D-2D calculations between Frame1 and Frame2 to obtain the camera pose R and unit length t of the current frame, and then calculate the 3D points through triangulation.
[0155] Then, 3D-2D calculations are performed between Frame2 and Frame3 to obtain the camera pose R of the current frame and t1 with a length of t. Then, 3D points are calculated by triangulation.
[0156] Through continuous iteration, the camera pose and 3D spatial points of the current image frame can be continuously added to the historical data, so that the camera pose can be confirmed and the 3D spatial points can be generated in the next image frame.
[0157] The image processing method provided in this application embodiment receives image data of all image frames sent by a smart wearable device, as well as matching relationship data between the current image frame and the previous image frame, through a terminal. Based on the matching relationship data, matching feature points in the image data of the current image frame are determined. Then, based on the matching feature points, the image data of the current image frame, and the image data of the previous image frame, the camera pose and 3D spatial points of the current image frame are determined. Finally, based on all received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames, a 3D model is determined. This application embodiment can utilize the computing resources of the smart wearable device and the terminal computing resources collaboratively to achieve effective computing power allocation, resulting in minimal computational latency and the most refined 3D image reconstruction effect possible.
[0158] Furthermore, by optimizing the camera pose and 3D spatial points, optimized data for each image frame is obtained to generate a 3D image. In other words, optimizing existing data further improves the generation quality of the 3D image, enhancing the user experience.
[0159] Reference Figure 4 , Figure 4 This is a flowchart illustrating an image processing method provided in one embodiment of this application, including the following steps:
[0160] Step 401: Start shooting;
[0161] Step 402: The camera of the smart wearable device begins shooting;
[0162] Step 403: Collect the current frame to be processed in real time at equal intervals;
[0163] Step 404: Extract image features using local resources of the smart wearable device;
[0164] Step 405: Calculate the matching relationship between the current frame to be processed and the previous frame to be processed using the local computing resources of the smart wearable device;
[0165] Step 406: Upload the matching relationship data and the compressed data of the current image frame to the terminal connected to the smart wearable device;
[0166] Step 407: Perform camera pose estimation for the current frame to be processed on the terminal;
[0167] Step 408: Triangulate the current frame to be processed on the terminal to generate three-dimensional spatial points;
[0168] Step 409: Perform BA optimization on all generated 3D spatial points and estimated poses at the terminal;
[0169] Step 410: Perform Business Optimization (BA) on all data in the terminal;
[0170] Step 411: Obtain the target pose, camera intrinsic parameters, and 3D sparse point cloud of all frames to be processed;
[0171] Step 412: The terminal uses the MVS algorithm or NeRF algorithm to calculate the three-dimensional model;
[0172] Step 413: Output a 3D image including the 3D model.
[0173] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute an image processing method, which includes:
[0174] Acquire multiple image frames at equal time intervals;
[0175] For each of the image frames, feature extraction is performed to determine the feature points of the current image frame, wherein the feature point data of the feature points includes a descriptor describing the image region surrounding the feature points;
[0176] Each of the aforementioned feature points is matched with all feature points of the previous image frame to determine the feature points of the previous image frame that match the aforementioned feature points.
[0177] The matching relationship data and image data are sent to the terminal, so that the terminal determines the three-dimensional model of the target based on the matching relationship data and the image data. The matching relationship data includes the matched feature points, the feature point ID and feature point data of the feature points of the previous image frame, and the image data includes the image data of the current image frame.
[0178] or,
[0179] The system receives matching relationship data and image data sent by a smart wearable device. The matching relationship data is determined by the smart wearable device based on feature matching of each feature point in the current image frame and all feature points in the previous image frame. The matching relationship data includes the matched feature points, the feature point IDs of the feature points in the previous image frame, and the feature point data. The image data includes the image data of the current image frame, and the feature point data of the feature points includes a descriptor describing the image region surrounding the feature point.
[0180] Based on the matching relationship data, the current image frame, and the previous image frame, the camera pose and three-dimensional spatial points of the current image frame are determined;
[0181] Perform BA optimization on all the determined camera poses and the three-dimensional space points;
[0182] If no new image frame is received, perform BA optimization on all the determined camera poses and 3D spatial points to obtain the final camera poses and 3D spatial points.
[0183] Based on all received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames, a 3D model is determined.
[0184] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0185] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to perform the image processing methods provided by the above methods, the method including:
[0186] Acquire multiple image frames at equal time intervals;
[0187] For each of the image frames, feature extraction is performed to determine the feature points of the current image frame, wherein the feature point data of the feature points includes a descriptor describing the image region surrounding the feature points;
[0188] Each of the aforementioned feature points is matched with all feature points of the previous image frame to determine the feature points of the previous image frame that match the aforementioned feature points.
[0189] The matching relationship data and image data are sent to the terminal, so that the terminal determines the three-dimensional model of the target based on the matching relationship data and the image data. The matching relationship data includes the matched feature points, the feature point ID and feature point data of the feature points of the previous image frame, and the image data includes the image data of the current image frame.
[0190] or,
[0191] The system receives matching relationship data and image data sent by a smart wearable device. The matching relationship data is determined by the smart wearable device based on feature matching of each feature point in the current image frame and all feature points in the previous image frame. The matching relationship data includes the matched feature points, the feature point IDs of the feature points in the previous image frame, and the feature point data. The image data includes the image data of the current image frame, and the feature point data of the feature points includes a descriptor describing the image region surrounding the feature point.
[0192] Based on the matching relationship data, the current image frame, and the previous image frame, the camera pose and three-dimensional spatial points of the current image frame are determined;
[0193] Perform BA optimization on all the determined camera poses and the three-dimensional space points;
[0194] If no new image frame is received, perform BA optimization on all the determined camera poses and 3D spatial points to obtain the final camera poses and 3D spatial points.
[0195] Based on all received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames, a 3D model is determined.
[0196] In another aspect, this application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the image processing methods provided by the methods described above, the method comprising:
[0197] Acquire multiple image frames at equal time intervals;
[0198] For each of the image frames, feature extraction is performed to determine the feature points of the current image frame, wherein the feature point data of the feature points includes a descriptor describing the image region surrounding the feature points;
[0199] Each of the aforementioned feature points is matched with all feature points of the previous image frame to determine the feature points of the previous image frame that match the aforementioned feature points.
[0200] The matching relationship data and image data are sent to the terminal, so that the terminal determines the three-dimensional model of the target based on the matching relationship data and the image data. The matching relationship data includes the matched feature points, the feature point ID and feature point data of the feature points of the previous image frame, and the image data includes the image data of the current image frame.
[0201] or,
[0202] The system receives matching relationship data and image data sent by a smart wearable device. The matching relationship data is determined by the smart wearable device based on feature matching of each feature point in the current image frame and all feature points in the previous image frame. The matching relationship data includes the matched feature points, the feature point IDs of the feature points in the previous image frame, and the feature point data. The image data includes the image data of the current image frame, and the feature point data of the feature points includes a descriptor describing the image region surrounding the feature point.
[0203] Based on the matching relationship data, the current image frame, and the previous image frame, the camera pose and three-dimensional spatial points of the current image frame are determined;
[0204] Perform BA optimization on all the determined camera poses and the three-dimensional space points;
[0205] If no new image frame is received, perform BA optimization on all the determined camera poses and 3D spatial points to obtain the final camera poses and 3D spatial points.
[0206] Based on all received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames, a 3D model is determined.
[0207] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0208] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0209] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An image processing method, characterized in that, Applied to terminals, including: The system receives matching relationship data and image data sent by a smart wearable device. The matching relationship data is determined by the smart wearable device based on feature matching of each feature point in the current image frame and all feature points in the previous image frame. The matching relationship data includes the matched feature points, the feature point IDs of the feature points in the previous image frame, and the feature point data. The image data includes the image data of the current image frame, and the feature point data of the feature points includes a descriptor describing the image region surrounding the feature point. Based on the matching relationship data, the current image frame, and the previous image frame, the camera pose and three-dimensional spatial points of the current image frame are determined; The camera pose of the current image frame is set as an identity matrix, and the target matrix is determined based on the identity matrix and the matching relationship data. The target matrix is decomposed to obtain the camera pose of the current image frame; Triangulation is performed based on the camera pose of the current image frame, the camera pose of the previous image frame, and the matching relationship data to determine the three-dimensional spatial points; For the image data of the current image frame that is continuously added, 2D-2D matching is performed between the image data of the current image frame and the image data of the previous image frame to obtain the 2D-3D matching between the two. Then, the camera pose of the current frame is calculated using RANSAC-PNP. Then, triangulation is used to generate three-dimensional spatial points in the current frame, and the newly added three-dimensional spatial points are added to the historical data. Perform BA optimization on all the determined camera poses and the three-dimensional space points; If no new image frame is received, perform BA optimization on all the determined camera poses and 3D spatial points to obtain the final camera poses and 3D spatial points. Based on all received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames, a 3D model is determined.
2. The image processing method according to claim 1, characterized in that, Also includes: All received image frames, the final camera pose, camera intrinsic parameters, and 3D spatial points of all image frames are input into the target algorithm to output the 3D model. The target algorithm is either the MVS algorithm or the NeRF algorithm.
3. The image processing method according to any one of claims 1-2, characterized in that, Also includes: Based on the 3D model, output multi-view rendered videos or 3D data files.
4. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the image processing method as described in any one of claims 1 to 3.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image processing method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
AR / VR scene map collection method
CN110855601A
Peripheral map establishing method and device of mobile equipment, equipment and storage medium
CN115239902A