Method and device for coding multi-view image
Through three-dimensional reconstruction technology, the image coding process is automated, and the problems of cumbersome and inconsistent manual coding in the existing technology are solved, and efficient and automated multi-view image batch coding is achieved.
Patent Information
- Application Number
- CN202311465967.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, image coding is usually done manually, resulting in large workloads, repetitive and cumbersome workloads, and the consistency of coding areas in different pictures cannot be guaranteed.
Through a three-dimensional reconstruction method, in response to the object image coding request, the three-dimensional model is reconstructed, and the coding area model is selected from the three-dimensional model, and it is projected into a multi-view image. Each image obtains the corresponding coding area and performs batch coding processing.
It effectively reduces manual participation, reduces the workload of manual processing, realizes batch coding processing at the same location of multiple images, and ensures the consistency of coding areas in different images.
Smart Images

Figure CN119942035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method and device for coding multi-view images. Background Art
[0002] In order to better display item information, multi-view images are generally used for item display. In the displayed item images, some sensitive information that is inconvenient to expose (such as production date, barcode, QR code, etc.) needs to be coded, which is usually manually processed by art designers using image processing software (for example, PhotoShop). In multi-view image display, there are often many pictures that need to be processed, so they need to be processed manually one by one.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0004] Image coding is usually done manually. The workload of manually processing multiple images is large and the work is repetitive and tedious. In multi-view images, it is necessary to manually determine whether each image needs to be coded, and the consistency of the coded areas in different images cannot be guaranteed. Summary of the invention
[0005] In view of this, the embodiments of the present invention provide a method and device for coding multi-view images, which can effectively reduce manual participation, reduce the workload of manual processing, realize batch coding processing of the same position of multiple images, and ensure the consistency of coding areas in different images.
[0006] To achieve the above object, according to one aspect of an embodiment of the present invention, a method for coding a multi-view image is provided, comprising:
[0007] In response to a request for coding an object image, performing three-dimensional reconstruction based on multi-view images of the object to obtain a three-dimensional model of the object;
[0008] In response to receiving the coding area model selected from the three-dimensional model, projecting the coding area model onto each image in the multi-view images to obtain the coding area corresponding to each image;
[0009] The coding area of each image is subjected to coding processing to code the multi-view images.
[0010] Optionally, performing three-dimensional reconstruction based on multi-view images of the object to obtain a three-dimensional model of the object includes: using a neural radiation field algorithm to perform three-dimensional reconstruction based on multi-view images of the object to obtain the three-dimensional model of the object.
[0011] Optionally, three-dimensional reconstruction is performed based on multi-view images of the object to obtain a three-dimensional model of the object, including: obtaining at least one key pixel point in the multi-view images of the object; obtaining a camera position, camera direction and actual color value corresponding to each key pixel point; inputting the camera position, camera direction and actual color value corresponding to each key pixel point into a pre-trained three-dimensional reconstruction model to obtain the coordinates of points on the model surface; and constructing the three-dimensional model of the object according to the coordinates of the points on the model surface.
[0012] Optionally, the three-dimensional reconstruction model is trained in the following manner: for each object in the training set, obtaining at least one key pixel point in the multi-view image of the object; using the camera position and camera direction corresponding to each key pixel point as the input of the initial model, and calculating the color value corresponding to each key pixel point; for each key pixel point, comparing the color value corresponding to the key pixel point with the actual color value for supervised learning to obtain the three-dimensional reconstruction model.
[0013] Optionally, a signed distance function is used in the three-dimensional reconstruction model to define the positional relationship between a point in the three-dimensional space and the model surface, and the point on the model surface is a point in the three-dimensional space whose closest distance to the model surface is 0.
[0014] Optionally, the coding area model is projected onto each image in the multi-view images to obtain the coding area corresponding to each image, including: for each image in the multi-view images, obtaining a projection matrix according to a camera intrinsic parameter matrix corresponding to the image; and calculating the coordinates of the coding area corresponding to the image according to the projection matrix, the posture matrix corresponding to the image, and the matrix corresponding to the coding area model.
[0015] Optionally, the coding area of each image is coded to code the multi-view image, including: down-sampling each image to obtain a down-sampled image of each image; mean blurring the down-sampled image to obtain a blurred image; up-sampling the blurred image to code each image.
[0016] According to another aspect of an embodiment of the present invention, there is provided a device for coding a multi-view image, comprising:
[0017] A three-dimensional model reconstruction module, configured to respond to an object image coding request and perform three-dimensional reconstruction based on multi-view images of the object to obtain a three-dimensional model of the object;
[0018] a coding area acquisition module, configured to, in response to receiving a coding area model selected from the three-dimensional model, project the coding area model onto each image in the multi-view images to obtain a coding area corresponding to each image;
[0019] The coding processing module is used to perform coding processing on the coding area of each image to code the multi-view images.
[0020] According to another aspect of an embodiment of the present invention, there is provided an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for coding multi-view images provided in an embodiment of the present invention.
[0021] According to another aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for coding a multi-view image provided by an embodiment of the present invention is implemented.
[0022] An embodiment of the above invention has the following advantages or beneficial effects: by responding to a coding request for an object image, a three-dimensional model of the object is obtained by performing three-dimensional reconstruction based on the multi-view images of the object; in response to receiving a coding area model selected from the three-dimensional model, the coding area model is projected onto each image in the multi-view images to obtain a coding area corresponding to each image; coding the coding area of each image to code the multi-view images can effectively reduce manual participation, reduce the workload of manual processing, realize batch coding processing of the same position of multiple images, and ensure the consistency of coding areas in different images.
[0023] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.
[0025] Figure 1 is a schematic diagram of main steps of a method for coding a multi-view image according to an embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of a three-dimensional reconstruction effect of a multi-view image according to an embodiment of the present invention;
[0027] Figure 3 It is a schematic diagram of the processing process of the coding area model according to one embodiment of the present invention;
[0028] Figure 4This is a schematic diagram of a coding area projection process according to an embodiment of the present invention;
[0029] Figure 5 This is a schematic diagram of an image coding result according to an embodiment of the present invention;
[0030] Figure 6 is a schematic diagram of main modules of a device for coding a multi-view image according to an embodiment of the present invention;
[0031] Figure 7 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;
[0032] Figure 8 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0033] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0034] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, storage and other aspects of user personal information involved in the technical solution disclosed in the present invention are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information, network security and national security.
[0035] In order to solve the technical problems existing in the prior art, the present invention proposes a method for batch coding of multi-view images based on three-dimensional reconstruction, which realizes batch coding processing of the same position of multiple images of the same object from different viewpoints through an algorithm. Thus, the following technical effects are achieved: reducing manual repetitive operations; no need to manually judge whether each image needs coding processing; ensuring that the coding area of each image is consistent.
[0036] The present invention proposes a method for batch coding of multi-view images based on three-dimensional reconstruction. The overall process includes: first, three-dimensional reconstruction of the object is performed based on images of multiple viewpoints; second, the area to be coded is selected in the reconstructed model, and only the three-dimensional model corresponding to the coding area is retained; then, the three-dimensional model of the coding area is projected to each image to obtain the coding area of each image; finally, the selected coding area is coded.
[0037] Figure 1 FIG. 1 is a schematic diagram of the main steps of a method for coding a multi-view image according to an embodiment of the present invention. Figure 1 As shown, the method for coding a multi-view image according to an embodiment of the present invention mainly includes the following steps S101 to S103.
[0038] Step S101: In response to an object image coding request, a 3D reconstruction is performed based on the multi-view images of the object to obtain a 3D model of the object. According to an embodiment of the present invention, when reconstructing the 3D model of the object, a traditional algorithm based on SFM (Structure from motion, a 3D reconstruction method) or MVS (Multi View Stereo, a 3D reconstruction method of multi-view stereo) can be used for 3D reconstruction, or other methods can be used.
[0039] However, traditional 3D reconstruction algorithms cannot obtain relatively good 3D models for objects with weak textures, repeated textures, and complex structures, so the 3D reconstruction effect of many objects is poor.
[0040] According to an embodiment of the present invention, when performing 3D reconstruction based on multi-view images of an object to obtain a 3D model of the object, it may specifically include: using a neural radiation field algorithm to perform 3D reconstruction based on multi-view images of the object to obtain a 3D model of the object. The method used in the present invention is to perform 3D reconstruction of an object based on a neural radiation field, and can reconstruct a high-quality 3D model of an object with weak texture, repeated texture, or complex structure.
[0041] According to an embodiment of the present invention, when a three-dimensional model of the object is obtained by three-dimensional reconstruction based on a multi-view image of the object, the method may specifically include: obtaining at least one key pixel in the multi-view image of the object; obtaining the camera position, camera direction and actual color value corresponding to each key pixel; inputting the camera position, camera direction and actual color value corresponding to each key pixel into a pre-trained three-dimensional reconstruction model to obtain the coordinates of a point on the surface of the model; and constructing the three-dimensional model of the object according to the coordinates of the point on the surface of the model. Among them, the key pixel of the image can be pre-set when the image is taken. For example, when the image is taken, a virtual frame can be generated on the photo interface through the Simultaneous Localization and Mapping (SLAM) technology to select the projection position point of the key position of the object, and then take the photo, so as to avoid focusing on the background of the object, so as to better determine the key pixel. By performing three-dimensional reconstruction of the model only based on the key pixel in the image, the time and computing resources of the three-dimensional reconstruction of the model can be greatly saved, and the coding position selection based on the key pixel can be better performed. Parameters such as the camera position and camera direction corresponding to each key pixel point can be associated with the image and saved when the image is captured, and the actual color value corresponding to each key pixel point can be obtained from the image.
[0042] According to an embodiment of the present invention, the three-dimensional reconstruction model is obtained by training in the following manner: for each object in the training set, obtain at least one key pixel in the multi-view image of the object; use the camera position and camera direction corresponding to each key pixel as the input of the initial model, and calculate the color value corresponding to each key pixel; for each key pixel, compare the color value corresponding to the key pixel with the actual color value for supervised learning to obtain the three-dimensional reconstruction model. In the three-dimensional reconstruction model, a signed distance function is used to define the positional relationship between a point in three-dimensional space and a model surface, and a point on the model surface is a point in three-dimensional space whose closest distance to the model surface is 0.
[0043] The following describes the training process of the 3D reconstruction model according to an embodiment of the present invention in conjunction with an embodiment.
[0044] The 3D reconstruction based on neural radiation field is essentially to use neural network to represent the 3D model, obtain the image through neural network rendering, and then use the input image for supervision. In this 3D reconstruction model, SDF (sign distance function) is used to describe the surface of the 3D model.
[0045] Before model training, an initial model must first be established. In an embodiment of the present invention, the initial model includes two functions defined as follows:
[0046] (1) The signed distance function value from a point in space to a surface Input a point in space, and output the shortest distance from the point to a surface (which can be a curved surface). The sign is positive outside the surface and negative inside.
[0047] (2) The color value from a point in space to a certain viewing angle Its input is the camera position and camera orientation Output as color information Among them, x, y, z are the three-dimensional coordinate values of the camera, and θ is the rotation angle of the camera. is the pitch angle of the camera.
[0048] Both functions are represented by neural networks. These two functions are interrelated. The positional relationship between a point in space and the model surface directly determines the actual color value of the point in the image. In theory, if the point is on the model surface, the color information c calculated for the point should be equal to the actual color value of the point.
[0049] The model surface of the 3D model to be reconstructed can be represented as the zero-level set of SDF, that is, the point corresponding to f = 0. The core idea is to use volume rendering to train the SDF network. The larger the S density value is, the greater the S density value is. Input a position (x, y, z) in space and output an SDF value corresponding to the S density.
[0050] When training a 3D reconstruction model, you can select a group of key pixels each time based on the key areas or key pixels of the object when the image was taken, use the camera position and camera direction corresponding to these key pixels when shooting as the input of the initial model, calculate the corresponding color values, compare them with the actual color values of these key pixels in the actual image, supervise the network, and use L1 (L1_loss, mean absolute error) as the loss function to train the 3D reconstruction model. When a certain number of iterations is reached, the training is completed and the 3D reconstruction model is obtained. Then the points corresponding to f=0 can be extracted from the 3D reconstruction model as the surface of the 3D model, thereby obtaining the corresponding 3D model. Figure 2 As shown, Figure 2 The figure is a schematic diagram of the effect of three-dimensional reconstruction of multi-view images according to an embodiment of the present invention. By performing three-dimensional reconstruction on multi-angle images of an object, a three-dimensional model of the object is obtained.
[0051] Step S102: In response to receiving the coding area model selected from the three-dimensional model, the coding area model is projected onto each image in the multi-view image to obtain the coding area corresponding to each image. After the object is three-dimensionally reconstructed to obtain the three-dimensional model, the user can manually select the area to be coded from the three-dimensional model, and delete the remaining unselected areas from the three-dimensional model to obtain the coding area model.
[0052] Figure 3 FIG. 1 is a schematic diagram of the processing process of the coding area model according to an embodiment of the present invention. Figure 3 As shown, after selecting the coding area from the 3D model on the left, the models corresponding to other unselected areas are deleted to obtain the coding area model.
[0053] Afterwards, the coding area model is projected into each image in the multi-view images to determine the coding area corresponding to each image.
[0054] According to one embodiment of the present invention, when the coding area model is projected onto each image in the multi-view image to obtain the coding area corresponding to each image, it specifically includes: for each image in the multi-view image, a projection matrix is obtained according to the camera intrinsic parameter matrix corresponding to the image; according to the projection matrix, the posture matrix corresponding to the image and the matrix corresponding to the coding area model, the coordinates of the coding area corresponding to the image are calculated. When the image is captured, the camera's intrinsic parameter matrix, extrinsic parameter matrix, posture matrix and other information at the time of shooting will be saved synchronously for use in subsequent processing, and the camera's intrinsic parameter matrix, posture matrix and other information can also be obtained by post-processing the image (for example: through a camera calibration method, a posture calculation method, etc.). Among them, the camera's posture matrix and the camera position and camera direction at the time of image capture can be converted to each other, and the specific conversion method is not described in detail in this invention.
[0055] The projection process of projecting the coding area model onto each image in the multi-view image is consistent. Here, the projection process is introduced by taking one image as an example.
[0056] Suppose the camera intrinsic parameter matrix corresponding to a certain image is Among them, f x and f y represents the focal length of the camera, c x and c y represents the principal point coordinates of the camera; the attitude matrix is Among them, r0, r1, ..., r8 are the rotation matrix elements of the camera in space, which can be obtained according to the camera position and camera direction conversion, t0, t1, t2 are the positions of the camera in space; the reconstructed coding area model is recorded as M (essentially a series of three-dimensional coordinates), and in the embodiment of the present invention, OpenGL can be used for rendering to project the three-dimensional coding area model onto a two-dimensional image. Among them, the projection matrix P can be obtained from the camera intrinsic parameter matrix:
[0057]
[0058] Among them, width and height are the width and height of the image respectively, far is the far plane, and near is the near plane. Since the imaging space is in the shape of a trumpet, the area between the near plane and the far plane is the valid area. Therefore, the far plane and the near plane are the two planes selected to limit the effective projection area. The near plane is closer to the camera, and the far plane is farther away from the camera.
[0059] Finally, according to the projection matrix P, the posture matrix V corresponding to the image, and the matrix M corresponding to the coding area model, the coordinates of the coding area corresponding to the image can be calculated. Specifically, it can be obtained by calculating P×V×M. By calculating V×M, the result of projecting the coding area model to the camera space can be obtained, and then by multiplying P with the result, the result of projecting the coding area model to the camera space and then projecting it to the image space can be obtained, that is, the coordinates of the coding area corresponding to the image.
[0060] Figure 4 FIG. 1 is a schematic diagram of the coding area projection process according to an embodiment of the present invention. Figure 4 As shown, after performing the above-mentioned projection calculation based on the coding area model and the multi-angle images, the coding area corresponding to each image can be obtained.
[0061] Step S103: coding the coding area of each image to code the multi-view image. Coding the coding area may be performed in various forms such as blurring the coding area, adding mosaic effect, etc.
[0062] In one embodiment of the present invention, the coding area of each image is coded to code the multi-view image, which may specifically include: downsampling each image to obtain a downsampled image of each image; performing mean blurring on the downsampled image to obtain a blurred image; upsampling the blurred image to code each image. The image can be scaled by downsampling the image; then the downsampled image is fuzzy by mean blurring the scaled image; and then the blurred image is upsampled to restore the image to its original resolution, thereby improving the image coding efficiency.
[0063] Figure 5 FIG. 1 is a schematic diagram of an image coding result according to an embodiment of the present invention. Figure 5 As shown, the result obtained after coding an image is shown.
[0064] According to the above steps S101 to S103, firstly, a three-dimensional reconstruction is performed based on the multi-view images of the object to obtain a three-dimensional model of the object. Secondly, the area to be coded is selected in the reconstructed model, and only the three-dimensional model corresponding to the coding area is retained; then, the coding area model is projected to each picture to obtain the coding area of each picture; finally, the selected area is coded. This image coding method effectively reduces the manual participation and the workload of manual processing, realizes batch coding processing of the same position of multiple images, and can ensure the consistency of coding areas in different pictures.
[0065] Figure 6 2 is a schematic diagram of the main modules of the device for coding a multi-view image according to an embodiment of the present invention. Figure 6 As shown, the device 600 for coding a multi-view image according to an embodiment of the present invention mainly includes a three-dimensional model reconstruction module 601 , a coding area acquisition module 602 and a coding processing module 603 .
[0066] A three-dimensional model reconstruction module 601 is used to perform three-dimensional reconstruction based on multi-view images of the object to obtain a three-dimensional model of the object in response to an object image coding request;
[0067] A coding area acquisition module 602 is used for, in response to receiving a coding area model selected from the three-dimensional model, projecting the coding area model onto each image in the multi-view image to obtain a coding area corresponding to each image;
[0068] The coding processing module 603 is used to perform coding processing on the coding area of each image to code the multi-view images.
[0069] According to an embodiment of the present invention, the three-dimensional model reconstruction module 601 may also be used to: use a neural radiation field algorithm to perform three-dimensional reconstruction based on multi-view images of the object to obtain a three-dimensional model of the object.
[0070] According to another embodiment of the present invention, the three-dimensional model reconstruction module 601 can also be used to: obtain at least one key pixel point in the multi-view image of the object; obtain the camera position, camera direction and actual color value corresponding to each key pixel point; input the camera position, camera direction and actual color value corresponding to each key pixel point into a pre-trained three-dimensional reconstruction model to obtain the coordinates of the points on the model surface; and construct a three-dimensional model of the object according to the coordinates of the points on the model surface.
[0071] According to another embodiment of the present invention, the three-dimensional reconstruction model is obtained by training in the following manner: for each object in the training set, obtaining at least one key pixel point in the multi-view image of the object; using the camera position and camera direction corresponding to each of the key pixel points as inputs of the initial model, and calculating the color value corresponding to each of the key pixel points; for each of the key pixel points, comparing the color value corresponding to the key pixel point with the actual color value for supervised learning to obtain the three-dimensional reconstruction model.
[0072] According to another embodiment of the present invention, the three-dimensional reconstruction model uses a signed distance function to define the positional relationship between a point in the three-dimensional space and the model surface, and the point on the model surface is a point in the three-dimensional space whose closest distance to the model surface is 0.
[0073] According to another embodiment of the present invention, the coding area acquisition module 602 can also be used to: for each image in the multi-view image, obtain a projection matrix according to the camera intrinsic parameter matrix corresponding to the image; and calculate the coordinates of the coding area corresponding to the image according to the projection matrix, the posture matrix corresponding to the image and the matrix corresponding to the coding area model.
[0074] According to another embodiment of the present invention, the coding processing module 603 can also be used to: downsample each image to obtain a downsampled image of each image; perform mean blur processing on the downsampled image to obtain a blurred image; upsample the blurred image to perform coding processing on each image.
[0075] According to the technical solution of the embodiment of the present invention, by responding to a coding request for an object image, three-dimensional reconstruction is performed based on the multi-view images of the object to obtain a three-dimensional model of the object; in response to receiving a coding area model selected from the three-dimensional model, the coding area model is projected to each image in the multi-view images to obtain a coding area corresponding to each image; and coding is performed on the coding area of each image to code the multi-view images, which can effectively reduce manual participation, reduce the workload of manual processing, realize batch coding processing of the same position of multiple images, and ensure the consistency of coding areas in different images.
[0076] Figure 7 An exemplary system architecture 700 is shown to which the method for coding a multi-view image or the apparatus for coding a multi-view image according to the embodiments of the present invention can be applied.
[0077] like Figure 7 As shown, system architecture 700 may include terminal devices 701, 702, 703, network 704 and server 705. Network 704 is used to provide a medium for communication links between terminal devices 701, 702, 703 and server 705. Network 704 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0078] Users can use terminal devices 701, 702, and 703 to interact with server 705 through network 704 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 701, 702, and 703, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only examples).
[0079] The terminal devices 701 , 702 , and 703 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0080] The server 705 may be a server that provides various services, such as a background management server that provides support for websites browsed by users using terminal devices 701, 702, and 703 (for example only). The background management server may respond to the received image coding request and other data, perform three-dimensional reconstruction based on the multi-view images of the object to obtain a three-dimensional model of the object; in response to receiving a coding area model selected from the three-dimensional model, project the coding area model to each image in the multi-view images to obtain a coding area corresponding to each image; perform coding processing on the coding area of each image to perform coding and other processing on the multi-view images, and feed back the processing result (for example, image coding result - for example only) to the terminal device.
[0081] It should be noted that the method for coding the multi-view image provided in the embodiment of the present invention is generally executed by the server 705 , and accordingly, the device for coding the multi-view image is generally arranged in the server 705 .
[0082] It should be understood that Figure 7 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0083] Reference below Figure 8 , which shows a schematic diagram of the structure of a computer system 800 of a terminal device or server suitable for implementing an embodiment of the present invention. Figure 8 The terminal device or server shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0084] like Figure 8 As shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the system 800 are also stored. The CPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0085] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed, so that a computer program read therefrom is installed into the storage section 808 as needed.
[0086] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, the above-mentioned functions defined in the system of the present invention are executed.
[0087] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0088] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0089] The units or modules involved in the embodiments of the present invention may be implemented by software or hardware. The units or modules described may also be arranged in a processor, for example, they may be described as: a processor including a three-dimensional model reconstruction module, a coding area acquisition module, and a coding processing module. The names of these units or modules do not, in some cases, constitute limitations on the units or modules themselves, for example, the coding processing module may also be described as "a module for coding the coding area of each image to code the multi-view image".
[0090] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: in response to a coding request for an object image, performing three-dimensional reconstruction based on a multi-view image of the object to obtain a three-dimensional model of the object; in response to receiving a coding area model selected from the three-dimensional model, projecting the coding area model to each image in the multi-view image to obtain a coding area corresponding to each image; and coding the coding area of each image to code the multi-view image.
[0091] According to the technical solution of the embodiment of the present invention, by responding to a coding request for an object image, three-dimensional reconstruction is performed based on the multi-view images of the object to obtain a three-dimensional model of the object; in response to receiving a coding area model selected from the three-dimensional model, the coding area model is projected to each image in the multi-view images to obtain a coding area corresponding to each image; and coding is performed on the coding area of each image to code the multi-view images, which can effectively reduce manual participation, reduce the workload of manual processing, realize batch coding processing of the same position of multiple images, and ensure the consistency of coding areas in different images.
[0092] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for coding a multi-view image, characterized in that: include: In response to a request for coding an object image, performing three-dimensional reconstruction based on multi-view images of the object to obtain a three-dimensional model of the object; In response to receiving the coding area model selected from the three-dimensional model, projecting the coding area model onto each image in the multi-view images to obtain the coding area corresponding to each image; The coding area of each image is subjected to coding processing to code the multi-view images.
2. The method according to claim 1, characterized in that Performing three-dimensional reconstruction based on multi-view images of an object to obtain a three-dimensional model of the object includes: Using a neural radiation field algorithm, three-dimensional reconstruction is performed based on multi-view images of the object to obtain a three-dimensional model of the object.
3. The method according to claim 1 or 2, characterized in that: Performing three-dimensional reconstruction based on multi-view images of an object to obtain a three-dimensional model of the object includes: Acquire at least one key pixel point in the multi-view images of the object; Get the camera position, camera direction and actual color value corresponding to each key pixel; Inputting the camera position, camera direction and actual color value corresponding to each key pixel point into a pre-trained 3D reconstruction model to obtain the coordinates of the point on the surface of the model; A three-dimensional model of the object is constructed according to the coordinates of the points on the surface of the model.
4. The method according to claim 3, characterized in that The three-dimensional reconstruction model is obtained by training in the following way: For each object in the training set, obtaining at least one key pixel point in the multi-view images of the object; The camera position and camera direction corresponding to each key pixel point are used as inputs of the initial model to calculate the color value corresponding to each key pixel point; For each of the key pixel points, the color value corresponding to the key pixel point is compared with the actual color value to perform supervised learning to obtain the three-dimensional reconstruction model.
5. The method according to claim 3, characterized in that: In the three-dimensional reconstruction model, a signed distance function is used to define the positional relationship between a point in the three-dimensional space and the model surface. The point on the model surface is a point in the three-dimensional space whose closest distance to the model surface is 0.
6. The method according to claim 1, characterized in that Projecting the coding area model onto each image in the multi-view images to obtain a coding area corresponding to each image includes: For each image in the multi-view images, obtaining a projection matrix according to a camera intrinsic parameter matrix corresponding to the image; The coordinates of the coding area corresponding to the image are calculated according to the projection matrix, the posture matrix corresponding to the image and the matrix corresponding to the coding area model.
7. The method according to claim 1, characterized in that Performing coding processing on the coding area of each image to code the multi-view images includes: Performing down-sampling processing on each of the images to obtain a down-sampled image of each of the images; Performing mean blur processing on the downsampled image to obtain a blurred image; The blurred image is up-sampled to perform coding processing on each of the images.
8. A device for coding a multi-view image, characterized in that: include: A three-dimensional model reconstruction module, configured to respond to an object image coding request and perform three-dimensional reconstruction based on multi-view images of the object to obtain a three-dimensional model of the object; a coding area acquisition module, configured to, in response to receiving a coding area model selected from the three-dimensional model, project the coding area model onto each image in the multi-view images to obtain a coding area corresponding to each image; The coding processing module is used to perform coding processing on the coding area of each image to code the multi-view images.
9. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.