An image synthesis method, device, system, electronic device, and storage medium
By generating and transmitting target location data, the problems of large data volume and high bandwidth consumption in the image synthesis process are solved, and more efficient image synthesis is achieved.
Patent Information
- Application Number
- CN202310267115.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-14
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-03-14
AI Technical Summary
In existing technologies, since the size of the mask image is the same as that of the first image, a large amount of data needs to be transmitted during the image synthesis process, which consumes bandwidth resources and is inefficient.
The first device generates target location data, which represents the position of the edge pixels of the area occupied by the target object in the first image, and sends it to the second device. The second device then acquires and synthesizes the image based on this data.
This reduces the amount of data transmitted during image synthesis, lowers bandwidth requirements, and improves the efficiency of image synthesis.
Smart Images

Figure CN116245777B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to an image synthesis method, device, system, electronic device and storage medium. BACKGROUND
[0002] With the rapid development of image processing technology, in some scenarios, it is necessary to synthesize the image region occupied by a target object (for example, a person or a vehicle) in a first image and a second image as background to obtain a synthesized image.
[0003] In the related art, when a second device (such as a mobile phone, a PC (Personal Computer), or other smart terminal device located locally) needs to synthesize the image region occupied by a target object in a first image and a second image, the second device can send the first image to a first device (such as a server). Correspondingly, the first device can generate a mask image for the target object in the first image and send the mask image to the second device. The pixel value of each pixel point in the mask image indicates whether the pixel position corresponding to the pixel point belongs to the image region occupied by the target object in the first image. Then, for each pixel position, the second device can calculate the pixel value of the pixel position in the synthesized image based on the pixel value of the pixel position in the mask image, the pixel value of the pixel position in the first image, and the pixel value of the pixel position in the second image.
[0004] Based on the above method, in the process of generating the synthesized image, the first device needs to send the mask image to the second device. Since the size of the mask image is consistent with that of the first image, if the size of the first image is large, the amount of data that needs to be transmitted between the first device and the second device will also be large, which consumes a large amount of bandwidth resources and may take a long time, resulting in low efficiency of image synthesis. SUMMARY
[0005] The embodiments of the present application aim to provide an image synthesis method, device, system, electronic device and storage medium to reduce the amount of data to be transmitted, reduce the bandwidth required for transmitting data, and improve the efficiency of image synthesis. The specific technical solutions are as follows:
[0006] In a first aspect, the embodiments of the present application provide a method for image synthesis. The method is applied to a first device, and includes: receiving a first image containing a target object sent by a second device; generating target position data of the target object based on the first image, wherein the target position data represents a position of an edge pixel point of a first image region occupied by the target object in the first image; and sending the target position data to the second device, so that the second device obtains pixel values of each pixel point contained in the first image region based on the target position data after receiving the target position data, and obtains a synthesized image by combining pixel values of each pixel point contained in a second image region; wherein the second image region represents an image region other than a corresponding image region of the first image region in a second image.
[0007] In some embodiments, the generating of the target position data of the target object based on the first image includes: generating a mask image of the target object based on the first image, wherein a pixel value of each pixel point in the mask image represents whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; and generating the target position data of the target object based on the pixel value of each pixel point in the mask image.
[0008] In some embodiments, the generating of the target position data of the target object based on the pixel value of each pixel point in the mask image includes: determining, for each row in the mask image, a pixel point at which a pixel value changes in the row as a first edge pixel point in the row; obtaining a pixel coordinate of the first edge pixel point in the row to obtain first position data corresponding to the row; and obtaining the target position data of the target object based on the first position data corresponding to each row.
[0009] In some embodiments, the obtaining of the first position data corresponding to the row includes: obtaining a pixel coordinate of the first edge pixel point in the row and a pixel coordinate of a pixel point at a specified end point of the row to obtain the first position data corresponding to the row; wherein a pixel value of the pixel point at the specified end point represents that a corresponding pixel position in the first image belongs to the first image region.
[0010] In some embodiments, the pixel coordinate of the first edge pixel point in the row includes a vertical coordinate of the row and a horizontal coordinate of each first edge pixel point in the row.
[0011] In some embodiments, the method further comprises: for each row in the mask image, if the pixel values of the pixels in the row all indicate that the corresponding pixel positions in the first image belong to the first image region, obtaining the pixel coordinates of the pixels at the two end points of the row as the first position data corresponding to the row; or, obtaining a first preset identifier and the vertical coordinate of the row as the first position data corresponding to the row; wherein the first preset identifier indicates that the corresponding pixel positions of the pixels in a row all belong to the first image region in the first image.
[0012] In some embodiments, the generating the target position data for the target object based on the pixel values of each pixel in the mask image comprises: for each column in the mask image, determining the pixels in the column where the pixel values change as the first edge pixels in the column; obtaining the pixel coordinates of the first edge pixels in the column to obtain the first position data corresponding to the column; and obtaining the target position data for the target object based on the first position data corresponding to each column.
[0013] In some embodiments, the obtaining the pixel coordinates of the first edge pixels in the column to obtain the first position data corresponding to the column comprises: obtaining the pixel coordinates of the first edge pixels in the column and the pixel coordinates of the pixels at the specified end points of the column to obtain the first position data corresponding to the column; wherein the pixel values of the pixels at the specified end points indicate that the corresponding pixel positions in the first image belong to the first image region.
[0014] In some embodiments, the pixel coordinates of the first edge pixels in the column include the horizontal coordinates of the column and the vertical coordinates of the first edge pixels in the column.
[0015] In some embodiments, the method further comprises: for each column in the mask image, if the pixel values of the pixels in the column all indicate that the corresponding pixel positions in the first image belong to the first image region, obtaining the pixel coordinates of the pixels at the two end points of the column as the first position data corresponding to the column; or, obtaining a second preset identifier and the horizontal coordinate of the column as the first position data corresponding to the column; wherein the second preset identifier indicates that the corresponding pixel positions of the pixels in a column all belong to the first image region in the first image.
[0016] In some embodiments, the generating the target position data of the target object based on the pixel value of each pixel in the mask image comprises: determining, for each row in the mask image, a pixel point at which the pixel value changes in the row as a first edge pixel point in the row; obtaining the pixel coordinates of the first edge pixel point in the row to obtain first position data corresponding to the row; obtaining first candidate data representing the position of the edge pixel point of the first image region based on the first position data corresponding to each row; determining, for each column in the mask image, a pixel point at which the pixel value changes in the column as a first edge pixel point in the column; obtaining the pixel coordinates of the first edge pixel point in the column to obtain first position data corresponding to the column; obtaining second candidate data representing the position of the edge pixel point of the first image region based on the first position data corresponding to each column; if the data amount of the first candidate data is less than the data amount of the second candidate data, determining the first candidate data as the target position data of the target object; if the data amount of the first candidate data is not less than the data amount of the second candidate data, determining the second candidate data as the target position data of the target object.
[0017] In a second aspect, the embodiment of the present application provides an image synthesis method, the method is applied to a second device, and the method comprises the following steps: sending a first image containing a target object to a first device, so that the first device generates target position data of the target object based on the first image and sends the target position data to the second device; wherein the target position data represents the position of an edge pixel point of a first image region occupied by the target object in the first image; after receiving the target position data, obtaining pixel values of each pixel point contained in the first image region based on the target position data; and combining the pixel values of each pixel point contained in the first image region and the pixel values of each pixel point contained in a second image region to obtain a synthesis image; wherein the second image region represents other image regions in a second image except for an image region corresponding to the first image region.
[0018] In some embodiments, the target position data comprises: first position data corresponding to a row in the mask image; a pixel value of each pixel point in the mask image representing whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; and the obtaining of the pixel value of each pixel point included in the first image region based on the target position data comprises: for each first position data in the target position data, if the first position data comprises pixel coordinates of a plurality of pixel points, obtaining, based on the pixel coordinates of the plurality of pixel points, the pixel value of each pixel point between each pair of pixel points in the first image to obtain the pixel value of the pixel point corresponding to the first image region in the row to which the first position data belongs; wherein each pair of pixel points represents: in the first image, pixel points corresponding to two adjacent pixel coordinates in the pixel coordinates of the plurality of pixel points; any two pairs of pixel points do not include the same pixel point; if the first position data comprises a first preset identifier and a vertical coordinate, obtaining the pixel value of each pixel point in the row corresponding to the vertical coordinate in the first image to obtain the pixel value of the pixel point corresponding to the first image region in the row to which the first position data belongs.
[0019] In some embodiments, the target position data comprises: first position data corresponding to a column in the mask image; a pixel value of each pixel point in the mask image representing whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; and the obtaining of the pixel value of each pixel point included in the first image region based on the target position data comprises: for each first position data in the target position data, if the first position data comprises pixel coordinates of a plurality of pixel points, obtaining, based on the pixel coordinates of the plurality of pixel points, the pixel value of each pixel point between each pair of pixel points in the first image to obtain the pixel value of the pixel point corresponding to the first image region in the column to which the first position data belongs; wherein each pair of pixel points represents: in the first image, pixel points corresponding to two adjacent pixel coordinates in the pixel coordinates of the plurality of pixel points; any two pairs of pixel points do not include the same pixel point; if the first position data comprises a second preset identifier and a horizontal coordinate, obtaining the pixel value of each pixel point in the column corresponding to the horizontal coordinate in the first image to obtain the pixel value of the pixel point corresponding to the first image region in the column to which the first position data belongs.
[0020] In some embodiments, the combining of the pixel value of each pixel point included in the first image region and the pixel value of each pixel point included in the second image region to obtain a synthesis image comprises: for each pixel point in the first image region, replacing the pixel value of the pixel point in the second image with the pixel value of the pixel point in the first image to obtain a synthesis image.
[0021] In a third aspect, the embodiment of the present application provides an image synthesis system, the system comprising a first device and a second device, wherein: the second device is configured to send a first image containing a target object to the first device; the first device is configured to generate target position data of the target object based on the first image when the first image is received, and send the target position data to the second device; wherein the target position data indicates a position of an edge pixel point of a first image region occupied by the target object in the first image; and the second device is further configured to obtain pixel values of each pixel point contained in the first image region based on the target position data after the target position data is received, and obtain a synthesis image by combining the pixel values of each pixel point contained in the first image region and pixel values of each pixel point contained in a second image region; wherein the second image region represents other image regions in a second image except for an image region corresponding to the first image region.
[0022] In a fourth aspect, the embodiment of the present application provides an image synthesis device, the device being applied to a first device, and the device comprising: a first receiving module configured to receive a first image containing a target object sent by a second device; a target position data generation module configured to generate target position data of the target object based on the first image; wherein the target position data indicates a position of an edge pixel point of a first image region occupied by the target object in the first image; and a first sending module configured to send the target position data to the second device, so that the second device obtains pixel values of each pixel point contained in the first image region based on the target position data after the target position data is received, and obtains a synthesis image by combining pixel values of each pixel point contained in a second image region; wherein the second image region represents other image regions in a second image except for an image region corresponding to the first image region.
[0023] In some embodiments, the target position data generation module comprises: a mask image generation submodule configured to generate a mask image of the target object based on the first image; wherein a pixel value of each pixel point in the mask image indicates whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; and a target position data generation submodule configured to generate target position data of the target object based on the pixel value of each pixel point in the mask image.
[0024] In some embodiments, the target position data generation submodule comprises: a first determination unit configured to determine, for each row in the mask image, a pixel point at which a pixel value changes in the row as a first edge pixel point in the row; a first acquisition unit configured to acquire a pixel coordinate of the first edge pixel point in the row to obtain first position data corresponding to the row; and a first generation unit configured to obtain target position data for the target object based on the first position data corresponding to each row.
[0025] In some embodiments, the first acquisition unit is specifically configured to acquire the pixel coordinate of the first edge pixel point in the row and a pixel coordinate of a pixel point at a specified end point of the row to obtain the first position data corresponding to the row; and the pixel value of the pixel point at the specified end point indicates that a corresponding pixel position in the first image belongs to the first image region.
[0026] In some embodiments, the pixel coordinate of the first edge pixel point in the row comprises a vertical coordinate of the row and a horizontal coordinate of each first edge pixel point in the row.
[0027] In some embodiments, the apparatus further comprises: a first acquisition module configured to, for each row in the mask image, if the pixel value of each pixel point in the row indicates that a corresponding pixel position in the first image belongs to the first image region, acquire a pixel coordinate of a pixel point at each end point of the row as first position data corresponding to the row; or a second acquisition module configured to acquire a first preset identifier and a vertical coordinate of the row as the first position data corresponding to the row; and the first preset identifier indicates that each pixel point in a row corresponds to a pixel position in the first image that belongs to the first image region.
[0028] In some embodiments, the target position data generation submodule comprises: a second determination unit configured to determine, for each column in the mask image, a pixel point at which a pixel value changes in the column as a first edge pixel point in the column; a second acquisition unit configured to acquire a pixel coordinate of the first edge pixel point in the column to obtain first position data corresponding to the column; and a second generation unit configured to obtain target position data for the target object based on the first position data corresponding to each column.
[0029] In some embodiments, the second acquisition unit is specifically configured to acquire the pixel coordinate of the first edge pixel point in the column and a pixel coordinate of a pixel point at a specified end point of the column to obtain the first position data corresponding to the column; and the pixel value of the pixel point at the specified end point indicates that a corresponding pixel position in the first image belongs to the first image region.
[0030] In some embodiments, the pixel coordinates of the first edge pixel points in the column include: the horizontal coordinates of the column, and the vertical coordinates of the first edge pixel points in the column.
[0031] In some embodiments, the device further includes: a third obtaining module, configured to, for each column in the mask image, if the pixel values of all pixel points in the column represent that the corresponding pixel positions in the first image belong to the first image region, obtain the pixel coordinates of the pixel points at the two endpoints of the column as the first position data corresponding to the column; or a fourth obtaining module, configured to obtain a second preset identifier and the horizontal coordinates of the column as the first position data corresponding to the column; wherein the second preset identifier represents that the corresponding pixel positions of all pixel points in a column in the first image belong to the first image region.
[0032] In some embodiments, the target position data generation submodule includes: a first candidate data generation unit, configured to, for each row in the mask image, determine the pixel points at which the pixel values change in the row as the first edge pixel points in the row; obtain the pixel coordinates of the first edge pixel points in the row to obtain the first position data corresponding to the row; and obtain first candidate data representing the positions of the edge pixel points of the first image region based on the first position data corresponding to each row; a second candidate data generation unit, configured to, for each column in the mask image, determine the pixel points at which the pixel values change in the column as the first edge pixel points in the column; obtain the pixel coordinates of the first edge pixel points in the column to obtain the first position data corresponding to the column; and obtain second candidate data representing the positions of the edge pixel points of the first image region based on the first position data corresponding to each column; a third determination unit, configured to, if the data amount of the first candidate data is less than the data amount of the second candidate data, determine the first candidate data as the target position data for the target object; and a fourth determination unit, configured to, if the data amount of the first candidate data is not less than the data amount of the second candidate data, determine the second candidate data as the target position data for the target object.
[0033] In a fifth aspect, the embodiment of the present application provides an image synthesis device, which is applied to a second device, and comprises: a second sending module, configured to send a first image containing a target object to a first device, so that the first device generates target position data of the target object based on the first image, and sends the target position data to the second device; wherein the target position data indicates a position of an edge pixel point of a first image region occupied by the target object in the first image; a second receiving module, configured to acquire pixel values of each pixel point contained in the first image region based on the target position data after receiving the target position data; and a synthesis module, configured to combine the pixel values of each pixel point contained in the first image region and pixel values of each pixel point contained in a second image region to obtain a synthesis image; wherein the second image region indicates other image regions in a second image except for an image region corresponding to the first image region.
[0034] In some embodiments, the target position data contains first position data corresponding to a row in a mask image; pixel values of each pixel point in the mask image indicate whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; and the second receiving module is specifically configured to: for each first position data in the target position data, if the first position data contains pixel coordinates of a plurality of pixel points, acquire pixel values of each pixel point between the pixel points in the first image based on the pixel coordinates of the plurality of pixel points, to obtain pixel values of pixel points corresponding to the first image region in a row to which the first position data belongs; wherein each pixel point pair indicates pixel points corresponding to two adjacent pixel coordinates in the pixel coordinates of the plurality of pixel points in the first image; any two pixel point pairs do not contain the same pixel point; and if the first position data contains a first preset identifier and a vertical coordinate, acquire pixel values of each pixel point in a row corresponding to the vertical coordinate in the first image, to obtain pixel values of pixel points corresponding to the first image region in the row to which the first position data belongs.
[0035] In some embodiments, the target position data comprises: first position data corresponding to columns in a mask image; a pixel value of each pixel point in the mask image representing whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; the second receiving module is specifically configured to: for each first position data in the target position data, if the first position data comprises pixel coordinates of a plurality of pixel points, based on the pixel coordinates of the plurality of pixel points, obtaining pixel values of pixel points between each pixel point pair in the first image to obtain pixel values of pixel points corresponding to the first image region in a column to which the first position data belongs in the first image; wherein each pixel point pair represents: in the first image, pixel points corresponding to two adjacent pixel coordinates in the pixel coordinates of the plurality of pixel points; any two pixel point pairs do not comprise the same pixel point; if the first position data comprises a second preset identifier and an abscissa, obtaining pixel values of pixel points in a column corresponding to the abscissa in the first image to obtain pixel values of pixel points corresponding to the first image region in the column to which the first position data belongs.
[0036] In some embodiments, the synthesizing module is specifically configured to: for each pixel point in the first image region, replace a pixel value of a pixel point in the second image that is consistent with a position of the pixel point with a pixel value of the pixel point in the first image to obtain a synthesized image.
[0037] A sixth aspect of the embodiments of the present application provides an electronic device, comprising:
[0038] a memory for storing a computer program;
[0039] a processor for executing the program stored on the memory to implement the image synthesis method described above.
[0040] A seventh aspect of the embodiments of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the image synthesis method described above.
[0041] The embodiments of the present application further provide a computer program product comprising instructions which, when executed on a computer, cause the computer to perform the image synthesis method described above.
[0042] The embodiments of the present application have the following beneficial effects:
[0043] The embodiment of the present application provides a kind of image synthesis method, the method is applied to first equipment, method includes: receiving the first image containing target object that second equipment sends;First image is generated based on the target position data of target object;Wherein, target position data indicates: the position of the edge pixel point of the first image area occupied by target object in first image;Target position data is sent to second equipment, so that second equipment receives target position data, based on target position data, obtains the pixel value of each pixel point contained in first image area, and combines the pixel value of each pixel point contained in second image area, obtains synthesis image;Wherein, second image area indicates: the other image area except the image area corresponding to first image area in second image.
[0044] Based on the above processing, first equipment only needs to send target position data to second equipment, so that second equipment can generate synthesis image. Since the size of mask image is consistent with first image, target position data indicates the position of the edge pixel point of first image area, and the data quantity is much smaller than that of mask image corresponding to first image, and then, relative to the mode in prior art, the data quantity required in image synthesis process is reduced, the bandwidth required for transmitting data is reduced, and the efficiency of image synthesis is improved.
[0045] Of course, implementing any product or method of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other embodiments can also be obtained by those skilled in the art based on these drawings.
[0047] Figure 1 A schematic diagram of obtaining a synthesis image in related art;
[0048] Figure 2 An architecture diagram of an image synthesis system provided by the embodiment of the present application;
[0049] Figure 3 A flowchart of an image synthesis method provided by the embodiment of the present application;
[0050] Figure 4 A flowchart of another image synthesis method provided by the embodiment of the present application;
[0051] Figure 5 A schematic diagram of a first image provided by the embodiment of the present application;
[0052] Figure 6 A method for generating a mask image based on a first image is provided in an embodiment of the present application. Figure 5 A mask image generated according to a first image is shown in the figure.
[0053] Figure 7 A schematic diagram of a mask image generated according to a target position data in a first mode is provided in an embodiment of the present application.
[0054] Figure 8 A schematic diagram of a mask image generated according to a target position data in a second mode is provided in an embodiment of the present application.
[0055] Figure 9 A flowchart of a method for generating a composite image is provided in an embodiment of the present application.
[0056] Figure 10 A structural diagram of an image synthesis device is provided in an embodiment of the present application.
[0057] Figure 11 A structural diagram of another image synthesis device is provided in an embodiment of the present application.
[0058] Figure 12 A structural diagram of an electronic device is provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art based on the present application are within the scope of protection of the present application.
[0060] With the rapid development of image processing technology, in some scenarios, it is necessary to synthesize the image region occupied by a target object (e.g., a person, a vehicle) in a first image containing the target object with a second image as a background to obtain a composite image.
[0061] As shown in the figure, Figure 1 Figure 1 A schematic diagram of a method for generating a composite image in the related art is shown. The first device can be a server, and correspondingly, the second device can be a terminal device used by a user; or the first device can be a GPU (Graphics Processing Unit) in an electronic device, and correspondingly, the second device can be a CPU (Central Processing Unit) in the electronic device.
[0062] In the related art, when a second device needs to synthesize an image region occupied by a target object in an original image (i.e., a first image) and a background image (i.e., a second image), the second device can send the first image to a first device. Correspondingly, the first device can generate a mask image (also referred to as an alpha image) for the target object in the first image, and send the mask image to the second device. A pixel value of each pixel point in the mask image indicates whether a pixel position corresponding to the pixel point belongs to the image region occupied by the target object in the first image. Further, for each pixel position, the second device can calculate a pixel value of the pixel position in a synthesized image based on a pixel value of the pixel position in the mask image, a pixel value of the pixel position in the first image, and a pixel value of the pixel position in the second image, as shown in formula (1):
[0063] Image = alpha * foreground + (1-alpha) * background (1)
[0064] wherein Image represents a pixel value of any pixel position in the synthesized image, foreground represents a pixel value of the pixel position in the first image, background represents a pixel value of the pixel position in the second image, and alpha represents a weight of the pixel value of the pixel position in the synthesized image, which is the pixel value of the pixel position in the first image. The value of alpha is in the interval [0, 1], for example, alpha can be 0.5.
[0065] Based on the above manner, in the process of generating the synthesized image, the first device needs to send the mask image to the second device. Since the size of the mask image is consistent with that of the first image, if the size of the first image is large, the amount of data to be transmitted between the first device and the second device will also be large, which consumes a lot of time and leads to low efficiency of image synthesis.
[0066] In addition, in the process of generating the synthesized image, the second device obtains the pixel value of each pixel point in the synthesized image based on the above formula (1), that is, in the related art, the process of generating the synthesized image also involves a large amount of floating-point calculation, which also leads to low efficiency of image synthesis.
[0067] To solve the above problems, an embodiment of the present application provides an image synthesis system, which is described below with reference to Figure 2 , Figure 2 An architecture diagram of an image synthesis system provided by an embodiment of the present application. The image synthesis system includes a first device 201 and a second device 202, wherein:
[0068] The second device 202 is configured to send a first image containing a target object to the first device 201.
[0069] The first device 201 is configured to, when the first image is received, generate target position data of a target object based on the first image, and send the target position data to the second device 202.
[0070] The target position data indicates a position of an edge pixel point of a first image region occupied by the target object in the first image.
[0071] The second device 202 is further configured to, after receiving the target position data, acquire pixel values of each pixel point contained in the first image region based on the target position data, and obtain a composite image by combining the pixel values of each pixel point contained in the first image region and pixel values of each pixel point contained in a second image region.
[0072] The second image region indicates an image region other than the first image region in the second image.
[0073] The image synthesis system provided by the embodiments of the present application can make the second device 202 generate a composite image only by sending the target position data from the first device 201 to the second device 202. Since the mask image has the same size as the first image, and the target position data indicates the position of the edge pixel point of the first image region, the data amount of the target position data is much smaller than that of the mask image corresponding to the first image. Therefore, compared with the prior art, the data amount required to be transmitted in the image synthesis process is reduced, the bandwidth required for transmitting data is reduced, and the efficiency of image synthesis is improved.
[0074] In order to protect part of the function code in the image synthesis process, reduce the load of the second device 202, and further improve the efficiency of image synthesis, the first device 201 and the second device 202 can be two independent electronic devices, and the first device 201 and the second device 202 can communicate through a network. For example, the first device 201 can be a server providing an image synthesis function, and the second device 202 can be a user terminal, such as a mobile phone, a computer, etc., which needs to acquire a composite image. Based on this, the technical personnel can deploy the related code for generating the target position data in the first device 201 with better performance, so as to reduce the load of the second device 202, and at the same time, protect part of the function code (i.e. the code deployed in the first device 201) in the image synthesis process.
[0075] Alternatively, the first device 201 and the second device 202 can also be integrated in the same electronic device (which can be referred to as an integrated device), and the first device 201 and the second device 202 can communicate through a hardware connection. For example, the first device 201 can be a GPU (Graphics Processing Unit) in the integrated device, and the second device 202 can be a CPU (Central Processing Unit) in the integrated device. Since the first device has higher parallel processing capability and is more suitable for processing image data, the efficiency of generating target position data can be improved, and thus the efficiency of image synthesis can be further improved.
[0076] In addition, since the amount of data required for transmission in the image synthesis process is reduced, the bandwidth required for data transmission can also be reduced, and the synthesized image can be obtained in a low-bandwidth situation, that is, the low-bandwidth scenario between the first device and the second device can be adapted.
[0077] Based on the same inventive concept, the embodiment of the present application provides an image synthesis method, which can be applied to a first device and a second device. The first device can be the first device 201 in the image synthesis system described above, and the second device can be the second device 202 in the image synthesis system described above. Referring to Figure 3 , Figure 3 A flowchart of an image synthesis method provided by the embodiment of the present application is shown in the figure, which can include the following steps:
[0078] S301: The second device sends a first image containing a target object to the first device.
[0079] S302: The first device generates target position data for the target object based on the first image.
[0080] The target position data indicates the position of the edge pixel point of the first image region occupied by the target object in the first image.
[0081] S303: The first device sends the target position data to the second device.
[0082] S304: The second device obtains the pixel values of each pixel point contained in the first image region based on the target position data.
[0083] S305: The second device obtains a synthesized image by combining the pixel values of each pixel point contained in the first image region and the pixel values of each pixel point contained in the second image region.
[0084] The second image region indicates other image regions in the second image except for the image region corresponding to the first image region.
[0085] Based on the above processing, the first device only needs to send the target position data to the second device, so that the second device can generate the composite image. Since the size of the mask image is consistent with the first image, the target position data represents the positions of the edge pixel points of the first image region, and the data amount is much smaller than that of the mask image corresponding to the first image. Therefore, compared with the manner in the prior art, the data amount required in the image composition process is reduced, the bandwidth required for transmitting data is reduced, and the efficiency of image composition is improved.
[0086] For step S301, the target object represents an object that needs to be displayed in the composite image. Specifically, the type of the target object can be pre-set by a technical person according to business requirements. For example, the target object can be a foreground object in the first image. For example, the target object can also be a person or a vehicle in the image.
[0087] The second device can pre-acquire an image containing the target object. For example, the image containing the target object can be a single picture, or can also be a video frame in a video image. For example, for each video frame in the video image, the second device can determine the video frame as the first image and send the first image to the first device. In this way, the entire video image can be processed.
[0088] For steps S302-S303, the first image region is an image region occupied by the target object in the first image. The edge pixel points of the first image region represent the pixel points at the edge positions of the first image region in the first image. Specifically, the process of determining the first image region will be described in subsequent embodiments. For each pixel point, the pixel position of the pixel point can be represented in the form of a pixel coordinate.
[0089] In an implementation manner, the first device can determine the pixel points at the edge positions based on a target detection algorithm and an edge detection algorithm. For example, the target detection algorithm can be a target detection network based on R-CNN (Region-Convolutional Neural Networks, region detection-convolutional neural network). The edge detection algorithm can be a Canny multi-level edge detection algorithm or an edge detection algorithm based on a Roberts operator.
[0090] In another implementation manner, referring to Figure 4 , Figure 4 Another flowchart of an image composition method provided by the embodiment of the present application is shown in Figure 3 Based on the above, step S302 includes:
[0091] S3021: generating a mask image for the target object based on the first image.
[0092] The pixel value of each pixel point in the mask image indicates whether a pixel position corresponding to the pixel point in the first image belongs to the first image region.
[0093] S3022: Generate target position data for the target object based on the pixel value of each pixel point in the mask image.
[0094] In the embodiment of the present application, the first device can generate a mask image for the target object based on a matting algorithm. For example, the matting algorithm can be a Background Matting algorithm or a Deep Image Matting algorithm.
[0095] The size of the mask image generated by the first device is consistent with the size of the first image. For example, if the size of the first image is 2560 (width) * 1440 (height), the size of the mask image is also 2560 * 1440.
[0096] For each pixel point in the mask image, if the pixel position corresponding to the pixel point in the first image belongs to the first image region, the pixel value of the pixel point in the generated mask image is a first value; if the pixel position corresponding to the pixel point in the first image does not belong to the first image region, the pixel value of the pixel point in the generated mask image is a second value. For example, the first value can be 1 and the second value can be 0.
[0097] Referring to Figure 5 and Figure 6 , Figure 5 is a schematic diagram of a first image provided by an embodiment of the present application, Figure 6 is a mask image generated based on the first image shown in Figure 5 . Figure 5 In the embodiment, the target object is a person, and the first image region is an image region occupied by the person. Figure 6 In the embodiment, the white image region contains pixel points indicating that the pixel position corresponding to the pixel point in the first image belongs to the first image region, and the black image region contains pixel points indicating that the pixel position corresponding to the pixel point in the first image does not belong to the first image region.
[0098] The first device can generate the target position data of the target object in different manners. For example, the first device can take a row of pixels in the mask image as a unit for generating the target position data, i.e., determine the position of the pixel corresponding to the edge pixel of the first image region in each row of pixels in the mask image, and generate the target position data of the target object; or the first device can also take a column of pixels in the mask image as a unit for generating the target position data, i.e., determine the position of the pixel corresponding to the edge pixel of the first image region in each column of pixels in the mask image, and generate the target position data of the target object. The specific process of generating the target position data of the target object will be described in subsequent embodiments.
[0099] For each pixel in the mask image, the pixel value of the pixel can represent whether the corresponding position in the first image belongs to the first image region, and thus, based on the pixel value of the pixel in the mask image, the position of the edge pixel of the first image region can be determined. Based on this, the target position data of the target object obtained can more accurately reflect the position of the edge pixel of the first image region occupied by the target object in the first image.
[0100] For step S304, after receiving the target position data, the second device can determine the positions of the edge pixels, and then the second device can determine the image region (i.e., the first image region) with the pixels at the positions of the edge pixels as edges, and can obtain the pixel values of the pixels contained in the first image region.
[0101] For step S305, the second device can obtain the second image used for synthesis in advance. For example, the second image can be an image containing only the background, such as the background image shown in FIG. 6. Figure 1
[0102] The size of the second image is consistent with the size of the first image, and the size of the synthesized image obtained by the second device is consistent with the size of the first image and the size of the second image. For example, if the size of the first image is 2560*1440 and the size of the second image is 2560*1440, the size of the synthesized image obtained is also 2560*1440.
[0103] For each pixel in the synthesized image, the corresponding pixel position of the pixel in the first image belongs to the first image region, or the corresponding pixel position of the pixel in the second image belongs to the second image region. That is, the synthesized image is composed of the pixels contained in the first image region and the pixels contained in the second image region.
[0104] In one implementation, the second device can determine the size of the composite image according to the size of the second image, and for each pixel point in the composite image, if the pixel position corresponding to the pixel point in the first image belongs to the first image region, the pixel value of the pixel point in the composite image is determined as the pixel value of the pixel point at the same pixel position in the first image region; if the pixel position corresponding to the pixel point in the second image belongs to the second image region, the pixel value of the pixel point in the composite image is determined as the pixel value of the pixel point at the same pixel position in the second image region.
[0105] In another implementation, the second device can directly replace the pixel values of a part of the pixel points in the second image to obtain the composite image. Specifically, step S305 includes: for each pixel point in the first image region, replacing the pixel value of the pixel point in the second image with the pixel value of the pixel point in the first image to obtain the composite image.
[0106] Since the pixel values of the pixel points included in the second image region in the second image are unchanged, the second device can replace the pixel values of a part of the pixel points in the second image to obtain the second image after replacement as the composite image. Compared with calculating the pixel value of each pixel point in the composite image based on the above formula (1) in the related art, the method provided in the embodiments of the present application does not need to perform a large number of floating point calculations, can further reduce the calculation amount of generating the composite image, reduce the amount of data required to be transmitted in the image synthesis process, reduce the bandwidth required for transmitting data, and improve the efficiency of image synthesis.
[0107] The first device can generate the target position data for the target object in different ways. The above step S3022 can be implemented based on any one of the following multiple ways:
[0108] Way one: the first device can generate the target position data for the target object in a row scanning manner, that is, the first device can take the pixel points in a row of the mask image as a unit for generating the target position data, and correspondingly, step S3022 includes:
[0109] Step 1: for each row in the mask image, determine the pixel point at which the pixel value changes in the row as the first edge pixel point in the row.
[0110] Step 2: obtain the pixel coordinates of the first edge pixel point in the row to obtain the first position data corresponding to the row.
[0111] Step 3: based on the first position data corresponding to each row, obtain the target position data for the target object.
[0112] In the embodiment of the present application, each row in the mask image represents the pixel points in the row. For each row in the mask image, the first device can traverse the pixel values of the pixel points in the row to determine the first edge pixel points in the row. For example, the first device can traverse the pixel points in the row in the order of the horizontal coordinates of the pixel points from small to large, or can traverse the pixel points in the row in the order of the horizontal coordinates of the pixel points from large to small.
[0113] For each pixel point in the row, if the pixel value of the pixel point is different from the pixel values of the other pixel points adjacent to the pixel point in the row, and the pixel value of the pixel point indicates that the pixel position corresponding to the pixel point in the first image belongs to the first image region, the pixel point can be determined as the first edge pixel point in the row. That is, since the size of the mask image is consistent with the size of the first image, the pixel position of each first edge pixel point in the mask image is consistent with the pixel position of an edge pixel point of the first image region in the first image.
[0114] For example, the size of the mask image is 2560 (width) * 1440 (height), in the row with the vertical coordinate of 200, the pixel position of the pixel point D0 is (99, 200), and the pixel value is 0; the pixel position of the pixel point D1 is (100, 200), and the pixel value is 1; the pixel position of the pixel point D2 is (101, 200), and the pixel value is 1; the pixel position of the pixel point D3 is (102, 200), and the pixel value is 0. Then, the first device can determine that the pixel point D1 and the pixel point D2 are both the first edge pixel points in the row.
[0115] In the related art, since each row in the mask image contains 2560 pixel points, for the row with the vertical coordinate of 200 in the mask image, the first device needs to send the pixel values of each pixel point in the row, i.e., the pixel values of 2560 pixel points, to the second device. In the embodiment of the present application, the first device can send the pixel coordinates of the pixel point D1 and the pixel point D2 to the second device, and the data amount of the two pixel coordinates is much smaller than the data amount of the pixel values of a row of pixel points. That is, based on this, the first device can reduce the data amount of the data sent to the second device, reduce the bandwidth required for transmitting the data, reduce the time length consumed for transmitting the data, and improve the efficiency of the synthesized image.
[0116] Alternatively, for any pixel point in the row, if the pixel value of the pixel point is different from the pixel values of the other pixel points adjacent to the pixel point in the row, and the pixel value of the pixel point indicates that the pixel position corresponding to the pixel point in the first image does not belong to the first image region, the pixel point can be determined as the first edge pixel point in the row.
[0117] For example, for the pixel points D0-D3, the first device can determine the pixel point D0 and the pixel point D3 as the first edge pixel points in the row.
[0118] For each first edge pixel point in the row, the pixel coordinate of the first edge pixel point is composed of the horizontal coordinate and the vertical coordinate of the first edge pixel point. It can be understood that the vertical coordinates of the first edge pixel points in a row are the same. For example, for the pixel point D1 and the pixel point D2, the horizontal coordinate of the pixel point D1 is 100, and the vertical coordinate is 200; the horizontal coordinate of the pixel point D2 is 101, and the vertical coordinate is 200.
[0119] In addition, when the edge of the first image region overlaps with the left edge or the right edge of the first image, since the pixel value of the pixel point at the position in the mask image is the same as the pixel value of the adjacent pixel point in the row, the pixel point at the position in the mask image will not be determined as the first edge pixel point.
[0120] Correspondingly, the above step 2 comprises:
[0121] Step 21: obtaining the pixel coordinates of the first edge pixel points in the row and the pixel coordinates of the pixel points at the specified endpoints of the row, to obtain the first position data corresponding to the row.
[0122] The pixel value of the pixel point at the specified endpoint indicates that the corresponding pixel position in the first image belongs to the first image region.
[0123] For each row in the mask image, there are two endpoints in the row, which are the pixel point with the smallest horizontal coordinate in the row and the pixel point with the largest horizontal coordinate in the row.
[0124] For the pixel point at any endpoint in the row, there is only one other pixel point adjacent to it in the row. If the pixel value of the pixel point at the endpoint indicates that the corresponding pixel position in the first image belongs to the first image region, the endpoint is the specified endpoint, and the pixel point at the endpoint is the pixel point at the specified endpoint.
[0125] For example, the size of the mask image is 2560*1440, in the row with the vertical coordinate of 200, the pixel position of the pixel point D4 is (0, 200), and the pixel value is 1; the pixel position of the pixel point D5 is (1, 200), and the pixel value is 1; the pixel position of the pixel point D6 is (2558, 200), and the pixel value is 0; the pixel position of the pixel point D7 is (2559, 200), and the pixel value is 0. Then, the pixel point D4 is the pixel point at the specified endpoint, and the pixel point D7 is the pixel point at the endpoint, but not the pixel point at the specified endpoint. Correspondingly, the first position data corresponding to the row also includes the pixel coordinate (0, 200) of the pixel point D4.
[0126] For each row in the mask image, the first position data corresponding to the row can represent the edge of the first image region corresponding to the row. That is, for any row in the mask image, if there is a pixel point corresponding to the first image region in the row, the pixel coordinates determined based on step 21 above are even numbers.
[0127] Based on the above processing, when the edge of the first image region occupied by the target object overlaps with the edge of the first image, the first device can determine the pixel position of the pixel point at the overlapping position, so that the target position data obtained based on the first position data corresponding to each row can represent the position of the edge pixel point of the complete first image region.
[0128] In an implementation manner, the pixel coordinates of the first edge pixel points in the row include the horizontal coordinates and the vertical coordinates of the first edge pixel points in the row.
[0129] For example, for a row in the mask image, if the first edge pixel points in the row include pixel point E1 and pixel point E2, the pixel coordinates of the first edge pixel points in the row include the horizontal coordinate x1 and the vertical coordinate y1 of pixel point E1, and the horizontal coordinate x2 and the vertical coordinate y1 of pixel point E2.
[0130] In another implementation manner, the pixel coordinates of the first edge pixel points in the row include the vertical coordinate of the row, and the horizontal coordinates of the first edge pixel points in the row.
[0131] That is, for multiple first edge pixel points in the same row in the mask image, the pixel coordinates of the first edge pixel points in the row can be represented by the vertical coordinate of the row and the horizontal coordinates of the first edge pixel points in the row.
[0132] For example, for a row in the mask image, if the first edge pixel points in the row include pixel point E1 and pixel point E2, the pixel coordinates of the first edge pixel points in the row include the vertical coordinate y1 of the row, and the horizontal coordinate x1 of pixel point E1 and the horizontal coordinate x2 of pixel point E2.
[0133] Based on the above processing, in the case that there are multiple first edge pixel points in a row in the mask image, the same vertical coordinate in the pixel coordinates of the multiple first edge pixel points can be avoided from being repeatedly recorded in the obtained target position data, and further, the data amount of the generated target position data can be further reduced, the bandwidth required for transmitting data can be reduced, and the efficiency of the synthesized image can be improved.
[0134] In addition, since the first position data corresponding to each row of the mask image includes the pixel coordinates of the first edge pixel points in the row and the pixel coordinates of the pixel points at the specified endpoints of the row, the pixel coordinates of the first edge pixel points in each row of the mask image can be obtained based on the first position data corresponding to each row, and then the second device can determine the positions of the edge pixel points of the first image region in the first image based on the target position data.
[0135] In some embodiments, the method further includes:
[0136] Step 4: For each row of the mask image, if the pixel values of the pixel points in the row all indicate that the corresponding pixel positions in the first image belong to the first image region, the pixel coordinates of the pixel points at the two endpoints of the row are obtained as the first position data corresponding to the row.
[0137] Alternatively,
[0138] Step 5: The first preset identifier and the longitudinal coordinate of the row are obtained as the first position data corresponding to the row.
[0139] The first preset identifier indicates that the corresponding pixel positions of the pixel points in a row in the first image all belong to the first image region.
[0140] The description of the endpoints in each row of the mask image can refer to the above embodiments.
[0141] For each row of the mask image, if the pixel values of the pixel points in the row all indicate that the corresponding pixel positions in the first image belong to the first image region, the first device can set the value of the corresponding identifier bit (Flag) for the row (i.e., the first preset identifier), for example, the value of the identifier bit can be 0.
[0142] For example, the size of the mask image is 2560 (width) * 1440 (height), and if the pixel values of the pixel points in the row with a longitudinal coordinate of 300 in the mask image are all 1, the first position data corresponding to the row includes (0, 300) and (2559, 300). Alternatively, the first position data corresponding to the row includes the longitudinal coordinate (300) of the row and the first preset identifier (Flag = 0).
[0143] Based on the above processing, in the case that the corresponding pixel positions of each pixel point in a row in the mask image in the first image all belong to the first image region, the first position data corresponding to the row only contains the pixel coordinates of the pixel points at the two endpoints of the row, or the first preset identifier and the longitudinal coordinate of the row. Compared with the prior art which needs to transmit the pixel values of each pixel point in the mask image, the application can further reduce the data amount of the generated target position data, reduce the bandwidth required for transmitting data, and further improve the efficiency of the synthesized image.
[0144] In some embodiments, the method further comprises:
[0145] For each row in the mask image, if the pixel values of each pixel point in the row all indicate that the corresponding pixel position in the first image does not belong to the first image region, the third preset identifier and the longitudinal coordinate of the row are obtained as the first position data corresponding to the row.
[0146] The third preset identifier indicates that the corresponding pixel positions of each pixel point in a row in the first image all do not belong to the first image region.
[0147] In the embodiments of the application, if the pixel values of each pixel point in the row all indicate that the corresponding pixel position in the first image does not belong to the first image region, the first device can set the value of the corresponding identifier bit (Flag) for the row (i.e. the third preset identifier), for example, the value of the identifier bit can be -1.
[0148] For example, the size of the mask image is 2560 (width) * 1440 (height), and if the pixel values of each pixel point in the row with the longitudinal coordinate of 301 in the mask image are all 0, the first position data corresponding to the row is: the longitudinal coordinate of the row: 301, and the third preset identifier: Flag = -1.
[0149] Referring to Figure 7 , Figure 7 A schematic diagram of the mask image when generating the target position data in mode one is provided for the embodiments of the application. Figure 7 In the embodiments of the application, the mask image can be divided into three image regions A1, B1 and C1. For the rows contained in the image region A1, the pixel values of each pixel point in each row indicate that the corresponding pixel position in the first image does not belong to the first image region, that is, the first position data corresponding to each row in the image region A1 includes the third preset identifier and the longitudinal coordinate of the row. For the rows contained in the image regions B1 and C1, the first position data corresponding to each row includes the pixel coordinates of the first edge pixel point in the row.
[0150] In a case where each pixel in a row in the mask image corresponds to a pixel position in the first image that does not belong to the first image region, the first position data corresponding to the row includes the third preset identifier and the vertical coordinate of the row. Compared with the prior art in which the pixel value of each pixel in the mask image needs to be transmitted, the application can further reduce the data amount of the generated target position data, reduce the bandwidth required for transmitting the data, and further improve the efficiency of the synthesized image.
[0151] It can be understood that, in a case where large-size and high-resolution images are synthesized, or in a case where wide-angle shooting results in a relatively small proportion of the target object in the first image, compared with the prior art in which the complete mask image is transmitted, the method provided in the embodiments of the application can further improve the efficiency of obtaining the synthesized image and reduce the bandwidth required for transmitting the target position data. In addition, since the mask image has the same size as the first image and the first device can transmit the vertical coordinate of each row in the mask image to the second device, the second device can determine whether data loss occurs in the transmission of the target position data based on the received vertical coordinate, and thus can verify the integrity of the data received by itself.
[0152] The second device can obtain the pixel values of the pixels included in the first image region in a manner corresponding to the generation of the target position data by the first device. If the target position data includes the first position data corresponding to a row in the mask image, the second device can use the pixels in a row as a unit to obtain the synthesized image, and correspondingly, the step S304 includes:
[0153] Step 1: For each first position data in the target position data, if the first position data includes the pixel coordinates of multiple pixels, the pixel values of the pixels between each pair of pixels in the first image region corresponding to the row to which the first position data belongs are obtained from the first image based on the pixel coordinates of the multiple pixels.
[0154] Each pair of pixels indicates that, in the first image, the pixels corresponding to two adjacent pixel coordinates in the pixel coordinates of the multiple pixels; any two pairs of pixels do not include the same pixel.
[0155] Step 2: If the first position data includes the first preset identifier and the vertical coordinate, the pixel values of the pixels in the row corresponding to the vertical coordinate in the first image are obtained, and the pixel values of the pixels in the first image region corresponding to the row to which the first position data belongs are obtained.
[0156] The first position data contains pixel coordinates of a plurality of pixels, and indicates that there are pixels in the row to which the first position data belongs and which have the same positions as the edge pixels of the first image region. That is, the first position data contains pixel coordinates of the first edge pixels determined based on step 2, and can also contain pixel coordinates of the pixels at the specified end points of the row determined based on step 21. The second device can obtain pixel values of the pixels in the first image region corresponding to the row to which the first position data belongs based on the plurality of pixel coordinates contained in the first position data.
[0157] If the first position data contains pixel coordinates of a plurality of pixels, the second device can sort the pixel coordinates of the plurality of pixels in ascending or descending order of the horizontal coordinates, determine two pixel coordinates adjacent to each other as pixel coordinates of two pixels in a pixel pair, and each pixel coordinate corresponds to only one pixel pair.
[0158] For example, a certain first position data contains: a vertical coordinate of 50, and horizontal coordinates of 0, 5, 30 and 90. The pixel pairs determined based on the first position data have two groups, the first group of pixel pairs contains (0, 50) and (5, 50), and the second group of pixel pairs contains (30, 50) and (90, 50).
[0159] The pixels between each pixel pair represent that among the pixels of the row to which the pixel pair belongs, the horizontal coordinates are between the smaller value of the horizontal coordinates of the pixels contained in the pixel pair and the larger value of the horizontal coordinates of the pixels contained in the pixel pair, and contain all the pixels of the pixel pair. Then, the second device can obtain pixel values of the pixels between each pixel pair from the first image, that is, obtain pixel values of the pixels in the first image region corresponding to the row to which the first position data belongs.
[0160] For example, the first group of pixel pairs contains (0, 50) and (5, 50), and the pixels between the pixel pairs contain (0, 50), (1, 50), (2, 50), (3, 50), (4, 50) and (5, 50). Further, the second device can obtain pixel values of the pixels at pixel positions (0, 50), (1, 50), (2, 50), (3, 50), (4, 50) and (5, 50) in the first image.
[0161] In addition, if the first position data corresponding to a row contains only one pixel coordinate, the second device can obtain the pixel value of the pixel at the pixel coordinate in the first image to obtain the pixel value of the pixel in the first image region corresponding to the row to which the first position data belongs.
[0162] If the first position data contains the first preset identifier and the ordinate, it indicates that each pixel point in the row corresponding to the ordinate in the first image belongs to the first image region. The second device can obtain the pixel value of each pixel point in the row corresponding to the ordinate from the first image, and obtain the pixel value of the pixel point corresponding to the row in the first image region to which the first position data belongs.
[0163] Based on the above processing, the pixel value of each row of pixels in the first image region can be obtained.
[0164] Correspondingly, if the first position data contains the third preset identifier and the ordinate, it indicates that each pixel point in the row corresponding to the ordinate in the first image does not belong to the first image region. Then, the second device can obtain the pixel value of each pixel point in the row corresponding to the ordinate from the second image, and obtain the pixel value of the pixel point corresponding to the row in the second image region to which the first position data belongs.
[0165] Based on the above processing, since the target position data contains the first position data corresponding to the row in the mask image, the second device can determine the pixel value of each pixel point in each row of the synthesized image in units of rows, and further obtain the synthesized image. After receiving the target position data, the second device can determine the pixel position of the pixel point contained in the corresponding first image region based on the position of the edge pixel point in each row, without obtaining the pixel value of the corresponding position for each pixel point in the synthesized image one by one. That is, the efficiency of the synthesized image can be further improved in the post-processing process.
[0166] Method two: the first device can generate target position data for the target object in a column scanning manner, that is, the first device can take the pixel points in a column of the mask image as a unit for generating target position data. Correspondingly, step S3022 comprises:
[0167] Step (1): for each column in the mask image, determine the pixel point at which the pixel value changes in the column as the first edge pixel point in the column.
[0168] Step (2): obtain the pixel coordinates of the first edge pixel point in the column to obtain the first position data corresponding to the column.
[0169] Step (3): based on the first position data corresponding to each column, obtain the target position data for the target object.
[0170] In the embodiment of the present application, each column in the mask image represents the pixel points in the column. For each column in the mask image, the first device can traverse the pixel values of the pixel points in the column to determine the first edge pixel points in the column. For example, the first device can traverse the pixel points in the column in the order of the vertical coordinates of the pixel points from small to large, or can traverse the pixel points in the column in the order of the vertical coordinates of the pixel points from large to small.
[0171] For each pixel point in the column, if the pixel value of the pixel point is different from the pixel values of the other pixel points adjacent to the pixel point in the column, and the pixel value of the pixel point indicates that the pixel position corresponding to the pixel point in the first image belongs to the first image region, the pixel point can be determined as the first edge pixel point in the column. That is, since the size of the mask image is consistent with the size of the first image, the pixel position of each first edge pixel point in the mask image is consistent with the pixel position of an edge pixel point of the first image region in the first image.
[0172] For example, the size of the mask image is 2560 (width) * 1440 (height), in the column with the horizontal coordinate of 200, the pixel position of the pixel point F0 is (200, 99), and the pixel value is 0; the pixel position of the pixel point F1 is (200, 100), and the pixel value is 1; the pixel position of the pixel point F2 is (200, 101), and the pixel value is 1; the pixel position of the pixel point F3 is (200, 102), and the pixel value is 0. Then, the first device can determine that the pixel points F1 and F2 are the first edge pixel points in the column.
[0173] In the related art, since each column in the mask image contains 1440 pixel points, for the column with the horizontal coordinate of 200 in the mask image, the first device needs to send the pixel values of each pixel point in the column, i.e., the pixel values of 1440 pixel points, to the second device. In the embodiment of the present application, the first device can send the pixel coordinates of the pixel points F1 and F2 to the second device, and the data amount of the two pixel coordinates is much smaller than the data amount of the pixel values of a column of pixel points. That is, based on this, the first device can reduce the data amount of the data sent to the second device, reduce the bandwidth required for transmitting the data, reduce the time length consumed for transmitting the data, and improve the efficiency of the synthesized image.
[0174] Alternatively, for any pixel point in the column, if the pixel value of the pixel point is different from the pixel values of the other pixel points adjacent to the pixel point in the column, and the pixel value of the pixel point indicates that the pixel position corresponding to the pixel point in the first image does not belong to the first image region, the pixel point can be determined as the first edge pixel point in the column.
[0175] For example, for the pixel points F0-F3, the first device can determine the pixel point F0 and the pixel point F3 as the first edge pixel points in the column.
[0176] For each first edge pixel point in the column, the pixel coordinate of the first edge pixel point is composed of the horizontal coordinate and the vertical coordinate of the first edge pixel point. It can be understood that the horizontal coordinates of the first edge pixel points in a column are the same. For example, for the pixel point F1 and the pixel point F2, the horizontal coordinate of the pixel point F1 is 200, and the vertical coordinate is 100; the horizontal coordinate of the pixel point F2 is 200, and the vertical coordinate is 101.
[0177] In addition, when the edge of the first image region overlaps with the upper edge or the lower edge of the first image, since the pixel value of the pixel point at the position in the mask image is the same as the pixel value of the adjacent pixel point in the column, the pixel point at the position in the mask image will not be determined as the first edge pixel point.
[0178] Correspondingly, the above step (2) comprises:
[0179] Step (21): obtaining the pixel coordinates of the first edge pixel points in the column and the pixel coordinates of the pixel points at the specified endpoints of the column, to obtain the first position data corresponding to the column.
[0180] The pixel value of the pixel point at the specified endpoint indicates that the corresponding pixel position in the first image belongs to the first image region.
[0181] For each column in the mask image, there are two endpoints in the column, which are the pixel point with the smallest vertical coordinate in the column and the pixel point with the largest vertical coordinate in the column.
[0182] For the pixel point at any endpoint in the column, there is only one other pixel point adjacent to it in the column. If the pixel value of the pixel point at the endpoint indicates that the corresponding pixel position in the first image belongs to the first image region, the endpoint is the specified endpoint, and the pixel point at the endpoint is the pixel point at the specified endpoint.
[0183] For example, the size of the mask image is 2560*1440, in the column with the horizontal coordinate of 200, the pixel position of the pixel point F4 is (200, 0), and the pixel value is 1; the pixel position of the pixel point F5 is (200, 1), and the pixel value is 1; the pixel position of the pixel point F6 is (200, 1438), and the pixel value is 0; the pixel position of the pixel point F7 is (200, 1439), and the pixel value is 0. Then, the pixel point F4 is the pixel point at the specified endpoint, the pixel point F7 is the pixel point at the endpoint, but not the pixel point at the specified endpoint. Correspondingly, the first position data corresponding to the column also includes the pixel coordinate (200, 0) of the pixel point F4.
[0184] For each column in the mask image, the first position data corresponding to the column can represent the edge of the first image region in the column. That is, for any column in the mask image, if there is a pixel point corresponding to the first image region in the column, the pixel coordinates determined based on the above step (21) are even numbers.
[0185] Based on the above processing, when the edge of the first image region occupied by the target object overlaps with the edge of the first image, the first device can determine the pixel position of the pixel point at the overlapping position, and further make the target position data represent the position of the edge pixel point of the complete first image region.
[0186] In an implementation manner, the pixel coordinates of the first edge pixel points in the column include the horizontal coordinates and the vertical coordinates of the first edge pixel points in the column.
[0187] For example, for a column in the mask image, if the first edge pixel points in the column include pixel point E3 and pixel point E4, the pixel coordinates of the first edge pixel points in the column include the horizontal coordinate x3 and the vertical coordinate y3 of the pixel point E3, and the horizontal coordinate x3 and the vertical coordinate y4 of the pixel point E4.
[0188] In another implementation manner, the pixel coordinates of the first edge pixel points in the column include the horizontal coordinate of the column, and the vertical coordinates of the first edge pixel points in the column.
[0189] That is, for the plurality of first edge pixel points in the same column in the mask image, the pixel coordinates of the first edge pixel points in the column can be represented by the horizontal coordinate of the column and the vertical coordinates of the first edge pixel points in the column.
[0190] For example, for a column in the mask image, if the first edge pixel points in the column include the above pixel point E3 and pixel point E4, the pixel coordinates of the first edge pixel points in the column include the horizontal coordinate x3 of the column, and the vertical coordinate y3 of the pixel point E3 and the vertical coordinate y4 of the pixel point E4.
[0191] Based on the above processing, in the case that there are a plurality of first edge pixel points in a column in the mask image, the same horizontal coordinates in the pixel coordinates of the plurality of first edge pixel points can be avoided to be repeatedly recorded in the obtained target position data, and further, the data amount of the generated target position data can be further reduced, the bandwidth required for transmitting data can be reduced, and the efficiency of the synthesized image can be improved.
[0192] In addition, since the first position data corresponding to each column of the mask image includes the pixel coordinates of the first edge pixel points in the column and the pixel coordinates of the pixel points at the specified endpoints of the column, the pixel coordinates of the first edge pixel points in each column of the mask image can be obtained based on the first position data corresponding to each column, and then the second device can determine the positions of the edge pixel points of the first image region in the first image based on the target position data.
[0193] In some embodiments, the method further includes:
[0194] Step (4): For each column of the mask image, if the pixel values of the pixel points in the column all indicate that the corresponding pixel positions in the first image belong to the first image region, the pixel coordinates of the pixel points at the two endpoints of the column are obtained as the first position data corresponding to the column.
[0195] Alternatively,
[0196] Step (5): The second preset identifier and the horizontal coordinate of the column are obtained as the first position data corresponding to the column.
[0197] The second preset identifier indicates that the corresponding pixel positions of the pixel points in a column in the first image all belong to the first image region.
[0198] The description of the endpoints in each column of the mask image can refer to the above embodiments.
[0199] For each column of the mask image, if the pixel values of the pixel points in the column all indicate that the corresponding pixel positions in the first image belong to the first image region, the first device can set the value of the corresponding identifier bit (Flag) for the column (i.e., the second preset identifier), for example, the value of the identifier bit can be 0.
[0200] For example, the size of the mask image is 2560 (width) * 1440 (height), and if the pixel values of the pixel points in the column with a horizontal coordinate of 300 in the mask image are all 1, the first position data corresponding to the column includes (300, 0) and (300, 1439). Alternatively, the first position data corresponding to the column includes the horizontal coordinate (300) of the column and the second preset identifier (Flag = 0).
[0201] Based on the above processing, in the case that the corresponding pixel positions of each pixel point in a column in the mask image in the first image all belong to the first image region, the first position data corresponding to the column only contains the pixel coordinates of the pixel points at the two endpoints of the column, or the second preset identifier and the horizontal coordinate of the column. Compared with the prior art which needs to transmit the pixel value of each pixel point in the mask image, the application can further reduce the data amount of the generated target position data, reduce the bandwidth required for transmitting data, and further improve the efficiency of the synthesized image.
[0202] In some embodiments, the method further comprises:
[0203] For each column in the mask image, if the pixel values of each pixel point in the column all indicate that the corresponding pixel position in the first image does not belong to the first image region, the fourth preset identifier and the horizontal coordinate of the column are obtained as the first position data corresponding to the column.
[0204] The fourth preset identifier indicates that the corresponding pixel positions of each pixel point in a column in the first image all do not belong to the first image region.
[0205] In the embodiments of the application, if the pixel values of each pixel point in the column all indicate that the corresponding pixel position in the first image does not belong to the first image region, the first device can set the value of the corresponding identifier bit (Flag) for the column (i.e. the fourth preset identifier), for example, the value of the identifier bit can be -1.
[0206] For example, the size of the mask image is 2560 (width) * 1440 (height), and if the pixel values of each pixel point in the column with a horizontal coordinate of 301 in the mask image are all 0, the first position data corresponding to the column is: the horizontal coordinate of the column: 301, and the fourth preset identifier: Flag = -1.
[0207] Referring to Figure 8 , Figure 8 A schematic diagram of the mask image when generating the target position data in the manner two is provided for the embodiments of the application. Figure 8 In the embodiments of the application, the mask image can be divided into two image regions A2, B2 and C2. For the columns contained in the image region A2 and the image region C2, the pixel values of each pixel point in each column indicate that the corresponding pixel position in the first image does not belong to the first image region, that is, the first position data corresponding to each column in the image region A2 and the image region C2 includes the fourth preset identifier and the horizontal coordinate of the column. For the columns contained in the image region B2, the first position data corresponding to each column includes the pixel coordinates of the first edge pixel point in the column.
[0208] In a case where each pixel in a column in the mask image corresponds to a pixel position in the first image that does not belong to the first image region, the first position data corresponding to the column comprises the fourth preset identifier and the horizontal coordinate of the column. Compared with the prior art in which the pixel value of each pixel in the mask image needs to be transmitted, the application can further reduce the data amount of the generated target position data, reduce the bandwidth required for transmitting the data, and further improve the efficiency of the synthesized image.
[0209] It can be understood that, in a case where large-size and high-resolution images are synthesized, or in a case where wide-angle shooting results in a relatively small proportion of the target object in the first image, compared with the prior art in which the complete mask image is transmitted, the method provided in the embodiments of the application can further improve the efficiency of obtaining the synthesized image and reduce the bandwidth required for transmitting the target position data. In addition, since the mask image has the same size as the first image and the first device can transmit the horizontal coordinate of each column in the mask image to the second device, the second device can determine whether data loss occurs in the transmission of the target position data based on the received horizontal coordinate, and thus can verify the integrity of the data received by itself.
[0210] The second device can obtain the pixel values of the pixels included in the first image region in a manner corresponding to the generation of the target position data by the first device. If the target position data comprises the first position data corresponding to a column in the mask image, the second device can use the pixels in a column as a unit to obtain the synthesized image, and correspondingly, the step S304 comprises:
[0211] Step (I): for each first position data in the target position data, if the first position data comprises pixel coordinates of a plurality of pixels, the pixel values of the pixels between each pixel pair in the first image region in the column to which the first position data belongs are obtained from the first image based on the pixel coordinates of the plurality of pixels, to obtain the pixel values of the pixels in the first image region corresponding to the column.
[0212] Each pixel pair indicates that, in the first image, two adjacent pixel coordinates in the pixel coordinates of the plurality of pixels correspond to a pixel, and any two pixel pairs do not include the same pixel.
[0213] Step (II): if the first position data comprises the second preset identifier and the horizontal coordinate, the pixel values of the pixels in the column corresponding to the horizontal coordinate in the first image are obtained to obtain the pixel values of the pixels in the first image region corresponding to the column to which the first position data belongs.
[0214] The first position data contains pixel coordinates of a plurality of pixels, indicating that there are pixels in the column to which the first position data belongs and which have the same position as the edge pixels of the first image region. That is, the first position data contains the pixel coordinates of the first edge pixels determined based on step (2), and can also contain the pixel coordinates of the pixels at the specified end points of the column determined based on step (21). The second device can obtain the pixel values of the pixels in the first image region corresponding to the column to which the first position data belongs based on the plurality of pixel coordinates contained in the first position data.
[0215] If the first position data contains pixel coordinates of a plurality of pixels, the second device can sort the pixel coordinates of the plurality of pixels in ascending or descending order of the vertical coordinates, determine two pixel coordinates adjacent to each other as the pixel coordinates of two pixels in a pixel pair, and each pixel coordinate corresponds to only one pixel pair.
[0216] For example, a certain first position data contains: horizontal coordinate: 50, and vertical coordinates: 0, 5, 30 and 90. The pixel pairs determined based on the first position data have two groups, the first group of pixel pairs contains (50, 0) and (50, 5), and the second group of pixel pairs contains (50, 30) and (50, 90).
[0217] The pixels between each pixel pair represent that among the pixels of the column to which the pixel pair belongs, the vertical coordinates are between the smaller value and the larger value of the vertical coordinates of the pixels contained in the pixel pair, and contain all the pixels of the pixel pair. Then, the second device can obtain the pixel values of the pixels between each pixel pair from the first image, that is, obtain the pixel values of the pixels in the first image region corresponding to the column to which the first position data belongs.
[0218] For example, the first group of pixel pairs contains (50, 0) and (50, 5), and the pixels between the pixel pairs contain: (50, 0), (50, 1), (50, 2), (50, 3), (50, 4) and (50, 5). Further, the second device can obtain the pixel values of the pixels at pixel positions (50, 0), (50, 1), (50, 2), (50, 3), (50, 4) and (50, 5) in the first image.
[0219] In addition, if the first position data corresponding to a column contains only one pixel coordinate, the second device can obtain the pixel value of the pixel at the pixel coordinate in the first image, and obtain the pixel value of the pixel in the first image region corresponding to the column to which the first position data belongs.
[0220] If the first position data contains the first preset identifier and the abscissa, it indicates that each pixel point in the column corresponding to the abscissa in the first image belongs to the first image region. The second device can obtain the pixel value of each pixel point in the column corresponding to the abscissa from the first image to obtain the pixel value of the corresponding pixel point of the column to which the first position data belongs in the first image region.
[0221] Based on the above processing, the pixel value of each column of pixels in the first image region can be obtained.
[0222] Correspondingly, if the first position data contains the fourth preset identifier and the abscissa, it indicates that each pixel point in the column corresponding to the abscissa in the first image does not belong to the first image region. Then the second device can obtain the pixel value of each pixel point in the column corresponding to the abscissa from the second image to obtain the pixel value of the corresponding pixel point of the column to which the first position data belongs in the second image region.
[0223] Based on the above processing, since the target position data contains the first position data corresponding to the column in the mask image, the second device can determine the pixel value of each pixel point in each column in the synthesized image in units of columns, and further obtain the synthesized image. Moreover, after receiving the target position data, the second device can determine the pixel position of the pixel point contained in the corresponding first image region based on the position of the edge pixel point in each column, without obtaining the pixel value of the corresponding position for each pixel point in the synthesized image one by one, that is, the efficiency of the synthesized image can be further improved in the post-processing process.
[0224] Method three: step S3022 includes:
[0225] Step 1: For each row in the mask image, determine the pixel point where the pixel value changes in the row as the first edge pixel point in the row; obtain the pixel coordinates of the first edge pixel point in the row to obtain the first position data corresponding to the row; based on the first position data corresponding to each row, obtain the first candidate data representing the position of the edge pixel point of the first image region.
[0226] Step 2: For each column in the mask image, determine the pixel point where the pixel value changes in the column as the first edge pixel point in the column; obtain the pixel coordinates of the first edge pixel point in the column to obtain the first position data corresponding to the column; based on the first position data corresponding to each column, obtain the second candidate data representing the position of the edge pixel point of the first image region.
[0227] Step 3: If the data amount of the first candidate data is less than the data amount of the second candidate data, the first candidate data is determined as the target position data for the target object.
[0228] Step 4: If the data amount of the first candidate data is not less than the data amount of the second candidate data, the second candidate data is determined as the target position data for the target object.
[0229] The first image region occupied by the target object in the first image can be a relatively complex irregular region. For example, referring to FIG. 1, if the target position data is generated in a manner one, the first position data corresponding to each row in the image region A1 at least includes one longitudinal coordinate and a third preset identifier, the first position data corresponding to each row in the image region B1 at least includes one longitudinal coordinate and two transverse coordinates, and the first position data corresponding to each row in the image region C1 at least includes one longitudinal coordinate and four transverse coordinates. Figure 7 and Figure 8 , Figure 7
[0230] Figure 8 If the target position data is generated in a manner two, the first position data corresponding to each column in the image region A2 and C2 at least includes one transverse coordinate and a fourth preset identifier, and the first position data corresponding to each column in the image region B2 at least includes one transverse coordinate and two longitudinal coordinates.
[0231] For the image regions such as the image regions A1, A2 and C2, since they do not include the pixel points corresponding to the first image region, that is, for each row or column in the image regions, the data amount of the first position data corresponding to the row or column is small. Therefore, if the proportion of the image region A1 in the mask image is less than the sum of the proportions of the image regions A2 and C2 in the mask image, the data amount of the first candidate data generated by the first device through the row scanning manner can be greater than the data amount of the second candidate data generated by the first device through the column scanning manner.
[0232] In view of the above, the first device can generate the first candidate data representing the positions of the edge pixel points of the first image region through the row scanning manner (i.e., the manner one), and generate the second candidate data representing the positions of the edge pixel points of the first image region through the column scanning manner (i.e., the manner two), and then compare the data amount of the first candidate data with the data amount of the second candidate data, and select the one with smaller data amount as the target position data.
[0233] In addition, the first device can also send an identifier representing the manner of generating the target position data to the second device, so that the second device acquires the pixel values of the pixel points included in the first image region according to the manner corresponding to the identifier.
[0234] Based on the above processing, the first device can generate the first candidate data and the second candidate data representing the positions of the edge pixel points of the first image region in different manners, and select the target position data with smaller data amount, thereby further reducing the data amount required to be transmitted between the first device and the second device, reducing the bandwidth required for transmitting the data, and improving the efficiency of the synthesized image.
[0235] That is, in the present application, the manner in which the first device generates the target position data and the manner in which the second device obtains the pixel values of the pixel points included in the first image region can be pre-agreed, for example, the first device can generate the target position data in a row scanning manner, and the second device can obtain the pixel values of the pixel points included in the first image region in a row unit.
[0236] Alternatively, the first device can also select the manner of generating the target position data according to the first image, and notify the second device so that the second device obtains the synthesized image in the corresponding manner. That is, the first device can select the target position data with smaller data amount according to the actual scene.
[0237] As shown in Figure 9 , a flowchart for obtaining a synthesized image provided by an embodiment of the present application includes the following steps: Figure 9
[0238] S901: Start.
[0239] S902: Transmit a first image.
[0240] That is, the second device transmits the first image including the target object to the first device.
[0241] S903: Judge the extraction direction, and obtain horizontal extraction information or vertical extraction information.
[0242] That is, the first device generates a mask image for the target object based on the first image.
[0243] Then, the first device can determine, for each row in the mask image, the pixel point at which the pixel value changes in the row as the first edge pixel point in the row, obtain the pixel coordinates of the first edge pixel point in the row to obtain the first position data corresponding to the row, and obtain the first candidate data representing the positions of the edge pixel points of the first image region based on the first position data corresponding to each row.
[0244] For each column in the mask image, the first device determines the pixel point at which the pixel value changes in the column as the first edge pixel point in the column, obtains the pixel coordinates of the first edge pixel point in the column to obtain the first position data corresponding to the column, and obtains the second candidate data representing the positions of the edge pixel points of the first image region based on the first position data corresponding to each column.
[0245] If the data amount of the first candidate data is less than the data amount of the second candidate data, the first candidate data is determined as the target position data for the target object; if the data amount of the first candidate data is not less than the data amount of the second candidate data, the second candidate data is determined as the target position data for the target object.
[0246] S904: Return the extracted alpha information.
[0247] That is, the first device sends the target position data to the second device.
[0248] S905: Post-processing to obtain the matting result.
[0249] That is, after receiving the target position data, the second device obtains the pixel values of each pixel point contained in the first image region based on the target position data, and obtains a composite image by combining the pixel values of each pixel point contained in the second image region.
[0250] The second image region represents other image regions in the second image except for the image region corresponding to the first image region.
[0251] It should be noted that the two-dimensional face image in the embodiment comes from a public data set.
[0252] Based on the same inventive concept, the embodiment of the present application also provides an image synthesis device. The device is applied to a first device. Referring to Figure 10 , Figure 10 A structural diagram of an image synthesis device provided by the embodiment of the present application is provided. The device includes:
[0253] The first receiving module 1001 is configured to receive the first image containing the target object sent by the second device.
[0254] The target position data generation module 1002 is configured to generate target position data for the target object based on the first image. The target position data represents the position of the edge pixel point of the first image region occupied by the target object in the first image.
[0255] The first sending module 1003 is configured to send the target position data to the second device, so that the second device obtains the pixel values of each pixel point contained in the first image region based on the target position data after receiving the target position data, and obtains a composite image by combining the pixel values of each pixel point contained in the second image region. The second image region represents other image regions in the second image except for the image region corresponding to the first image region.
[0256] In some embodiments, the target position data generation module 1002 includes: a mask image generation submodule configured to generate a mask image for the target object based on the first image; wherein a pixel value of each pixel point in the mask image indicates whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; and a target position data generation submodule configured to generate target position data for the target object based on the pixel value of each pixel point in the mask image.
[0257] In some embodiments, the target position data generation submodule includes: a first determination unit configured to determine, for each row in the mask image, a pixel point at which a pixel value changes in the row as a first edge pixel point in the row; a first acquisition unit configured to acquire a pixel coordinate of the first edge pixel point in the row to obtain first position data corresponding to the row; and a first generation unit configured to obtain the target position data for the target object based on the first position data corresponding to each row.
[0258] In some embodiments, the first acquisition unit is specifically configured to: acquire the pixel coordinate of the first edge pixel point in the row and a pixel coordinate of a pixel point at a specified end point of the row to obtain the first position data corresponding to the row; wherein the pixel value of the pixel point at the specified end point indicates that a corresponding pixel position in the first image belongs to the first image region.
[0259] In some embodiments, the pixel coordinate of the first edge pixel point in the row includes: a vertical coordinate of the row and a horizontal coordinate of each first edge pixel point in the row.
[0260] In some embodiments, the apparatus further includes: a first acquisition module configured to, for each row in the mask image, if the pixel value of each pixel point in the row indicates that a corresponding pixel position in the first image belongs to the first image region, acquire a pixel coordinate of a pixel point at each end point of the row as first position data corresponding to the row; or a second acquisition module configured to acquire a first preset identifier and a vertical coordinate of the row as the first position data corresponding to the row; wherein the first preset identifier indicates that the corresponding pixel positions of each pixel point in a row in the first image all belong to the first image region.
[0261] In some embodiments, the target position data generation submodule includes: a second determination unit configured to determine, for each column in the mask image, a pixel point at which a pixel value changes in the column as a first edge pixel point in the column; a second acquisition unit configured to acquire a pixel coordinate of the first edge pixel point in the column to obtain first position data corresponding to the column; and a second generation unit configured to obtain the target position data for the target object based on the first position data corresponding to each column.
[0262] In some embodiments, the second obtaining unit is specifically configured to: obtain the pixel coordinates of the first edge pixel points in the column and the pixel coordinates of the pixel points at the specified end points of the column, to obtain the first position data corresponding to the column; wherein the pixel value of the pixel point at the specified end point indicates that the corresponding pixel position in the first image belongs to the first image region.
[0263] In some embodiments, the pixel coordinates of the first edge pixel points in the column include: the horizontal coordinates of the column and the vertical coordinates of the first edge pixel points in the column.
[0264] In some embodiments, the device further comprises: a third obtaining module configured to, for each column in the mask image, if the pixel values of the pixel points in the column all indicate that the corresponding pixel positions in the first image belong to the first image region, obtain the pixel coordinates of the pixel points at the two end points of the column as the first position data corresponding to the column; or a fourth obtaining module configured to obtain the second preset identifier and the horizontal coordinates of the column as the first position data corresponding to the column; wherein the second preset identifier indicates that the corresponding pixel positions of the pixel points in a column all belong to the first image region.
[0265] In some embodiments, the target position data generation submodule comprises: a first candidate data generation unit configured to, for each row in the mask image, determine the pixel points at which the pixel values change in the row as the first edge pixel points in the row; obtain the pixel coordinates of the first edge pixel points in the row to obtain the first position data corresponding to the row; and obtain the first candidate data representing the positions of the edge pixel points of the first image region based on the first position data corresponding to each row; a second candidate data generation unit configured to, for each column in the mask image, determine the pixel points at which the pixel values change in the column as the first edge pixel points in the column; obtain the pixel coordinates of the first edge pixel points in the column to obtain the first position data corresponding to the column; and obtain the second candidate data representing the positions of the edge pixel points of the first image region based on the first position data corresponding to each column; a third determination unit configured to, if the data amount of the first candidate data is less than the data amount of the second candidate data, determine the first candidate data as the target position data for the target object; and a fourth determination unit configured to, if the data amount of the first candidate data is not less than the data amount of the second candidate data, determine the second candidate data as the target position data for the target object.
[0266] Based on the same inventive concept, the embodiments of the present application also provide an image synthesis device. The device is applied to a second device. Referring to Figure 11 , Figure 11 Another structural diagram of an image synthesis device provided by the embodiments of the present application is provided. The device comprises:
[0267] The second sending module 1101 is configured to send a first image containing a target object to the first device, so that the first device generates target position data for the target object based on the first image, and sends the target position data to the second device; wherein the target position data indicates a position of an edge pixel point of a first image region occupied by the target object in the first image.
[0268] The second receiving module 1102 is configured to, after receiving the target position data, acquire pixel values of each pixel point contained in the first image region based on the target position data.
[0269] The synthesizing module 1103 is configured to combine the pixel values of each pixel point contained in the first image region and the pixel values of each pixel point contained in the second image region to obtain a synthesized image; wherein the second image region represents other image regions in the second image except for an image region corresponding to the first image region.
[0270] In some embodiments, the target position data contains first position data corresponding to a row in a mask image; a pixel value of each pixel point in the mask image indicates whether a pixel position corresponding to the pixel point in the first image belongs to the first image region; the second receiving module 1102 is specifically configured to, for each first position data in the target position data, if the first position data contains pixel coordinates of a plurality of pixel points, acquire pixel values of each pixel point pair in the first image based on the pixel coordinates of the plurality of pixel points, to obtain pixel values of pixel points corresponding to the first image region in a row to which the first position data belongs; wherein each pixel point pair represents, in the first image, pixel points corresponding to two adjacent pixel coordinates in the pixel coordinates of the plurality of pixel points; any two pixel point pairs do not contain the same pixel point; if the first position data contains a first preset identifier and a vertical coordinate, acquire pixel values of each pixel point in a row corresponding to the vertical coordinate from the first image, to obtain pixel values of pixel points corresponding to the first image region in the row to which the first position data belongs.
[0271] In some embodiments, the target location data includes: first location data corresponding to columns in a mask image; the pixel value of each pixel in the mask image indicates whether the pixel position corresponding to the pixel in the first image belongs to the first image region; the second receiving module 1102 is specifically used for: for each first location data in the target location data, if the first location data contains pixel coordinates of multiple pixels, then based on the pixel coordinates of multiple pixels, obtaining the pixel value of each pair of pixels in the first image to obtain the pixel value of the first image region corresponding to the pixel in the column to which the first location data belongs; wherein, each pair of pixels indicates: the pixel corresponding to two adjacent pixel coordinates in the pixel coordinates of multiple pixels in the first image; any two pairs of pixels do not contain the same pixel; if the first location data contains a second preset identifier and an abscissa, then obtaining the pixel value of each pixel in the column corresponding to the abscissa in the first image to obtain the pixel value of the first image region corresponding to the pixel in the column to which the first location data belongs.
[0272] In some embodiments, the synthesis module 1103 is specifically used to: for each pixel in the first image region, replace the pixel value of the pixel in the second image that has the same position as the pixel with the pixel value in the first image, so as to obtain a synthesized image.
[0273] This application also provides an electronic device, such as... Figure 12 As shown, the device includes: a memory 1201 for storing computer programs; and a processor 1202 for executing the program stored in the memory 1201 to implement the steps of any of the image synthesis methods described in the above embodiments. Furthermore, the electronic device may also include a communication bus and / or a communication interface, through which the processor 1202, the communication interface, and the memory 1201 communicate with each other.
[0274] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0275] The communication interface is configured to communicate with other devices. The memory can include a random access memory (RAM) and can also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor. The aforementioned processor can be a general purpose processor, including a central processing unit (CPU), a network processor (NP), etc., and can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components.
[0276] In a further embodiment provided in the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps of any of the above image synthesis methods.
[0277] In a further embodiment provided in the present application, a computer program product containing instructions, which, when run on a computer, causes the computer to execute any of the image synthesis methods in the above embodiments.
[0278] In the embodiments described above, all or some of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or some of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded into and executed by a computer, all or some of the processes or functions according to the embodiments described in the specification are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a storage medium (for example, Solid State Disk (SSD)) and the like.
[0279] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0280] Each of the embodiments in the specification is described in a related manner, and the same or similar parts between each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for system, device, electronic device, computer readable storage medium, and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0281] The above merely provides the preferred embodiment of the present application, and not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An image synthesis method, characterized in that, The method is applied to a first device, and the method includes: Receive a first image containing the target object sent by the second device; Based on the first image, a mask image for the target object is generated; wherein, the pixel value of each pixel in the mask image indicates whether the pixel position corresponding to that pixel in the first image belongs to the region of the first image; Based on the pixel value of each pixel in the mask image, target position data for the target object is generated; wherein, the target position data represents the position of the edge pixel of the first image region occupied by the target object in the first image; The target location data is sent to the second device so that, upon receiving the target location data, the second device obtains the pixel values of each pixel in the first image region based on the target location data, and combines them with the pixel values of each pixel in the second image region to obtain a composite image; wherein, the second image region refers to other image regions in the second image besides the image region corresponding to the first image region; The step of generating target location data for the target object based on the pixel value of each pixel in the mask image includes: for each row in the mask image, determining the pixel point where the pixel value changes in that row as the first edge pixel point in that row; obtaining the pixel coordinates of the first edge pixel point in that row to obtain the first location data corresponding to that row; and obtaining the target location data for the target object based on the first location data corresponding to each row. Alternatively, generating target location data for the target object based on the pixel value of each pixel in the mask image includes: for each column in the mask image, determining the pixel point where the pixel value changes in that column as the first edge pixel point in that column; obtaining the pixel coordinates of the first edge pixel point in that column to obtain the first location data corresponding to that column; and obtaining the target location data for the target object based on the first location data corresponding to each column. Alternatively, generating target location data for the target object based on the pixel value of each pixel in the mask image includes: for each row in the mask image, determining the pixels where the pixel value changes in that row as first edge pixels in that row; obtaining the pixel coordinates of the first edge pixels in that row to obtain first location data corresponding to that row; obtaining first candidate data representing the location of edge pixels in the first image region based on the first location data corresponding to each row; for each column in the mask image, determining the pixels where the pixel value changes in that column as first edge pixels in that column; obtaining the pixel coordinates of the first edge pixels in that column to obtain first location data corresponding to that column; obtaining second candidate data representing the location of edge pixels in the first image region based on the first location data corresponding to each column; if the amount of data in the first candidate data is less than the amount of data in the second candidate data, then the first candidate data is determined as target location data for the target object; if the amount of data in the first candidate data is not less than the amount of data in the second candidate data, then the second candidate data is determined as target location data for the target object.
2. The method according to claim 1, characterized in that, The step of obtaining the pixel coordinates of the first edge pixel in the row to obtain the first position data corresponding to the row includes: Obtain the pixel coordinates of the first edge pixel in the row and the pixel coordinates of the pixel at the specified endpoint of the row to obtain the first position data corresponding to the row; wherein, the pixel value of the pixel at the specified endpoint indicates that the corresponding pixel position in the first image belongs to the first image region.
3. The method according to claim 2, characterized in that, The pixel coordinates of the first edge pixel in the row include: the vertical coordinate of the row, and the horizontal coordinate of each first edge pixel in the row.
4. The method according to claim 2, characterized in that, The method further includes: For each row in the mask image, if the pixel value of each pixel in the row represents the pixel position in the first image that belongs to the first image region, then the pixel coordinates of the two endpoints of the row are obtained as the first position data corresponding to the row. or, Obtain the first preset identifier and the vertical coordinate of the row as the first position data corresponding to the row; wherein, the first preset identifier indicates that the pixel position of each pixel in the row in the first image belongs to the first image region.
5. The method according to claim 1, characterized in that, The step of obtaining the pixel coordinates of the first edge pixel in the column to obtain the first position data corresponding to the column includes: Obtain the pixel coordinates of the first edge pixel in the column and the pixel coordinates of the pixel at the specified endpoint of the column to obtain the first position data corresponding to the column; wherein, the pixel value of the pixel at the specified endpoint indicates that the corresponding pixel position in the first image belongs to the first image region.
6. The method according to claim 1, characterized in that, The pixel coordinates of the first edge pixel in this column include: the horizontal coordinate of the column, and the vertical coordinate of each first edge pixel in the column.
7. The method according to claim 1, characterized in that, The method further includes: For each column in the mask image, if the pixel value of each pixel in the column represents the pixel position in the first image that belongs to the first image region, then the pixel coordinates of the two endpoints of the column are obtained as the first position data corresponding to the column. or, Obtain the second preset identifier and the horizontal coordinate of the column as the first position data corresponding to the column; wherein, the second preset identifier indicates that the pixel position of each pixel in the column in the first image belongs to the first image region.
8. An image synthesis method, characterized in that, The method is applied to a second device, and the method includes: A first image containing a target object is sent to a first device, so that the first device generates target location data for the target object based on the first image, and sends the target location data to a second device; wherein, the target location data represents the position of the edge pixels of the first image region occupied by the target object in the first image; After receiving the target location data, the pixel values of each pixel point contained in the first image region are obtained based on the target location data; By combining the pixel values of each pixel in the first image region and the pixel values of each pixel in the second image region, a composite image is obtained; wherein, the second image region refers to other image regions in the second image besides the image region corresponding to the first image region; The target location data includes: first location data corresponding to rows in the mask image; the pixel value of each pixel in the mask image indicates whether the pixel position corresponding to that pixel in the first image belongs to the first image region; The step of obtaining the pixel values of each pixel point contained in the first image region based on the target location data includes: For each first location data in the target location data, if the first location data contains the pixel coordinates of multiple pixels, then based on the pixel coordinates of the multiple pixels, the pixel value of each pixel pair is obtained from the first image to obtain the pixel value of the corresponding pixel in the row to which the first location data belongs in the first image region; wherein, each pixel pair represents: in the first image, the pixel corresponding to two adjacent pixel coordinates among the pixel coordinates of the multiple pixels; any two pixel pairs do not contain the same pixel; If the first position data includes a first preset identifier and a vertical coordinate, then the pixel value of each pixel in the row corresponding to the vertical coordinate is obtained from the first image, and the pixel value of the pixel in the row to which the first position data belongs is obtained. Alternatively, the target location data may include: first location data corresponding to columns in the mask image; The step of obtaining the pixel values of each pixel point contained in the first image region based on the target location data includes: For each first location data in the target location data, if the first location data contains the pixel coordinates of multiple pixels, then based on the pixel coordinates of the multiple pixels, the pixel value of each pair of pixels in the first image is obtained, and the pixel value of the corresponding pixel in the column to which the first image region belongs is obtained. If the first location data includes a second preset identifier and a horizontal coordinate, then the pixel value of each pixel in the column corresponding to the horizontal coordinate is obtained from the first image, and the pixel value of the pixel in the column to which the first location data belongs is obtained.
9. The method according to claim 8, characterized in that, The step of combining the pixel values of each pixel in the first image region and the pixel values of each pixel in the second image region to obtain a composite image includes: For each pixel in the first image region, the pixel value of the pixel in the second image that has the same position as that pixel is replaced with the pixel value of that pixel in the first image to obtain a composite image.
10. An image synthesis system for performing the method of claim 1 or 8, characterized in that, The system includes a first device and a second device, wherein: The second device is used to send a first image containing the target object to the first device; The first device is configured to, upon receiving the first image, generate target location data for the target object based on the first image, and send the target location data to the second device; wherein, the target location data represents the position of the edge pixels of the first image region occupied by the target object in the first image; The second device is further configured to, upon receiving the target location data, obtain the pixel values of each pixel point contained in the first image region based on the target location data; and combine the pixel values of each pixel point contained in the first image region with the pixel values of each pixel point contained in the second image region to obtain a composite image; wherein, the second image region refers to other image regions in the second image besides the image region corresponding to the first image region.
11. An image synthesis apparatus for performing the method of claim 1, characterized in that, The device is applied to the first apparatus, the device comprising: The first receiving module is used to receive a first image containing a target object sent by the second device; The target location data generation module is used to generate target location data for the target object based on the first image; wherein, the target location data represents the position of the edge pixels of the first image region occupied by the target object in the first image; The first sending module is used to send the target location data to the second device, so that after receiving the target location data, the second device can obtain the pixel values of each pixel in the first image region based on the target location data, and combine them with the pixel values of each pixel in the second image region to obtain a composite image; wherein, the second image region refers to other image regions in the second image besides the image region corresponding to the first image region.
12. An image synthesis apparatus for performing the method of claim 8, characterized in that, The device is applied to a second device, the device comprising: The second sending module is used to send a first image containing a target object to the first device, so that the first device generates target location data for the target object based on the first image, and sends the target location data to the second device; wherein, the target location data represents the position of the edge pixels of the first image area occupied by the target object in the first image; The second receiving module is used to obtain the pixel values of each pixel point contained in the first image region based on the target location data after receiving the target location data. The compositing module is used to combine the pixel values of each pixel point contained in the first image region and the pixel values of each pixel point contained in the second image region to obtain a composite image; wherein, the second image region refers to other image regions in the second image besides the image region corresponding to the first image region.
13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-7 or 8-9.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-7 or 8-9.
Citation Information
Patent Citations
Image synthesis method, device and equipment
CN112150398A
Image zooming method and device, electronic equipment and storage medium
CN112508793A