Method, apparatus and system for processing multi-view images, storage medium

By analyzing the occlusion of multi-view images and selecting or synthesizing unoccluded images, the problem of incomplete information captured by the camera from different perspectives is solved, achieving complete and accurate display of the target object image and improving the effect of image recognition and analysis.

CN114648483BActive Publication Date: 2025-11-07ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011520775.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-21
Publication Date
2025-11-07
Estimated Expiration
2040-12-21

AI Technical Summary

Technical Problem

When using multiple cameras mounted at different locations within a given area to capture images from multiple perspectives, the information captured by the cameras at different locations varies from perspective to perspective. In some perspectives, there may be obstruction, resulting in incomplete or inaccurate images of the target object displayed in the camera images.

Method used

By acquiring multiple images taken by multiple cameras from different perspectives, the region information and display coordinates of the target object are determined, the occlusion situation under different perspectives is analyzed, and unoccluded images are selected or synthesized for processing. Image fusion is then performed using a perspective transformation matrix.

Benefits of technology

It effectively solves the occlusion problem, ensures the integrity and accuracy of the target object image, and improves the accuracy of image recognition and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648483B_ABST
    Figure CN114648483B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-view image processing method, device and system, storage medium.Therein, the method includes: obtaining multiple cameras from different view angles and multiple images photographed, at least two camera viewfinder positions are different;From the multiple images under different view angles, obtain the target image of the target object existing;Determine the region of at least two target images where the target object exists, and obtain the region information of the region;Based on the region information of at least two target images, determine whether the target image under different view angles exists shielding.The application solves the prior art by mounting camera in multiple different positions in the region range, to realize the multi-view viewfinder of the region in all directions, since the information captured by the camera in different positions is different under different view angles, there is shielding phenomenon under some view angles, resulting in the target object image displayed in the camera image is not complete, not accurate technical problem.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, in particular to a multi-view image processing method, device and system, and a storage medium. BACKGROUND

[0002] In order to analyze some people or objects in a scene, multiple image acquisition devices are usually deployed in the scene to obtain images from multiple perspectives. However, occlusion may exist in images from any perspective, and the occlusion of images has a great impact on most algorithms in computer vision.

[0003] In order to reduce the impact of image occlusion on computer vision algorithms, the following two solutions are commonly used: (1) predicting multiple camera images separately and fusing the prediction results of multiple camera images; (2) performing joint prediction after performing operations such as image feature splicing on images from multiple perspectives. These two solutions are expected to reduce the impact of occlusion by combining the prediction information of non-occlusion angle-related algorithms and the prediction information of occlusion angle-related algorithms. However, in the case of severe occlusion, the model is difficult to correctly predict, and at this time, the above two solutions will spread this misjudgment, resulting in incorrect final results.

[0004] In the prior art, multiple cameras are mounted at different positions in a region to achieve all-around multi-view shooting of the region. Since the cameras at different positions capture different information at different perspectives, occlusion exists at some perspectives, resulting in incomplete and inaccurate target object images displayed in the camera images. However, no effective solution has been proposed to solve this problem. SUMMARY

[0005] Embodiments of the present application provide a multi-view image processing method, device and system, and a storage medium to at least solve the technical problem that in the prior art, multiple cameras are mounted at different positions in a region to achieve all-around multi-view shooting of the region. Since the cameras at different positions capture different information at different perspectives, occlusion exists at some perspectives, resulting in incomplete and inaccurate target object images displayed in the camera images.

[0006] According to an aspect of the embodiments of the present application, a method for processing multi-view images is provided, comprising: obtaining a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras are located at different positions; obtaining target images in which a target object exists from the plurality of images from different perspectives, wherein the target object is an object to be observed in the target images; determining a region in which the target object exists in at least two target images and obtaining region information of the region, wherein the region information comprises: a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; and determining whether the target images from different perspectives exist occlusion based on the region information of the at least two target images.

[0007] According to another aspect of the embodiments of the present application, a method for processing multi-view images is also provided, comprising: starting a plurality of cameras in a region to be observed, wherein at least two cameras are located at different positions; obtaining a plurality of images captured by the plurality of cameras from different perspectives, wherein the plurality of images form a camera image set; displaying target images in which a target object exists in the camera image set in a display interface, wherein the target object is an object to be observed in the target images; displaying a region in which the target object exists in at least two target images and obtaining region information of the region, wherein the region information comprises: a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; and displaying at least one target image in which no occlusion exists in the display interface, wherein whether the target images from different perspectives exist occlusion is determined by analyzing the region information of the at least two target images.

[0008] According to another aspect of the embodiments of the present application, a processing device for multi-view images is also provided, comprising: a first obtaining module, configured to obtain a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras are located at different positions; a second obtaining module, configured to obtain target images in which a target object exists from the plurality of images from different perspectives, wherein the target object is an object to be observed in the target images; a first determining module, configured to determine a region in which the target object exists in at least two target images and obtain region information of the region, wherein the region information comprises: a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; and a second determining module, configured to determine whether the target images from different perspectives exist occlusion based on the region information of the at least two target images.

[0009] According to another aspect of the embodiments of the present application, a multi-view image processing apparatus is also provided, which comprises: a starting module configured to start a plurality of cameras to capture a region to be observed, wherein at least two cameras have different shooting positions; an obtaining module configured to obtain a plurality of images captured by the plurality of cameras from different perspectives, wherein the plurality of images form a camera image set; a first displaying module configured to display a target image in which a target object exists in the camera image set in a display interface, wherein the target object is an object to be observed in the target image; a second displaying module configured to display a region in which the target object exists in at least two target images and obtain region information of the region, wherein the region information comprises: a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; and a third displaying module configured to display at least one target image in which no occlusion exists in the display interface, wherein whether the target image from different perspectives has occlusion is determined by analyzing the region information of the at least two target images.

[0010] According to another aspect of the embodiments of the present application, a storage medium is also provided, which comprises a stored program, wherein the program controls a device in which the storage medium is located to perform the following steps when the program is running: obtaining a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; obtaining a target image in which a target object exists from the plurality of images from different perspectives, wherein the target object is an object to be observed in the target image; determining a region in which the target object exists in at least two target images and obtaining region information of the region, wherein the region information comprises: a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; and determining whether the target image from different perspectives has occlusion based on the region information of the at least two target images.

[0011] According to another aspect of the embodiments of the present application, a virtual machine memory allocation system is also provided, which comprises: a processor; and a memory connected with the processor and configured to provide the processor with instructions to process the following steps: obtaining a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; obtaining a target image in which a target object exists from the plurality of images from different perspectives, wherein the target object is an object to be observed in the target image; determining a region in which the target object exists in at least two target images and obtaining region information of the region, wherein the region information comprises: a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; and determining whether the target image from different perspectives has occlusion based on the region information of the at least two target images.

[0012] In the embodiment of the present application, a plurality of images captured by a plurality of cameras from different perspectives are acquired, wherein at least two cameras have different shooting positions; target images in which target objects exist are acquired from the plurality of images from different perspectives, wherein the target objects are objects to be observed in the target images; regions in which the target objects exist in the at least two target images are determined, and region information of the regions is acquired, wherein the region information includes: a block image of the target object in the target image, and display coordinates of the block image in the corresponding target image; whether the target images from different perspectives exist occlusion is determined based on the region information of the at least two target images. The above scheme determines whether the target images exist occlusion by analyzing the plurality of images from different perspectives, so that how to process the target images from different perspectives can be determined according to the determination result, and then the occlusion problem can be effectively solved by using intelligent scheduling and fusing information of the plurality of perspectives (cameras), and the technical problem that the target object images displayed in the camera images are incomplete and inaccurate due to the occlusion phenomenon in some perspectives caused by mounting cameras at a plurality of different positions in a region to realize omnidirectional multi-perspective shooting of the region is solved. BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:

[0014] Figure 1 A hardware structure block diagram of a computing device (or a mobile device) for implementing a multi-perspective image processing method is shown;

[0015] Figure 2 is a flowchart of a multi-perspective image processing method according to Embodiment 1 of the present application;

[0016] Figure 3 is a schematic diagram of an image of a smart classroom acquired by a camera;

[0017] Figure 4 is a schematic diagram of a block image in Figure 3 according to an embodiment of the present application;

[0018] Figure 5 is a schematic diagram of a block image in Figure 3 according to another embodiment of the present application;

[0019] Figure 6 is a schematic diagram of an optional two-perspective image processing method according to an embodiment of the present application;

[0020] Figure 7 is a flowchart of a multi-view image processing method according to an embodiment of the present application;

[0021] Figure 8 is a schematic diagram of a multi-view image processing apparatus according to an embodiment of the present application;

[0022] Figure 9 is a schematic diagram of a multi-view image processing apparatus according to an embodiment of the present application;

[0023] Figure 10 is a structural block diagram of a computing device according to an embodiment of the present application;

[0024] Figure 11 is a flowchart of a multi-view image processing method according to an embodiment of the present application;

[0025] Figure 12 is a schematic diagram of a multi-view image processing apparatus according to an embodiment of the present application;

[0026] Figure 13 is a flowchart of a multi-view image processing method according to an embodiment of the present application;

[0027] Figure 14 is a schematic diagram of a multi-view image processing apparatus according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] First, some of the nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:

[0031] Smart classroom: The smart classroom is a materialization of a typical smart learning environment, which is a high-end form of multimedia and network classrooms. It is a new type of classroom built with Internet of Things technology, cloud computing technology and intelligent technology, which includes tangible physical space and intangible digital space. It is a new type of classroom that realizes situational awareness and environmental management functions through various types of intelligent equipment to assist teaching content presentation, facilitate learning resource acquisition, and promote classroom interaction.

[0032] Embodiment 1

[0033] According to the embodiments of the present application, an embodiment of a multi-view image processing method is also provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0034] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computing device or a similar computing device. Figure 1 A hardware structure block diagram of a computing device (or mobile device) for implementing a multi-view image processing method is shown. As shown in Figure 1 The computing device 10 (or mobile device 10) can include one or more processors 102 (the processor 102 can include but is not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, and does not limit the structure of the above-mentioned electronic device. For example, the computing device 10 can include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .

[0035] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any one of the other elements of the computing device 10 (or mobile device). As referred to in the embodiments of the present application, the data processing circuitry serves as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.

[0036] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the multi-view image processing method of the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implements the multi-view image processing method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the computing device 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0037] The transmission module 106 is configured to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computing device 10. In one example, the transmission module 106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a radio frequency (RF) module, which is configured to communicate with the Internet in a wireless manner.

[0038] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the computing device 10 (or mobile device).

[0039] It should be noted that in some optional embodiments, the above-mentioned Figure 1 The computer device (or mobile device) shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that Figure 1is merely one instance of a particular, specific example, and is intended to show the types of components that can be present in the above-described computer device (or mobile device).

[0040] In the above operating environment, the present application provides a multi-view image processing method as shown in Figure 2 Figure 2 is a flowchart of a multi-view image processing method according to an embodiment of the present application.

[0041] Step S21: Obtain multiple images captured by multiple cameras from different perspectives, wherein at least two cameras have different shooting positions.

[0042] Specifically, the above cameras can be pan-tilt-zoom cameras that capture image information according to control according to a preset perspective. The multiple images captured from different perspectives obtained in the above step can be images of an observation area captured by multiple cameras within a same time period, and more specifically, can be images of the observation area captured by multiple cameras at a same time.

[0043] Since there are blind spots in the shooting of at least two cameras, it is difficult to obtain images of the entire observation area, and multiple objects in the observation area are subject to occlusion, in the above solution, at least two cameras have different shooting positions, so that images of the observation area can be more comprehensively obtained.

[0044] In an optional embodiment, taking the scenario of a smart classroom as an example, since the classroom contains multiple students, it is difficult to capture images of all the students with only one camera, therefore, in order to obtain images of the students, the above solution installs cameras at multiple positions in the classroom to obtain images of the students from multiple different angles.

[0045] In another optional embodiment, taking the traffic scenario as an example, the number of vehicles contained in the road is also very large, and it is difficult to obtain images of all the vehicles on the road with only one camera, therefore, multiple cameras are arranged to comprehensively obtain images of the vehicles appearing on the road.

[0046] Step S23: Obtain a target image containing a target object from the multiple images from different perspectives, wherein the target object is an object to be observed in the target image.

[0047] Specifically, the above step can be implemented through an image recognition solution. For example, an image recognition model can be obtained through training, and the image recognition model is used to determine a target image containing a target object from the multiple images, so as to extract the target image.

[0048] ​In an optional embodiment, taking the scenario of a smart classroom as an example, the target object can be a student. After a plurality of images are acquired, images containing the student are extracted. In another optional embodiment, taking the scenario of a traffic scene as an example, the target object can be a vehicle, and after a plurality of images are acquired, images containing the vehicle are extracted.

[0049] In step S25, a region in which the target object exists in the at least two target images is determined, and region information of the region is acquired, wherein the region information includes a block image of the target object in the target image and display coordinates of the block image in the corresponding target image.

[0050] In an optional embodiment, the region of the target object in the target image can be acquired by detecting a tracking model, and region information of the region can be acquired.

[0051] Figure 3 FIG. 1 is a schematic diagram of an image of a smart classroom acquired by a camera, in combination with Figure 3 As shown in FIG. 1, in the image, the target object is a teacher, and the target region is a face region of the teacher. Therefore, the face region of the teacher can be detected and tracked based on a detection tracking model, to obtain display coordinates of the region of the target object in the target image. The detection tracking model can frame the detected region of the target object in the target image with a rectangle according to the display coordinates of the region of the target object in the target image, to indicate the position of the region of the target object in the target image. The region framed by the rectangle is the block image of the target object in the target image.

[0052] In step S27, whether occlusion exists in the target images under different perspectives is determined based on the region information of the at least two target images.

[0053] Specifically, the determination of whether occlusion exists in the target images under different perspectives can be a determination of whether the region in which the target object is located in the target images under different perspectives has occlusion.

[0054] In an optional embodiment, the detection tracking model can be used to determine whether occlusion exists in the block image when the determination of occlusion in the block image is performed. The detection tracking model can be used to track the target object in real time according to the images acquired by the camera. When the images acquired by the camera include a plurality of target objects, the detection tracking model detects and tracks the plurality of target objects in the images respectively. During the detection and tracking, if a target object cannot be detected, it can be determined that the block image of the target object has occlusion.

[0055] The above embodiments of the present application obtain multiple images captured by multiple cameras from different perspectives, wherein at least two cameras have different viewfinder positions; from the multiple images from different perspectives, a target image in which a target object exists is obtained, wherein the target object is an object to be observed in the target image; a region in which the target object exists in at least two target images is determined, and region information of the region is obtained, wherein the region information includes a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; whether the target images from different perspectives exist occlusion is determined based on the region information of the at least two target images. The above scheme analyzes multiple images from multiple perspectives to determine whether the target images exist occlusion, so that how to process the target images from multiple perspectives can be determined according to the determination result, and then the occlusion problem can be effectively solved by using intelligent scheduling and fusing information of multiple perspectives (cameras), which solves the technical problem that the target object image displayed in the camera image is incomplete and inaccurate due to occlusion in some perspectives in the prior art, in which multiple cameras are mounted at multiple different positions in a region to realize omnidirectional multi-perspective view of the region.

[0056] As an optional embodiment, after obtaining the multiple images captured by the multiple cameras from different perspectives, the above method further includes: obtaining coordinate system information of at least two images captured from different perspectives, wherein the coordinate system information is determined based on a shooting parameter of the camera and a positional relationship between the two cameras; performing perspective matching on the at least two images from different perspectives based on the coordinate system information of the at least two images to obtain a positional relationship of the target object in the two images; and saving the positional relationship in a matrix to obtain a perspective conversion matrix, wherein the perspective conversion matrix records a coordinate conversion relationship between the images from different perspectives.

[0057] Specifically, because the angles of the at least two cameras are different, the at least two cameras each have its own coordinate system, and the coordinate system is different from a world coordinate system, so the coordinate system information of the at least two images is the coordinate system information of the image collected in the coordinate system of the camera itself. The shooting parameter of the camera is an internal parameter of the camera, which can include a focal length, a resolution, a sensor size, etc. The positional relationship between the two cameras is an external parameter of the camera, and the perspective conversion matrix is an affine transformation matrix between the two images. A same target object in images from different angles can be converted based on the perspective conversion matrix.

[0058] It should be noted that the above steps can be performed after the position and shooting angle of the camera are determined, and the same set of perspective conversion matrix can be used when the position and shooting angle of the camera do not change. When the position or shooting angle of the camera changes, the above steps need to be re-executed to determine the perspective conversion matrix between the two again.

[0059] The perspective conversion matrix can be one-to-one. Specifically, in the case where two cameras are included in the scene, the perspective conversion matrix is the perspective conversion matrix of the two cameras, that is, the images obtained by the two cameras can be converted by the perspective conversion matrix. In the case where N (N > 2) cameras are included in the scene, the perspective conversion matrix between each two cameras needs to be obtained, that is, N-1 perspective conversion matrices need to be obtained.

[0060] In an optional embodiment, taking obtaining the transformation matrix between image A and image B as an example, a plurality of feature points can be extracted from image A and image B first, for example, 9 feature points are extracted from image A and image B respectively, 9 groups of feature points are obtained, and then the affine transformation matrix is solved based on the 9 groups of feature points to obtain the final perspective conversion matrix. In another optional embodiment, in the case where the target object is a person, face recognition or pedestrian re-identification technology can also be used for matching. The two methods are content matching, which can be realized by feature description or model.

[0061] In yet another optional embodiment, matching of coordinate position relationship can also be realized. Specifically, the relative position relationship between image coordinates is obtained based on matrix conversion through camera internal parameters (focal length, resolution, sensor size) and external parameters (position relationship between multiple cameras), and the content in the image can be matched through the position relationship. In this scheme, the input information is the internal parameters of the camera, the coordinates of the feature points (non-coplanar) in the target coordinate system (3D) and the plane coordinate system (2D) of the camera. Through image and camera coordinate conversion methods such as solvePnP, the output information, that is, the position of the target coordinate system relative to the camera coordinate system, is obtained, and then the relative position between the camera coordinate systems is obtained through the external parameters of the camera, that is, the relative position between the cameras.

[0062] As an optional embodiment, based on the perspective conversion matrix, the association relationship of the target views with the same target object under different perspectives is determined, so that the target images under different perspectives can be converted to each other based on the perspective conversion matrix therebetween.

[0063] As an optional embodiment, based on the region information of the at least two target images, it is determined whether the target images at different perspectives exist occlusion, comprising: based on the region information of the target images at different perspectives, block images of the target object at different perspectives are determined; based on the feature information of the target object, types of the block images of the target object at different perspectives are determined, wherein the types include: occlusion and no occlusion; based on display coordinates of the block images in the corresponding target images, it is determined whether the target images at different perspectives exist occlusion.

[0064] Specifically, since the region information includes the block images of the target object in the target images and the display coordinates of the block images in the target images, the block images corresponding to the target object can be extracted from the target images based on the region information of the target images, and the block images contain the target object.

[0065] In an optional embodiment, taking the scenario of a smart classroom as an example, the target object can be at least two persons in the classroom, and therefore the block images of the at least two persons in the target images can be extracted. Further, the block images can be face images of the at least two persons. Whether the block images are overall images of the target objects or only face images of the target objects can be determined according to actual conditions. If the target images are used for face analysis, the block images can be only face images of the target objects, and if the target images are used for behavior analysis, the block images can be overall images of the target objects.

[0066] In another optional embodiment, taking a traffic scenario as an example, the target object can be at least two vehicles in an observation area, and therefore the block images of the at least two vehicles in the target images can be extracted. Further, the block images can be license plate images of the at least two vehicles. Similarly, whether the block images are overall images of the target objects or only face images of the target objects can be determined according to actual conditions. If the target images are used for vehicle recording, the block images can be only license plate images of the target objects, and if the target images are used for driving analysis of the vehicles, the block images can be overall images of the target objects.

[0067] In combination Figure 3 As shown in the image, the target object is a teacher, and the image of the classroom obtained by the camera at the perspective contains the image of the teacher. The face region of the target object can be recognized, and the face region of the target object is extracted as the block image at the perspective, as shown in Figure 4 As shown in the image, the face region of the target object can also be extracted as the block image at the perspective, as shown in Figure 5

[0068] ​After obtaining the block image of the target object, it is necessary to distinguish the blocked block image and the unblocked block image. The above-mentioned blocking and unblocking are two types of block images, and do not represent the actual blocking state. For example, in an optional embodiment, the type of block image can be determined according to the proportion of the blocked area in the total area of the block image. If the blocked area is less than the preset value of the total area of the block image, it is determined to be an unblocked state. Only if the blocked area is greater than the preset value of the total area of the block image, it is determined to be a blocked state. For another example, in another optional embodiment, the type of block image can be determined according to the blocked part in the block image. Taking a face image as an example, if the five features in the face image are not blocked, it is determined to be an unblocked block image. If any of the five features is blocked, it is determined to be a blocked block image.

[0069] Based on the feature information of the target object, the type of the block image of the target object under different viewing angles is determined. The determination can be based on a detection tracking model. When the detection tracking model is used to track at least two target objects in a to-be-observed region, when the block image of the target object is blocked, part of the feature information of the target object may not be detected due to the blocking. Therefore, when the detection tracking model detects that the feature information of the target object is incomplete, it is determined that the block image of the target object is blocked.

[0070] In the case where the block images of the at least two target objects in the target image are not blocked, it is determined that the target image is not blocked. In the case where the block image of the target object is blocked, based on the display coordinates of the block image in the corresponding target image, it is determined whether the target image under different viewing angles is blocked. The determination can be based on the range of the block image in the target image according to the display coordinates of the block image in the target image. If the range of the blocked block image in the target image is greater than a preset range, it is determined that the target image is blocked. If the range of the blocked block image in the target image is less than the preset range, it is determined that the target image is not blocked.

[0071] As an optional embodiment, after determining whether the target image under different viewing angles is blocked, the above-mentioned method further comprises: if it is determined that the target image under at least one viewing angle is blocked, deleting the target image with blocking from the plurality of images.

[0072] If the target image with blocking is retained for subsequent algorithms, not only the data amount of the subsequent algorithms will be increased, but also the accuracy of the subsequent algorithms will be affected. In the above-mentioned scheme, the target image with blocking is discarded, and only the picture without blocking angle is selected for prediction of the subsequent algorithm, thereby fundamentally avoiding the influence of the target picture with blocking on the subsequent algorithm and reducing the data processing amount.

[0073] As an optional embodiment, after the target images with occlusion are removed from the plurality of images, the method further includes: selecting the target images without occlusion as an image sample set, and performing one of the following functions on the images in the image sample set: recognizing an expression, recognizing a face, recognizing a behavior, and training a neural network model.

[0074] In an optional embodiment, taking a smart classroom as an example, the target object can be a student in the classroom, face recognition can be performed based on the image of the student, and expression recognition can be performed based on the face recognition result, so as to obtain the reaction of the student to the teaching, and then the analysis of the class state of at least two students can be applied; the image of the student can also be used for behavior recognition to determine whether the student is listening carefully; and the image can also be used as sample data for training a neural network model.

[0075] It should be noted that after obtaining the new synthesized image, the new image can be given to a subsequent algorithm, and the subsequent algorithm can perform any analysis or operation on the image, which is not limited by the present application.

[0076] The above scheme selects the target images without occlusion as an image sample set for image recognition, which can avoid the influence of the images with occlusion on image recognition, thereby improving the accuracy of image recognition.

[0077] It should also be noted that in the case where none of the plurality of images has occlusion, the plurality of images under different perspectives can all be given to a subsequent algorithm for processing as an image sample set.

[0078] As an optional embodiment, after determining whether the target images under different perspectives have occlusion, the method further includes: if it is determined that the target images under all perspectives have occlusion, extracting the regions without occlusion from the target images under different perspectives, and synthesizing the regions without occlusion under different perspectives to synthesize a new image without occlusion.

[0079] Specifically, in the above scheme, in the case where the target images under all perspectives have occlusion, the target images are converted and synthesized in perspective by a perspective conversion matrix, thereby synthesizing a new image without occlusion.

[0080] When converting the perspective of the target image, the region where the target object is located in the target image with a larger occlusion region can be converted to the target image with a smaller occlusion region. When there are more than two target images, a target image with the smallest occlusion range can be selected, and the regions where the target objects are located in the other images are all converted to the perspective of the target image with the smallest occlusion range. Finally, all the conversion results and the target image with the smallest occlusion range are synthesized, and a new synthesized image without occlusion is obtained.

[0081] In an optional embodiment, taking the smart classroom including the left camera A and the right camera B as an example, for the image of a student, the image of the student appears in the images captured by the left and right cameras, and in this case, the region in which the image of the student is not blocked in the camera A can be obtained, and after conversion by the perspective conversion matrix, the image is combined with the image captured by the camera B, so as to synthesize the unblocked image of the student. Similarly, the region in which the image of the student is not blocked in the camera B can also be obtained, and after conversion by the perspective conversion matrix, the image is combined with the image captured by the camera A, so as to synthesize the unblocked image of the student.

[0082] As an optional embodiment, after synthesizing the new unblocked image, the method further includes: regarding the synthesized new unblocked image as an image sample set, and performing one of the following functions on the images in the image sample set: recognizing an expression, recognizing a face, recognizing a behavior, and training a neural network model.

[0083] The above scheme can perform image recognition on the new unblocked image synthesized by multiple images with occlusion under all perspectives, so as to perform image recognition on a target object for which there is no unblocked image, thereby improving the image recognition capability, and further solving the technical problem in the prior art that a camera is mounted at multiple different positions in a region to realize all-around multi-perspective shooting of the region, and the information captured by the cameras at different positions is different under different perspectives, and there is an occlusion phenomenon under some perspectives, resulting in that the image of the target object displayed in the camera image is incomplete and inaccurate.

[0084] In an optional embodiment, taking the smart classroom as an example, the target object can be a student in the classroom, face recognition can be performed based on the image of the student, expression recognition can be performed based on the face recognition result, so as to obtain the reaction of the student to the teaching, and then the analysis of the class state of at least two students can be applied; behavior recognition can also be performed on the image of the student to determine whether the student is listening carefully; and the image can also be used as sample data for training a neural network model.

[0085] It should be noted that after obtaining the synthesized new image, the new image can be given to a subsequent algorithm, and the subsequent algorithm performs what kind of analysis or operation on the image is not limited in the present application.

[0086] Figure 6 is a schematic diagram of an optional two-perspective image processing method according to an embodiment of the present application, which is combined with Figure 6As shown, the two views are a left view and a right view. Before data processing, first, left view and right view matching is performed according to the images obtained by the left view and the right view to obtain a left-right view conversion matrix. The left view image and the right view image are input into a detection and tracking model, region occlusion is judged by the detection and tracking model to determine whether the target object is occluded, and then view selection or synthesis is performed according to the judgment result and the left-right view conversion matrix. The following is described separately.

[0087] In the case that the block images of the target object in the left view image and the right view image are not occluded and the detection confidence is high, both can be used for subsequent algorithm prediction; in the case that the block image of the target object in any one view is occluded but the block image of the target object in the other view is not occluded, the image in the view in which occlusion occurs is discarded, and only the image in the view in which occlusion does not occur is selected for subsequent algorithm prediction; in the case that the block images of the target object in the two views are both occluded, the part of the block image of the target object in one view that is not occluded is converted to the other view by the left-right view conversion matrix, and is synthesized with the part of the block image of the target object in the image in the other view that is not occluded to obtain a new synthesized image, and the new synthesized image is used for subsequent algorithm prediction. The subsequent algorithm prediction can include face recognition, expression recognition, and behavior recognition.

[0088] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0089] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a necessary general hardware platform, and of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk), and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the method of each embodiment of the present application.

[0090] Embodiment 2

[0091] According to the embodiment of the present application, an embodiment of a multi-view image processing method is also provided, Figure 7 is a flowchart of a multi-view image processing method according to the embodiment 2 of the present application, which is combined with Figure 7 as shown, the method comprises:

[0092] In step S71, multiple cameras in the to-be-observed area are started, wherein at least two cameras have different shooting positions.

[0093] Specifically, the camera can be a pan-tilt-zoom camera, and the image information is captured according to the control according to the preset view. Since at least two cameras have shooting blind areas, it is difficult to obtain the image of the entire to-be-observed area, and multiple objects in the to-be-observed area are blocked, at least two cameras have different shooting positions in the above-mentioned solution, so that the image in the to-be-observed area is more comprehensive.

[0094] In an optional embodiment, taking the scene of a smart classroom as an example, since the classroom contains multiple students, it is difficult to capture the images of all the students by installing only one camera, therefore, in order to obtain the images of the students, multiple cameras are installed in multiple positions in the classroom to obtain the images of the students from multiple different angles.

[0095] In another optional embodiment, taking the traffic scene as an example, the number of vehicles contained in the road is also very large, and it is difficult to obtain the images of all the vehicles on the road by using only one camera, therefore, multiple cameras are arranged to comprehensively obtain the images of the vehicles appearing on the road.

[0096] In step S73, multiple images captured by the multiple cameras from different angles are obtained, wherein the multiple images constitute a camera image set.

[0097] The multiple images captured from different angles in the above-mentioned step can be the images of the to-be-observed area captured by the multiple cameras in the same time period, more specifically, can be the images of the to-be-observed area captured by the multiple cameras at the same time.

[0098] In step S75, a target image in which a target object exists in the camera image set is displayed in a display interface, wherein the target object is an object to be observed in the target image.

[0099] The display interface can be a display interface displayed by a smart terminal. The target images under multiple angles can be displayed on one display interface in a split-screen manner, or only the target image under one angle can be displayed on the display interface, and the target images under other angles can be displayed by page turning operation.

[0100] In step S77, a region in which the target object exists in at least two target images is displayed, and region information of the region is obtained, wherein the region information includes a block image of the target object in the target image and display coordinates of the block image in the corresponding target image.

[0101] Specifically, the region of the target object in the target image can be obtained by detecting the tracking model, and the region information of the region can be obtained.

[0102] In step S79, at least one target image without occlusion is displayed in the display interface, wherein whether the target image at different angles exists occlusion is determined by analyzing the region information of the at least two target images.

[0103] In the above scheme, the target image with occlusion is deleted, and the target image without occlusion is retained, thereby avoiding the influence of the image with occlusion on the display effect, and solving the technical problem that in the prior art, the cameras are mounted at multiple different positions in the region to realize omnibearing multi-view imaging of the region, and the information captured by the cameras at different positions is different at different angles, and there is occlusion at some angles, resulting in that the target object image displayed in the camera image is incomplete and inaccurate.

[0104] Embodiment 3

[0105] According to the embodiments of the present application, a multi-view image processing device for implementing the multi-view image processing method in Embodiment 1 is further provided, Figure 8 is a schematic diagram of a multi-view image processing device according to an embodiment of the present application, as Figure 8 shown, the device 800 includes:

[0106] The first acquisition module 802 is configured to acquire a plurality of images captured by a plurality of cameras at different angles, wherein at least two cameras are located at different positions.

[0107] The second acquisition module 804 is configured to acquire target images in which a target object exists from the plurality of images at different angles, wherein the target object is an object to be observed in the target image.

[0108] The first determination module 806 is configured to determine a region in which the target object exists in at least two target images, and obtain region information of the region, wherein the region information includes a block image of the target object in the target image and display coordinates of the block image in the corresponding target image.

[0109] The second determination module 808 is configured to determine whether the target image at different angles exists occlusion based on the region information of the at least two target images.

[0110] It should be noted that the first obtaining module 802, the second obtaining module 804, the first determining module 806 and the second determining module 808 correspond to steps S21 to S27 in Embodiment 1, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules can run in the computing device 10 provided in Embodiment 1 as part of the device.

[0111] As an optional embodiment, the device further includes a third obtaining module configured to obtain coordinate system information of at least two images captured at different perspectives, wherein the coordinate system information is determined based on a shooting parameter of a camera and a positional relationship between two cameras; a matching module configured to perform perspective matching on the at least two images at different perspectives based on the coordinate system information of the at least two images, to obtain a positional relationship of the target object in the two images; and a saving module configured to save the positional relationship in a matrix to obtain a perspective conversion matrix, wherein the perspective conversion matrix records a coordinate conversion relationship between images at different perspectives.

[0112] As an optional embodiment, the device further includes a determining module configured to determine, based on the perspective conversion matrix, an association relationship of a target view having the same target object at different perspectives.

[0113] As an optional embodiment, the second determining module includes a first determining submodule configured to determine, based on region information of target images at different perspectives, a block image of the target object at different perspectives; a second determining submodule configured to determine, based on feature information of the target object, a type of the block image of the target object at different perspectives, wherein the type includes occlusion and no occlusion; and a third determining submodule configured to determine, based on display coordinates of the block image in a corresponding target image, whether the target image at different perspectives has occlusion.

[0114] As an optional embodiment, the device further includes a deleting module configured to, after determining whether the target image at different perspectives has occlusion, delete, if it is determined that the target image at at least one perspective has occlusion, the target image having occlusion from the plurality of images.

[0115] As an optional embodiment, the device further includes a selecting module configured to, after deleting, from the plurality of images, the target image having occlusion, select, as an image sample set, a target image not having occlusion, and perform one of the following functions using images in the image sample set: recognizing an expression, recognizing a face, recognizing a behavior, and training a neural network model.

[0116] As an optional embodiment, the above-mentioned apparatus further includes: a synthesis module, configured to, after determining whether there is occlusion in the target images under different viewpoints, if it is determined that there is occlusion in the target images under all viewpoints, extract unoccluded regions from the target images under different viewpoints, and synthesize the unoccluded regions under different viewpoints to synthesize a new unoccluded image.

[0117] As an optional embodiment, the above apparatus further includes: an execution module, configured to, after synthesizing a new unoccluded image, use the synthesized new unoccluded image as an image sample set, and use the images in the image sample set to perform one of the following functions: facial expression recognition, face recognition, behavior recognition, and training a neural network model.

[0118] Example 4

[0119] According to embodiments of the present invention, a multi-view image processing apparatus is also provided for implementing the multi-view image processing method in Embodiment 1 described above. Figure 9 This is a schematic diagram of a multi-view image processing apparatus according to an embodiment of this application, as shown below. Figure 9 As shown, the device 900 includes:

[0120] The startup module 902 is used to start multiple cameras in the area to be observed, wherein at least two cameras have different framing positions.

[0121] The acquisition module 904 is used to acquire multiple images captured by multiple cameras from different perspectives, wherein the multiple images constitute a camera image set.

[0122] The first display module 906 is used to display a target image containing a target object in the camera image set on the display interface, wherein the target object is the object to be observed in the target image.

[0123] The second display module 908 is used to display the region where the target object exists in at least two target images and obtain the region information of the region. The region information includes: the block image of the target object in the target image and the display coordinates of the block image in the corresponding target image.

[0124] The third display module 9010 is used to display at least one target image without occlusion in the display interface, wherein the presence or absence of occlusion in the target image under different viewing angles is determined by analyzing the region information of at least two target images.

[0125] It should be noted that the starting module 902, the obtaining module 904, the first display module 906, the second display module 908 and the third display module 910 correspond to steps S71 to S79 in Embodiment 2, and the five modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules can run in the computing device 10 provided in Embodiment 1 as part of the device.

[0126] Embodiment 5

[0127] Embodiments of the present application can provide a computing device, which can be any one of the computing devices in the computing device group. Alternatively, in the present embodiment, the above computing device can also be replaced by a terminal device such as a mobile terminal.

[0128] Alternatively, in the present embodiment, the above computing device can be located in at least one of the network devices in the computer network.

[0129] In the present embodiment, the above computing device can execute program codes of the following steps in the multi-view image processing method: obtaining a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; obtaining target images with target objects from the plurality of images from different perspectives, wherein the target objects are objects to be observed in the target images; determining regions with target objects in the at least two target images and obtaining region information of the regions, wherein the region information includes: block images of the target objects in the target images, and display coordinates of the block images in the corresponding target images; and determining whether there is occlusion in the target images from different perspectives based on the region information of the at least two target images.

[0130] Alternatively, Figure 10 is a structural block diagram of a computing device according to Embodiment 6 of the present application. As shown in Figure 10 , the computing device A can include one or more (only one is shown in the figure) processors 1002, a memory 1004, and a peripheral interface 1006.

[0131] The memory can be configured to store software programs and modules, such as program instructions / modules corresponding to the multi-view image processing method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the multi-view image processing method described above. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal A through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0132] The processor can call information and applications stored in the memory through the transmission device to perform the following steps: acquiring a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; acquiring target images in which a target object exists from the plurality of images from different perspectives, wherein the target object is an object to be observed in the target image; determining a region in which the target object exists in at least two target images, and acquiring region information of the region, wherein the region information includes a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; and determining whether the target images from different perspectives exist occlusion based on the region information of the at least two target images.

[0133] Optionally, the processor can further execute program codes of the following steps: acquiring coordinate system information of at least two images captured from different perspectives, wherein the coordinate system information is determined based on shooting parameters of the cameras and a positional relationship between the two cameras; performing perspective matching on the at least two images from different perspectives based on the coordinate system information of the at least two images to acquire a positional relationship of the target object in the two images; the positional relationship is saved in a matrix to obtain a perspective conversion matrix, wherein the perspective conversion matrix records a coordinate conversion relationship between the images from different perspectives.

[0134] Optionally, based on the perspective conversion matrix, an association relationship of a target view having the same target object from different perspectives is determined.

[0135] Optionally, the processor can further execute program codes of the following steps: determining whether the target images in different perspectives exist occlusion based on the region information of the at least two target images, comprising: determining the block images of the target object in different perspectives based on the region information of the target images in different perspectives; determining the type of the block images of the target object in different perspectives based on the feature information of the target object, wherein the type comprises: occlusion and no occlusion; determining whether the target images in different perspectives exist occlusion based on the display coordinates of the block images in the corresponding target images.

[0136] Optionally, the processor can further execute program codes of the following steps: after determining whether the target images in different perspectives exist occlusion, if it is determined that the target images in at least one perspective exist occlusion, deleting the target images existing occlusion from the plurality of images.

[0137] Optionally, the processor can further execute program codes of the following steps: after deleting the target images existing occlusion from the plurality of images, selecting the target images without occlusion as an image sample set, and using the images in the image sample set to perform one of the following functions: recognizing an expression, recognizing a face, recognizing a behavior, and training a neural network model.

[0138] Optionally, the processor can further execute program codes of the following steps: after determining whether the target images in different perspectives exist occlusion, if it is determined that the target images in all perspectives exist occlusion, extracting the regions without occlusion from the target images in different perspectives, and synthesizing the regions without occlusion in different perspectives to synthesize a new non-occlusion image.

[0139] Optionally, the processor can further execute program codes of the following steps: after synthesizing the new non-occlusion image, using the synthesized new non-occlusion image as an image sample set, and using the images in the image sample set to perform one of the following functions: recognizing an expression, recognizing a face, recognizing a behavior, and training a neural network model.

[0140] The embodiment of the present application provides a multi-view image processing scheme. A plurality of images captured by a plurality of cameras from different view angles are acquired, wherein at least two cameras are located at different positions; target images in which target objects exist are acquired from the plurality of images from different view angles, wherein the target objects are objects to be observed in the target images; regions in which the target objects exist in the at least two target images are determined, and region information of the regions is acquired, wherein the region information comprises: a block image of the target object in the target image, and display coordinates of the block image in the corresponding target image; whether the target images from different view angles exist occlusion is determined based on the region information of the at least two target images. The above scheme determines whether the target images exist occlusion by analyzing the plurality of images from different view angles, so that how to process the target images from different view angles can be determined according to the determination result, and then the occlusion problem can be effectively solved by using intelligent scheduling and fusing information of the plurality of view angles (cameras), and the technical problem that the cameras are mounted at a plurality of different positions in a region range to realize all-around multi-view imaging of the region in the prior art is solved, because the information captured by the cameras at different positions in different view angles is different, and the target object images displayed in the camera images are incomplete and inaccurate in some view angles.

[0141] Those skilled in the art can understand that, Figure 10 The structure shown is only schematic, and the computing device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or the like. Figure 10 It does not limit the structure of the electronic device. For example, the computing device 10 can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 10 It does not limit the structure of the electronic device. For example, the computing device 10 can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 10 It does not limit the structure of the electronic device. For example, the computing device 10 can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.

[0142] Those skilled in the art can understand that all or part of the steps of the various methods in the above embodiments can be completed by programs instructing the related hardware of the terminal device, and the programs can be stored in a computer readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.

[0143] Embodiment 6

[0144] The embodiment of the present application further provides a storage medium. Optionally, in the embodiment, the storage medium can be used to save program codes executed by the multi-view image processing method provided in the embodiment 1.

[0145] Optionally, in the embodiment, the storage medium can be located in any one of the computing devices in the computing device group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0146] Optionally, in the embodiment, the storage medium is configured to store program codes for performing the following steps: obtaining a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; obtaining target images in which a target object exists from the plurality of images from different perspectives, wherein the target object is an object to be observed in the target images; determining regions in which the target object exists in the at least two target images, and obtaining region information of the regions, wherein the region information comprises: a block image of the target object in the target image, and display coordinates of the block image in the corresponding target image; and determining whether the target images from different perspectives exist occlusion based on the region information of the at least two target images. The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0147] Embodiment 7

[0148] According to the embodiment of the present application, an embodiment of a multi-view image processing method is further provided, Figure 11 is a flowchart of a multi-view image processing method according to the embodiment 7 of the present application, which, as shown in Figure 11 , comprises the following steps:

[0149] In step S111, a plurality of images captured by a plurality of cameras from different perspectives are obtained, wherein at least two cameras have different shooting positions.

[0150] The plurality of images captured from different perspectives obtained in the above steps can be images of an observation region captured by a plurality of cameras in a same time period, and more specifically, can be images of the observation region captured by a plurality of cameras at a same time.

[0151] In step S113, target images in which a target object exists are displayed from the plurality of images.

[0152] The target images are displayed on a smart terminal. The target images from different perspectives can be displayed on a display interface in a split-screen manner, or only the target image from one perspective can be displayed on the display interface, and the target images from other perspectives can be displayed by page turning operation.

[0153] In step S115, display the identification information, wherein the identification information is used to indicate the region of the target object in at least two of the target images.

[0154] The identification information can be displayed on the target image to indicate the region of the target object in the target image. For example, a rectangular frame can be used to frame the region of the target object. Figure 3 is a schematic diagram of an image of a smart classroom acquired by a camera, combined with Figure 3 In the image, the target object is a teacher, and the target region is the face region of the teacher. Therefore, the face region of the teacher can be detected and tracked based on the detection and tracking model to obtain the region of the target object in the target image. In the image, the target object is a teacher, and the image acquired by the camera at this angle is an image containing the teacher. The face region of the target object can be recognized, so that the face region of the target object is extracted as the region of the target object at this angle, as shown in Figure 4 The entire region of the target object in the image can also be extracted as the region of the target object at this angle, as shown in Figure 5

[0155] In step S117, when the region of the target object in the target image is all blocked, the region of the target object without blocking is extracted from the plurality of target images, and the region without blocking is synthesized to generate a new image without blocking. After generating the new image without blocking, the new image without blocking can also be displayed.

[0156] Specifically, in the above scheme, when the target image at all angles is blocked, the target image is converted and synthesized in terms of the angle of view by using the view conversion matrix, so as to synthesize a new image without blocking.

[0157] When the view of the target image is converted, the region of the target object in the target image with a larger blocking region can be converted to the region of the target object in the target image with a smaller blocking region. When there are more than two target images, a target image with the smallest blocking range can be selected, and the region of the target object in other images is converted to the view of the target image with the smallest blocking range. Finally, all the converted results and the target image with the smallest blocking range are synthesized, so as to obtain the synthesized new image without blocking.

[0158] ​In an alternative embodiment, taking the smart classroom including left camera A and right camera B as an example, for the image of a student, the image of the student appears in the images captured by the left and right cameras, and in this case, the non-occluded area of the image of the student in the camera A can be obtained, and after conversion by the perspective conversion matrix, the image is combined with the image captured by the camera B, so as to synthesize the non-occluded image of the student. Similarly, the non-occluded area of the image of the student in the camera B can also be obtained, and after conversion by the perspective conversion matrix, the image is combined with the image captured by the camera A, so as to synthesize the non-occluded image of the student.

[0159] Embodiment 8

[0160] According to the embodiments of the present application, a multi-view image processing device for implementing the multi-view image processing method in Embodiment 7 is also provided, Figure 12 is a schematic diagram of a multi-view image processing device according to Embodiment 8 of the present application, as Figure 12 shown, the device 1200 includes:

[0161] The acquisition module 1202 is configured to acquire a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions.

[0162] The first display module 1204 is configured to acquire a target image in which a target object exists from the plurality of images from different perspectives, wherein the target object is an object to be observed in the target image.

[0163] The second display module 1206 is configured to display identification information, wherein the identification information is used to indicate the area in which the target object exists in at least two target images.

[0164] The synthesis module 1208 is configured to extract the area in which the target object is not occluded from the plurality of target images in the case that the area in which the target object is not occluded exists in the target images, and to synthesize the area in which the target object is not occluded to generate a new non-occluded image.

[0165] It should be noted that the acquisition module 1202, the first display module 1204, the second display module 1206 and the synthesis module 1208 correspond to steps S111 to S117 in Embodiment 7, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules as part of the device can run in the computing device 10 provided in Embodiment 1.

[0166] Embodiment 9

[0167] According to the embodiment of the present application, an embodiment of a multi-view image processing method is also provided, Figure 13 is a flowchart of a multi-view image processing method according to the embodiment 9 of the present application, which is combined with Figure 13 The step includes:

[0168] In step S131, multiple images of the to-be-observed region under different view angles are displayed, and the multiple images include the target object.

[0169] The multiple images obtained in the above step can be images of the to-be-observed region captured by multiple cameras in the same time period, and more specifically, can be images of the to-be-observed region captured by multiple cameras at the same time. The target image is on the intelligent terminal. The target images under multiple view angles can be displayed on one display interface in a split-screen manner, or only the target image under one view angle can be displayed on the display interface, and the target images under other view angles can be displayed through page turning operation.

[0170] In step S133, identification information is displayed, wherein the identification information is used to indicate the region of the target object in the multiple images.

[0171] The identification information can be displayed on the target image, and is used to indicate the region of the target object in the target image. For example, a rectangular frame can be used to frame the region of the target object. Figure 3 is a schematic diagram of an image of a smart classroom captured by a camera, which is combined with Figure 3 In the image, the target object is a teacher, and the target region is the face region of the teacher, so the face region of the teacher can be detected and tracked based on a detection and tracking model to obtain the region of the target object in the target image. In the image, the target object is a teacher, and the image of the classroom captured by the camera at this angle is an image containing the teacher. The face region of the target object can be recognized, so that the face region of the target object is extracted as the region of the target object under this view angle, as shown in Figure 4 or the entire region of the target object in the image can be extracted as the region of the target object under this view angle, as shown in Figure 5

[0172] In step S135, selection prompt information is sent out in a case where the region of the target object in the multiple images all has occlusion, wherein the selection prompt information is used to prompt whether to synthesize the multiple images with occlusion.

[0173] ​In the above step, the prompt information can be sent by popping up a prompt box. In an optional embodiment, the prompt information is sent in the case where it is determined that all the images containing the target object have occlusion: all the images of XX have occlusion, do you want to synthesize the multiple images of XX. The prompt information also includes a control that can be selected by the user, and the user set selects according to the needs.

[0174] In step S137, the selection information is received, and in the case where the selection information is yes, the region without occlusion of the target object is extracted from the multiple images, and the region without occlusion is synthesized to generate a new image without occlusion.

[0175] In the case where the user selects to synthesize the multiple images of the target object, the application extracts the region without occlusion of the target object in the multiple images for synthesis. Since the multiple images are images of different angles, the complete image of the target object can be obtained by synthesis. In the case where the user selects not to synthesize the multiple images of the target object, the prompt information is sent, and the multiple images are continuously displayed.

[0176] Embodiment 10

[0177] According to the embodiments of the present application, a multi-view image processing device for implementing the multi-view image processing method in the above embodiment 9 is also provided, Figure 14 is a schematic diagram of a multi-view image processing device according to the embodiment 10 of the present application, as Figure 14 shown, the device 1400 includes:

[0178] The first display module 1402 is configured to display multiple images of a to-be-observed region under different angles, and the multiple images include a target object.

[0179] The second display module 1404 is configured to display identification information, where the identification information is used to indicate the region of the multiple images where the target object exists.

[0180] The prompt module 1406 is configured to send selection prompt information in the case where the region of the target object in the multiple images has occlusion, where the selection prompt information is used to prompt whether to synthesize the multiple images with occlusion.

[0181] The synthesis module 1408 is configured to receive selection information, and in the case where the selection information is yes, extract the region without occlusion of the target object from the multiple images, and synthesize the region without occlusion to generate a new image without occlusion.

[0182] It should be noted that the first display module 1402, the second display module 1404, the prompting module 1406 and the synthesizing module 1408 correspond to steps S131 to S137 in Embodiment 9, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules can run in the computing device 10 provided in Embodiment 1 as part of the device.

[0183] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0184] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.

[0185] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the present embodiment.

[0186] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0187] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0188] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. A method of processing a multi-view image, characterized by, The method comprises: obtaining a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; obtaining target images in which a target object exists from the plurality of images from different perspectives; determining a region in which the target object exists in at least two target images, and obtaining region information of the region, wherein the region information comprises a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; determining whether the target images from different perspectives exist occlusion based on the region information of the at least two target images; wherein the determination of whether the target images from different perspectives exist occlusion based on the region information of the at least two target images comprises: determining a range of the block image in the target image based on the display coordinates in the region information; if the range is greater than a preset range, it is determined that the target images from different perspectives exist occlusion; if the range is less than the preset range, it is determined that the target images from different perspectives do not exist occlusion.

2. The method of claim 1, wherein, After obtaining the plurality of images captured by the plurality of cameras from different perspectives, the method further comprises: obtaining coordinate system information of at least two images captured from different perspectives, wherein the coordinate system information is determined based on shooting parameters of the cameras and a positional relationship between two cameras; performing perspective matching on the at least two images from different perspectives based on the coordinate system information of the at least two images, to obtain a positional relationship of the target object in the two images; saving the positional relationship in a matrix to obtain a perspective conversion matrix, wherein the perspective conversion matrix records a coordinate conversion relationship between the images from different perspectives.

3. The method of claim 2, wherein, determining an association relationship of target views having the same target object from different perspectives based on the perspective conversion matrix.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: determining block images of the target object from different perspectives based on the region information of the target images from different perspectives; determining a type of the block images of the target object from different perspectives based on feature information of the target object, wherein the type comprises occlusion and no occlusion.

5. The method of claim 4, wherein, After determining whether the target images from different perspectives exist occlusion, the method further comprises: if it is determined that the target images from at least one perspective exist occlusion, deleting the target images in which the occlusion exists from the plurality of images.

6. The method of claim 5, wherein, After deleting the target images in which the occlusion exists from the plurality of images, the method further comprises: selecting target images in which no occlusion exists as an image sample set, and performing one of the following functions using images in the image sample set: recognizing an expression, recognizing a face, recognizing a behavior, and training a neural network model.

7. The method of claim 4, wherein, After determining whether the target images from different perspectives exist occlusion, the method further comprises: if it is determined that the target images from all perspectives exist occlusion, extracting a region in which no occlusion exists from the target images from different perspectives, and synthesizing the regions in which no occlusion exists from different perspectives to synthesize a new image in which no occlusion exists.

8. The method of claim 7, wherein, After synthesizing the new unoccluded image, the method further comprises: taking the synthesized new unoccluded image as an image sample set, and performing one of the following functions on images in the image sample set: identifying an expression, identifying a face, identifying a behavior, and training a neural network model.

9. A method of processing a multi-view image, characterized by, The method comprises: starting a plurality of cameras in a to-be-observed area, wherein at least two cameras are deployed at different positions in the to-be-observed area; obtaining a plurality of images captured by the plurality of cameras from different perspectives, wherein the plurality of images constitute a camera image set; displaying, in a display interface, target images in which a target object exists in the camera image set; displaying a region in which the target object exists in at least two target images, and obtaining region information of the region, wherein the region information comprises a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; displaying, in a display interface, at least one target image in which there is no occlusion, wherein whether the target image from different perspectives has occlusion is determined by analyzing the region information of at least two target images. The method comprises:

10. A multi-view image processing apparatus, characterized by comprising: starting a plurality of cameras in a to-be-observed area, wherein at least two cameras are deployed at different positions in the to-be-observed area; obtaining a plurality of images captured by the plurality of cameras from different perspectives, wherein the plurality of images constitute a camera image set; displaying, in a display interface, target images in which a target object exists in the camera image set; displaying a region in which the target object exists in at least two target images, and obtaining region information of the region, wherein the region information comprises a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; displaying, in a display interface, at least one target image in which there is no occlusion, wherein whether the target image from different perspectives has occlusion is determined by analyzing the region information of at least two target images. The method comprises:

11. A multi-view image processing apparatus, characterized by comprising: starting a plurality of cameras in a to-be-observed area, wherein at least two cameras are deployed at different positions in the to-be-observed area; obtaining a plurality of images captured by the plurality of cameras from different perspectives, wherein the plurality of images constitute a camera image set; displaying, in a display interface, target images in which a target object exists in the camera image set; displaying a region in which the target object exists in at least two target images, and obtaining region information of the region, wherein the region information comprises a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; displaying, in a display interface, at least one target image in which there is no occlusion, wherein whether the target image from different perspectives has occlusion is determined by analyzing the region information of at least two target images. The method comprises: starting a plurality of cameras in a to-be-observed area, wherein at least two cameras are deployed at different positions in the to-be-observed area; obtaining a plurality of images captured by the plurality of cameras from different perspectives, wherein the plurality of images constitute a camera image set; An acquisition module is configured to acquire a plurality of images captured by a plurality of cameras from different perspectives, wherein the plurality of images form a camera image set; A first display module is configured to display a target image in which a target object exists in the camera image set in a display interface; A second display module is configured to display a region in which the target object exists in at least two target images and acquire region information of the region, wherein the region information includes a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; A third display module is configured to display at least one target image in which no occlusion exists in the display interface, wherein whether the target image in different perspectives exists occlusion is determined by analyzing the region information of at least two target images; The third display module is configured to determine whether the target image in different perspectives exists occlusion by analyzing the region information of at least two target images by performing the following steps: determining a range of the block image in the target image based on the display coordinates in the region information; if the range is greater than a preset range, it is determined that the target image in different perspectives exists occlusion; if the range is less than the preset range, it is determined that the target image in different perspectives does not exist occlusion.

12. A storage medium, the storage medium comprising a stored program, characterized in that When the program is running, the device in which the storage medium is located is controlled to perform the following steps: acquiring a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; acquiring target images in which a target object exists from the plurality of images in different perspectives; determining a region in which the target object exists in at least two target images and acquiring region information of the region, wherein the region information includes a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; determining whether the target image in different perspectives exists occlusion based on the region information of at least two target images; wherein determining whether the target image in different perspectives exists occlusion based on the region information of at least two target images includes: determining a range of the block image in the target image based on the display coordinates in the region information; if the range is greater than a preset range, it is determined that the target image in different perspectives exists occlusion; if the range is less than the preset range, it is determined that the target image in different perspectives does not exist occlusion.

13. A multi-view image processing system, characterized by comprising: comprise: a processor; and a memory connected with the processor, configured to provide the processor with instructions for processing the following processing steps: acquiring a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; acquiring target images in which a target object exists from the plurality of images in different perspectives; determining a region in which the target object exists in at least two target images, and obtaining region information of the region, wherein the region information comprises a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; determining whether the target image under different perspectives exists occlusion based on the region information of the at least two target images; wherein the determining whether the target image under different perspectives exists occlusion based on the region information of the at least two target images comprises: determining a range of the block image in the target image based on the display coordinates in the region information; if the range is greater than a preset range, determining that the target image under different perspectives exists occlusion; if the range is less than the preset range, determining that the target image under different perspectives does not exist occlusion.

14. A method of processing a multi-view image, characterized by, comprising: obtaining a plurality of images captured by a plurality of cameras from different perspectives, wherein at least two cameras have different shooting positions; displaying target images in which a target object exists in the plurality of images; displaying identification information, wherein the identification information is used to indicate a region in which the target object exists in at least two target images; in a case where the region of the target object in the target images all exists occlusion, extracting a region in which the target object does not exist occlusion from the target images, and synthesizing the region in which the target object does not exist occlusion to generate a new occlusion-free image; wherein the method further comprises: determining a region in which the target object exists in at least two target images, and obtaining region information of the region, wherein the region information comprises a block image of the target object in the target image and display coordinates of the block image in the corresponding target image; determining a range of the block image in the target image based on the display coordinates in the region information; if the range is greater than a preset range, determining that the target image under different perspectives exists occlusion; if the range is less than the preset range, determining that the target image under different perspectives does not exist occlusion.

15. A method of processing a multi-view image, characterized by, comprising: displaying a plurality of images of a to-be-observed region under different perspectives, wherein the plurality of images comprise a target object; displaying identification information, wherein the identification information is used to indicate a region in which the target object exists in the plurality of images; in a case where the region of the target object in the plurality of images all exists occlusion, issuing selection prompt information, wherein the selection prompt information is used to prompt whether to synthesize the plurality of images in which occlusion exists; receiving selection information, and in a case where the selection information is yes, extracting a region in which the target object does not exist occlusion from the plurality of images, and synthesizing the region in which the target object does not exist occlusion to generate a new occlusion-free image; The method further comprises: determining a region in which the target object exists in at least two target images, and obtaining region information of the region, wherein the region information comprises: a block image of the target object in the target image, and display coordinates of the block image in the corresponding target image; determining a range of the block image in the target image based on the display coordinates in the region information; and determining that the target images under different viewing angles exist occlusion if the range is greater than a preset range, or determining that the target images under different viewing angles do not exist occlusion if the range is less than the preset range.

Citation Information

Patent Citations

  • Image collection control method and apparatus used for vehicle calibration

    CN107886544A

  • Face detection method and device, face recognition method and device and storage medium

    CN108710859A

  • Shielding evaluation method, device, equipment and system based on mutual view angle, and medium

    CN111489384A