Image Processing Method, Intelligent Camera Device, Intelligent Camera System and Storage Medium
Through point-to-point architecture and free space perspective conversion algorithm, the problems of bandwidth and high server performance in multi-camera systems are solved, and the computing power is dispersed and fast perspective selection is achieved, which is suitable for sports events, security and smart shopping malls.
Patent Information
- Application Number
- CN202410597551.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-05-14
AI Technical Summary
In the prior art, the bandwidth and server performance requirements of the multi-camera system are high, making it difficult for users to quickly and resiliently select the desired browsing screen.
Using a point-to-point architecture network transmission method, each camera has computing power as a central control camera device, and uses a free space perspective conversion algorithm to generate target picture data.
It reduces the communication bandwidth requirement and realizes dispersed computing power. Users can quickly select field of view pictures from specific perspectives, suitable for sports events, security and smart shopping malls.
Smart Images

Figure CN118474305B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent camera technology, and particularly to an image processing method, apparatus, terminal and storage medium. Background Art
[0002] The network transmission method between multiple cameras adopts a centralized architecture, that is, a centralized server is set up specifically for receiving and processing the acquisition data of all cameras, which often requires a huge amount of bandwidth and high-performance servers, and it is difficult for users to flexibly and quickly select the desired viewing screen.
[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an image processing method, apparatus, terminal and storage medium for the above-mentioned defects of the existing technology, aiming to solve the problems that the multi-camera system based on the centralized architecture in the existing technology has high requirements for bandwidth and server performance, and it is difficult for users to flexibly and quickly select the desired viewing screen.
[0005] The technical solution adopted by the present invention to solve the problem is as follows:
[0006] In the first aspect, an embodiment of the present invention provides an image processing method, which is applied to any intelligent camera device within a target area, and the method includes:
[0007] When a connection request from a user terminal is obtained, establish a data connection with the user terminal and determine itself as the central control camera device;
[0008] Obtain a target position, and determine whether to interact with other intelligent camera devices within the target area according to the target position, wherein a point-to-point data transmission method is adopted between any two intelligent camera devices;
[0009] If an interaction is performed, receive the image data of several intelligent camera devices through the interaction, and generate target screen data based on its own image data and the received image data through a preset free space view conversion algorithm, wherein the target screen data is used to reflect the field of view screen corresponding to the target position.
[0010] In an implementation manner, the determining whether to interact with other intelligent camera devices within the target area according to the target position includes:
[0011] Determine whether there are other intelligent camera devices whose fields of view corresponding to the target position have overlapping areas except itself according to the target position;
[0012] If such an intelligent camera device exists, determine to interact with the intelligent camera device.
[0013] In one embodiment, the method further includes:
[0014] If such an intelligent camera device does not exist, determine not to interact with other intelligent camera devices in the target area.
[0015] In one embodiment, each piece of image data received through the interaction is local image data of the overlapping area collected by the corresponding intelligent camera device, and each piece of image data received through the interaction has been converted to the coordinate system corresponding to the target position on the corresponding intelligent camera device based on the free space perspective conversion algorithm. Generating target picture data based on its own image data and the received image data through the preset free space perspective conversion algorithm includes:
[0016] Converting the local image data of the overlapping area collected by itself to the coordinate system corresponding to the target position through the free space perspective conversion algorithm;
[0017] Generating the target picture data according to the local image data of itself after coordinate system conversion and the received local image data.
[0018] In one embodiment, each piece of image data received through the interaction is local image data of the overlapping area collected by the corresponding intelligent camera device. Generating target picture data based on its own image data and the received image data through the preset free space perspective conversion algorithm includes:
[0019] Converting the local image data of the overlapping area collected by itself and the received local image data to the coordinate system corresponding to the target position through the free space perspective conversion algorithm;
[0020] Generating the target picture data according to all the local image data after coordinate system conversion.
[0021] In one embodiment, the method further includes:
[0022] If no interaction is performed, generate the target picture data based on its own image data through the free space perspective conversion algorithm.
[0023] In one embodiment, generating the target picture data based on its own image data through the free space perspective conversion algorithm includes:
[0024] Converting the local image data of the overlapping area collected by itself to the coordinate system corresponding to the target position through the free space perspective conversion algorithm;
[0025] Generate the target screen data based on the local image data of itself after coordinate system conversion.
[0026] In a second aspect, an embodiment of the present invention further provides an intelligent camera device, and the device includes:
[0027] A data connection module, configured to establish a data connection with the user terminal when a connection request from the user terminal is obtained, and determine itself as a central control camera device;
[0028] A data acquisition module, configured to acquire the target position of the user terminal, and determine whether to interact with other intelligent camera devices in the target area according to the target position, where the target position is the geographical location information of the user terminal or the position information input by the user on the user terminal, and a point-to-point data transmission method is adopted between two intelligent camera devices;
[0029] A screen generation module, configured to, if an interaction is performed, receive the image data of a plurality of intelligent camera devices through the interaction, and generate target screen data based on its own image data and the received image data through a preset free space perspective conversion algorithm, where the target screen data is used to reflect the field of view screen corresponding to the target position.
[0030] In a third aspect, an embodiment of the present invention further provides an intelligent camera system, and the system includes a plurality of intelligent camera devices as described above.
[0031] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a plurality of instructions are stored, and the instructions are suitable for being loaded and executed by a processor to implement the steps of any of the above-mentioned image processing methods.
[0032] Advantages of the present invention: In the embodiments of the present invention, the network transmission method between multiple cameras adopts a point-to-point architecture. Each camera has computing power and can be used as a central control camera device without a centralized server, realizing the dispersion of computing power and reducing the requirement for communication bandwidth. And combined with a preset free space perspective conversion algorithm, it provides a field of view screen with a specific perspective for users, and can be applied to multiple technical fields such as sports events, security fields, and smart shopping mall fields. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1It is a schematic flowchart of the image processing method provided by an embodiment of the present invention.
[0035] Figure 2 It is a schematic connection diagram of a user terminal and an intelligent camera device provided by an embodiment of the present invention.
[0036] Figure 3 It is a schematic architecture diagram of centralization and decentralization provided by an embodiment of the present invention.
[0037] Figure 4 It is a schematic architecture diagram of centralization and P2P provided by an embodiment of the present invention.
[0038] Figure 5 It is a schematic diagram of an observer screen provided by an embodiment of the present invention.
[0039] Figure 6 It is a schematic installation diagram of an intelligent camera device provided by an embodiment of the present invention.
[0040] Figure 7 It is a schematic diagram of the screen of the image obtained by the intelligent camera device provided by an embodiment of the present invention.
[0041] Figure 8 It is a schematic diagram of the free space perspective conversion algorithm provided by an embodiment of the present invention.
[0042] Figure 9 It is a schematic diagram of an overlapping area provided by an embodiment of the present invention.
[0043] Figure 10 It is a schematic diagram of the user's feedback screen provided by an embodiment of the present invention.
[0044] Figure 11 It is a schematic diagram of the entire process from the generation to the feedback of the target screen data provided by an embodiment of the present invention.
[0045] Figure 12 It is a schematic diagram of a sports event application scenario provided by an embodiment of the present invention.
[0046] Figure 13 It is a schematic diagram of a camera array provided by an embodiment of the present invention.
[0047] Figure 14 It is a schematic diagram of different space perspectives provided by an embodiment of the present invention.
[0048] Figure 15 It is a schematic module diagram of an intelligent camera device provided by an embodiment of the present invention.
[0049] Figure 16 It is the principle framework of the terminal provided by an embodiment of the present invention. Detailed implementation manners
[0050] The present invention discloses an image processing method, an intelligent camera device, an intelligent camera system and a storage medium. To make the objectives, technical solutions and effects of the present invention clearer and more explicit, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0051] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.
[0052] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.
[0053] In view of the above defects of the prior art, the present invention provides an image processing method, which is applied to any intelligent camera device within a target area. The method includes: when a connection request from a user terminal is received, establishing a data connection with the user terminal and determining itself as the central control camera device; obtaining a target position, and judging whether to interact with other intelligent camera devices within the target area according to the target position, wherein point-to-point data transmission is adopted between any two intelligent camera devices; if an interaction is to be performed, receiving image data of a number of intelligent camera devices through the interaction, and generating target picture data based on its own image data and the received image data through a preset free space view conversion algorithm, wherein the target picture data is used to reflect the field of view picture corresponding to the target position. In the present invention, the network transmission mode between multiple cameras adopts a point-to-point architecture, and each camera has computing power and can be used as the central control camera device without a centralized server, realizing decentralized computing power and reducing the requirement for communication bandwidth. Moreover, the present invention also combines a preset free space view conversion algorithm, which can provide a field of view picture with a specific view for users and is applicable to multiple technical fields such as sports events, security fields, and smart shopping malls.
[0054] As Figure 1 shown, the method specifically includes:
[0055] Step S100: When a connection request from a user terminal is received, establish a data connection with the user terminal and determine itself as the central control camera device.
[0056] Specifically, the target area in this embodiment can be an area selected by the user or the area where the user terminal is currently located. The user terminal can be the user's mobile phone or other devices with networking functions. The intelligent camera device in this embodiment has an independent computing unit. The user terminal can be connected to the independent computing unit of any intelligent camera device within the target area through the network, and the connected intelligent camera device becomes the current central control camera device in the target area and can dispatch and coordinate the nearby intelligent camera devices.
[0057] For example, as Figure 2 shown, multiple intelligent camera devices are pre-set within the target area, and each intelligent camera device corresponds to a different mobile hotspot (WIFI). The user selects an intelligent camera device by himself / herself and connects to the independent computing unit of this intelligent camera device through the network, and this intelligent camera device becomes the central control camera device.
[0058] Step S200: Obtain a target position, and judge whether to interact with other intelligent camera devices within the target area according to the target position, wherein point-to-point data transmission is adopted between any two intelligent camera devices.
[0059] Specifically, the technical solution of this embodiment is mainly used to generate a field of view screen corresponding to a target position. The target position can be the position information input by the user on the terminal, or the position information input by the user through the terminal in combination with other devices, or the geographical location information of the user terminal itself. Usually, the central control camera device is the intelligent camera device closest to the target position selected by the user, and the overlapping area between the field of view of the central control camera device and the field of view of the target position is relatively large. However, other intelligent camera devices within the target area may also have an overlapping area with the field of view of the target position. Therefore, the central control camera device needs to determine whether to interact with other intelligent camera devices according to the target position to obtain the image data they collect. To reduce the requirement for communication bandwidth, this embodiment sets that the intelligent camera devices adopt a peer-to-peer data transmission method (Peer-to-Peer, P2P), and the user terminal and the intelligent camera device also adopt a direct communication method without forwarding information through a central server.
[0060] As Figure 3 , 4 shown, in terms of the design mode, P2P breaks the traditional client / server mode (Client / Server, i.e., C / S). Each node in the P2P network (i.e., the intelligent camera device and its corresponding independent computing unit) has an equal status and can act as a server (i.e., the central control camera device) to provide services for other nodes, and at the same time enjoy the services provided by other nodes. What is provided / enjoyed between each node is camera data and the target position, where the target position can include parameters such as coordinates and angles. The P2P architecture can use small-area computing power hardware to complete the required computing work, reduce the burden of unnecessary data transmission load, and solve the requirements for a large amount of computing power, storage space, and communication bandwidth of the centralized server.
[0061] In one implementation, the image data collected by each intelligent camera device is image data with a distance value, and a binocular Camera or a ToF Camera, etc., with a distance value can be used.
[0062] In one implementation, the determining whether to interact with other intelligent camera devices within the target area according to the target position includes:
[0063] Determining whether there are intelligent camera devices other than itself that have an overlapping area with the field of view corresponding to the target position according to the target position;
[0064] If there are such intelligent camera devices, then determine to interact with such intelligent camera devices.
[0065] Generally speaking, after each intelligent camera device is connected to a user terminal and a service request is made, the respective video data can start to be exchanged. Specifically, the central control camera device will, according to the received target position, determine whether there are other intelligent camera devices around whose fields of view overlap with the field of view of the target position. If there are such intelligent camera devices, the central control camera device needs to obtain the video data they collect through interaction. The video data of such intelligent camera devices will be transmitted to the independent computing unit of the central control camera device in a P2P manner to generate a more comprehensive field of view picture of the target position.
[0066] In one implementation, the method further includes:
[0067] If there are no such intelligent camera devices, it is determined not to interact with other intelligent camera devices within the target area.
[0068] Specifically, if there are no such intelligent camera devices, it means that only the field of view of the central control camera device overlaps with the field of view of the target position, and the central control camera device does not need to interact with other intelligent camera devices.
[0069] In one implementation, each piece of video data received through interaction is the local video data of the overlapping area collected by the corresponding intelligent camera device, and each piece of video data received through interaction has been converted to the coordinate system corresponding to the target position on the corresponding intelligent camera device based on the free space perspective conversion algorithm. Generating the target picture data based on the preset free space perspective conversion algorithm for its own video data and the received video data includes:
[0070] Converting the local video data of the overlapping area collected by itself to the coordinate system corresponding to the target position through the free space perspective conversion algorithm;
[0071] Generating the target picture data according to the local video data of itself after coordinate system conversion and the received local video data.
[0072] Specifically, since there are differences between the coordinate systems of each intelligent camera device and the target location, in order to accurately generate the field of view image of the target location, a free space perspective conversion algorithm is preset in this embodiment. Each intelligent camera device contains this algorithm, which can perform coordinate system conversion on the local image data belonging to the overlapping area in the image data collected by itself through this algorithm, and uniformly convert it to the coordinate system of the target location, and display the 2D image corresponding to the real space. In addition, in order to reduce the computing power burden of the central control intelligent camera device, this embodiment adopts a computing power dispersion method, and distributes the algorithm process of coordinate system conversion to each intelligent camera device interacting with the central control intelligent camera device for execution. That is, each intelligent camera device interacting with the central control intelligent camera device needs to execute the coordinate system conversion step locally, and then send the locally image data after coordinate system conversion to the central control intelligent camera device. The central control intelligent camera device only needs to perform coordinate system conversion on the local image data of the overlapping area collected by itself, and then fuse all the locally image data after coordinate system conversion to obtain the target image data for reflecting the field of view image of the target location. As Figure 5 shown, the target image data can actually include the image data of multiple intelligent camera devices. Finally, the target image data can be generated to display the images captured by the intelligent camera devices in the nearby area according to the requirements of the user (i.e., the observer), and can include the images captured by multiple intelligent camera devices. For example, the user terminal is connected to the current intelligent camera device, and the images of the three nearest intelligent camera devices are displayed.
[0073] In one implementation, if the positions of two intelligent camera devices do not change, the calculation method of the image overlapping area specifically includes:
[0074] Pre-shoot multiple groups of images of the calibration checkerboard through two intelligent camera devices to obtain an image set;
[0075] Calculate the internal and external parameters between the two intelligent camera devices according to the image set, and calculate the fixed image overlapping area according to the internal and external parameters.
[0076] Specifically, taking two devices as an example (the device can be an intelligent camera device or a user terminal), assuming that the two devices are intelligent camera devices, after the intelligent camera devices are installed at the beginning, the calibration of the camera will be performed, and the overlapping area between the two intelligent camera devices will be calculated. The calculation method of the overlapping area can use the checkerboard as the target. First, install two intelligent camera devices, take multiple groups of images for the two intelligent camera devices, and use their images to obtain the internal and external parameters between the two intelligent camera devices, and calculate the image overlapping area of the two intelligent camera devices through the internal and external parameters. Figure 6 is a schematic diagram of the installation of the intelligent camera device, Figure 7It is a schematic diagram of the image obtained by the corresponding intelligent camera device. Taking the checkerboard in the figure as the target, the overlapping area of the two intelligent camera devices is calculated and the position is recorded. If the position of the intelligent camera device remains unchanged, the image overlapping area is fixed.
[0077] In another implementation, if the positions of the two intelligent camera devices change, the calculation method of the image overlapping area specifically includes:
[0078] Re-obtain the images of the two intelligent camera devices, extract a number of feature points from the images, and calculate the feature descriptors of each feature point;
[0079] Compare the feature descriptors of the feature points in the two images through a preset matching algorithm to obtain the corresponding relationship between the feature points;
[0080] Take the corresponding relationship between the feature points as the matching result, and perform image registration on the two images according to the matching result;
[0081] Judge whether the two images overlap. If the two images overlap, calculate the image overlapping area.
[0082] Generally speaking, if the positions of the two intelligent camera devices change, the feature point method is used to find the image overlapping area and calculate the position of the image overlapping area. As Figure 8 shown, the free space view transformation algorithm uses the overlapping area of the two intelligent camera devices to find the corresponding features, and calculates the relative difference positions of the two intelligent camera devices, and finds the corresponding features, which can be analogous to the checkerboard method. Specifically, first re-obtain the images through the two intelligent camera devices, extract the feature points in the images, and for each feature point, calculate the feature descriptor of the feature point for subsequent feature point matching. The matching process uses the specified matching algorithm to compare the feature points in the two images, so as to determine the corresponding relationship between the feature points and obtain the matching result. Perform subsequent image registration according to the matching result, judge whether the two images overlap, and calculate the image overlapping area when the two images overlap. If there is no overlapping area in the images of the two intelligent camera devices, a specified picture (such as a black picture) can be used as the output, or the image of one intelligent camera device can be output.
[0083] In one implementation, the calculation method of the overlapping area of the two images specifically includes:
[0084] Calculate the bounding boxes of the two images respectively, where the bounding box of each image is determined based on the minimum and maximum horizontal and vertical coordinates of the image on the plane;
[0085] Judge whether there is an intersection area between the bounding boxes of the two images;
[0086] If it exists, calculate the intersection coordinates based on the bounding boxes of the two images;
[0087] Calculate the area of the intersection region based on the intersection coordinates, and use the area of the intersection region as the area of the overlapping region of the two images.
[0088] To quickly determine whether there is an overlapping part between two images, in this embodiment, the bounding box in computational geometry is used for image analysis and calculation. Specifically, each image has a bounding box defined by its minimum abscissa, minimum ordinate, maximum abscissa, and maximum ordinate on the plane. Determine whether the bounding box meets the preset overlapping conditions. If it is satisfied, it is determined that there is an overlapping part between the two images, that is, the intersection region. The intersection region is usually rectangular, so the area of the overlapping region can be calculated by calculating the four coordinate points of the rectangle corresponding to the intersection region.
[0089] Illustrate with an example, bounding box overlap check: Assume that each image has a bounding box defined by its minimum and maximum coordinates on the plane: (x_min, y_min) and (x_max, y_max). For two images (A) and (B), the bounding boxes are defined as follows:
[0090] Image (A): (A_min = (A_xmin, A_ymin), A_max = (A_xmax, A_ymax);
[0091] Image (B): (B_min = (B_xmin, B_ymin), B_max = (B_xmax, B_ymax);
[0092] The conditions for image overlap are:
[0093] 1. (A_xmax >= B_xmin) and (A_xmin <= B_xmax);
[0094] 2. (A_ymax >= B_ymin) and (A_ymin <= B_ymax);
[0095] These conditions ensure that the two bounding boxes have at least shared space in the horizontal and vertical directions.
[0096] If the images overlap, calculate the area of the overlapping region by determining the coordinates of the intersection rectangle. The calculation method of the intersection coordinates is as follows:
[0097] I_xmin = max(A_xmin, B_xmin);
[0098] I_xmax = min(A_xmax, B_xmax);
[0099] I_ymin = max(A_ymin, B_ymin);
[0100] I_ymax = min(A_ymax, B_ymax);
[0101] If (I_xmax >= I_xmin) && (I_ymax >= I_ymin), then the area of the intersection region can be calculated as:
[0102] Area = (I_xmax - I_xmin) * (I_ymax - I_ymin).
[0103] In one implementation, the method further includes:
[0104] Construct a sphere equation and a line equation respectively based on the position and angle input by the observer;
[0105] Calculate the intersection coordinates of the line and the sphere according to the sphere equation and the line equation, and determine a new viewing angle according to the intersection coordinates;
[0106] If the new viewing angle is a blind area of the central control intelligent camera device and there is no overlapping area with other intelligent camera devices, then return a specified screen.
[0107] Specifically, the user can arbitrarily select an intelligent camera device to log in, and can move forward, backward, and pan according to the first-person viewing angle. The terminal changes its viewing angle according to the input angle and position. To reduce the computational overhead of the terminal, in this embodiment, a sphere equation and a line equation are respectively constructed based on the observer's position and the input angle, and the intersection coordinates of the line and the sphere can be calculated by solving the two equations simultaneously. The intersection coordinates can reflect the direction and position that the observer is looking at. When the viewing angle is an area that the intelligent camera device cannot see and there is no overlapping area with another intelligent camera device, then a specified screen (such as a black screen) can be used as the output result.
[0108] For example, (x1, y1, z1) is the observer's position, (x2, y2, z2) is the direction and position being looked at, and (x3, y3, z3) is the upward orientation.
[0109] The 3D line passing through (x1, y1, z1) is:
[0110] α(x - x1) = β(y - y1) = γ(z - z1);
[0111] Use the sphere formula for simplification (the general formula for an ellipsoid is too complex):
[0112] (x 2 + y 2 + z 2 = r 2 )
[0113] α = sin(angle) * cos(pitch), β = cos(angle) * cos(pitch), γ = sin(pitch);
[0114] The intersection coordinates of the straight line and the sphere can be obtained by solving the above equations simultaneously as (x2, y2, z2).
[0115] In another implementation, each piece of image data received through interaction is local image data of the overlapping area collected by the corresponding intelligent camera device. The target picture data is generated based on its own image data and the received image data through a preset free-space perspective conversion algorithm, including:
[0116] Converting the local image data of the overlapping area collected by itself and each piece of received local image data to the coordinate system corresponding to the target position through the free-space perspective conversion algorithm;
[0117] Generating the target picture data according to all the local image data after coordinate system conversion.
[0118] Specifically, in this embodiment, the method of concentrating computing power can also be adopted, and the algorithm process of coordinate system conversion is uniformly executed on the central control intelligent camera device. That is, the central control intelligent camera device needs to perform coordinate system conversion on its own local image data and other local image data obtained through interaction, and then fuse all the local image data after coordinate system conversion to obtain the target picture data for reflecting the field of view of the target position.
[0119] In one implementation, the method further includes:
[0120] If no interaction is performed, the target picture data is generated based on its own image data through the free-space perspective conversion algorithm.
[0121] Specifically, if the central control camera device determines that there is no need to interact with other intelligent camera devices, indicating that only the field of view of the central control camera device overlaps with the field of view of the target position, the central control camera device only needs to use the image data collected by itself to generate the target picture data for reflecting the field of view of the target position.
[0122] In one implementation, generating the target picture data based on its own image data through the free-space perspective conversion algorithm includes:
[0123] Converting the local image data of the overlapping area collected by itself to the coordinate system corresponding to the target position through the free-space perspective conversion algorithm;
[0124] Generate the target picture data based on the local image data of itself after coordinate system conversion.
[0125] Specifically, if the central control camera device only needs to use its own image data to generate the target picture data, the central control camera device only needs to input the local image data corresponding to the overlapping area in its own image data into the free space perspective conversion algorithm for calculation, and then the target picture data can be obtained through coordinate system conversion.
[0126] For example, as Figure 9 shown, P1 is the user and P2 is the intelligent camera device. Determine the shooting data of the dark area in P2 according to the field of view overlapping area of P1 and P2, and then transmit the shooting data to P1 after coordinate conversion. As Figure 10 shown, the picture obtained by P1 only has the picture corresponding to the overlapping area, and only a black picture can be seen in the non-overlapping area.
[0127] In one implementation, the method further includes:
[0128] Send the target picture data to the user terminal.
[0129] Specifically, as Figure 11 shown, the central control camera device can send the generated target picture data back to the user terminal.
[0130] For example, as Figure 12 、 13 shown, when this embodiment is applied to a sports event, the user can connect their terminal to any intelligent camera device pre-set on the field and input the target position for expected immersive viewing on their terminal. The connected intelligent camera device is the central control camera device, which will receive the target position sent by the user. After the above solution, the target picture data for reflecting the field of view picture of the target position on the field is generated, and the target picture data is sent back to the user terminal, and the user can immerse themselves in watching the field picture of the target position.
[0131] In another implementation, the method further includes:
[0132] Send the target picture data to other devices selected by the user.
[0133] Specifically, the central control camera device can send the generated target picture data to other devices selected by the user.
[0134] For example, when this embodiment is applied to the field of smart shopping malls, users can pre-set a anti-loss program on the user terminal. That is, when the location information of the user terminal does not change within a preset time period, the anti-loss program is automatically started to establish a connection with an intelligent camera device in the shopping mall; or the user uses other devices that have logged in to the user terminal to remotely control the user terminal to establish a connection with an intelligent camera device in the shopping mall. The connected intelligent camera device is the central control camera device, which will receive the current location of the user terminal, that is, the target location. The central control camera device generates target picture data for reflecting the field of view picture of the user terminal through the above solution, and sends the target picture data to other devices that have logged in to the user terminal. The user can receive the target picture data through other devices and then identify where their terminal is left in the shopping mall.
[0135] In summary, the advantages of the present invention are as follows:
[0136] 1. P2P architecture: It can reduce the burden of data transmission load and solve the requirements for a large amount of computing power, storage space, and communication bandwidth of the centralized server;
[0137] 2. Free space perspective conversion algorithm: After the intelligent camera device obtains the visual image of the real world and then matches the input real space coordinates, through the free space perspective conversion algorithm, a 2D image corresponding to the real space is displayed (as Figure 14 shown).
[0138] 3. P2P architecture + free space perspective conversion algorithm: It can arbitrarily increase or decrease the number of terminal devices, achieve the user's goal, reduce a lot of computing power, and maintain the flexibility of use.
[0139] Based on the above embodiments, the present invention also provides an intelligent camera device, as Figure 15 shown, the device includes:
[0140] A data connection module 01, configured to establish a data connection with the user terminal when a connection request of the user terminal is obtained, and determine itself as the central control camera device;
[0141] A data acquisition module 02, configured to acquire the target location of the user terminal, and determine whether to interact with other intelligent camera devices in the target area according to the target location, where the target location is the geographical location information of the user terminal or the location information input by the user on the user terminal, and a point-to-point data transmission method is adopted between each pair of intelligent camera devices;
[0142] A screen generation module 03, configured to, if an interaction is performed, receive image data of a plurality of intelligent camera devices through the interaction, and generate target screen data based on its own image data and the received image data through a preset free space perspective conversion algorithm, wherein the target screen data is used to reflect the field of view screen corresponding to the target position.
[0143] Based on the above embodiments, the present invention further provides an intelligent camera system, which includes a plurality of intelligent camera devices as shown above.
[0144] Specifically, the intelligent camera device in this embodiment can be regarded as a terminal with computing capabilities, and its principle block diagram can be as Figure 16 shown. The terminal includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes an image processing method.
[0145] Those skilled in the art can understand that Figure 16 the principle block diagram shown in
[0146] merely shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0147] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0148] In summary, the present invention discloses an image processing method, an intelligent camera device, an intelligent camera system, and a storage medium. The method is applied to any intelligent camera device within a target area. The method includes: when a connection request from a user terminal is obtained, establishing a data connection with the user terminal and determining itself as a central control camera device; obtaining a target position, and judging whether to interact with other intelligent camera devices within the target area according to the target position, wherein a point-to-point data transmission method is adopted between any two intelligent camera devices; if an interaction is to be performed, receiving image data of a plurality of intelligent camera devices through the interaction, and generating target picture data based on its own image data and the received image data through a preset free space view conversion algorithm, wherein the target picture data is used to reflect the field of view picture corresponding to the target position. In the present invention, the network transmission method between multiple cameras adopts a point-to-point architecture. Each camera has computing power and can be used as a central control camera device without a centralized server, realizing the decentralization of computing power and reducing the requirement for communication bandwidth. Moreover, the present invention also combines a preset free space view conversion algorithm, which can provide a field of view picture with a specific view for users and is applicable to multiple technical fields such as sports events, the security field, and the intelligent shopping mall field.
[0149] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or modifications can be made according to the above description, and all such improvements and modifications shall fall within the protection scope of the appended claims of the present invention.
Claims
1. An image processing method, characterized in that, The method is applied to any intelligent camera device within a target area, and the method includes: When a connection request from a user terminal is received, establish a data connection with the user terminal and determine itself as the central control camera device; Obtain a target position, and determine whether to interact with other intelligent camera devices within the target area according to the target position, wherein a point-to-point data transmission method is adopted between any two intelligent camera devices; If interaction is to be performed, receive image data of a number of intelligent camera devices through the interaction, and generate target picture data based on its own image data and the received image data through a preset free-space perspective conversion algorithm, wherein the target picture data is used to reflect the field-of-view picture corresponding to the target position; Determining whether to interact with other intelligent camera devices within the target area according to the target position includes: Determine whether there is an intelligent camera device whose field of view corresponding to the target position has an overlapping area except itself according to the target position; If there is such an intelligent camera device, determine to interact with the intelligent camera device; If the positions of two intelligent camera devices do not change, the calculation method of the image overlapping area specifically includes: pre-shooting multiple groups of images of a calibration checkerboard through the two intelligent camera devices to obtain an image set; calculating the internal parameters and external parameters between the two intelligent camera devices according to the image set, and calculating a fixed image overlapping area according to the internal parameters and external parameters; If the positions of two intelligent camera devices change, the calculation method of the image overlapping area specifically includes: re-obtain the images of the two intelligent camera devices, extract a number of feature points according to the images, and calculate the feature descriptors of each feature point; compare the feature descriptors of the feature points in the two images through a preset matching algorithm to obtain the corresponding relationship between the feature points; use the corresponding relationship between the feature points as the matching result, and perform image registration on the two images according to the matching result; determine whether the two images overlap, if the two images overlap, calculate the image overlapping area; the calculation method of the overlapping area of the two images specifically includes: calculate the bounding boxes of the two images respectively, wherein the bounding box of each image is determined based on the minimum and maximum horizontal and vertical coordinates of the image on the plane; determine whether there is an intersection area between the bounding boxes of the two images; if there is, calculate the intersection coordinates according to the bounding boxes of the two images; calculate the area of the intersection area according to the intersection coordinates, and use the area of the intersection area as the area of the overlapping area of the two images.
2. The image processing method according to claim 1, wherein The method further includes: If there is no such intelligent camera device, determine not to interact with other intelligent camera devices within the target area.
3. The image processing method according to claim 1, wherein Each piece of image data received through the interaction is local image data of the overlapping area collected by the corresponding intelligent camera device, and each piece of image data received through the interaction has been converted to the coordinate system corresponding to the target position on the corresponding intelligent camera device based on the free-space perspective conversion algorithm. Generating target picture data based on its own image data and the received image data through a preset free-space perspective conversion algorithm includes: Convert the local image data of the overlapping area collected by itself to the coordinate system corresponding to the target position through the free space perspective conversion algorithm; Generate the target picture data according to the local image data of itself after coordinate system conversion and the received local image data.
4. The image processing method according to claim 1, wherein Each piece of image data received through interaction is the local image data of the overlapping area collected by the corresponding intelligent camera device. The method of generating target picture data based on its own image data and the received image data through a preset free space perspective conversion algorithm includes: Convert the local image data of the overlapping area collected by itself and the received local image data to the coordinate system corresponding to the target position through the free space perspective conversion algorithm; Generate the target picture data according to all the local image data after coordinate system conversion.
5. The image processing method according to claim 1, wherein The method further includes: If no interaction is performed, generate the target picture data based on its own image data through the free space perspective conversion algorithm.
6. The image processing method according to claim 5, wherein The method of generating the target picture data based on its own image data through the free space perspective conversion algorithm includes: Convert the local image data of the overlapping area collected by itself to the coordinate system corresponding to the target position through the free space perspective conversion algorithm; Generate the target picture data according to the local image data of itself after coordinate system conversion.
7. An intelligent camera device, characterized in that, The device includes: A data connection module, configured to establish a data connection with the user terminal when a connection request from the user terminal is obtained, and determine itself as the central control camera device; A data acquisition module, configured to acquire the target position of the user terminal, and determine whether to interact with other intelligent camera devices in the target area according to the target position, where the target position is the geographical location information of the user terminal or the position information input by the user on the user terminal, and a point-to-point data transmission method is adopted between two intelligent camera devices; A picture generation module, configured to, if interaction is performed, receive the image data of a number of intelligent camera devices through interaction, and generate target picture data based on its own image data and the received image data through a preset free space perspective conversion algorithm, where the target picture data is used to reflect the field of view picture corresponding to the target position; Determining whether to interact with other intelligent camera devices in the target area according to the target position includes: Determine whether there is an intelligent camera device whose field of view corresponding to the target position has an overlapping area except itself according to the target position; If there is such an intelligent camera device, determine to interact with the intelligent camera device; If the positions of two intelligent camera devices do not change, the calculation method of the image overlapping area specifically includes: pre-shooting multiple groups of images of a calibration checkerboard by two intelligent camera devices to obtain an image set; calculating the internal and external parameters between the two intelligent camera devices according to the image set, and calculating the fixed image overlapping area according to the internal and external parameters; If the positions of two intelligent camera devices change, the method for calculating the overlapping area of images specifically includes: re-acquiring the images of the two intelligent camera devices, extracting a number of feature points from the images, and calculating the feature descriptors of each feature point; comparing the feature descriptors of the feature points in the two images through a preset matching algorithm to obtain the corresponding relationship between the feature points; using the corresponding relationship between the feature points as the matching result, and performing image registration on the two images according to the matching result; determining whether the two images overlap, and if the two images overlap, calculating the overlapping area of the images; the method for calculating the overlapping area of the two images specifically includes: respectively calculating the bounding boxes of the two images, where the bounding box of each image is determined based on the minimum and maximum horizontal and vertical coordinates of the image on the plane; determining whether there is an overlapping area between the bounding boxes of the two images; if there is, calculating the intersection coordinates according to the bounding boxes of the two images; calculating the area of the overlapping area according to the intersection coordinates, and using the area of the overlapping area as the area of the overlapping area of the two images.
8. An intelligent camera system, characterized in that, The system includes a number of intelligent camera devices as described in claim 7.
9. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that, The instructions are suitable for being loaded and executed by a processor to implement the steps of the image processing method described in any one of claims 1-6 above.
Citation Information
Patent Citations
Method and device for mobile terminal synergistic shooting
CN104426588A
Video monitoring system and method
CN107231547A
Recovery machine built-in camera shooting control method, device and system
CN114945071A