A method and apparatus for unfolding a fisheye image
By adjusting the starting position during the fisheye image unfolding process, the problem of object segmentation during fisheye image unfolding is solved, improving image quality and target detection accuracy, simplifying operation and reducing resource consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2022-08-11
- Publication Date
- 2026-05-05
AI Technical Summary
If a suitable starting position is not determined during the unfolding process of a fisheye image, the image of the object to be detected may be segmented at both ends of the unfolded image, affecting the image quality and the accuracy of object detection.
The original fisheye image is unfolded based on the starting position to obtain the first unfolded image. The tail end of the first unfolded image is then stitched together with the head end of the copied image to form a stitched image. Target detection is then performed to determine whether there are segmented detection objects. If so, the starting position is adjusted and the image is unfolded again until there are no segments.
It improves the unfolding effect of fisheye images and the accuracy of target detection, simplifies the operation process, and reduces resource consumption.
Smart Images

Figure CN115471643B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method and apparatus for unfolding a fisheye image. Background Technology
[0002] A fisheye camera is an ultra-wide-angle camera that simulates the effect of a fish looking up at the water's surface. Fisheye cameras offer advantages such as a large field of view, the ability to capture diverse scenes, and suitability for shooting in confined spaces. By capturing images with a fisheye camera, the field of view is significantly larger than that of a conventional lens, resulting in richer information in the captured image. Therefore, fisheye cameras are widely used in fields such as virtual reality technology, robot navigation, visual monitoring, and intelligent assisted driving.
[0003] Fisheye images captured by fisheye cameras exhibit distortion at their edges. Furthermore, the information content is highest and the distortion lowest near the center of the image, decreasing with increasing radius and gradually increasing with decreasing information content and distortion. Therefore, information obtained directly from a fisheye image by the naked eye is inaccurate, and image recognition based on fisheye images is relatively poor. Consequently, in practical use, fisheye images are typically processed by unfolding them according to the camera's characteristics.
[0004] However, if a suitable starting position is not determined during the unfolding process of a fisheye image, the image of the object to be detected may be segmented at both ends of the unfolded image, resulting in poor image quality of the unfolded fisheye image and affecting the accuracy of target detection. Summary of the Invention
[0005] This application provides a method and apparatus for unfolding a fisheye image, which can be used to adjust the starting position of the unfolded fisheye image to improve the image quality based on the unfolded fisheye image and the accuracy of target detection.
[0006] To achieve the above technical objectives, this application adopts the following technical solution:
[0007] Firstly, this application provides a method for unfolding a fisheye image. The method specifically includes: firstly, unfolding the original fisheye image based on a starting position to obtain a first unfolded image of the original fisheye image, wherein the unfolding start end of the original fisheye image corresponds to the beginning end of the first unfolded image, and the unfolding end end of the original fisheye image corresponds to the end end of the first unfolded image. Then, stitching the end end of the first unfolded image with the beginning end of a copied image to obtain a stitched image of the original fisheye image, wherein the copied image is an image obtained by copying the first unfolded image. Next, performing object detection on the stitched image to determine whether there is an image of a segmented detection object in the first unfolded image. If there is an image of a segmented detection object in the first unfolded image, adjusting the starting position and unfolding the original fisheye image based on the adjusted position to obtain a second unfolded image, wherein there is no image of a segmented detection object in the second unfolded image.
[0008] The fisheye image unfolding method provided in this application has at least the following beneficial effects: First, the method pre-unfolds the original fisheye image based on a preset initial position to obtain a preliminary unfolded result. Then, based on the circular characteristics of a 360° circular fisheye image during unfolding, the images at the beginning and end of the first unfolded image can be stitched together to obtain the image at the starting position in the original fisheye image. Therefore, a stitched image of the preliminary unfolded image can be obtained through a copying and stitching process. This stitched image is then subjected to detection processing to determine whether the unfolded image based on the initial position will segment the image of the detected object at both ends of the unfolded image, i.e., whether there is a segmented image of the detected object in the first unfolded image. This determination process, combining image copying and stitching, is simple to operate, easy to implement, and more practical. Furthermore, when it is determined that there is a segmented image of the detected object in the first unfolded image, the starting position of the image unfolding operation can be adjusted according to the target detection result, and the initial fisheye image can be unfolded based on the adjusted starting position. This results in a better image quality in the unfolded image, improving the user's viewing experience and increasing the accuracy of target detection on the unfolded image.
[0009] In one possible implementation, the above-mentioned stitching of the tail end of the first unfolded image with the head end of the copied image includes: completely copying the first unfolded image to obtain a first copied image; and stitching the tail end of the first unfolded image with the head end of the first copied image to obtain a stitched image of the original fisheye image.
[0010] As can be understood, this implementation first completely copies the first unfolded image, thus obtaining a stitched image of the original fisheye image. Based on this stitched image, it can reflect the image information of the target object to be detected relatively completely and accurately.
[0011] In another possible implementation, the above-mentioned stitching of the tail end of the first unfolded image with the head end of the copied image includes: starting from the head end of the first unfolded image, partially copying the first unfolded image to obtain a second copied image, wherein the width of the second copied image is greater than or equal to a first preset width threshold; and stitching the tail end of the first unfolded image with the head end of the second copied image to obtain a stitched image of the original fisheye image.
[0012] It is understandable that this implementation method obtains the stitched image by partially copying the first unfolded image, which can reduce the resource consumption of the image copying process and reduce resource waste.
[0013] Optionally, the above-mentioned stitching of the tail end of the first unfolded image with the head end of the copied image may further include: starting from the tail end of the first unfolded image, partially copying the first unfolded image to obtain another second copied image, wherein the width of the second copied image is greater than or equal to a first preset width threshold; and stitching the head end of the first unfolded image with the tail end of the second copied image to obtain a stitched image of the original fisheye image.
[0014] In another possible implementation, the above-mentioned object detection of the stitched image to determine whether there is an image of a segmented detection object in the first unfolded image includes: performing object detection on all images of the stitched image to obtain one or more object detection boxes, each object detection box including an image of a detection object; if the area of any object detection box in the one or more object detection boxes on the stitched image overlaps with the stitching position, then it is determined that there is an image of a segmented detection object in the first unfolded image; the stitching position is the stitching position between the stitched image, the first unfolded image, and the copied image.
[0015] It is understandable that this implementation performs target detection on all images in the obtained stitched image, thus resulting in a high accuracy rate for the target detection results.
[0016] In another possible implementation, the above-mentioned stitched image is subjected to object detection to determine whether there is an image of a segmented detection object in the first expanded image, including: performing object detection on the first sub-image in the stitched image to obtain one or more object detection boxes, each object detection box including an image of a detection object, wherein the first sub-image includes the stitching position of the first expanded image and the copied image, and the first sub-image includes the entire image of the first expanded image and the second sub-image in the copied image, the width of the second sub-image being greater than or equal to a first preset width threshold; if the area of any object detection box in the one or more object detection boxes on the stitched image overlaps with the stitching position, then it is determined that there is an image of a segmented detection object in the first expanded image; the stitching position is the stitching position of the first expanded image and the copied image in the stitched image.
[0017] It is understandable that this implementation method performs object detection on a portion of the first sub-image, i.e., the stitched image, to determine if there is an image of the segmented detection object in the first unfolded image. In this way, the resource consumption of the image object detection process can be reduced, and the waste of resources can be minimized.
[0018] In another possible implementation, the first preset width threshold is the preset maximum width value of the target detection box, which is a width value determined based on the maximum imaging size of the detected object in the original fisheye image.
[0019] In another possible implementation, the above-mentioned object detection of the stitched image to determine whether there is a segmented detection object in the first expanded image includes: dividing the stitched image into one or more third sub-images based on a preset width; wherein the width of the repeated image between any two adjacent third sub-images is greater than or equal to a second preset width threshold; performing object detection on the one or more third sub-images respectively to obtain one or more object detection boxes, each object detection box including an image of a detection object; if the area of any object detection box in the one or more object detection boxes on the third sub-image overlaps with the stitching position, then it is determined that there is a segmented detection object in the first expanded image; the stitching position is the stitching position between the first expanded image and the copied image in the stitched image.
[0020] In another possible implementation, the aspect ratio of the third sub-image is 3:2, 4:3, 16:9, or 1:1.
[0021] It is understood that this implementation method can divide the stitched image into a third sub-image, that is, an image with a common aspect ratio such as 3:2, 4:3, 16:9 or 1:1. In this way, the aspect ratio of the image for target detection is the same as that of common images, so that various common image recognition devices can also perform the detection process in the above method, making the application scope of the method provided in this application wider.
[0022] In another possible implementation, the above method further includes: determining a first detection object; locating the first detection object in real time and displaying an image of the first detection object.
[0023] Optionally, the first detection target is the person who is speaking, or the participants in the meeting.
[0024] In another possible implementation, the original fisheye image is an image captured by an image acquisition device, which includes a fisheye lens and multiple audio acquisition components arranged in an array. The center point of the fisheye lens coincides with the center point of the multiple audio acquisition components, which are used to acquire audio information in the environment in which the image acquisition device is placed.
[0025] Optionally, multiple audio acquisition components can be arranged in a circular array.
[0026] Understandably, based on this implementation method, if the first detection target is a person who is speaking, multiple audio acquisition components can assist in locating the first detection target based on the location of the acquired audio source.
[0027] Secondly, this application provides a fisheye image unfolding device, the device comprising: an unfolding module, configured to unfold an original fisheye image based on a starting position to obtain a first unfolded image of the original fisheye image; the unfolding starting end of the original fisheye image corresponds to the first end of the first unfolded image, and the unfolding ending end of the original fisheye image corresponds to the tail end of the first unfolded image; a stitching module, configured to stitch the tail end of the first unfolded image with the first end of a copied image to obtain a stitched image of the original fisheye image, wherein the copied image is an image obtained by copying the first unfolded image; a detection module, configured to perform target detection on the stitched image to determine whether there is an image of a segmented detection object in the first unfolded image; and a processing module, configured to, if there is an image of a segmented detection object in the first unfolded image, adjust the starting position and unfold the original fisheye image based on the adjusted position to obtain a second unfolded image, wherein there is no image of a segmented detection object in the second unfolded image.
[0028] In one possible implementation, the above-mentioned stitching module is specifically used to: completely copy the first unfolded image to obtain a first copied image; and stitch the tail end of the first unfolded image with the head end of the first copied image to obtain a stitched image of the original fisheye image.
[0029] In another possible implementation, the above-mentioned stitching module is further specifically used to: partially copy the first unfolded image starting from the beginning of the first unfolded image to obtain a second copied image, wherein the width of the second copied image is greater than or equal to a first preset width threshold. The end of the first unfolded image is then stitched together with the beginning of the second copied image to obtain a stitched image of the original fisheye image.
[0030] In another possible implementation, the detection module described above is specifically used to: perform target detection on all images of the stitched image to obtain one or more target detection boxes, each target detection box including an image of a detected object; if any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the stitched image, then it is determined that there is an image of the segmented detected object in the first unfolded image; the stitching position is the stitching position between the first unfolded image and the copied image in the stitched image.
[0031] In another possible implementation, the detection module is further specifically used to: perform target detection on the first sub-image in the stitched image to obtain one or more target detection boxes, each target detection box including an image of a detected object, wherein the first sub-image includes the stitching position of the first expanded image and the copied image, and the first sub-image includes the entire image of the first expanded image and the second sub-image in the copied image, the width of the second sub-image being greater than or equal to a first preset width threshold; if any target detection box in the one or more target detection boxes has an overlapping area with the stitching position in the stitched image, then it is determined that there is an image of the segmented detected object in the first expanded image; the stitching position is the stitching position of the first expanded image and the copied image in the stitched image.
[0032] In another possible implementation, the first preset width threshold is the preset maximum width value of the target detection box, which is a width value determined based on the maximum imaging size of the detected object in the original fisheye image.
[0033] In another possible implementation, the detection module is further specifically used to: divide the stitched image into one or more third sub-images based on a preset width; wherein the width of the repeating image between any two adjacent third sub-images is greater than or equal to a second preset width threshold; perform target detection on the one or more third sub-images respectively to obtain one or more target detection boxes, each target detection box including an image of a detected object; if the area of any target detection box in the one or more target detection boxes on the third sub-image overlaps with the stitching position, then it is determined that there is an image of the segmented detected object in the first expanded image; the stitching position is the stitching position between the first expanded image and the copied image in the stitched image.
[0034] In another possible implementation, the aspect ratio of the third sub-image is 3:2, 4:3, 16:9 or 1:1;
[0035] In another possible implementation, the above processing module is further configured to: determine the first detection object; locate the first detection object in real time and display the image of the first detection object.
[0036] In another possible implementation, the original fisheye image is an image captured by an image acquisition device, which includes a fisheye lens and multiple audio acquisition components arranged in an array. The center point of the fisheye lens coincides with the center point of the multiple audio acquisition components, which are used to acquire audio information in the environment in which the image acquisition device is placed.
[0037] Thirdly, this application provides an image acquisition device, comprising: a fisheye lens for capturing raw fisheye images; a processor for executing the fisheye image unfolding method as described in the first aspect and any possible implementation thereof; and a display for displaying the fisheye image processed by the processor in real time.
[0038] In one possible implementation, the image acquisition device further includes an array of multiple audio acquisition components for acquiring audio information. A processor is also used to perform sound source localization processing on the audio information to determine the location of a first detected object, which is a person speaking. A display is used to display, in real time, an image of the first detected object in a fisheye image processed by the processor.
[0039] Fourthly, this application provides an electronic device including a memory and a processor. The memory and processor are coupled. The memory stores computer program code, which includes computer instructions. When the processor executes the computer instructions, it causes the electronic device to perform the fisheye image unfolding method as described in the first aspect and any possible implementation thereof.
[0040] Fifthly, this application provides a chip system applied to a fisheye image unfolding apparatus; the chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the memory of the fisheye image unfolding apparatus and send signals to the processors, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, it causes the electronic device to perform the fisheye image unfolding method as described in the first aspect and any possible implementation thereof.
[0041] In a sixth aspect, this application provides a computer-readable storage medium storing computer instructions that, when executed on an electronic device, cause the electronic device to perform the fisheye image unfolding method as described in the first aspect and any possible implementation thereof.
[0042] In a seventh aspect, this application provides a computer program product comprising computer instructions that, when executed on an electronic device, cause the electronic device to perform the fisheye image unfolding method as described in the first aspect and any possible implementation thereof.
[0043] For a detailed description of aspects two through seven and their various implementations in this application, please refer to the detailed description in aspect one and its various implementations; and for a detailed description of the beneficial effects of aspects two through seven and their various implementations, please refer to the beneficial effect analysis in aspect one and its various implementations, which will not be repeated here. Attached Figure Description
[0044] Figure 1 A schematic diagram of a monitoring system provided in an embodiment of this application;
[0045] Figure 2 A schematic diagram of a monitoring scenario provided in an embodiment of this application;
[0046] Figure 3 This application provides a schematic diagram of the hardware structure of a fisheye image unfolding device.
[0047] Figure 4 A flowchart illustrating a method for unfolding a fisheye image as provided in an embodiment of this application;
[0048] Figure 5 A schematic diagram of an unfolded image provided for an embodiment of this application;
[0049] Figure 6 A schematic diagram of a stitched image provided in an embodiment of this application;
[0050] Figure 7 A schematic diagram of a target detection box in a stitched image provided in an embodiment of this application;
[0051] Figure 8 A schematic diagram illustrating a first preset width threshold provided in an embodiment of this application;
[0052] Figure 9 This is a schematic diagram of an image acquisition device provided in an embodiment of this application;
[0053] Figure 10 A flowchart illustrating a method for unfolding a fisheye image as provided in an embodiment of this application;
[0054] Figure 11 A schematic diagram of a third sub-image provided in an embodiment of this application;
[0055] Figure 12 This is a schematic diagram of the structure of a fisheye image unfolding device provided in an embodiment of this application. Detailed Implementation
[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0057] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The words "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply difference. It should be noted that in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0058] To determine a suitable starting position for fisheye image unfolding, this application provides a method for unfolding a fisheye image. The method specifically includes: firstly, unfolding the original fisheye image based on the starting position to obtain a first unfolded image of the original fisheye image, wherein the unfolding start end of the original fisheye image corresponds to the beginning end of the first unfolded image, and the unfolding end end of the original fisheye image corresponds to the end end of the first unfolded image. Then, stitching the end end of the first unfolded image with the beginning end of a copied image to obtain a stitched image of the original fisheye image, where the copied image is a copy of the first unfolded image. Next, performing object detection on the stitched image to determine whether there is an image of a segmented detection object in the first unfolded image. If there is an image of a segmented detection object in the first unfolded image, adjusting the starting position and unfolding the original fisheye image based on the adjusted position to obtain a second unfolded image, wherein there is no image of a segmented detection object in the second unfolded image.
[0059] Based on this, the starting position of the image unfolding operation can be adjusted according to the target detection results, and the initial fisheye image can be unfolded based on the adjusted starting position, resulting in higher accuracy of target detection based on the unfolded image.
[0060] like Figure 1As shown in the figure, this application provides a monitoring system. The monitoring system 100 includes an image acquisition device 10, an image processing device 20, and an image display device 30. The image acquisition device 10 and the image display device 30 are respectively connected to the image processing device 20.
[0061] It should be understood that the above connection methods can be wireless, such as Bluetooth or Wi-Fi; or wired, such as fiber optic, etc., without limitation.
[0062] The image acquisition device 10 may include one or more fisheye cameras for acquiring fisheye images within the field of view of the image acquisition device 10.
[0063] Furthermore, in practical use, the image acquisition device 10 needs to be deployed according to the specific usage scenario and the user's specific needs. For example, when monitoring an office area, the image acquisition device 10 can be ceiling-mounted in the center of the ceiling of the office area. Alternatively, when monitoring doorways or windows, the image acquisition device 10 can be wall-mounted on the wall directly opposite the door or window. Or, the image acquisition device 10 can also be a desktop camera, which can be placed on a desktop.
[0064] For example, the monitoring system 100 can be deployed in a conference setting, such as... Figure 2 As shown, the image acquisition device 10 can be placed in the center of the conference table 21 so that the image acquisition device 10 can capture 360° images of all participants located around the conference table.
[0065] The image processing device 20 can be used to receive and process image information acquired by the image acquisition device 10. Furthermore, the image processing device 20 can also be used to transmit the processed image information to the image display device 30, so that a user can view the image information acquired by the image acquisition device 10 through the image display device 30.
[0066] Optionally, the image processing device 20 can be the management server of the monitoring system. It can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data servers.
[0067] The image display device 30 can be used to display received image information captured in the monitored area to the user. For example, in a conference scenario, the terminal device 20 can be placed in the same conference room as the image acquisition device 10. The terminal device 20 can display images of the participants around the conference table, or it can display close-up images of the person speaking. Optionally, the monitoring system 100 may include multiple image display devices 30.
[0068] The image display device 30 can be a user device with video or image playback capabilities. For example, the image display device 30 can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, etc., with a display or other similar structure.
[0069] In some embodiments, the image processing device 20 may be integrated into the image acquisition device 10, or the image processing device 20 may be set up independently from the image acquisition device 10.
[0070] In some embodiments, the image processing device 20 and the image display device 30 are integrated into one device, or the image processing device 20 and the image display device 30 are two independent devices. In some embodiments, the image processing device 20, the image acquisition device 10, and the image display device 30 can be integrated into one device. Optionally, embodiments of this application also provide an image acquisition device that can be applied in a conference scenario. The image acquisition device includes: a fisheye lens for capturing a raw fisheye image of the conference scenario where the conference system is located; a processor for the fisheye image unfolding method; and a display for displaying the fisheye image processed by the processor in real time.
[0071] Optionally, the image acquisition device also includes multiple audio acquisition components arranged in an array, which are used to acquire audio information in the conference scene. The processor is also used to perform sound source localization processing on the audio information to determine the location of a first detected object, which is the person speaking. The display is also used to display the image of the first detected object in a fisheye image processed by the processor in real time.
[0072] This application embodiment also provides a fisheye image unfolding device (hereinafter referred to as the unfolding device for ease of description), which is the executing entity of the above-described fisheye image unfolding method. The unfolding device can be an electronic device with data processing capabilities, or a functional module within that electronic device; there is no limitation on this. For example, the electronic device can be the image processing device 20 in the above-described monitoring system 100.
[0073] The following example uses the unfolding device as an electronic device, combined with... Figure 3 One hardware structure of the deployment device 200 is described.
[0074] like Figure 3 As shown, the deployment device 200 includes a processor 210, a communication line 220, and a communication interface 230.
[0075] Optionally, the deployment device 200 may also include a memory 240. The processor 210, memory 240, and communication interface 230 can be connected via a communication line 220.
[0076] The processor 210 can be a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 210 can also be any other device with processing capabilities, such as a circuit, device, or software module, without limitation.
[0077] In one example, processor 210 may include one or more CPUs, for example Figure 3 CPU0 and CPU1 in the CPU.
[0078] As an optional implementation, the deployment device 200 may include multiple processors, for example, in addition to processor 210, it may also include processor 270. A communication line 220 is used to transmit information between the components included in the deployment device 200.
[0079] Communication interface 230 is used for communicating with other devices or other communication networks. These other communication networks can be Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc. Communication interface 230 can be a module, circuit, transceiver, or any device capable of enabling communication.
[0080] The memory 240 is used to store instructions. These instructions can be computer programs.
[0081] The memory 240 can be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions; it can also be a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions; it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, etc., without limitation.
[0082] It should be noted that the memory 240 can exist independently of the processor 210, or it can be integrated with the processor 210. The memory 240 can be used to store instructions, program code, or some data, etc. The memory 240 can be located inside or outside the deployment device 200, without restriction.
[0083] The processor 210 is configured to execute instructions stored in the memory 240 to implement the communication method provided in the following embodiments of this application. For example, when the unfolding device 200 is a terminal or a chip or system-on-a-chip in the terminal, the processor 210 can execute instructions stored in the memory 240 to implement the fisheye image unfolding method provided in this application.
[0084] As an optional implementation, the unfolding device 200 also includes an output device 250 and an input device 260. The output device 250 can be a display screen, speaker, or other device capable of outputting data from the unfolding device 200 to the user. The input device 260 can be a keyboard, mouse, microphone, joystick, or other device capable of inputting data into the unfolding device 200.
[0085] It should be pointed out that, Figure 3 The structure shown does not constitute a limitation on the computing device, except Figure 3 In addition to the components shown, the computing device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0086] The embodiments provided in this application will now be described in detail with reference to the accompanying drawings.
[0087] like Figure 4 As shown in the embodiment of this application, a method for unfolding a fisheye image is provided. Optionally, this method is... Figure 3 The deployment device 200 shown is executed, and the method includes the following steps:
[0088] S101, The unfolding device unfolds the original fisheye image based on the starting position to obtain the first unfolded image of the original fisheye image.
[0089] Here, the raw fisheye image refers to, for example, the unprocessed fisheye image information acquired by the image acquisition device 10 in the monitoring system 100 described above. For example, using... Figure 2 Taking the image acquisition device 10 placed on the conference table as an example, Figure 5 (a) shows a raw fisheye image that the image acquisition device 10 may acquire, wherein the raw fisheye image is a circular image and includes personnel 1, 2, 3 and 4 near the conference table.
[0090] Furthermore, the aforementioned starting position refers to the default position at which the unfolding device begins the unfolding process when unfolding the original fisheye image. Optionally, the starting position can refer to a dividing line in the original fisheye image, for example... Figure 5 The dividing line 51 in (a) of the text.
[0091] The first expanded image is obtained by expanding the original fisheye image based on a default starting position. Furthermore, the starting point of the expanded original fisheye image corresponds to the beginning of the first expanded image, and the ending point of the expanded original fisheye image corresponds to the end of the first expanded image.
[0092] For example, with Figure 5 Taking the original fisheye image shown in (a) as an example, the unfolding device can use the dividing line 51 as the starting position for unfolding and unfold the circular original fisheye image in a clockwise direction, such as... Figure 5 As shown in (b) above, the first unfolded image of the rectangle is obtained. Wherein, as... Figure 5 As shown in (b), the two sides 52 and 53 of the first unfolded image are... Figure 5 This corresponds to the dividing line 51 in (a). Furthermore, the end where edge 52 of the first unfolded image is located is the beginning of the first unfolded image, and the end where edge 53 of the first unfolded image is located is the end of the first unfolded image.
[0093] It should be noted that due to the shooting characteristics of fisheye cameras, the fisheye images captured by fisheye cameras will produce a certain degree of distortion. Therefore, the unfolding device can first unfold the original fisheye image to reduce the impact of image distortion on the accuracy of subsequent image recognition.
[0094] However, if the expanded fisheye image is used for image recognition or target detection, to improve the accuracy of target detection, it is necessary to ensure the integrity of the image of the target being detected in the expanded fisheye image. Therefore, the expanding device needs to determine a suitable starting position for expanding. Thus, the expanding device can first expand the original fisheye image based on a default starting position, and then determine whether the starting position is reasonable based on the obtained first expanded image.
[0095] S102, The unfolding device splices the tail end of the first unfolded image with the head end of the copied image to obtain a spliced image of the original fisheye image.
[0096] The aforementioned copied image is an image obtained by copying the first unfolded image. Optionally, the copied image can be an image obtained by completely copying the first unfolded image. Alternatively, the copied image can also be an image obtained by copying a portion of the first unfolded image.
[0097] In one possible implementation, the unfolding device can completely replicate the first unfolded image to obtain a first replicated image.
[0098] For example, with Figure 5 Taking the first unfolded image in (b) as an example, the unfolding device can completely replicate the first unfolded image, such as... Figure 6 As shown in (a) of the diagram, a first copied image 61 is obtained. Wherein, the first copied image 61 and... Figure 5 The first unfolded image in (b) is exactly the same.
[0099] Furthermore, the unfolding device can stitch the tail end of the first unfolded image with the head end of the first copied image to obtain a stitched image of the original fisheye image.
[0100] For example, such as Figure 6 As shown in (a), the first end of the first copied image 61 can point to one end of the side 62 of the rectangular image. Therefore, the last end of the first copied image 61 can point to one end of the side 62 of the rectangular image. Furthermore, the image at the first end of the first copied image 61 is the same as the image at the first end of the first unfolded image. The image at the last end of the first copied image 61 is the same as the image at the last end of the first unfolded image.
[0101] like Figure 6 As shown in (a), in this stitched image, the tail end of the first unfolded image is stitched with the head of the first copied image 61, that is, the edge 53 of the first unfolded image coincides with the edge 62 of the first copied image, so as to stitch the first unfolded image and the first copied image together. The edge 53 of the first unfolded image and the edge 62 of the first copied image both correspond to the starting position in the original unfolded image, i.e., the dividing line 51.
[0102] In another possible implementation, the unfolded image can start from the beginning of the first unfolded image, partially copy the first unfolded image, and obtain the second copied image.
[0103] The first image of the second copied image is identical to the first image of the expanded image. Furthermore, the width of the second copied image is greater than or equal to a first preset width threshold. The width value of the rectangular image refers to the length of its horizontal side.
[0104] Furthermore, the aforementioned first preset width threshold is the minimum width value of the second copied image.
[0105] For example, with Figure 5 Taking the first unfolded image in (b) as an example, the unfolding device can copy a portion of the image in the first unfolded image, such as... Figure 6 As shown in (b), a second copied image 64 is obtained. The image at the beginning of the second copied image 64 is... Figure 5 The first image in (b) is identical to the image at the beginning. Furthermore, the width of the second copied image 64 is greater than or equal to the aforementioned preset threshold.
[0106] Furthermore, the unfolding device stitches the tail end of the first unfolded image with the head end of the second copied image to obtain a stitched image of the original fisheye image.
[0107] For example, such as Figure 6 As shown in (b), the beginning of the second copied image 64 can refer to one end of the side 65 of the rectangular image. Furthermore, the image at the beginning of the second copied image 64 is similar to... Figure 5 The first unfolded image shown in (b) is identical to the image at the beginning of the first unfolded image. The image at the end of the second copied image 64 is different from the image at the end of the first unfolded image.
[0108] Correspondingly, such as Figure 6 As shown in (b) of the image, in this stitched image, the tail end of the first unfolded image is stitched with the head end of the second copied image, that is, the edge 53 of the first unfolded image is aligned with the edge 65 of the second copied image, so as to stitch the first unfolded image and the first copied image together.
[0109] Alternatively, the unfolded image can start from the tail end of the first unfolded image and partially copy it to obtain a second copied image. Thus, the tail end of the second copied image is the same as the tail end of the first unfolded image. Furthermore, the width of the second copied image is greater than or equal to a first preset width threshold. The width value of the rectangular image refers to the length of its horizontal side.
[0110] Furthermore, the unfolding device can stitch the first end of the unfolded image with the last end of the second copied image to obtain a stitched image of the original fisheye image.
[0111] It should be noted that in this implementation, by copying a second image with a smaller area for stitching, the resource consumption of the image copying process can be reduced, thus reducing resource waste.
[0112] S103. The unfolding device performs target detection on the stitched image to determine whether there is an image of the segmented detection object in the first unfolded image.
[0113] The detection targets can be human bodies, faces, animals, vehicles, or other possible targets. Therefore, the unfolding device can perform target detection on the stitched image based on preset detection targets, such as human bodies.
[0114] Optionally, the unfolding device can perform target detection on all images of the stitched image, or it can perform target detection on a portion of the stitched image.
[0115] In one possible implementation, the unfolding device performs target detection on all images of the stitched image.
[0116] Optionally, the unfolding device can detect one or more objects in the stitched image and determine the corresponding target detection box for each object on the stitched image. A target detection box on the stitched image includes an image of one object.
[0117] For example, with Figure 7 Taking the stitched image shown in (a) as an example, if the object to be detected is a human body, the unfolding device can determine target detection boxes 71, 72, 73, 74, 75, 76, 77, and 78 in the stitched image. Each of these target detection boxes includes images corresponding to person 1, person 2, person 3, and person 4, respectively.
[0118] Therefore, if any of the target detection boxes obtained above coincides with the splicing position of the first unfolded image and the copied image, the unfolding device can determine that there is an image of the segmented detection object in the first unfolded image.
[0119] Optionally, the unfolding device may determine whether a target detection box coincides with the splicing position of the first unfolded image and the copied image based on the following formula (1).
[0120] x1 < w2 < (x1 + w1) (1)
[0121] Where x1 is the x-coordinate of the top-left corner of the target detection box, w1 is the side length of the target detection box along the horizontal axis, and w2 is the side length of the first expanded image along the horizontal axis. Optionally, if the copied image in the stitched image is the same as the first expanded image, then w2 is... W is the side length of the stitched image along the horizontal axis.
[0122] like Figure 7 As shown in (a), the unfolding device can establish a rectangular coordinate system XOY with the lower left corner of the stitched image as the origin. The X-axis direction is the aforementioned horizontal axis direction, and the Y-axis direction is the vertical axis direction. Furthermore, the coordinates of a target detection box can be represented as (x1, y1, w1, h1). Here, x1 is the horizontal coordinate of the upper left corner of the target detection box in this coordinate system, y1 is the vertical coordinate of the upper left corner of the target detection box in this coordinate system, w1 is the side length of the target detection box along the horizontal axis, and h1 is the side length of the target detection box along the vertical axis.
[0123] It should be understood that if the coordinates of a target detection box satisfy Formula 1 above, it means that the stitching position of the first expanded image and the copied image coincides with the target detection box, that is, the image of the detected object corresponding to the target detection box is segmented. For example... Figure 7 As shown in (b), in this stitched image, the target image frame 79 of person 4 partially overlaps with the stitching position of the stitched image. Furthermore, in this first unfolded image, the image of person 4 is segmented.
[0124] Specifically, the unfolding device can sequentially determine the coordinates of each target detection box, and based on the above formula (1), determine whether each target detection box coincides with the stitching position of the stitched image. Furthermore, if there is a target detection box that coincides with the stitching position, the unfolding device can determine that there is an image of the segmented detection object in the first unfolded image.
[0125] It should be noted that if the first unfolded image contains images of the detected object that have been segmented, that is, if the first unfolded image does not show the complete image of the detected object, the detection accuracy of the detected object will be affected during the target detection process of the first unfolded image.
[0126] In another possible implementation, the unfolding device can perform target detection on the first sub-image of the stitched image.
[0127] The first sub-image may include the splicing position of the first expanded image and the copied image, and the first sub-image includes a second sub-image in the first expanded image and the copied image, wherein the width of the second sub-image is greater than or equal to the first preset width threshold mentioned above.
[0128] Optionally, the first preset width threshold can be the preset maximum width value of the target detection box.
[0129] Optionally, the unfolding device can determine the first preset width threshold based on the type of the target being detected and the image size required for target detection in the current monitoring scene.
[0130] like Figure 8 As shown, the detected object at a distance d1 from the image acquisition device corresponds to image frame 81 in the expanded image, the detected object at a distance d2 from the image acquisition device corresponds to image frame 82 in the expanded image, and the detected object at a distance d3 from the image acquisition device corresponds to image frame 83 in the expanded image. At this time, if the image size of the target that the expanding device can detect is as shown in image frame 82, the width of image frame 82 in the expanded image can be determined to be the aforementioned first preset width threshold.
[0131] The target detection process of the unfolding device on the first sub-image can refer to the target detection process of the above implementation method, and will not be repeated here.
[0132] In some embodiments, if the unfolding device determines that all target detection boxes do not overlap with the splicing position of the first unfolded image and the copied image, the unfolding device may determine that there is no image of the segmented detection object in the first unfolded image.
[0133] S104. If the first unfolded image contains an image of the segmented detection object, the unfolding device adjusts its starting position and unfolds the original fisheye image based on the adjusted position to obtain the second unfolded image.
[0134] In the second expanded image, there is no image of the detected object that has been segmented.
[0135] If the unfolding device determines that there is no segmented detection object in the first unfolded image, then the unfolding device adjusts the starting position to a position that does not coincide with any of the target boxes in the first unfolded image, based on the determined target detection boxes.
[0136] In some embodiments, the unfolding device can also identify a first detection object among a plurality of detection objects, locate the first detection object in real time, and display an image of the first detection object.
[0137] Optionally, the first detection target mentioned above includes people who are speaking, participants in a meeting, etc.
[0138] Optionally, the aforementioned original fisheye image is an image acquired by an image acquisition device, which includes a fisheye lens and multiple audio acquisition components arranged in an array, such as... Figure 9As shown, the center point of the fisheye lens coincides with the center point of the arrangement of multiple audio acquisition components, which are used to acquire audio information from the environment in which the image acquisition device is placed. Optionally, the multiple audio acquisition components can be arranged in a circular array.
[0139] For example, taking an image acquisition device deployed in a meeting setting as an example, the first detection object can be a person who is speaking. Thus, multiple audio acquisition components can also be used to determine the position of the first detection object and assist in locating the first detection object.
[0140] Optionally, when multiple audio acquisition components determine the position of the first detection object based on the sound source, the initial position of the above-mentioned image unfolding can be taken as 0°, and the x-coordinate of the first detection object in the unfolded image can be determined according to the following formula 2.
[0141] x=w×α / 360 (2)
[0142] Where w is the side length of the unfolded image along the horizontal axis, and α is the sound source localization angle.
[0143] The fisheye image unfolding method provided in this application has at least the following beneficial effects: First, the method pre-unfolds the original fisheye image based on a preset initial position to obtain a preliminary unfolded result. Then, based on the circular characteristics of a 360° circular fisheye image during unfolding, the images at the beginning and end of the first unfolded image can be stitched together to obtain the image at the starting position in the original fisheye image. Therefore, a stitched image of the preliminary unfolded image can be obtained through a copying and stitching process. This stitched image is then subjected to detection processing to determine whether the unfolded image based on the initial position will segment the image of the detected object at both ends of the unfolded image, i.e., whether there is a segmented image of the detected object in the first unfolded image. This determination process, combining image copying and stitching, is simple to operate, easy to implement, and more practical. Furthermore, when it is determined that there is a segmented image of the detected object in the first unfolded image, the starting position of the image unfolding operation can be adjusted according to the target detection result, and the initial fisheye image can be unfolded based on the adjusted starting position. This results in a better image quality in the unfolded image, improving the user's viewing experience and increasing the accuracy of target detection on the unfolded image.
[0144] In some embodiments, based on Figure 4 The provided method, to facilitate target detection, allows the unfolding device to divide the stitched image obtained in step S102 into one or more third sub-images based on commonly used image aspect ratios, and then determine the image containing the segmented detection object in the first unfolded image. For example... Figure 10 As shown, step S103 above can be specifically implemented as follows:
[0145] S1031, The unfolding device divides the stitched image into one or more third sub-images based on a preset image size.
[0146] Among them, the width of the repeating image between any two adjacent third sub-images is greater than or equal to the second preset width threshold.
[0147] Optionally, the second preset width threshold can also be the preset maximum width value of the target detection box. The second preset width threshold can be equal to or different from the first preset width threshold.
[0148] Optionally, the aspect ratio of the third sub-image is 3:2, 4:3, 16:9, or 1:1.
[0149] S1032, the unfolding device performs target detection on one or more third sub-images respectively, and obtains one or more target detection boxes, wherein one target detection box includes an image of a detected object.
[0150] The specific process of object detection in the third sub-image can be referred to the relevant description in step S103 above, and will not be repeated here.
[0151] In some embodiments, the unfolding device can also detect whether two adjacent third sub-images have the same target detection box. If so, the unfolding device can merge the two third sub-images, retaining only one of the identical target detection boxes.
[0152] For example, such as Figure 11 As shown, the third sub-image 1 includes target detection boxes for personnel 4 and personnel 1, and the third sub-image 2 includes target detection boxes for personnel 1 and personnel 2. Both third sub-images include the target detection box for personnel 1. At this time, the unfolding device can merge these two third sub-images, retaining only the target detection box for personnel 1 in the third sub-image 1, or the unfolding device can retain only the target detection box for personnel 1 in the third sub-image 2.
[0153] S1033. If any target detection box in one or more target detection boxes coincides with the splicing position of the first unfolded image and the copied image, the unfolding device determines that there is an image of the segmented detection object in the first unfolded image.
[0154] It should be noted that this method can divide the stitched image into a third sub-image, that is, an image with a common aspect ratio such as 3:2, 4:3, 16:9 or 1:1. In this way, the aspect ratio of the image for target detection is the same as that of common images, so various common image recognition devices can also perform the detection process in the above method, making the application scope of the method provided in this application wider.
[0155] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the aforementioned functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0156] like Figure 12 The diagram shown is a structural schematic of a fisheye image unfolding device 300 provided in an embodiment of this application. The device 300 may include: an unfolding module 301, a stitching module 302, a detection module 303, and a processing module 304.
[0157] The expansion module 301 is used to expand the original fisheye image based on the starting position to obtain a first expanded image of the original fisheye image; the starting end of the expansion of the original fisheye image corresponds to the beginning end of the first expanded image, and the ending end of the expansion of the original fisheye image corresponds to the end end of the first expanded image.
[0158] The stitching module 302 is used to stitch the tail end of the first unfolded image with the head end of the copied image to obtain a stitched image of the original fisheye image. The copied image is the image obtained by copying the first unfolded image.
[0159] The detection module 303 is used to perform target detection on the stitched image to determine whether there is an image of the segmented detection object in the first unfolded image.
[0160] The processing module 304 is used to adjust the starting position and expand the original fisheye image based on the adjusted position to obtain a second expanded image if the first expanded image contains an image of the segmented detection object. The second expanded image does not contain an image of the segmented detection object.
[0161] In one possible implementation, the stitching module 302 is specifically used to: completely copy the first unfolded image to obtain a first copied image; and stitch the tail end of the first unfolded image with the head end of the first copied image to obtain a stitched image of the original fisheye image.
[0162] In another possible implementation, the stitching module 302 is further configured to: partially copy the first unfolded image, starting from the beginning of the first unfolded image, to obtain a second copied image, wherein the width of the second copied image is greater than or equal to a first preset width threshold. The end of the first unfolded image is then stitched together with the beginning of the second copied image to obtain a stitched image of the original fisheye image.
[0163] In another possible implementation, the detection module 303 is specifically used to: perform target detection on all images of the stitched image to obtain one or more target detection boxes, each target detection box including an image of a detected object; if any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the stitched image, then it is determined that there is an image of the segmented detected object in the first unfolded image; the stitching position is the stitching position between the first unfolded image and the copied image in the stitched image.
[0164] In another possible implementation, the detection module 303 is further specifically used to: perform target detection on the first sub-image in the stitched image to obtain one or more target detection boxes, each target detection box including an image of a detected object, wherein the first sub-image includes the stitching position of the first expanded image and the copied image, and the first sub-image includes the entire image of the first expanded image and the second sub-image in the copied image, the width of the second sub-image being greater than or equal to a first preset width threshold; if any target detection box in the one or more target detection boxes has an overlapping area with the stitching position in the stitched image, then it is determined that there is an image of a segmented detected object in the first expanded image; the stitching position is the stitching position of the first expanded image and the copied image in the stitched image.
[0165] In another possible implementation, the first preset width threshold is the preset maximum width value of the target detection box, which is a width value determined based on the maximum imaging size of the detected object in the original fisheye image.
[0166] In another possible implementation, the detection module 303 is further specifically used to: divide the stitched image into one or more third sub-images based on a preset width; wherein the width of the repeating image between any two adjacent third sub-images is greater than or equal to a second preset width threshold; perform target detection on the one or more third sub-images respectively to obtain one or more target detection boxes, each target detection box including an image of a detected object; if the area of any target detection box in the one or more target detection boxes on the third sub-image overlaps with the stitching position, then it is determined that there is an image of the segmented detected object in the first expanded image; the stitching position is the stitching position between the first expanded image and the copied image in the stitched image.
[0167] In another possible implementation, the aspect ratio of the third sub-image is 3:2, 4:3, 16:9 or 1:1;
[0168] In another possible implementation, the processing module 304 is further configured to: determine the first detection object; locate the first detection object in real time and display an image of the first detection object.
[0169] In another possible implementation, the original fisheye image is an image captured by an image acquisition device, which includes a fisheye lens and multiple audio acquisition components arranged in an array. The center point of the fisheye lens coincides with the center point of the multiple audio acquisition components, which are used to acquire audio information in the environment in which the image acquisition device is placed.
[0170] For a detailed description of the above-mentioned optional methods, please refer to the foregoing method embodiments, which will not be repeated here. Furthermore, the explanation of any of the fisheye image unfolding devices 300 provided above, as well as the description of their beneficial effects, can be found in the corresponding method embodiments described above, and will not be repeated here.
[0171] As an example, combined Figure 3 The functions implemented by the processing module 304 of the fisheye image unfolding device 300 can be achieved through... Figure 3 Processor 210 or processor 270 in the middle executes Figure 3 The program code implementation in memory 240 is, of course, not limited to this.
[0172] Those skilled in the art will readily recognize that, based on the units and algorithm steps described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is implemented in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0173] It should be noted that, Figure 12 The module division shown is illustrative and represents only one logical functional division; in actual implementation, other division methods are possible. For example, two or more functions can be integrated into a single processing module. These integrated modules can be implemented either in hardware or as software functional modules.
[0174] This application also provides a computer-readable storage medium including computer-executable instructions that, when run on a computer, cause the computer to perform any of the methods provided in the above embodiments. For example, Figure 4 One or more features in S101 to S104 can be performed by one or more computer-executable instructions stored in the computer-readable storage medium.
[0175] This application also provides a computer program product containing computer execution instructions, which, when run on a computer, causes the computer to perform any of the methods provided in the above embodiments.
[0176] This application also provides a chip, including a processor and an interface. The processor is coupled to a memory through the interface. When the processor executes a computer program in the memory or computer execution instructions, any of the methods provided in the above embodiments are executed.
[0177] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).
[0178] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
[0179] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for unfolding a fisheye image, characterized in that, The method includes: The original fisheye image is unfolded based on the starting position to obtain the first unfolded image of the original fisheye image; the unfolding start end of the original fisheye image corresponds to the first end of the first unfolded image, and the unfolding end end of the original fisheye image corresponds to the tail end of the first unfolded image. The tail end of the first unfolded image is stitched together with the head end of the copied image to obtain the stitched image of the original fisheye image, and the copied image is the image obtained by copying the first unfolded image; The stitched image is subjected to target detection to obtain one or more target detection boxes, and each target detection box includes an image of a detected object; If any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the stitched image, then it is determined that there is an image of a segmented detection object in the first expanded image; the stitching position is the stitching position between the first expanded image and the copied image in the stitched image. If the first expanded image contains an image of the segmented detection object, the starting position is adjusted, and the original fisheye image is expanded based on the adjusted position to obtain a second expanded image, wherein the second expanded image does not contain an image of the segmented detection object.
2. The method according to claim 1, characterized in that, The step of stitching the tail end of the first unfolded image with the head end of the copied image includes: Completely copy the first unfolded image to obtain the first copied image; The tail end of the first unfolded image is stitched together with the head end of the first copied image to obtain the stitched image of the original fisheye image.
3. The method according to claim 1, characterized in that, The step of splicing the tail end of the first unfolded image with the head end of the copied image includes: Starting from the first end of the first expanded image, the first expanded image is partially copied to obtain a second copied image, wherein the width of the second copied image is greater than or equal to a first preset width threshold. The tail end of the first unfolded image is stitched together with the head end of the second copied image to obtain the stitched image of the original fisheye image.
4. The method according to any one of claims 1 to 3, characterized in that, The step of performing target detection on the stitched image to obtain one or more target detection boxes includes: Target detection is performed on all images of the stitched image to obtain one or more target detection boxes.
5. The method according to claim 1 or 2, characterized in that, The step of performing target detection on the stitched image to obtain one or more target detection boxes includes: Target detection is performed on the first sub-image in the stitched image to obtain one or more target detection boxes; wherein, the first sub-image includes the stitching position of the first expanded image and the copied image, and the first sub-image includes the entire image of the first expanded image and the second sub-image in the copied image, and the width of the second sub-image is greater than or equal to a first preset width threshold.
6. The method according to claim 5, characterized in that, The first preset width threshold is the preset maximum width value of the target detection box, and the preset maximum width value is the width value determined based on the maximum imaging size of the detected object in the original fisheye image.
7. The method according to any one of claims 1 to 3, characterized in that, The step of performing target detection on the stitched image to obtain one or more target detection boxes includes: The stitched image is divided into one or more third sub-images based on a preset width; wherein the width of the repeating image between any two adjacent third sub-images is greater than or equal to a second preset width threshold. Target detection is performed on the one or more third sub-images respectively to obtain one or more target detection boxes; If any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the stitched image, then determining that the first unfolded image contains an image of a segmented detection object includes: If any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the third sub-image, then it is determined that there is an image of a segmented detection object in the first expanded image.
8. The method according to claim 7, characterized in that, The aspect ratio of the third sub-image is 3:2, 4:3, 16:9 or 1:
1.
9. The method according to any one of claims 1 to 3, 6 or 8, characterized in that, The method further includes: Identify the primary target for testing; The first detected object is located in real time, and an image of the first detected object is displayed.
10. The method according to any one of claims 1 to 3, 6 or 8, characterized in that, The original fisheye image is an image acquired by an image acquisition device, which includes a fisheye lens and multiple audio acquisition components arranged in an array. The center point of the fisheye lens coincides with the center point of the multiple audio acquisition components. The multiple audio acquisition components are used to acquire audio information in the environment in which the image acquisition device is placed.
11. A device for unfolding a fisheye image, characterized in that, The device includes: An unfolding module is used to unfold the original fisheye image based on a starting position to obtain a first unfolded image of the original fisheye image; the unfolding start end of the original fisheye image corresponds to the first end of the first unfolded image, and the unfolding end end of the original fisheye image corresponds to the last end of the first unfolded image. The stitching module is used to stitch the tail end of the first unfolded image with the head end of the copied image to obtain the stitched image of the original fisheye image, wherein the copied image is an image obtained by copying the first unfolded image; The detection module is used to perform target detection on the stitched image to obtain one or more target detection boxes, wherein each target detection box includes an image of a detected object; The processing module is configured to determine that the first expanded image contains the image of the segmented detection object if the area of any one of the target detection boxes in the one or more target detection boxes on the stitched image overlaps with the stitching position; the stitching position is the stitching position of the first expanded image and the copied image in the stitched image. The processing module is further configured to, if the first expanded image contains an image of the segmented detection object, adjust the starting position and expand the original fisheye image based on the adjusted position to obtain a second expanded image, wherein the second expanded image does not contain an image of the segmented detection object.
12. The apparatus according to claim 11, characterized in that, The splicing module is specifically used for: Completely copy the first unfolded image to obtain the first copied image; The tail end of the first unfolded image is stitched together with the head end of the first copied image to obtain the stitched image of the original fisheye image; The splicing module is also specifically used for: Starting from the first end of the first expanded image, the first expanded image is partially copied to obtain a second copied image, wherein the width of the second copied image is greater than or equal to a first preset width threshold. The tail end of the first unfolded image is stitched together with the head end of the second copied image to obtain the stitched image of the original fisheye image; The detection module is specifically used for: Target detection is performed on all images of the stitched image to obtain one or more target detection boxes, and each target detection box includes an image of a detected object. If any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the stitched image, then it is determined that there is an image of a segmented detection object in the first expanded image; the stitching position is the stitching position between the first expanded image and the copied image in the stitched image. The detection module is also specifically used for: Target detection is performed on the first sub-image in the stitched image to obtain one or more target detection boxes. Each target detection box includes an image of a detected object. The first sub-image includes the stitching position of the first expanded image and the copied image, and the first sub-image includes the entire image of the first expanded image and the second sub-image in the copied image. The width of the second sub-image is greater than or equal to a first preset width threshold. If any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the stitched image, then it is determined that there is an image of a segmented detection object in the first expanded image; the stitching position is the stitching position between the first expanded image and the copied image in the stitched image. The first preset width threshold is the preset maximum width value of the target detection box, and the preset maximum width value is the width value determined based on the maximum imaging size of the detected object in the original fisheye image; The detection module is also specifically used for: The stitched image is divided into one or more third sub-images based on a preset width; wherein the width of the repeating image between any two adjacent third sub-images is greater than or equal to a second preset width threshold. Target detection is performed on the one or more third sub-images respectively to obtain one or more target detection boxes, and each target detection box includes an image of a detected object; If any of the target detection boxes in the one or more target detection boxes overlaps with the stitching position in the third sub-image, then it is determined that there is an image of a segmented detection object in the first expanded image; the stitching position is the stitching position between the first expanded image and the copied image in the stitched image; The aspect ratio of the third sub-image is 3:2, 4:3, 16:9 or 1:1; The processing module is further configured to: Identify the primary target for testing; The first detected object is located in real time, and an image of the first detected object is displayed. The original fisheye image is an image acquired by an image acquisition device, which includes a fisheye lens and multiple audio acquisition components arranged in an array. The center point of the fisheye lens coincides with the center point of the multiple audio acquisition components. The multiple audio acquisition components are used to acquire audio information in the environment in which the image acquisition device is placed.
13. An image acquisition device, characterized in that, The image acquisition device includes: A fisheye lens, used to capture raw fisheye images; A processor for executing the fisheye image unfolding method according to any one of claims 1 to 10; A display for real-time display of fisheye images processed by the processor.
14. The image acquisition device according to claim 13, characterized in that, The image acquisition device also includes multiple audio acquisition components arranged in an array, which are used to acquire audio information; The processor is further configured to perform sound source localization processing on the audio information to determine the location of the first detection object, wherein the first detection object is the person who is speaking; The display is also used to display, in real time, the image of the first detected object in the fisheye image processed by the processor.
15. An electronic device, characterized in that, The electronic device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code including computer instructions; When the processor executes the computer instructions, the electronic device performs the fisheye image unfolding method as described in any one of claims 1-10.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on an electronic device, cause the electronic device to perform the fisheye image unfolding method as described in any one of claims 1-10.
Citation Information
Patent Citations
Target detection method, target detection device and computer readable storage medium
CN114419428A
Interactive conference device based on vision and sound fusion
CN213213667U
Minimizing dead zones in panoramic images
US20050151837A1
Video conference panoramic image spreading method
US20210176444A1