Active focusing method and panoramic image capturing device

TWI935470BActive Publication Date: 2026-08-11CHICONY ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
TW113133968
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2026-08-11
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

Conventional panoramic image capturing devices using multiple lenses face issues with image stitching due to varying focal lengths, as they typically use fixed-focus lenses, and incorporating active focusing lenses results in un-stitchable images.

Method used

An active focusing method that utilizes a microphone array to locate sound sources, selects a reference lens module based on autofocus values, adjusts all lens modules to a uniform focal length, and stitches images with synchronized focus for a clear panoramic image.

Benefits of technology

Ensures that images from multiple lenses with different focal lengths can be stitched together effectively, preventing distortion or blurring in the final panoramic image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001905405_001
    Figure TWG2TB001905405_001
  • Figure TWG2TB001905405_002
    Figure TWG2TB001905405_002
  • Figure TWG2TB001905405_003
    Figure TWG2TB001905405_003
Patent Text Reader

Abstract

An active focusing method is applied to a panoramic image capturing device. The active focusing method includes the following steps: receiving multiple sound signals in a space from a microphone array; obtaining the location of at least one sound source based on the sound signals; acquiring at least one image image close to the sound source location; analyzing and comparing the sound signals and the image images to select one of the multiple lens modules as a reference lens module, the field of view of which covers a main speaker; obtaining an autofocus value of the reference lens module for the main speaker; adjusting the lens modules according to the autofocus value so that the lens modules have the same focal length; and the processor acquiring an adjusted image image from each of the lens modules and stitching them together to form a panoramic image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to an active focusing method, and more particularly to an active focusing method applied in a panoramic image capturing device. [Previous Technology]

[0002] With the advancement of optical and image processing technologies, image capturing devices with multiple lenses have emerged. For example, panoramic image capturing devices (or 360-degree panoramic cameras) use two to four lenses arranged in a ring to capture multiple images with different field of view ranges, and then use image processing technology to stitch the multiple images into a panoramic image. In addition, some video conferencing systems also use panoramic image capturing devices to obtain panoramic images of the meeting space.

[0003] However, conventional panoramic image capturing devices can only use multiple fixed-focus lenses, so the focal length cannot be adjusted, which sometimes results in out-of-focus images. If the panoramic image capturing device uses multiple active focusing lenses, the images cannot be stitched together because each active focusing lens has a different focal length, so there is a need for improvement. [Summary of the Invention]

[0004] In view of the above-mentioned problems, the main objective of the present invention is to provide an active focusing method and a panoramic image capturing device using the active focusing method, which selects a reference lens module by analyzing and comparing the sound signal and the sound signal of the image screen, and adjusts the focal length of the lens module according to an autofocus value of the reference lens module, so as to solve the problem that conventional panoramic image capturing devices are difficult to use active focusing lenses.

[0005] To achieve the above objective, the present invention provides an active focusing method applied to a panoramic image capturing device. The panoramic image capturing device includes a plurality of lens modules, a microphone array, and a processor. The active focusing method includes the following steps: the processor receives a plurality of sound signals in a space from the microphone array; the processor obtains the location of at least one sound source based on the sound signals; the processor obtains at least one image frame of the lens modules near the location of the at least one sound source, which includes at least one speaker; the processor analyzes and compares the sound signals with the at least one image frame to select one of the lens modules as a reference lens module, the field of view of the reference lens module covering a main speaker; the processor obtains an autofocus value of the reference lens module for the main speaker; the processor adjusts the lens modules according to the autofocus value so that the lens modules have the same focal length; and the processor obtains an adjusted image frame from each of the lens modules and stitches them together to form a panoramic image frame.

[0006] According to one embodiment of the present invention, the microphone array includes a plurality of microphones, which are spaced apart in the panoramic image capturing device. The active focusing method further includes the following steps: the processor divides the space into a plurality of spatial directions according to the field of view corresponding to the lens modules; and the microphones respectively pick up sound from the spatial directions to obtain the sound signals.

[0007] According to one embodiment of the present invention, each sound signal includes directional information. When the average volume of each of the sound signals is greater than a threshold value, the processor obtains the directional information corresponding to each of the sound signals and marks it as the directional information of the at least one sound source.

[0008] According to one embodiment of the present invention, the active focusing method further includes the following steps: the processor analyzes and compares the sound signals with the at least one image frame to calculate an integral value of the at least one speaker; and the processor selects the main speaker based on a highest integral value.

[0009] According to one embodiment of the present invention, when the average volume of the sound signals is greater than a threshold value, the processor adds a first weight score to obtain an integral value.

[0010] According to one embodiment of the present invention, the processor performs image recognition on a facial movement in the at least one image frame. When the facial movement matches the sound signal, the processor adds a second weight score to obtain an integral value.

[0011] According to one embodiment of the present invention, the processor includes a memory unit that stores a focus parameter table. When the reference lens module executes an autofocus program, the processor obtains the autofocus value based on the focus parameter table.

[0012] According to one embodiment of the present invention, the microphone array includes a plurality of microphones. These microphones are at least partially located between two adjacent lens modules.

[0013] To achieve the above objectives, the present invention further provides a panoramic image capturing device, which includes a plurality of lens modules, a microphone array, and a processor. The processor is electrically connected to the lens modules and the microphone array. The processor performs the following steps according to an active focusing method: receiving a plurality of sound signals in a space from the microphone array; obtaining the orientation of at least one sound source based on the sound signals; obtaining at least one image frame of the lens modules near the orientation of the at least one sound source, which includes at least one speaker; analyzing and comparing the sound signals and the at least one image frame to select one of the lens modules as a reference lens module, the field of view of the reference lens module covering a main speaker; obtaining an autofocus value of the reference lens module for the main speaker; adjusting the lens modules according to the autofocus value so that the lens modules have the same focal length; and obtaining an adjusted image frame from each of the lens modules and stitching them together to form a panoramic image frame.

[0014] As described above, according to the active focusing method of the present invention and the panoramic image capturing device thereof, by analyzing and comparing the sound signal and the image, one of the multiple lens modules of the panoramic image capturing device is selected as a reference lens module, and the field of view corresponding to the reference lens module covers a main speaker. An autofocus value of the reference lens module for the main speaker is obtained, and all lens modules are adjusted according to the autofocus value so that all lens modules have the same focal length. Finally, the lens modules respectively acquire adjusted image images with the same focal length and stitch them together to form a panoramic image. Since the adjusted image images have the same focal length, the situation where images cannot be stitched due to different focal lengths can be avoided, and the resulting panoramic image will not be distorted or blurred due to stitching. Therefore, the panoramic image capturing device using the active focusing method of the present invention can use an active focusing lens, that is, the aforementioned lens module can be an active focusing lens module.

Implementation Method

[0016] In order to better understand the technical content of the present invention, preferred embodiments are described below.

[0017] Figure 1 is a top view of a panoramic image capturing device according to an embodiment of the present invention. Figures 2A and 2B are schematic diagrams of the usage environment of the panoramic image capturing device shown in Figure 1. Figure 3 is a block diagram of the panoramic image capturing device shown in Figure 1. Figures 4A and 4B are flowcharts of an active focusing method according to an embodiment of the present invention. Please refer to Figures 1, 2A, 3, 4A, and 4B. First, the panoramic image capturing device 1 of this embodiment is applied in a video conferencing system to acquire a panoramic image of a space S (e.g., a conference room space), as shown in Figures 2A and 2B. Also, Figures 2A and 2B are schematic diagrams showing the same panoramic image capturing device 1 applied in different sound (speaking) environments. In this embodiment, the panoramic image capturing device 1 includes a plurality of lens modules 11-14, a microphone array 20, and a processor 30. The panoramic image capturing device 1 of this embodiment has four lens modules 11-14, which are respectively oriented towards different directions within space S to capture images of different fields of view. As shown in Figures 1, 2A, and 2B, lens module 11 is located on the upper side, lens module 12 on the right side, lens module 13 on the lower side, and lens module 14 on the left side. Among them, lens modules 11-14 are active focusing lens modules, which have an autofocus (AF) module, enabling lens modules 11-14 to perform autofocus procedures.

[0018] The microphone array 20 includes a plurality of microphones. In this embodiment, the microphone array 20 has eight microphones 21-28, which are spaced apart on the panoramic image capturing device 1. Preferably, the microphone array 20 in this embodiment can be an omnidirectional microphone array, with microphones 21-28 arranged at equal intervals on the panoramic image capturing device 1, as shown in Figures 1, 2A, and 2B. Preferably, the number of microphones 21-28 in this embodiment is greater than the number of lens modules 11-14, with a ratio of 2:1. Preferably, microphones 21-28 are at least partially located between two adjacent lens modules 11-14. For example, microphones 21, 23, 25, and 27 correspond to lens modules 11, 12, 13, and 14, respectively, while microphones 22, 24, 26, and 28 are located between lens modules 11, 12, 13, and 14, respectively. Specifically, microphone 22 is located between two adjacent lens modules 11 and 12; microphone 24 is located between two adjacent lens modules 12 and 13; microphone 26 is located between two adjacent lens modules 13 and 14; and microphone 28 is located between lens modules 11 and 14.

[0019] As shown in FIG3, the processor 30 is electrically connected to the lens modules 11-14 and the microphone array 20 to receive and transmit related parameters, image signals, sound signals, or commands. The processor 30 may be, for example, but not limited to, a central processing unit (CPU) or a microprocessor. In this embodiment, the processor 30 also includes a memory unit 31, which stores the active focusing method shown in FIG4A and FIG4B, and executes the active focusing method to achieve the function of the present invention. Furthermore, this embodiment only illustrates a preferred embodiment of the present invention; to avoid redundancy, not all possible variations and combinations are described in detail. However, those skilled in the art should understand that not all of the above-mentioned modules or components are necessarily necessary. Other more detailed conventional modules or components may also be included to implement the present invention. Each module or component may be omitted or modified as needed, and there may be other modules or components between any two modules. The operation of each module and component is further explained below according to the steps of the active focusing method. It should be noted that the following step numbers are arranged sequentially for ease of explanation only, and are not intended to limit the order of the steps. Multiple steps can be interchanged or performed simultaneously, and this invention is not limited thereto.

[0020] Step S01: Divide a space S into multiple directional spaces S1, S2, S3, and S4 according to the field of view corresponding to the lens modules 11 to 14.

[0021] As described above, the four lens modules 11, 12, 13, and 14 in this embodiment face different directions within space S to capture images with different field of view ranges. In this embodiment, the processor 30 may preset four directional spaces S1, S2, S3, and S4. Among them, the directional spaces S1, S2, S3, and S4 are the field of view ranges corresponding to the lens modules 11, 12, 13, and 14, as shown in Figures 2A and 2B.

[0022] Furthermore, the parameters or image signals transmitted by the lens modules 11-14 each have different device information, enabling the processor 30 to determine which orientation space S1, S2, S3, or S4 the image is from. For example, when the processor 30 receives the image signal from the lens module 11, it can determine from the device information that it is an image from orientation space S1.

[0023] It should also be noted that the orientation spaces S1, S2, S3, and S4 shown in Figures 2A and 2B are examples for the convenience of describing the active focusing method of the present invention. The processor 30 can directly know its relative relationship with the microphones 21 to 28 based on the parameters transmitted by the lens modules 11 to 14 or the device information carried by the image signals, so as to perform subsequent analysis and comparison.

[0024] Step S02: Receive multiple sound signals in space S from the microphone array 20.

[0025] In this embodiment, the microphone array 20 has eight microphones 21 to 28, each of which can pick up sound signals from space S. Therefore, the microphone array 20 can acquire up to eight sets of sound signals at the same time. Preferably, the microphones 21 to 28 are spaced at the same distance, so that the eight microphones 21 to 28 can pick up sound from the directional spaces S1, S2, S3, and S4 on an average basis, thereby acquiring multiple sound signals. After receiving multiple sound signals from space S from the microphone array 20, the processor 30 can determine the location of the sound source by the average volume (e.g., amplitude) of the sound signals from different microphones 21 to 28, for example, in step S03.

[0026] Step S03: Obtain the location of at least one sound source based on these sound signals.

[0027] In this embodiment, the processor 30 analyzes the average volume of each sound signal from different microphones 21-28 to calculate the location of the sound source, referred to as the sound source location in this embodiment. Specifically, the memory unit 31 of the processor 30 also stores the relative positions of the four lens modules 11-14 and the eight microphones 21-28. Similarly, the sound signals transmitted by the eight microphones 21-28 each have different device information, referred to as location information in this embodiment. In other words, each sound signal in this embodiment includes location information. Therefore, the processor 30 can determine which directional space S1, S2, S3, S4 the microphones 21-28 transmitting the sound signal are close to by using the location information contained in the sound signal.

[0028] Furthermore, the processor 30 obtains the location of at least one sound source based on a plurality of sound signals. Specifically, there may be multiple sound sources in space S, allowing the processor 30 to obtain multiple sound source locations. In this embodiment, the processor 30 can obtain the sound source location based on an average volume (i.e., amplitude) of each sound signal. Taking the sound (speaking) environment shown in Figure 2A as an example, the main speaker P1 (i.e., the speaker in the meeting) shown in Figure 2A is located in directional space S1, while there are two other speakers P2 in directional space S2 who are chatting quietly with lower volume. Microphones 21, 22, 23, 24, and 25 near the main speaker P1 and the quiet speakers P2 can all receive sound and generate sound signals. Because microphone 22 is closer to the main speaker P1 and microphone 23 is closer to the soft-speaking speaker P2, processor 30 determines, based on the average volume (i.e., amplitude) and location information of each sound signal, that the average volume of the sound signals acquired by microphones 22 and 23 is larger, while the average volume of the sound signals acquired by microphones 21 and 24 is smaller. Therefore, it can be determined that one sound source is located in azimuth space S1 and is closer to microphone 22, and the other sound source is located in azimuth space S2 and is closer to microphone 23.

[0029] Taking the sound (speaking) environment shown in Figure 2B as an example, the main speaker P1 (i.e., the person speaking in the meeting) and two quiet speakers P2 are both located in the directional space S2. Microphones 22, 23, and 24, which are close to the main speaker P1 and the quiet speakers P2, can receive sound and generate sound signals. Since microphone 23 is closest to the main speaker P1 and the quiet speakers P2, the processor 30 determines, based on the average volume (i.e., amplitude) of each sound signal and the directional information, that the average volume of the sound signal obtained by microphone 23 is larger, while the average volume of the sound signals obtained by microphones 22 and 24 is smaller. Therefore, it can be determined that the sound source is located in the directional space S2 and is close to microphone 23.

[0030] It should be noted that if the sound source is located between two adjacent microphones (e.g., microphones 21 and 22), the sound source location can be determined based on the volume and angle of the sound signals obtained by the two microphones 21 and 22.

[0031] Step S04: Obtain at least one image from the lens modules 11-14 that are close to the at least one sound source, which contains at least one speaker.

[0032] Next, the processor 30 acquires at least one image from the lens modules 11-14 located near the sound source. Taking Figure 2A as an example, the sound source locations are in azimuth space S1 and near microphone 22, and in azimuth space S2 and near microphone 23, respectively, and are adjacent to lens modules 11 and 12. Therefore, in step S04, the processor 30 acquires two image images from lens modules 11 and 12 in azimuth space S1 and azimuth space S2. Since the processor 30 acquires the image images based on the sound source location, the image images contain at least one speaker.

[0033] Taking Figure 2B as an example, the sound source is located in the azimuth space S2 and is close to the microphone 23, which is adjacent to the lens module 12. Therefore, in step S04, the processor 30 obtains an image of the azimuth space S2 from the lens module 12. Since the processor 30 obtains the image based on the sound source location, the image contains at least one speaker.

[0034] It should be noted that after the panoramic image capturing device 1 is powered on, the processor 30 continuously receives image frames from the lens modules 11-14 in the azimuth spaces S1, S2, S3, and S4, and continuously receives sound signals from the microphones 21-28 in the space S. If there is a speaker in the space S (e.g., a main speaker P1 or a soft-spoken speaker P2), an image frame containing at least one speaker can be obtained in step S04. In other words, step S04 defines obtaining an image frame containing at least one speaker, but in practical applications, the processor 30 continuously receives image frames from the azimuth spaces S1, S2, S3, and S4 and sound signals from the space S.

[0035] Step S05: Analyze and compare the sound signals with the at least one image to calculate an integral value of the speaker.

[0036] In this embodiment, the processor 30 can analyze all the image frames to obtain each person's image in each image frame. Taking Figure 2A as an example, the image frame in orientation space S1 contains five people's images. Furthermore, the processor 30 can perform image recognition on the facial movements of the people's images and can simply determine that there is one speaker (i.e., the main speaker P1) and four ordinary people P3 without mouth movements. Also, the image frame in orientation space S2 also contains five people's images, but one of them (i.e., the main speaker P1) only has a body part and his / her facial movements cannot be identified. The processor 30 can simply determine that the image frame in orientation space S2 contains two speakers (i.e., soft-spoken speakers P2) and two ordinary people P3 without mouth movements. The image frame in orientation space S3 contains three people's images, all of whom are ordinary people P3 without mouth movements. The image frame in orientation space S4 contains four people's images, all of whom are ordinary people P3 without mouth movements.

[0037] In this embodiment, each speaker (i.e., those identified as having mouth movements) can be assigned an initial score, and weighted scoring conditions can be set. Therefore, a primary speaker P1 in the directional space S1 image and two soft-spoken speakers P2 in the directional space S2 image each have an initial score. Regarding the weighted scoring conditions, the processor 30 can first analyze the sound signal, assess whether the average volume of the sound signal is greater than a threshold value, and the processor 30 can compare the image and the sound signal, and determine whether the facial movements of the speaker in the image match the sound signal. When the average volume of the sound signal is greater than a threshold value, the processor 30 can add a first weight score. Furthermore, the processor 30 performs image recognition on the facial movements in the image. When the facial movements match the sound signal, the processor 30 can also add a second weight score. In this embodiment, the sum of the initial score, the first weight score (if any), and the second weight score (if any) is used as the speaker's score.

[0038] Specifically, the processor 30 can determine, based on the configuration of the lens modules 11-14 and the microphone array 20, whether the sound signals generated by the microphones 21, 22, and 23 contain the voice of the speaker (i.e., the main speaker P1) in the directional space S1, and determine whether the average volume of the sound signals from the microphones 21, 22, and 23 is greater than a threshold value. For example, if the microphone 22 is closest to the speaker (i.e., the main speaker P1) in the directional space S1, the average volume of the sound signal from the microphone 22 is greater than a threshold value, and a first weight score can be added. However, if the average volume of the sound signals from the microphones 21 and 23 is less than the threshold value, no first weight score is added. Next, the processor 30 compares the facial movements of the speaker (i.e., the main speaker P1) in the image frame of the directional space S1 with the sound signal from the microphone 22 with the highest volume. For example, if the timing of the sound signal from the microphone 22 matches the mouth movements of the speaker (i.e., the main speaker P1) in the image, a second weight score can be added. Taking Figure 2A as an example, the integral value of the speaker (i.e. the main speaker P1) in the image of the azimuth space S1 is the sum of the original score, the first weight score and the second weight score.

[0039] Similarly, the processor 30 can determine, based on the configuration of the lens modules 11-14 and the microphone array 20, whether the sound signals generated by the microphones 22, 23, and 24 contain the voice of the speaker in the directional space S2 (i.e., the soft-spoken speaker P2), and then determine whether the average volume of the sound signals from the microphones 22, 23, and 24 is greater than a threshold value. Since the speaker in the directional space S2 is the soft-spoken speaker P2 who is chatting with each other at a low volume, the average volume of the sound signals from the microphones 22, 23, and 24 is still less than the threshold value, and the first weight score cannot be increased. Similarly, the processor 30 compares the facial movements of the speaker (i.e., the soft-spoken speaker P2) in the image frame of the directional space S2 with the sound signal from the microphone 22 with the highest volume. Since the sound signal from the microphone 22 does not match the facial movements of the speaker (i.e., the soft-spoken speaker P2) in the image frame of the directional space S2, the second weight score cannot be increased. Taking Figure 2A as an example, the integral value of the speaker (i.e., the soft-spoken speaker P2) in the image of the azimuth space S2 is the original score.

[0040] Taking Figure 2B as an example, the processor 30 can analyze the image frames of directional spaces S1, S2, S3, and S4, and perform simple image recognition on the facial movements of the people in the image frames. Specifically, the image frame of directional space S1 contains four ordinary people P3 who do not make mouth movements. The image frame of directional space S2 contains three speakers (i.e., the main speaker P1 and the soft-spoken speaker P2) and two ordinary people P3 who do not make mouth movements. The image frame of directional space S3 contains three ordinary people P3 who do not make mouth movements. The image frame of directional space S4 contains four ordinary people P3 who do not make mouth movements.

[0041] Similarly, each speaker (i.e., those identified with mouth movements) is assigned an initial score, and weighted scoring conditions are set. Therefore, the three speakers in the directional space S2 image (i.e., the main speaker P1 and the soft-spoken speaker P2) each have an initial score. Next, the processor 30 analyzes the sound signals to evaluate whether to increase the first weight score and the second weight score (i.e., the weighted scoring conditions). Specifically, the processor 30 can determine from the configuration of the lens modules 11-14 and the microphone array 20 that the sound signals generated by the microphones 22, 23, and 24 contain the voices of the speakers in directional space S2 (i.e., the main speaker P1 and the soft-spoken speaker P2), and determine whether the average volume of the sound signals from the microphones 22, 23, and 24 is greater than a threshold value. For example, if the microphone 23 is closest to the speaker in directional space S2 (i.e., the soft-spoken speaker P2 and the main speaker P1), the average volume of the sound signal from the microphone 23 is greater than the threshold value, and the first weight score can be increased. The average volume of the sound signals from microphones 22 and 24 is less than the threshold value, so the first weight score is not increased.

[0042] Next, the processor 30 compares the facial movements of the speakers (i.e., the main speaker P1 and the soft-spoken speaker P2) in the image frame of the azimuth space S2 with the sound signal from the microphone 23 with the highest volume. Taking Figure 2B as an example, the timing of the sound signal from the microphone 23 matches the mouth movements of one speaker behind in the image (i.e., the main speaker P1), but does not match the mouth movements of the two speakers in front (i.e., the soft-spoken speaker P2). Therefore, the speaker behind (i.e., the main speaker P1) can gain a second weight score; while the speaker in front (i.e., the soft-spoken speaker P2) cannot gain a second weight score. Taking Figure 2B as an example, the integral value of one speaker behind in the image frame of the azimuth space S2 (the main speaker P1) is the sum of the original score, the first weight score, and the second weight score, while the integral value of the two speakers in front (i.e., the soft-spoken speaker P2) is the sum of the original score and the first weight score.

[0043] Preferably, regarding weighted scoring, the processor 30 in this embodiment first determines the average volume of the sound signal, and then compares the image and the sound signal. In other embodiments, the processor 30 may also first compare the image and the sound signal, and then determine the average volume of the sound signal; this invention is not limited in this respect. In other words, this invention does not limit the order in which the first weighted score and the second weighted score are calculated, and they can also be performed simultaneously.

[0044] Step S06: Select a main speaker based on a highest score.

[0045] The processor 30 selects the highest score from the scores of each speaker calculated in step S05, which is referred to as the highest score in this embodiment. Taking Figure 2A as an example, the score of the speaker in the image frame of azimuth space S1 is the sum of the original score, the first weighted score, and the second weighted score, which is the highest score. In this embodiment, the speaker corresponding to the highest score is also determined by the processor 30 to be the speaker who is giving an opinion in the meeting, i.e., the main speaker P1.

[0046] Taking Figure 2B as an example, the score of a speaker in the image of the azimuth space S2 is the sum of the original score, the first weight score and the second weight score, which is the highest score, and is determined by the processor 30 to be the main speaker P1.

[0047] Step S07: Select one of the lens modules 11 to 14 as a reference lens module, the field of view of the reference lens module covering the main speaker.

[0048] Next, taking Figure 2A as an example, the processor 30 selects lens module 11, whose field of view is azimuth space S1, as the reference lens module from the plurality of lens modules 11 to 14. Taking Figure 2B as an example, the processor 30 selects lens module 12, whose field of view is azimuth space S2, as the reference lens module from the plurality of lens modules 11 to 14. Therefore, the field of view corresponding to the reference lens module (i.e., lens module 11) covers the main speaker P1 (i.e., the speaker corresponding to the highest integral value).

[0049] It should be noted that in the examples of Figures 2A and 2B, only the field of view of lens module 11 or lens module 12 covers the entirety of the main speaker P1, so lens module 11 or lens module 12 is directly selected. In other embodiments, if the field of view of multiple lens modules (e.g., lens modules 11 and 12) covers the entirety of the main speaker P1, the distance between lens modules 11 and 12 and the main speaker P1 can be further analyzed, and the lens module closest to the main speaker P1 (i.e., one of lens modules 11 and 12) can be selected.

[0050] Step S08: Obtain an autofocus value for the main speaker from the reference lens module.

[0051] As mentioned above, the lens modules 11 to 14 in this embodiment are all active focusing lens modules. Therefore, the reference lens module (i.e., lens module 11 or lens module 12) can perform an autofocus program on the main speaker P1 (i.e. the speaker corresponding to the highest score) to obtain a focus value and use it as the autofocus value of this embodiment.

[0052] Specifically, the memory unit 31 of the processor 30 in this embodiment stores a focus parameter table. The focus parameter table includes multiple focus distances and their corresponding multiple focus values. Furthermore, the lens modules 11-14 in this embodiment each include a pair of focus motors (not shown). The focus distance is the distance from the lens modules 11-14 to the object being focused on, and the focus value can be a number of drive steps of the focus motors.

[0053] After the panoramic image capturing device 1 of this embodiment is powered on, the lens modules 11-14 acquire image images with different field of view ranges, and simultaneously acquire the focusing distance between each lens module 11-14 and each person in space S. Taking lens module 11 as an example, the focusing distance between lens module 11 and the main speaker P1 and four ordinary persons P3 in the directional space S1 can be acquired respectively. In step S06, the main speaker P1 is found based on the highest integral value. In step S07, a reference lens module (i.e., lens module 11 or lens module 12) is selected. Then, in step S08, the reference lens module (i.e., lens module 11 or lens module 12) performs an autofocus program on the main speaker P1, and the processor 30 obtains the focusing value of the main speaker P1 (e.g., the number of drive steps of the focusing motor) according to the focusing parameter table, and defines it as the autofocus value.

[0054] Step S09: Adjust the lens modules 11 to 14 according to the autofocus value so that the lens modules 11 to 14 have the same focal length.

[0055] The processor 30 adjusts the focal length of all lens modules 11 to 14 according to the autofocus value obtained in step S07, so that the lens modules 11 to 14 respectively obtain image images with the same focal length.

[0056] Step S10: Obtain an adjusted image from each of the lens modules 11 to 14 and stitch them together to form a panoramic image.

[0057] The processor 30 acquires an adjusted image from each of the lens modules 11-14, meaning an image adjusted to have the same focal length. Finally, the processor 30 stitches the adjusted image from the lens modules 11-14 into a panoramic image. Since the adjusted image has the same focal length, the inability to stitch images due to different focal lengths is avoided, and the resulting panoramic image will not be distorted or blurry due to stitching.

[0058] In summary, according to the active focusing method of the present invention and the panoramic image capturing device thereof, by analyzing and comparing the sound signal and the image, one of the multiple lens modules of the panoramic image capturing device is selected as a reference lens module, and the field of view corresponding to the reference lens module covers a main speaker. An autofocus value of the reference lens module for the main speaker is obtained, and all lens modules are adjusted according to the autofocus value so that all lens modules have the same focal length. Finally, the lens modules respectively acquire adjusted image images with the same focal length and stitch them together to form a panoramic image. Since the adjusted image images have the same focal length, the situation where images cannot be stitched due to different focal lengths can be avoided, and the resulting panoramic image will not be distorted or blurred due to stitching. Therefore, the panoramic image capturing device using the active focusing method of the present invention can use an active focusing lens, that is, the aforementioned lens module can be an active focusing lens module.

[0059] It should be noted that the above embodiments are examples for ease of explanation, and the scope of the claims made by the present invention should be determined by the claims of the patent application, and not limited to the above embodiments. [Simplified Explanation of the Diagram]

[0015] Figure 1 is a top view of a panoramic image capturing device according to an embodiment of the present invention. Figures 2A and 2B are schematic diagrams of the usage environment of the panoramic image capturing device shown in Figure 1. Figure 3 is a block diagram of the panoramic image capturing device shown in Figure 1. Figures 4A and 4B are flowcharts of an active focusing method according to an embodiment of the present invention.

Claims

1. An active focusing method applied to a panoramic image capturing device, comprising a plurality of lens modules, a microphone array, and a processor, the active focusing method comprising the following steps: the processor receives a plurality of sound signals in a space from the microphone array; the processor obtains the orientation of at least one sound source based on the sound signals; the processor obtains at least one image frame of the lens modules near the orientation of the at least one sound source, the at least one image frame containing at least one speaker; the processor analyzes and compares the sound signals with the at least one image frame to calculate an integral value of the at least one speaker and selects a main speaker based on a highest integral value, and selects one of the lens modules as a reference lens module, the field of view of the reference lens module covering the main speaker; the processor obtains an autofocus value of the reference lens module for the main speaker; the processor adjusts the lens modules according to the autofocus value so that the lens modules have the same focal length; and the processor obtains an adjusted image frame from each of the lens modules and stitches them together to form a panoramic image frame.

2. The active focusing method as described in claim 1, wherein the microphone array includes a plurality of microphones spaced apart in the panoramic image capturing device, wherein the active focusing method further includes the following steps: the processor divides the space into a plurality of directional spaces according to the field of view corresponding to the lens modules; and the microphones respectively pick up sound in the directional spaces to obtain the sound signals.

3. The active focusing method as described in claim 1, wherein each of the sound signals includes azimuth information, and the processor obtains the azimuth of the at least one sound source based on an average volume of each of the sound signals and the azimuth information.

4. The active focusing method as described in claim 1, wherein when an average volume of the sound signals is greater than a threshold value, the processor adds a first weight score to obtain an integral value.

5. The active focusing method as described in claim 1, wherein the processor performs image recognition on a facial movement in at least one image frame, and when the facial movement matches the sound signal, the processor adds a second weight score to obtain an integral value.

6. The active focusing method as described in claim 1, wherein the processor includes a memory unit storing a focusing parameter table, the reference lens module executes an autofocus program, and the processor obtains the autofocus value based on the focusing parameter table.

7. The active focusing method as described in claim 1, wherein the microphone array includes a plurality of microphones, which are at least partially located between two adjacent lens modules.

8. A panoramic image capturing device, comprising: A plurality of lens modules; a microphone array; The system also includes a processor electrically connected to the lens modules and the microphone array, and performs the following steps according to an active focusing method: receiving multiple sound signals from the microphone array within a space; obtaining the location of at least one sound source based on the sound signals; obtaining at least one image frame from the lens modules near the location of the at least one sound source, the at least one image frame containing at least one speaker; analyzing and comparing the sound signals and the at least one image frame to calculate an integral value for the at least one speaker and selecting a primary speaker based on a highest integral value, and selecting one of the lens modules as a reference lens module, the field of view of which covers the primary speaker; obtaining an autofocus value for the primary speaker from the reference lens module; adjusting the lens modules according to the autofocus value so that the lens modules have the same focal length; and obtaining an adjusted image frame from each of the lens modules and stitching them together to form a panoramic image frame.

Citation Information

Patent Citations

  • Panoramic camera

    CN107105135A

  • Conference video processing method and system

    CN116016836A